The bench · five models sit the frozen exams on a 24 GB card
The Instrument Travels
exhibit thirty-one The bench
Published 2026-08-29 (UTC)
a small (human) team and a fleet of AI agents — we set the exams, ruled the calls and signed the numbers; the harness ran the legs itself
main measurement window 2026-08-26, 07:00–15:13Z (UTC): the main run stopped at 09:58Z, a gap-fill lane ran on to 10:14Z when the trigger fired again, and scoring resumed at 13:55Z
A rules question comes in — a table of friends, mid-game, one of them typing 'can I build on a card I already played this turn?' — and a model we picked by exam answers it. Two weeks ago we picked on a 96 GB workstation card. This time we ran the very same frozen exam — same items, same seeds, same pass marks — on a 24 GB consumer card, to find out whether it would say the same thing in a smaller room. Five models sat it: two we had examined before, re-pulled at the tags as they stood on 2026-08-26; two newcomers; one we had blocked and wanted to re-check. One reproduced its published verdict on a different build and different silicon, which is the whole point of freezing an exam. The other four returned four different verdicts — a seat blocked by a stray byte, a judge that caught every planted error and failed anyway, two models that would not board the card, and a row that owes a re-run. Nothing earned a seat, and the bench added a fifth result by stopping itself at a disk floor rather than gambling on the margin.
The chair, and who gets to sit in it
Somewhere a rules question comes in — a table of friends, mid-game, one of them typing "can I build on a card I already played this turn?" — and the answer comes from a model sitting in a chair we gave it after an exam. Every few weeks the chairs get contested: a new model ships, a familiar one gets a new build under the same name, and someone asks whether the newcomer would answer that question better. The only fair way to find out is to make it sit the same exam the incumbent sat, and to trust nothing the exam did not see.
The question is never "is it good?" — it is "does it earn a chair, on hardware a normal person owns, under the same exams everything else passed?" What is different this time is the room. Two weeks ago we published our ranked verdicts, measured on a 96 GB workstation card; this battery ran on a 24 GB consumer card of the kind sitting in a gaming PC today. So it asked two questions at once: do any newcomers earn a chair — and when models we examined before are re-pulled and re-examined in a smaller room, does the frozen instrument say the same thing? If it does not, every ranking we publish is a photograph of a moment. If it does, the instrument travels.
What we ran, and the rule behind every number
Five candidates on a single 24 GB RTX 3090, one at a time, that card emptied for the run:
- two repro rows — qwen3.6:27b and nemotron-3.5-lightning:30b, examined on 2026-08-12 and re-pulled at the tags as they stood on 2026-08-26. Their earlier verdicts were not the same: qwen cleared both judge floors and was RANKED, nemotron missed both (9 of 12 kills, 6 of 9 preservations) and its toolbench did not carry. Only one is a new build — qwen's manifest digest changed since 2026-08-12, nemotron's is byte-identical, so a nemotron row is a harness-and-hardware check, not a model comparison.
- two new kids — gemma4:31b and laguna-xs-2.1:latest.
- one re-probe — muse-glimmer:30b, first probed two weeks ago and seat-blocked then. The earlier chair trials ranked glimmer's Q8 build; this battery pulled the default tag.
The box's second card kept serving live traffic throughout — a separate card with its own memory, and the checking is in the kit (the receipts published beside this page) rather than in our word for it: the lane log records a read-only probe of all three live seats at nine checkpoints, all five fit receipts carry prod_untouched: true, and the teardown receipt carries the last probe of all.
Candidates run as shipped: the default published tag, the vendor's own quantisation — a fit verdict is a verdict on that shipped artifact, not on the architecture. Every count below is copied from that leg's own scorer file — a leg is one exam run for one model — and the page computes nothing. The fit gate loads each model at a 32,768-token context: residency is a fact about weights plus a working context, and 32k is what our seats run.
One rule governs every comparison here. A count — kill-recall, preservation, task success, items right — is a property of the model, and travels across cards. A rate — tokens a second — is a property of the model and the silicon, and does not. So the counts below are compared to the 96 GB run and the speeds are not, ever.
The run itself is part of the data, told plainly. The battery stopped itself at 09:58Z, mid-leg, when free space on the bench box's disk fell below the harness's 60 GiB safety floor. The stop was not idle time, and the page has to say so because three published cells were measured inside it: with the battery's own units standing down, a gap-fill lane re-ran the re-probe's outstanding legs — glimmer's schema probe, assistant trial and toolbench, finishing at 10:00Z, 10:06Z and 10:14Z, and those three runs are the glimmer cells below. One second later, at 10:14:59Z, the speed leg tried to start, the trigger fired again, and that was the end of it.
We cleared the disk, restarted the bench process (the lane log records the new process ID, so a reproducibility claim spanning the halt spans a restart too), re-verified all five weight digests — 5 of 5 MATCH — resumed at 13:55Z, and time-boxed the remainder: qwen ran to completion, nemotron ran its fit gate only, which turns out to be the only result it needed, and we dropped laguna's remaining legs. When the battery sealed we tore its model store down to 16 KB under four guards: one host and never production, results verified present before any delete, named tags only and no wildcards, before/after receipts. Weights are re-pullable by design; measurements are not. That teardown ran at 18:13Z, after the window closed — bookkeeping rather than measurement.
The words these tables use, and the floors they score against
- A seat, which we also call a chair — a model slot our products actually serve from: RuleSage, our rules assistant, and its sibling games and tools. Every seat here runs on Ollama 0.32.13 at think:false. A seat is the prize; the things a model sits to win one are exams.
- Posture — whether the model's hidden reasoning channel is off (think:false, the seat contract) or on (think:true). "Held both ways" means the probe was run in both.
- The exams — a frozen battery: a schema-discipline probe (C0), a judge trial with pre-registered floors (C1), an assistant trial (C2), a tool-use trial (C3), speed (C5), a long-context filing test (C7), a 60-item field exam in three 20-item tasks, and the judge seat's own trial (43 cases × 3 calls = 129). The C-numbers are the harness's own labels, kept because they are the filenames in the kit; the numbering has gaps because the harness is older than this roster, and C4, the narrator's chair, is excluded by operator ruling. Frozen means same items, same seeds, same floors, same reply budgets as the published runs.
- Kill-recall and preservation — the two counts a judge trial scores. We plant known errors in the material, and kill-recall is how many the model caught. We also leave correct material alone in there, and preservation is how much of that it left standing. A judge that flags everything scores a perfect kill-recall and a terrible preservation, which is exactly why both floors have to clear.
- RANKED — a clean, comparable run.
- EXPLORATORY — real numbers, fenced: printed, but walled off from any ranking.
- DESCRIPTIVE — numbers reported, with no registered floor to pass or fail.
- UNMEASURABLE — valid attempts exist, but the measurement-failure rate crosses 10% and the exam refuses to convert what remains into a verdict. The judge-seat exam refuses the counts outright; the tools trial prints them fenced as inadmissible.
- SEAT-BLOCKED — the model works, but something in what it returns breaks the contract our seats require.
- Strict parse binds — a caller who requested a JSON schema and got unparseable bytes has a failed contract whatever the reason. Diagnosis can be sympathetic; the verdict is not.
- The empty cells — four reasons leave a cell without a score and the tables keep them apart: NOT-CARRIED (valid attempts came back zero, nothing to score), NOT-RUN (spill) — the model did not fit the card whole, and we do not run speed or long-context through a half-CPU model, which is a rule rather than a judgment call — NOT-RUN (stop) for the disk halt and NOT-RUN (trim) for the time-box we set. One cell is none of the four: laguna's judge trial was cut mid-leg, with real calls already made, and it says so rather than pretending the leg never started. A row's first empty reason propagates rightward.
The floors, pre-registered before any scoring. The judge trial passes at kill-recall ≥ 11 of 12 AND preservation ≥ 8 of 9; the judge seat's trial at kill-recall ≥ 23 of 27 AND preservation ≥ 13 of 16. A cross-run difference under 2 items reads TIED. Past a 10% measurement-failure rate an exam is UNMEASURABLE. The registration fixing all of this ships in the kit with its sha256, 423ed960…, and so does the 21-item judge fixture it counts, fede4e32….
One exam runs at settings that are not this battery's, and it changes how a column reads. The judge seat's 43-case trial is an older, separately frozen exam and runs at its own registered settings — a 16,384-token context and no think flag at all — where every other instrument here runs 32,768 with think:false explicit. We ran it as published rather than re-tuned, because its floors of 23/27 and 13/16 only mean anything against the run they were set on. Its counting rule differs too: this battery's judge trial puts a truncation against the reply budget in its own run-quality column, outside the 10% ceiling, while the seat exam's own frozen rule counts a truncation as a response failure. Where the two treat the same failure differently below, the difference is the exam, not the model.
The standings
Every cell is copied from that leg's own scorer file. Sizes are the fit receipts' own MiB, read with a 32,768-token context already loaded — a loaded footprint, not a shipped weight file.
The schema probe (C0) comes first because it gates the rest: qwen3.6:27b and gemma4:31b held the schema both ways, muse-glimmer:30b is SEAT-BLOCKED, laguna-xs-2.1:latest returned the battery's one inverted trap, and nemotron-3.5-lightning:30b never sat it. Both findings have their own section below.
Table A — the door: the fit gate, at a 32,768-token context
| candidate | role | boards whole | on GPU (%) | loaded footprint (MiB) | on GPU (MiB) | spilled to CPU (MiB) |
|---|---|---|---|---|---|---|
| qwen3.6:27b | repro | yes | 100 | 16,469 | 16,469 | 0 |
| muse-glimmer:30b | re-probe | yes | 100 | 15,831 | 15,831 | 0 |
| laguna-xs-2.1:latest | new kid | yes | 100 | 19,440 | 19,440 | 0 |
| gemma4:31b | new kid | no | 92.5 | 20,552 | 19,011 | 1,541 |
| nemotron-3.5-lightning:30b | repro | no | 83.1 | 24,434 | 20,298 | 4,136 |
Table B — the judge trial (C1): 21 items × 3 repeats = 63 calls
| candidate | posture | kill-recall (of 12, floor 11) | preservation (of 9, floor 8) | calls scored clean (of 63) | calls truncated | verdict |
|---|---|---|---|---|---|---|
| qwen3.6:27b | think:false | 11 | 8 | 60 | 3 | RANKED |
| muse-glimmer:30b | reasoning-strength medium | 11 | 7 | 63 | 0 | EXPLORATORY |
| muse-glimmer:30b | reasoning-strength high | 12 | 6 | 52 | 10 | EXPLORATORY |
| gemma4:31b | think:false | not scored | not scored | 0 | 0 | NOT-CARRIED |
| laguna-xs-2.1:latest | think:false | not scored | not scored | 19 | 0 | cut mid-leg |
| nemotron-3.5-lightning:30b | not sat | not scored | not scored | 0 | 0 | NOT-RUN (trim) |
Glimmer's high row also lost one call to a transport failure (52 + 10 + 1 = 63), and both its rows are fenced by the C0 verdict whatever the counts say. gemma4's 63 calls were all response failures. Laguna's 19 calls landed clean across 7 of the 21 items before the stop cut the leg; they ship in the kit, and they are not a result.
Table C — the judge seat's own 43-case exam: 129 calls per candidate
| candidate | kills (of 27, floor 23) | preservation (of 16, floor 13) | calls that failed to score (of 129) | replies unwrapped from a code fence | verdict |
|---|---|---|---|---|---|
| gemma4:31b | 27 | 9 | 12 | 117 | FAIL |
| muse-glimmer:30b | refused by the exam | refused by the exam | 18 | 0 | UNMEASURABLE |
| qwen3.6:27b | refused by the exam | refused by the exam | 93 | 0 | UNMEASURABLE |
laguna and nemotron never reached this exam — NOT-RUN (stop) and NOT-RUN (trim). The two UNMEASURABLE rows do have kill and preservation counts in the scorer file; the exam refused them at the ceiling, so we do not print them as results.
Table D — the exams with no registered floor: counts only, no pass or fail
| candidate | assistant, C2 (of 20) | tools task success (of 19) | tools outcome | field exam (of 60) | filing recall (of 18) | filing abstentions (of 18) | filing fabrications |
|---|---|---|---|---|---|---|---|
| qwen3.6:27b | 15 | 14 | RANKED | 58 | 16 | 18 | 0 |
| muse-glimmer:30b | 16 | 13 | UNMEASURABLE | 55 | 18 | 18 | 0 |
| gemma4:31b | 16 | 12 | RANKED | 58 | NOT-RUN (spill) | NOT-RUN (spill) | NOT-RUN (spill) |
laguna and nemotron sat none of these — NOT-RUN (stop) and NOT-RUN (trim). The tools trial runs 19 frozen tasks twice each, and glimmer failed 6 of those 38 attempts (15.8%), over the 10% ceiling: its 13 is fenced as inadmissible, and so are all 19 of its tasks, not 6 of them. The filing exam asks 36 filings twice each, 72 calls, and neither candidate that sat it failed a single response.
Table E — speed (C5): median decode rate, one real prompt loaded at each context size
| candidate | prompt size (tokens) | median decode (tok/s) | calls counted | calls with null counters |
|---|---|---|---|---|
| qwen3.6:27b | 1,000 | 62.0 | 13 | 0 |
| qwen3.6:27b | 8,000 | 60.8 | 10 | 0 |
| qwen3.6:27b | 32,000 | 59.1 | 10 | 0 |
| muse-glimmer:30b | 1,000 | 41.1 | 13 | 1 |
| muse-glimmer:30b | 8,000 | NOT-CARRIED | 0 | 10 |
| muse-glimmer:30b | 32,000 | 39.5 | 9 | 1 |
gemma4 is NOT-RUN (spill) at every tier, laguna NOT-RUN (stop), nemotron NOT-RUN (trim). The measured prompts run 996 to 31,996 tokens, a 32× spread, so the near-flat curve (62.0 → 59.1 tok/s) is a finding about attention cost on this card, not an unfilled tier.
What the battery was for: the instrument travels
On 2026-08-12, two weeks before this battery, qwen3.6:27b earned RANKED on the judge trial — kill-recall 11/12, preservation 9/9 — on the 96 GB card. This battery re-pulled the same tag and got a different build — a different manifest digest, 9d5803d4… where 2026-08-12 recorded a50eda8e… — on different silicon, and the frozen exam read: kill-recall 11/12, preservation 8/9 — both floors cleared, RANKED again.
The margin earns a plain sentence, because the 8/9 is not a judgment the model got wrong. Of 63 calls, 60 scored clean, none failed the response contract, and three hit the 1,024-token reply budget — and those three were all three repeats of a single preservation item, so that item carries no verdict at all. qwen preserved eight of the eight items it was scored on, and the ninth was unscorable rather than lost. Against 2026-08-12's 63 clean calls and 9 of 9, the preservation delta is one item, under the pre-registered two-item tie band, and reads TIED — the whole distance between "cleared at the floor" and "clean repeat" is those three truncations. Self-consistency across repeats moved the same single item: 21 of 21 on 2026-08-12, 20 of 21 here.
Honesty about the claim itself: that is one row, and the second repro row was time-boxed to its fit gate — and its 2026-08-12 judge row had failed both floors anyway, so the sibling this claim awaits is a second passing row, not nemotron's. "Travels" is an n=1 result, not yet a law.
The detail that makes this measurement rather than memory: the kill-recall count is identical but the missed item moved — the earlier build missed c1-k-bez-c1, this battery's build missed c1-k-bez-h1: a different case in the same family, landing on the same score. A frozen instrument catching a moving model in a slightly different place, at the same reading, is what reproducibility actually looks like.
Two other instruments sat both runs, and counts travel, so both are fair to compare. The assistant trial — descriptive, but frozen — reproduced exactly: 15 of 20, 11 of the 13 items with an exact check, 4 of the 7 scored by the registered proxy rule, zero response failures, the same in both runs on a different build and a different card. The toolbench went the other way: 16/19 task success on 2026-08-12, 14/19 in this battery, grounded — passed and made every pre-registered call — 14/19 → 12/19, both RANKED and admissible, the same 19 frozen tasks. A two-item move is not under the tie band, so we print it as a real move rather than noise, and it is a build-to-build move rather than a card-to-card one, because counts do not care about silicon. One instrument reproduced exactly, one reproduced at the floor, one slipped. That is what "travels" honestly looks like at n=1.
Schema-perfect, and still blocked
We first wrote muse-glimmer:30b's schema-probe finding down wrong, and the correction earns its paragraph. On the long-context filing exam glimmer was flawless — 18/18 recall, 18/18 correct abstentions, zero fabrications across 72 calls. And it is SEAT-BLOCKED, because under an enforced JSON schema it emits schema-perfect JSON — and then the runtime leaks a trailing end-of-text sentinel into the bytes, which breaks a strict parser. Our probe first said "can't hold a schema" while the field exam's own 20-item schema-extraction task scored it 19/20; two instruments, same model, opposite verdicts — so we read the raw bytes, and the bytes overruled both summaries. The verdict stands (strict parse binds: a caller who asked for a schema got unparseable bytes), but the diagnosis inverts — a one-line fix at the caller, not a model to discard. A blocked seat and a broken model are different facts, and the exam records which one it saw.
The judge's chair scored one candidate of three
Three candidates sat the judge seat's 43-case trial; the exam scored exactly one. gemma4:31b caught all 27 planted kills — clearing that floor — and failed anyway: preservation 9/16 against a floor of 13. A judge that fails almost everything will of course catch every planted error; the trial exists precisely to price that.
One thing about that verdict belongs on the page rather than only in the kit. 117 of gemma4's 129 replies arrived wrapped in a code fence and were unwrapped before scoring. That exam asks for JSON in the prompt rather than through a schema grammar, so unwrapping is what its own frozen rule does, and strict parse binds where a schema was sent — which this exam never sends. It is still worth printing that the one candidate the exam scored is the only one of the three that needed it: the other two arrived at zero fenced replies each.
The other two came back UNMEASURABLE: glimmer at 18 of 129 calls failing to score (14%), qwen at 93 of 129 (72%) — truncations against that exam's own 1,024-token reply budget, in its own 16,384-token context. A verbose model meeting a short cap, refused by the exam rather than converted into a score. Under this battery's judge trial the same failure mode sits outside the ceiling entirely; under the seat exam's older rule it crosses it. The difference is the exam.
The same trade appeared inside glimmer's judge trial, where raising its reasoning-strength dial from medium to high bought a perfect kill sweep (11/12 → 12/12) and paid for it in preservation (7/9 → 6/9). That dial is a line in the system prompt, not Ollama's think flag: both rows went out at think:false, the seat contract, and the kit's per-call records carry the wire settings.
Nothing was seated this battery. qwen3.6:27b boarded the card whole and returned a ranked or clean result on seven of the eight exams it sat; the eighth — the judge seat — was UNMEASURABLE under the frozen budget. We intend a re-run, and the frozen-instrument rule binds us too: the reply budget is part of the exam, gemma4's FAIL was measured under the same 1,024 tokens, so a changed budget would be a new exam version applied to every candidate. We would publish such a re-run fenced from the published floors — not as a seating run.
What a 24 GB door actually filters
Three of five board the card whole at 32k — 15,831, 16,469 and 19,440 MiB resident, 100% on GPU. Two do not: gemma4:31b runs 92.5% on GPU and spills 1,541 MiB, and nemotron-3.5-lightning spills 4,136 MiB, running 83.1% on GPU. Nemotron's shipped weights are 25.43 GB (23.68 GiB) against a card holding 24.0 GiB, so once a 32,768-token context is loaded beside them there is nowhere left to sit and 4,136 MiB goes to the CPU. That shipped size is read from the runtime's own tag listing, and the same runtime reports different sizes for byte-identical builds across versions — so the digest, not the size, identifies a build here. The kit records that anomaly as an open flag, unresolved as it ships.
Both spills are published as the finding, speed and long-context rows skipped rather than measured through them — and both are verdicts on the artifact as shipped: repacked smaller, either might board, and that would be a different candidate. gemma4:31b still posted the joint-best field exam of the battery (58/60, tied with qwen) — a good model this card cannot hold whole is a different fact from a bad model, and the tables keep them apart.
The row that owes a re-run
laguna-xs-2.1:latest fits whole (19,440 MiB) and showed the battery's only inverted trap. Under think:false — the exact posture our seats contract — its schema output parses clean; under think:true it drops the schema constraint. That is the inversion: glimmer, and the documented majority, fail the other way round, so a caller who "fixed" a schema problem here by turning thinking on would be breaking it. Whether that lives in the model's schema discipline or simply in the runtime's leak not firing on laguna's template is exactly what its full battery would have said.
It lost that battery twice, first to the disk stop mid-leg, then to the time-box we set. What survives is 19 valid calls across 7 of the judge trial's 21 items, made and never scored. Its row stays honestly empty, and it is first in line for the next window.
What this page does not say
- No seats were awarded. qwen's claim covers seven exams and is incomplete by the eighth; the re-run rule above says what would make it complete, for every candidate equally.
- One tier is a build-specific anomaly, resolved by a pre-named test: glimmer's 8k speed tier returned empty counters on all ten attempts while qwen's returned ten clean readings on the same harness, card and prompt construction. The anomaly belongs to the glimmer build, and it is recorded NOT-CARRIED.
- Both speed rows are recoveries, and the timed work is intact. Both C5 summary passes crashed after their timed calls — glimmer on a manifest-shape bug of ours, qwen on a hardcoded byte-compare 404 in the harness. The medians here are recomputed from the per-call records under the frozen counting rules, and the kit ships the file that says so.
- The speed numbers were measured on a box that was also serving. The live seats ran on the other card, but the host is shared, so the timed calls could carry contention. The check is in the raw: each candidate's 1k tier opens with a three-call clean baseline, and it lands within 0.3% of the ten timed calls that follow for qwen (62.1 against 62.0 tok/s) and 2.4% for glimmer (41.9 against 40.9). The published 1k figure is the median of all thirteen.
- One battery, one card, build-pinned digests. Every verdict names the build it measured; a tag is not a model — which is exactly why the repro rows exist.
- Speeds are not compared to the 96 GB run, by the count-travels rule above; a reader who diffs them against the earlier page is measuring silicon, not models.
What to take with you
Five things, in the page's own words:
- Counts travel; rates don't. A frozen exam's correctness counts are a property of the model and can be compared across cards; tokens-a-second belong to the model and the silicon and are new measurements every time.
- The instrument travelled at n=1, and all three readings are printed. One trial reproduced exactly, one reproduced at the floor with the missed item moved to a sibling case, and one slipped by two items, which the tie band does not cover.
- A 24 GB door filters by artifact, not architecture. Two of five spilled as shipped; one of those posted the joint-best field exam. Repacked smaller they might board — and that would be a different candidate.
- Frozen means frozen, including the exam you inherited. The judge seat's trial runs at its own registered settings and its own counting rule, both disclosed here, because "fixing" it to match the battery would make it a different exam with the same name.
- The bench stopped itself rather than fill its disk. It halted at a conservative floor, resumed only after every digest re-verified, and we tore its store down under four guards — weights are re-pullable; measurements are not.
How to check our work
The kit is beside this page, and what it is for is re-deriving every number printed here from the calls we actually made: the five fit receipts, every scorer file these tables read, the per-call records (including the judge seat's rows with their own num_ctx and think settings left exactly as sent), the build digests, the stop/resume log, the four-guard teardown receipt, the pre-registration with its golden hashes — 423ed960… for the judge trial's registration, fede4e32… for the 21 judge items — and a counting-rules note that states the count-travels rule, the 10% ceiling, the two-item tie band and the judge-seat deviation in one place. Where a scorer computes a rate it carries a Wilson 95% interval beside it; where it counts, it prints counts of N.
Every timestamp of the run itself in the kit is UTC. There is one exception and the kit's README names it: the tool-use trial's 19 frozen tasks are set in a fictional harbour with its own scenario clock, dated 2026-08-01, -08-02 and -08-12 rather than the run day, published as captured because the published checker matches the clock reading itself. None of them is a reading of any real clock.
If a number here does not reproduce from the kit, say so at the contact desk, where a person reads every message.
The rest of the seminar
The chairs these candidates contested were set in the chair trials on the 96 GB card, with the seat trials and the outside judges behind them. The consumer-card side of the workshop is A Rig Your Friend Already Owns, and what a smaller room costs in watts rather than chairs is What 150 Watts Buys. Thanks are owed to the people behind the five builds measured here; we wrote none of them.
The whole shelf holds the benches behind the claims we publish, failures included. If there is a piece of the machinery you want opened next, say so — the suggestion box is read.
Released under the house licence; main measurement window 2026-08-26, 07:00–15:13Z (UTC), with a stop at 09:58Z and a resume at 13:55Z. One 24 GB consumer card, Ollama 0.32.13, every candidate at its default published tag and the vendor's own quantisation; the 96 GB card of the earlier run is named by class throughout.