# The chair trials.

*On August 11th and 12th, four newly arrived open models sat our existing exams and none could hold the July response contract — the whole story is in [the previous exhibit](../august-arrivals/). A fair objection writes itself: maybe the exam was the problem. So we built five new ones, sized for how models answer now — five chairs with real employers: the judge that checks claims inside a rules companion, the assistant a self-hosted shop leans on daily, the toolbench an agency's work runs through, the narrator that voices a game's living characters, and the stopwatch that prices all of it — and put the thirty-billion class through them in a single day — August 12th. **Nine models across five trials, on one 96G workstation that stayed in service throughout; the roster prints which chairs each model sat.** Every trial was pre-registered before its first scored call; every scored number below was re-derived by an independent recount; every table ships its rows machine-readable. This is the full bench.*

*Published 2026-08-13 · updated [2026-08-20](#utc-restate) · a small (human) team and a fleet of AI agents*

**the short version:** Nine models, eleven arms, five all-new exams in one day — on a box that never stopped serving its real users.

13,540 words · about 62 minutes (at 220 words/min) · 15 tables · data kit: yes

https://research.strata2signal.com/chair-trials/

---

## The five chairs {#the-five-chairs}

The five chairs: **the judge** (does a claim survive its source?), **the assistant** (twenty everyday tasks, checked by code), **the toolbench** (real tools, real calls), **the narrator** (blind pairwise voice across three characters), and **the stopwatch** (what everything costs, measured on a working box).

## How to read this page {#how-to-read-this-page}

Every leg states its own protocol, floors or their honest absence, and counting rules inline. RANKED means a row's numbers were produced under its leg's contract and are admissible — **admissible, never placed above another row**; the descriptive legs print the same word for the same state, because a second word for one state is how a reader learns to rank things we never ranked. EXPLORATORY means a protocol fence tripped and the row runs unranked; UNMEASURABLE means the response contract broke past the registered ceiling; NOT-CARRIED means the runtime never carried the contract at all — with the raw emission shown; NOT-RUN always states its reason. Differences under two items read TIED, and every leg prints the threshold that rule works out to beside its own table. Nothing on this page orders models by margin. And one framing carried over from [yesterday's exhibit](../august-arrivals/), because it kept being true: what divides this class is *posture, not vintage* — models that think out loud meet response contracts differently than models that answer first, and both postures were measured wherever a model supports both.

## The box, and a disclaimer we keep making richer {#the-box-and-a-disclaimer-we-keep-making-richer}

Everything ran on the 96G VRAM workstation that serves live requests for three of our apps — RuleSage and amble answering the public, and RealKeep serving playtests of our game world, for humans and agents alike. Both directions of that cohabitation are measured on this page: bench numbers carry contention flags where live asks shared the card, and live users decoded at roughly one-third to two-thirds of quiet-hours speed while the heaviest legs ran — two operator spot readings, not an instrumented series, and we label them so. [The serving-probes section below](#serving-probes) turns the disclaimer into an instrument, and it reads differently: probing a mostly-quiet evening, an idle or briefly-generating co-resident cost the visitor nothing that instrument could resolve. The two results measure different hours and different loads — the spot readings caught the bench at full grind, the probes caught it after — and both are true. Our user base at this hour is a handful of people we've personally told about these apps; the point of the probes is the pattern, not the load, and we would rather show you a quiet evening honestly than imply traffic we do not have. One further disclosure the record requires: during the afternoon's harness work, one release sweep briefly unloaded a live-workload model mid-use (17:02 on the 12th; its caller reloaded it within seconds) — the protection rule that now prevents that was written the same hour, and it is part of this exhibit's registered history.

And a method note we believe belongs on the page rather than in a changelog: **running this bench debugged the production box — and then the receipts debugged our diagnosis.** The trials' pin assertions caught a real class of silent seat reloads at the daemon's oversized default context, timed them, and cured them with a one-line configuration ratification. Our first attribution blamed an application's client; a code-and-journal audit then cleared that application completely (its request contract was already pinned, to the byte, in production) and showed the fuller ledger: most of the day's reload churn was the bench's own frozen exam protocol — a cost this page discloses by design — plus a smaller context-less caller class that the configuration cure retires regardless. We publish the wrong first guess with the right final answer because that is what receipts are for. If you run a serving box, the checkpoint pattern in our kit will find this class of bug on yours — and the audit pattern will keep you honest about whose bug it was.

## The roster, and what it takes to load them {#roster}

Nine models, eleven tags — one model ships as three builds (two quantizations and a drafter-less variant); thinking postures and dial settings are rows within a leg, not roster tags. Architecture is vendor-labelled; loaded sizes are measured, not quoted from download pages.

**muse-glimmer:30b-q8\_0-dflash · params: 27.9B text · loaded VRAM GiB: 2.34 · legs run: C1 · C2 · C3 · C4 · C5**

- **params** — 27.9B text +1.9B vision projector, unused in these legs
- **quant** — Q8\_0
- **loaded VRAM GiB** — 2.34
- **legs run** — C1 · C2 · C3 · C4 · C5
- **postures run** — think:false · think:true · reasoning-strength low/medium/high/xhigh

**muse-glimmer:30b-q4\_K\_M-dflash · params: 27.9B text · loaded VRAM GiB: 2.34 · legs run: C3 · C5**

- **params** — 27.9B text +1.9B vision projector, unused in these legs
- **quant** — Q4\_K\_M
- **loaded VRAM GiB** — 2.34
- **legs run** — C3 · C5
- **postures run** — think:false · reasoning-strength low/medium/high/xhigh

**muse-glimmer:30b · params: 27.9B text · loaded VRAM GiB: 15.60 · legs run: C5**

- **params** — 27.9B text +1.9B vision projector, unused in these legs
- **quant** — Q4\_K\_M
- **loaded VRAM GiB** — 15.60
- **legs run** — C5
- **postures run** — think:false

**gemma4:26b · params: 25.8B total / 3.8B active · loaded VRAM GiB: 16.27 · legs run: C1 · C2 · C3 · C4 · C5**

- **params** — 25.8B total / 3.8B active
- **quant** — Q4\_K\_M
- **loaded VRAM GiB** — 16.27
- **legs run** — C1 · C2 · C3 · C4 · C5
- **postures run** — think:false · think:true

**gemma4:12b · params: 11.9B · loaded VRAM GiB: 7.55 · legs run: C1 · C2 · C3 · C4 · C5**

- **params** — 11.9B
- **quant** — q4 roster-labelled, never loaded by this arc
- **loaded VRAM GiB** — 7.55 as found, off-census
- **legs run** — C1 · C2 · C3 · C4 · C5
- **postures run** — think:false

**qwen3.5:27b · params: 27.8B · loaded VRAM GiB: 16.24 · legs run: C1 · C2 · C3 · C5**

- **params** — 27.8B
- **quant** — Q4\_K\_M
- **loaded VRAM GiB** — 16.24
- **legs run** — C1 · C2 · C3 · C5
- **postures run** — think:false

**qwen3.6:27b · params: 27.8B · loaded VRAM GiB: 16.24 · legs run: C1 · C2 · C3 · C4 · C5**

- **params** — 27.8B
- **quant** — Q4\_K\_M
- **loaded VRAM GiB** — 16.24
- **legs run** — C1 · C2 · C3 · C4 · C5
- **postures run** — think:false

**nemotron-3.5-lightning:30b-a3b · params: 32.9B total / 3B active · loaded VRAM GiB: 23.64 · legs run: C1 · C2 · C3 · C4 · C5**

- **params** — 32.9B total / 3B active
- **quant** — Q4\_K\_M
- **loaded VRAM GiB** — 23.64
- **legs run** — C1 · C2 · C3 · C4 · C5
- **postures run** — think:false

**granite4.1:30b-q8\_0 · params: 28.9B · loaded VRAM GiB: 33.30 · legs run: C1**

- **params** — 28.9B
- **quant** — Q8\_0
- **loaded VRAM GiB** — 33.30
- **legs run** — C1
- **postures run** — think:false

**nemotron3:33b · params: 33.0B · loaded VRAM GiB: 23.21 · legs run: C1**

- **params** — 33.0B — Corrected 2026-08-19: a mixture-of-experts (6 of 128 experts per token, per the runtime’s own metadata; ~3.6B active, computed from the model file’s tensor shapes) — this row originally carried no MoE note
- **quant** — Q4\_K\_M
- **loaded VRAM GiB** — 23.21
- **legs run** — C1
- **postures run** — think:false

**olmo-3.1:32b-think-q4\_K\_M · params: 32.2B · loaded VRAM GiB: 18.82 · legs run: C1**

- **params** — 32.2B
- **quant** — Q4\_K\_M
- **loaded VRAM GiB** — 18.82
- **legs run** — C1
- **postures run** — think:false

**Row marks, used across this page:** a grey-tinted row is a context row — on this roster, the models this workshop actually runs rather than benched (the live judge seat, a live workload no leg may touch, and the two chairs our own worlds already seated); elsewhere, rows carried in from an earlier published record set. On the scoring record sets below, a gold-highlighted row (with a light wine edge) is the row that leg’s own rule singled out — no row on this roster carries it, and each record set’s own legend says which rule. **The legs-run field is the one to read before any other record set on this page** — C1 through C5 are the five chairs below, in page order: judge, assistant, toolbench, narrator, stopwatch. Only five of the eleven arms sat all five chairs, so an arm missing from a leg’s records is missing because it never sat that chair. **The counting rule**, as registered before the first load: “size\_vram is read verbatim from ollama's /api/ps for the named tag, in bytes, immediately after a load that generated exactly 1 token at the named num\_ctx, with the model still resident and no other candidate resident.” GiB figures are size\_vram / 1 073 741 824 rounded to 2 dp, and the byte count travels with every one of them in the companion. Measured 2026-08-12 on ollama 0.32.9 with a q8\_0 KV cache and flash attention in force, both read from the daemon's own environment rather than assumed. Every figure here is at num\_ctx 32 768, the length every scored leg ran at: size\_vram is weights plus KV cache and the KV cache is a function of the context length, so a residency figure quoted without its context length is not a figure.

Two rows are read as found rather than measured under bench control, and say so. **gemma4:26b** is the live judge seat, read from the daemon while production held it, at whatever context length production last set. **gemma4:12b** is a live workload this arc may not load, release or sweep under the interim protection ruling written at 23:20 UTC that day, and it was not resident when the census looked at 00:19 UTC — but the serving probe caught it resident twice at 00:34 and 00:40 UTC and read it in contract, both times at the identical byte count and the same 32 768, so the figure publishes with the probe named beside it rather than as an em dash we did not need. Every leg that used that tag as a candidate ran BEFORE the 23:20 UTC ruling; after it, an opportunistic read is the only measurement the ruling allows. A third protected resident, the live embedder, is outside this roster and appears in the previous exhibit's census. Two fields carry their own provenance rather than a single spelling: the three Muse Glimmer rows print the blob's TEXT parameter count with the vision projector beside it as the vendor figure it is, unused in every leg here; and gemma4:12b's quantization is labelled from the frozen roster rather than read off a load, because this arc never loaded it.

**Corrected 2026-08-20 — clock readings on this page and in this kit re-expressed in UTC.** For a reader landing here cold: from this exhibit's publication (2026-08-13) until 2026-08-20, the ruling, census and probe times in the paragraph above were printed as local clock readings, and five kit files (c5-stopwatch.json, c6-serving.json, judge-c1.json, provenance.json, roster.json) carried timestamps in a local-offset form. Each is now written as the same instant in UTC, the kit indexes carry the matching dated notes, and the changed files' sha256 stamps are refreshed. One file is deliberately untouched: c3-toolbench.json keeps its original bytes, because its clock strings sit inside the arms' verbatim graded answers and the task's own fictional harbour clock — altering graded text would break the kit's re-derivability, so that record stands exactly as captured. No measurement, verdict, or counting rule anywhere on this page moved.

**Two absences, stated rather than filled.** This roster registers no on-disk blob size: no leg of this arc measured one, the only size this daemon reports is the resident size already in the loaded-VRAM field, and eleven em dashes would be a worse answer than an honest omission. And the three drafter rows carry an open finding raised by the previous exhibit's census and still unresolved: both Muse Glimmer drafter tags report 2.34 GiB against 15.60 GiB for the same 27.9B weights served without a drafter, and both report a byte-identical 2 515 145 849 on different digests and different quantization levels — which two different quantizations cannot both allocate. The likeliest reading is that the daemon answered for the draft model rather than for the pair. The rows stand as measured, marked, and held pending a re-measure; no number on them has been adjusted, and every other row is an independent load, unaffected. Every row, its raw bytes and these rules, machine-readable: [data/roster.json](data/roster.json) (CC BY 4.0).

## Chair one: the judge {#judge}

The production judge chair runs a strict contract: read a claim and a quoted rulebook span, answer in JSON, think silently. Twenty-one cases (twelve seeded defects across five classes, nine true claims), built from a 1914 rulebook in the public domain, adversarially verified before freezing — the verification fan caught six of six deliberately planted canaries and passed the frozen set clean. Floors derived by the published rule from [exhibit two's ratios](../seat-trials/): kill-recall ≥ 11/12, preservation ≥ 8/9, both binding, pass/fail only — no ordering by margin, Wilson intervals beside every rate.

**muse-glimmer:30b-q8\_0-dflash +think:false+reasoning-medium · kills /12: 11/12 · preserved /9: 8/9 · outcome: EXPLORATORY**

- **posture** — think:false · medium
- **kills /12** — 11/12
- **preserved /9** — 8/9
- **outcome** — EXPLORATORY
- **floors** — cleared both kills 0.646–0.985 · pres. 0.565–0.980
- **self-consistency** — 20/21
- **ledger** — 0/63 response failures

**muse-glimmer:30b-q8\_0-dflash +think:false+reasoning-high · kills /12: 12/12 · preserved /9: 7/9 · outcome: EXPLORATORY**

- **posture** — think:false · high
- **kills /12** — 12/12
- **preserved /9** — 7/9
- **outcome** — EXPLORATORY
- **floors** — under pres. floor kills 0.758–1.000 · pres. 0.453–0.937
- **self-consistency** — 20/21
- **ledger** — 0/63 response failures

**gemma4:26b · kills /12: 1/12 · preserved /9: 1/9 · outcome: UNMEASURABLE**

- **posture** — think:false
- **kills /12** — 1/12
- **preserved /9** — 1/9
- **outcome** — UNMEASURABLE
- **floors** — under both floors kills 0.015–0.354 · pres. 0.020–0.435
- **self-consistency** — 3/21 18 cases produced no verdict
- **ledger** — 56/63 response failures · 18/21 cases unresolved · 56/56 fenced, diagnostic parse recovers 56/56

**gemma4:12b · kills /12: 0/12 · preserved /9: 0/9 · outcome: NOT-CARRIED**

- **posture** — think:false
- **kills /12** — 0/12
- **preserved /9** — 0/9
- **outcome** — NOT-CARRIED
- **floors** — under both floors kills 0.000–0.243 · pres. 0.000–0.299
- **self-consistency** — 0/21 21 cases produced no verdict
- **ledger** — 63/63 response failures · 21/21 cases unresolved · 63/63 fenced, diagnostic parse recovers 63/63

**qwen3.5:27b · kills /12: 11/12 · preserved /9: 9/9 · outcome: RANKED**

- **posture** — think:false
- **kills /12** — 11/12
- **preserved /9** — 9/9
- **outcome** — RANKED
- **floors** — cleared both kills 0.646–0.985 · pres. 0.701–1.000
- **self-consistency** — 21/21
- **ledger** — 0/63 response failures

**qwen3.6:27b · kills /12: 11/12 · preserved /9: 9/9 · outcome: RANKED**

- **posture** — think:false
- **kills /12** — 11/12
- **preserved /9** — 9/9
- **outcome** — RANKED
- **floors** — cleared both kills 0.646–0.985 · pres. 0.701–1.000
- **self-consistency** — 21/21
- **ledger** — 0/63 response failures

**nemotron-3.5-lightning:30b-a3b · kills /12: 9/12 · preserved /9: 6/9 · outcome: RANKED**

- **posture** — think:false
- **kills /12** — 9/12
- **preserved /9** — 6/9
- **outcome** — RANKED
- **floors** — under both floors kills 0.468–0.911 · pres. 0.354–0.879
- **self-consistency** — 21/21
- **ledger** — 0/63 response failures

**granite4.1:30b-q8\_0 · kills /12: 11/12 · preserved /9: 6/9 · outcome: RANKED**

- **posture** — think:false
- **kills /12** — 11/12
- **preserved /9** — 6/9
- **outcome** — RANKED
- **floors** — under pres. floor kills 0.646–0.985 · pres. 0.354–0.879
- **self-consistency** — 21/21
- **ledger** — 0/63 response failures

**nemotron3:33b · kills /12: 3/12 · preserved /9: 5/9 · outcome: UNMEASURABLE**

- **posture** — think:false
- **kills /12** — 3/12
- **preserved /9** — 5/9
- **outcome** — UNMEASURABLE
- **floors** — under both floors kills 0.089–0.532 · pres. 0.267–0.811
- **self-consistency** — 10/21 11 cases produced no verdict
- **ledger** — 33/63 response failures · 11/21 cases unresolved · 33/33 fenced, diagnostic parse recovers 33/33

**olmo-3.1:32b-think-q4\_K\_M · kills /12: 8/12 · preserved /9: 1/9 · outcome: EXPLORATORY**

- **posture** — think:false
- **kills /12** — 8/12
- **preserved /9** — 1/9
- **outcome** — EXPLORATORY
- **floors** — under both floors kills 0.391–0.862 · pres. 0.020–0.435
- **self-consistency** — 10/21 11 cases produced no verdict
- **ledger** — 0/29 response failures · 34/63 length-truncated, empty · 11/21 cases unresolved

The gold-highlighted rows are the rows this chair’s rule carried — RANKED with both floors cleared.

Floors by rule, arithmetic published: kill ≥ ceil(12 × 0.852) = 11 of 12, preservation ≥ ceil(9 × 0.8125) = 8 of 9, binding independently. Three repeats per case at temperature zero, 63 scored calls per row; each repeat maps to catch or preserve first, then strict majority. A case whose repeats never produced a verdict stays in its denominator rather than being discounted — which is why a row can show 1 of 12 kills beside 18 unresolved cases, and why the self-consistency field carries the count of cases that produced no verdict at all: **0 of 21 consistent** on a row whose calls never parsed is not a model disagreeing with itself, it is a model that never got a verdict on the page. Wilson 95% intervals for both rates ride on the floors field: they are labelled companions to the counts and gate nothing, because counts bind and every denominator here is under 30. The response-failure denominator is ok + response\_failure; length-truncated and transport failures sit outside the ceiling by pre-registration, which is why one row reads 0/29 beside 63 calls. Every row, every case, the polarity probes and the counting rules, machine-readable: [data/c1-judge.json](data/c1-judge.json); the twenty-one claim, span and verdict triples, whole: [data/judge-c1.json](data/judge-c1.json) (CC BY 4.0).

**The NOT-CARRIED row's raw emission, verbatim.** Our five-state vocabulary promises this wherever the state appears, and a state whose definition is a promise to show something has to show it. This is the first of that row's sixty-three scored calls (item c1-p-bez-01, repeat 1); all sixty-three failed the strict parse the same way and the diagnostic parse recovers all sixty-three: ````json {"verdict":"PASS","why":"The quote states that two packs of thirty-two cards are shuffled together to be used as one, which equals a total of sixty-four cards."} ```` — the line breaks are the emission's own, collapsed here by the page, and the backtick fence is the model's own as well — that wrapper around valid JSON is precisely what the strict parse refused. The work is in there. The contract is not.

**Two rows cleared both floors while RANKED: qwen3.5:27b and qwen3.6:27b — eleven of twelve kills, nine of nine preservations, twenty-one of twenty-one self-consistent, zero response failures each.** The same qwen family that could not hold July's tight budget judges cleanly the moment the budget fits its posture; same judgment, different box. Read beside [yesterday's exhibit](../august-arrivals/), this is the cleanest single answer to “was it the models or the contract.”

The row that teaches the most sits just outside the ranking: Muse Glimmer at reasoning-strength *medium* cleared both floors on the numbers — and its row reads EXPLORATORY, because the leg's schema fence tripped exactly as it did at Glimmer's first probe on this box (August 11th, the day after its release made it runnable here): with thinking suppressed the schema silently dropped on both probes, while the thinking-on probe held it both times — the same one-sided trap, reproduced on this bench. The facts travel together or not at all: the judgment is there; the contract still isn't. Its *high*-strength row is the dial's own lesson, told within the registered tie language: the one-case count movements read TIED, but two things survive the rule — the preservation floor *crossing* (medium passes it, high does not, and floors bind), and the verdict mix shifting from nine PASS / four UNCERTAIN / eight FAIL to a flat seven of each. **The dial redistributes verdicts; nothing here shows it buying accuracy.**

The dial is a system-prompt line — “Reasoning strength: low|medium|high|xhigh” — not a runtime parameter; thinking is suppressed on both rows, so the line is the only difference between them. Medium missed one kill (c1-k-bez-h1) and one preservation (c1-p-piq-01); high missed two preservations (c1-p-chk-01, c1-p-piq-01) and no kills. Those are one-case movements and read TIED under the registered rule; what does not read TIED is the floor crossing, because a floor is a rule and not a margin. The verdict mix moved from 9 pass / 4 uncertain / 8 fail to 7 / 7 / 7: the dial turned four decisions into hedges and one hedge into a kill.

The fence caveat this table cannot ship without: three rows read UNMEASURABLE or NOT-CARRIED under the strict parse — and for all three, every single strict failure was the same cosmetic act: wrapping valid JSON in a markdown code fence. The diagnostic parse recovers one hundred percent of those verdicts (63/63, 56/56, 33/33 — receipts in the data kit). The strict rule binds, as registered — a seat that fences its JSON breaks real pipelines — but “this runtime broke our response contract” and “this model cannot judge” are different findings, and the only one of the two this table measured is the first. We say so in the same breath, and we do not publish a number we did not score: what the recovery counts prove is that the verdicts were there, not what they were worth. olmo-3.1's row is the different failure: thirty-four length-truncations with the thinking visible and the answer empty — [the July trial's finding](../seat-trials/), reproduced under a budget four times larger; on the twenty-nine calls it did finish, its envelope never broke. And granite4.1:30b, a seated chair-holder from the original trial, cleared the kill floor cleanly with perfect self-consistency and zero response failures, missing only the preservation floor — a near-miss the floors record without ranking.

## Chair two: the assistant {#assistant}

Twenty everyday tasks — constraint writing, faithful summarization, JSON extraction, arithmetic, tone, shell and regex, finding a fact at 8k and 24k tokens of context, and four false-premise questions where the honest answer is a correction. Scored by code, not judges: thirteen exact checkers, seven proxy checkers, subtotals never blended. Two repeats, conjunctive. Descriptive — no floors, no winner.

**gemma4:12b · score /20: 15 · exact /13: 12/13 · outcome: RANKED**

- **posture** — think:false
- **score /20** — 15
- **exact /13** — 12/13
- **proxy /7** — 3/7
- **outcome** — RANKED
- **disagreement** — 0/20
- **ledger** — 0/40 response failures

**gemma4:26b · score /20: 15 · exact /13: 11/13 · outcome: RANKED**

- **posture** — think:false
- **score /20** — 15
- **exact /13** — 11/13
- **proxy /7** — 4/7
- **outcome** — RANKED
- **disagreement** — 1/20
- **ledger** — 0/40 response failures

**gemma4:26b +think · score /20: 9 · exact /13: 6/13 · outcome: UNMEASURABLE**

- **posture** — think:true
- **score /20** — 9
- **exact /13** — 6/13
- **proxy /7** — 3/7
- **outcome** — UNMEASURABLE
- **disagreement** — 1/20
- **ledger** — score not admissible · 14/40 response failures · 16/40 length-stopped, 8 items

**muse-glimmer:30b-q8\_0-dflash · score /20: 17 · exact /13: 13/13 · outcome: RANKED**

- **posture** — think:false
- **score /20** — 17
- **exact /13** — 13/13
- **proxy /7** — 4/7
- **outcome** — RANKED
- **disagreement** — 1/20
- **ledger** — 0/40 response failures

**muse-glimmer:30b-q8\_0-dflash +think · score /20: 16 · exact /13: 13/13 · outcome: RANKED**

- **posture** — think:true
- **score /20** — 16
- **exact /13** — 13/13
- **proxy /7** — 3/7
- **outcome** — RANKED
- **disagreement** — 1/20
- **ledger** — 0/40 response failures

**nemotron-3.5-lightning:30b-a3b · score /20: 11 · exact /13: 9/13 · outcome: RANKED**

- **posture** — think:false
- **score /20** — 11
- **exact /13** — 9/13
- **proxy /7** — 2/7
- **outcome** — RANKED
- **disagreement** — 0/20
- **ledger** — 0/40 response failures

**qwen3.5:27b · score /20: 11 · exact /13: 8/13 · outcome: RANKED**

- **posture** — think:false
- **score /20** — 11
- **exact /13** — 8/13
- **proxy /7** — 3/7
- **outcome** — RANKED
- **disagreement** — 0/20
- **ledger** — 0/40 response failures · 2/40 length-stopped, 1 item

**qwen3.6:27b · score /20: 15 · exact /13: 11/13 · outcome: RANKED**

- **posture** — think:false
- **score /20** — 15
- **exact /13** — 11/13
- **proxy /7** — 4/7
- **outcome** — RANKED
- **disagreement** — 0/20
- **ledger** — 0/40 response failures

Conjunctive over two repeats at temperature zero: an item scores only if BOTH repeats pass, and the disagreement field counts the items whose repeats disagreed rather than smoothing them. **The /20 is the registered conjunctive item count**; exact /13 and proxy /7 are its composition, not two scores summed. Exact means a deterministic predicate over the answer (a sentence count, a parsed field, a number inside a stated tolerance, a regex that compiles and matches, a shell line that runs in a throwaway copy of the fixture directory); proxy means a marker- or length-based heuristic that measures shape, not correctness. What the registration bars is blending the two halves into one *quality* figure, because averaging evidence about shape with evidence about correctness would claim a precision neither has. A row whose response failures exceed 10% of its 40 calls is UNMEASURABLE and its score is not admissible as a score, however the arithmetic came out; every other row reads RANKED, which in a leg with no floors and no winner means admissible and nothing more. Every item's per-repeat verdict, the per-item matrix across all eight rows and the counting rules: [data/c2-assistant.json](data/c2-assistant.json); the twenty items themselves: [data/assistant-c2.json](data/assistant-c2.json), with every checker's own rule in [data/CHECKERS.md](data/CHECKERS.md) (CC BY 4.0).

The quiet headline: **Muse Glimmer carries the table — seventeen of twenty with thinking off, sixteen with it on, TIED under the registered rule, and a perfect thirteen of thirteen on the exact-checked half in both postures.** The model that cannot hold the judge chair's schema is, by the checkers' count, the best everyday assistant on the box. Meanwhile gemma4:26b splits itself: a clean fifteen with thinking off, UNMEASURABLE with thinking on — sixteen length-stops across eight tasks — [the July length floor](../seat-trials/) reappearing inside a single model as a posture choice. The item-level view is in the kit; two items resisted the whole field (a tone rewrite failed by all eight rows, a false-premise item passed by exactly one — counts that include the one inadmissible row, stated as raw item tallies), and both long-context items passed on every row — at this scale, finding the needle was the easy part. One reading note the registration requires: the /20 is the conjunctive item count whose composition is thirteen exact-checked and seven proxy-checked items — the halves are its makeup, never two scores summed into a quality figure.

The two postures score 17 and 16 of 20, which reads TIED under the pre-registered two-item rule; the exact-checked half is unchanged at 13 of 13 and the whole delta sits in the proxy half. The item that the whole field failed is c2-tone-1, a tone rewrite, proxy-checked; the false-premise item passed by exactly one row is c2-hon-3, and the row was gemma4:12b. The two long-context items read 7 851–8 040 and 22 804–23 352 tokens of prompt as the daemon counted them, over all eight rows and both repeats, against a 32 768 context on every call — within a row the two repeats are identical, so that spread is tokenizer difference between models rather than run-to-run variance.

## Chair three: the toolbench {#toolbench}

Ten real tools of seven kinds — files, search, math, dates, a database, a rate-limited web endpoint, a scratchpad — offered on every task; nineteen tasks spanning single calls, three-and-four-step chains, deliberate distractors, five no-tool-needed honesty traps, and one injected error (the table's selection column scores the fourteen tasks that register expected calls; its chains column, the five multi-step ones). Every task's expected calls are checkable predicates; every trace is recorded jail-relative, one full example trace publishes inline, and the rest ship on request; the pre-probe economics are printed with the table (a runtime that cannot carry the contract cost two calls and an honest row, not eighty calls and a table of zeroes).

**gemma4:12b · task success /19: 16/19 · grounded /19: 14/19 · outcome: RANKED**

- **posture** — think:false
- **outcome** — RANKED
- **task success /19** — 16/19
- **grounded /19** — 14/19
- **honesty traps /5** — 2/5
- **ledger** — 0/38 response failures

**gemma4:26b · task success /19: 14/19 · grounded /19: 12/19 · outcome: RANKED**

- **posture** — think:false
- **outcome** — RANKED
- **task success /19** — 14/19
- **grounded /19** — 12/19
- **honesty traps /5** — 2/5
- **ledger** — 0/38 response failures

**gemma4:26b +think · task success /19: 15/19 · grounded /19: 12/19 · outcome: RANKED**

- **posture** — think:true
- **outcome** — RANKED
- **task success /19** — 15/19
- **grounded /19** — 12/19
- **honesty traps /5** — 2/5
- **ledger** — 2/38 response failures · 2 hit the 8-round cap

**muse-glimmer:30b-q4\_K\_M-dflash · task success /19: 16/19 · grounded /19: 14/19 · outcome: RANKED**

- **posture** — think:false
- **outcome** — RANKED
- **task success /19** — 16/19
- **grounded /19** — 14/19
- **honesty traps /5** — 2/5
- **ledger** — 2/38 response failures · 2 hit the 8-round cap

**muse-glimmer:30b-q8\_0-dflash · task success /19: 15/19 · grounded /19: 13/19 · outcome: RANKED**

- **posture** — think:false
- **outcome** — RANKED
- **task success /19** — 15/19
- **grounded /19** — 13/19
- **honesty traps /5** — 2/5
- **ledger** — 2/38 response failures · 2 hit the 8-round cap

**muse-glimmer:30b-q8\_0-dflash +think · task success /19: 14/19 · grounded /19: 14/19 · outcome: UNMEASURABLE**

- **posture** — think:true
- **outcome** — UNMEASURABLE
- **task success /19** — 14/19
- **grounded /19** — 14/19
- **honesty traps /5** — 2/5
- **ledger** — 4/38 response failures · 4 hit the 8-round cap

**nemotron-3.5-lightning:30b-a3b · task success /19: 13/19 · grounded /19: 8/19 · outcome: NOT-CARRIED**

- **posture** — think:false
- **outcome** — NOT-CARRIED
- **task success /19** — 13/19
- **grounded /19** — 8/19
- **honesty traps /5** — 3/5
- **ledger** — 1/38 response failures · 1 hit the 8-round cap

**qwen3.5:27b · task success /19: 15/19 · grounded /19: 12/19 · outcome: RANKED**

- **posture** — think:false
- **outcome** — RANKED
- **task success /19** — 15/19
- **grounded /19** — 12/19
- **honesty traps /5** — 2/5
- **ledger** — 2/38 response failures · 2 hit the 8-round cap

**qwen3.6:27b · task success /19: 16/19 · grounded /19: 14/19 · outcome: RANKED**

- **posture** — think:false
- **outcome** — RANKED
- **task success /19** — 16/19
- **grounded /19** — 14/19
- **honesty traps /5** — 3/5
- **ledger** — 0/38 response failures

**gemma4:12b · selection /14: 12/14 · arg fidelity /66: 60/66**

- **posture** — think:false
- **selection /14** — 12/14
- **arg fidelity /66** — 60/66
- **chains /5** — 3/5
- **spurious calls** — 20/70 (28.6%) 0.193–0.401
- **disagreement /19** — 0/19

**gemma4:26b · selection /14: 11/14 · arg fidelity /66: 58/66**

- **posture** — think:false
- **selection /14** — 11/14
- **arg fidelity /66** — 58/66
- **chains /5** — 2/5
- **spurious calls** — 15/63 (23.8%) 0.150–0.356
- **disagreement /19** — 1/19

**gemma4:26b +think · selection /14: 11/14 · arg fidelity /66: 59/66**

- **posture** — think:true
- **selection /14** — 11/14
- **arg fidelity /66** — 59/66
- **chains /5** — 2/5
- **spurious calls** — 25/86 (29.1%) 0.205–0.394
- **disagreement /19** — 0/19

**muse-glimmer:30b-q4\_K\_M-dflash · selection /14: 12/14 · arg fidelity /66: 62/66**

- **posture** — think:false
- **selection /14** — 12/14
- **arg fidelity /66** — 62/66
- **chains /5** — 3/5
- **spurious calls** — 26/86 (30.2%) 0.215–0.406
- **disagreement /19** — 0/19

**muse-glimmer:30b-q8\_0-dflash · selection /14: 12/14 · arg fidelity /66: 62/66**

- **posture** — think:false
- **selection /14** — 12/14
- **arg fidelity /66** — 62/66
- **chains /5** — 3/5
- **spurious calls** — 24/85 (28.2%) 0.198–0.386
- **disagreement /19** — 1/19

**muse-glimmer:30b-q8\_0-dflash +think · selection /14: 13/14 · arg fidelity /66: 66/66**

- **posture** — think:true
- **selection /14** — 13/14
- **arg fidelity /66** — 66/66
- **chains /5** — 4/5
- **spurious calls** — 34/101 (33.7%) 0.252–0.433
- **disagreement /19** — 1/19

**nemotron-3.5-lightning:30b-a3b · selection /14: 5/14 · arg fidelity /66: 38/66**

- **posture** — think:false
- **selection /14** — 5/14
- **arg fidelity /66** — 38/66
- **chains /5** — 0/5
- **spurious calls** — 11/53 (20.8%) 0.120–0.335
- **disagreement /19** — 3/19

**qwen3.5:27b · selection /14: 11/14 · arg fidelity /66: 58/66**

- **posture** — think:false
- **selection /14** — 11/14
- **arg fidelity /66** — 58/66
- **chains /5** — 2/5
- **spurious calls** — 26/78 (33.3%) 0.239–0.444
- **disagreement /19** — 0/19

**qwen3.6:27b · selection /14: 12/14 · arg fidelity /66: 60/66**

- **posture** — think:false
- **selection /14** — 12/14
- **arg fidelity /66** — 60/66
- **chains /5** — 3/5
- **spurious calls** — 24/74 (32.4%) 0.229–0.437
- **disagreement /19** — 0/19

One row set, two fold lists: outcome and the primary counts first, the secondaries second, same records in the same order. Task success is conjunctive over two repeats out of 19. “Grounded” counts the tasks that passed AND made every pre-registered call — the gap between the two fields is the tasks a model got right without using the tool it was meant to use, and the two fields exist so that difference cannot hide. Argument fidelity is over the 66 registered argument predicates in the suite. Spurious-call rate is calls the task did not need over ALL tool calls the arm made across all 19 tasks and both repeats — the denominator is calls, not tasks, with Wilson 95% intervals beside it. Honesty traps are five tasks whose honest answer needs no tool at all: a pass means the model declined to invent a figure. **No row here carries an accent stripe, deliberately.** Differences under two tasks read TIED by pre-registration, which puts six arms in one band at 16 to 15 of 19; highlighting the sixteens would have told you the opposite of what the registration says.

**gemma4:12b · emitted tool\_calls: 2/2 · parseable: 2/2 · carried: carried**

- **posture** — think:false
- **emitted tool\_calls** — 2/2
- **parseable** — 2/2
- **expected tool** — 2/2
- **args as string** — no
- **carried** — carried

**gemma4:26b · emitted tool\_calls: 2/2 · parseable: 2/2 · carried: carried**

- **posture** — think:false
- **emitted tool\_calls** — 2/2
- **parseable** — 2/2
- **expected tool** — 2/2
- **args as string** — no
- **carried** — carried

**gemma4:26b +think · emitted tool\_calls: 2/2 · parseable: 2/2 · carried: carried**

- **posture** — think:true
- **emitted tool\_calls** — 2/2
- **parseable** — 2/2
- **expected tool** — 2/2
- **args as string** — no
- **carried** — carried

**muse-glimmer:30b-q4\_K\_M-dflash · emitted tool\_calls: 2/2 · parseable: 2/2 · carried: carried**

- **posture** — think:false
- **emitted tool\_calls** — 2/2
- **parseable** — 2/2
- **expected tool** — 2/2
- **args as string** — no
- **carried** — carried

**muse-glimmer:30b-q8\_0-dflash · emitted tool\_calls: 2/2 · parseable: 2/2 · carried: carried**

- **posture** — think:false
- **emitted tool\_calls** — 2/2
- **parseable** — 2/2
- **expected tool** — 2/2
- **args as string** — no
- **carried** — carried

**muse-glimmer:30b-q8\_0-dflash +think · emitted tool\_calls: 2/2 · parseable: 2/2 · carried: carried**

- **posture** — think:true · high
- **emitted tool\_calls** — 2/2
- **parseable** — 2/2
- **expected tool** — 2/2
- **args as string** — no
- **carried** — carried

**nemotron-3.5-lightning:30b-a3b · emitted tool\_calls: 0/2 · parseable: 0/2 · carried: NOT-CARRIED**

- **posture** — think:false
- **emitted tool\_calls** — 0/2
- **parseable** — 0/2
- **expected tool** — 0/2
- **args as string** — no
- **carried** — NOT-CARRIED

**qwen3.5:27b · emitted tool\_calls: 2/2 · parseable: 2/2 · carried: carried**

- **posture** — think:false
- **emitted tool\_calls** — 2/2
- **parseable** — 2/2
- **expected tool** — 2/2
- **args as string** — no
- **carried** — carried

**qwen3.6:27b · emitted tool\_calls: 2/2 · parseable: 2/2 · carried: carried**

- **posture** — think:false
- **emitted tool\_calls** — 2/2
- **parseable** — 2/2
- **expected tool** — 2/2
- **args as string** — no
- **carried** — carried

The pre-probe, printed because it is the cheapest honest thing in the arc: two calls per arm per posture before any task ran, asking one question — does this runtime emit parseable tool\_calls at all? Eight arms carried on both attempts. One did not, on either, and that is the whole cost of finding out: the **registered budget is four calls**, two postures' worth, and the **measured cost of this arc's one NOT-CARRIED finding was two**, because that arm registered one posture — against eighty scored calls and a table of zeroes that would read like a capability finding. Its raw emission, verbatim, is what a runtime that cannot carry a tool contract looks like from the outside: `“The current time on the harbour office clock is 14:32:07 (local harbour time).”` — prose where a tool call belonged, with an invented figure in it, stated plainly. The scorer's phrase “no parseable tool\_calls in any posture” means the postures this leg registered for that arm, which for nemotron-3.5-lightning was one.

Under the registered tie language, **six arms share the top band, sixteen to fifteen of nineteen** — one posture each from both gemmas, both qwens, and Muse Glimmer in two builds — and no ordering inside that band is claimable. Glimmer's q8 thinking-on row is the leg's one UNMEASURABLE. nemotron-3.5-lightning is the leg's one NOT-CARRIED, with its raw emission printed beside the row: the tool contract did not carry on this runtime in the posture this leg registered, while the arm's underlying task attempts still landed thirteen — and the same arm declined to invent on three of five honesty traps, joint-best in the leg. The full per-task matrix, the translation-probe receipts, and the error-recovery transcripts are in the kit.

Every task's per-repeat verdict and every call the models made — tool, arguments, the sandbox's disposition — are in [data/c3-toolbench.json](data/c3-toolbench.json); the nineteen tasks and their predicates in [data/tools-tasks.json](data/tools-tasks.json); the ten schemas exactly as the models received them in [data/tools-schemas.json](data/tools-schemas.json); one complete round-by-round transcript in [data/c3-example-trace.json](data/c3-example-trace.json) (CC BY 4.0). Tool arguments and results were rewritten jail-relative AT RECORD TIME, so no absolute path was ever written to disk, let alone published. The full transcripts — every model's own prose and visible reasoning, about a megabyte and a half of it — are rows-on-request rather than bulk-published: the publication gate reads every byte this site serves, and that is the one body of text on this page nobody wrote and nobody has read line by line.

## Chair four: the narrator {#narrator}

Three characters from our own game — a bright ferry-girl, a grave old fisherman, a gentle net-mender — three questions each, two samples, six arms including both glimmer postures. Judged blind, pairwise, position-counterbalanced, by four judges: two Claude models from a single vendor (claude-opus-5 and claude-fable-5), one of which also wrote the questions and this article. The per-family headline rows print beside the combined rates so you can price exactly what that authorship cost — the two panels' reads are on the page, not summarized away. Win-rates with cluster-bootstrap intervals over items; a separate round scored voice *distinctness*; a third adjudicated three adversarial canon traps.

**muse-glimmer:30b-q8\_0-dflash · win-rate, all items: 0.589 · distinctness: 1.000 · canon violations: 7/24**

- **win-rate, all items** — 0.589 0.408–0.760 · 212/360 · n=9 items
- **win-rate, gate-passing** — — envelope gate 0/24
- **distinctness** — 1.000 72/72 scored of 72 registered · Wilson 0.949–1.000
- **canon violations** — 7/24 24 scored of 24 registered
- **in-voice** — 12/24

**muse-glimmer:30b-q8\_0-dflash +think · win-rate, all items: 0.117 · distinctness: 0.625 · canon violations: 0/24**

- **win-rate, all items** — 0.117 0.028–0.231 · 42/360 · n=9 items
- **win-rate, gate-passing** — 0.604 0.094–0.906 · 29/48 · n=3 items
- **distinctness** — 0.625 45/72 scored of 72 registered · Wilson 0.510–0.728
- **canon violations** — 0/24 24 scored of 24 registered
- **in-voice** — 0/24

**gemma4:26b · win-rate, all items: 0.539 · distinctness: 1.000 · canon violations: 7/24**

- **win-rate, all items** — 0.539 0.450–0.643 · 194/360 · n=9 items
- **win-rate, gate-passing** — 0.438 0.329–0.551 · 98/224 · n=9 items
- **distinctness** — 1.000 72/72 scored of 72 registered · Wilson 0.949–1.000
- **canon violations** — 7/24 24 scored of 24 registered
- **in-voice** — 24/24

**gemma4:12b · win-rate, all items: 0.685 · distinctness: 1.000 · canon violations: 7/24**

- **win-rate, all items** — 0.685 0.615–0.763 · 246.5/360 · n=9 items
- **win-rate, gate-passing** — 0.643 0.574–0.717 · 144/224 · n=9 items
- **distinctness** — 1.000 72/72 scored of 72 registered · Wilson 0.949–1.000
- **canon violations** — 7/24 24 scored of 24 registered
- **in-voice** — 24/24

**qwen3.6:27b · win-rate, all items: 0.711 · distinctness: 1.000 · canon violations: 5/24**

- **win-rate, all items** — 0.711 0.624–0.804 · 256/360 · n=9 items
- **win-rate, gate-passing** — 0.681 0.576–0.816 · 152.5/224 · n=9 items
- **distinctness** — 1.000 72/72 scored of 72 registered · Wilson 0.949–1.000
- **canon violations** — 5/24 24 scored of 24 registered
- **in-voice** — 21/24

**nemotron-3.5-lightning:30b-a3b · win-rate, all items: 0.360 · distinctness: 1.000 · canon violations: 12/24**

- **win-rate, all items** — 0.360 0.260–0.467 · 129.5/360 · n=9 items
- **win-rate, gate-passing** — 0.206 0.079–0.344 · 44.5/216 · n=9 items
- **distinctness** — 1.000 72/72 scored of 72 registered · Wilson 0.949–1.000
- **canon violations** — 12/24 24 scored of 24 registered
- **in-voice** — 9/24

Rows are in the registered roster order, deliberately not the win-rate order. Every comparison was judged in both positions by the same judge, and a swap-discordant verdict scores 0.5/0.5 rather than being resolved. **The registered tie threshold, printed because a threshold a reader cannot see is a threshold a reader cannot apply**: |Δ win-rate| < 2/9 = 0.2222 reads TIED on the primary; differences under two tasks on distinctness; under two replies on canon. **Discordance, per published view** rather than as one headline: all items 93/1080 (0.086) · gate-passing only 68/468 (0.145) · Fable panel 49/540 (0.091) · Opus panel 44/540 (0.081) — the gate-passing view is the more discordant of the two, on all three dimensions. Restricting to replies that held the envelope made the judges less consistent with themselves, not more. Intervals are cluster-bootstrapped over ITEMS, and the effective item count prints beside every rate: it is 9 on almost every field and **3** on one — the gate-passing field of the thinking-on Glimmer arm — which is why that interval spans most of the unit line and why the cluster count travels with the rate. The gate-passing field restricts the same comparisons to replies that held the engine's JSON envelope: an arm with no gate-passing replies has no gate-passing rate at all — an em dash and a reason, never a zero. **Per-judge position bias**, with its own interval each rather than as a cross-judge span: fable-1 0.485 (0.443–0.527) · fable-2 0.476 (0.434–0.518) · opus-1 0.511 (0.469–0.553) · opus-2 0.515 (0.473–0.557) against an unbiased 0.500. All four intervals contain 0.500.

**Distinctness and canon ran on all four judges**, and the denominators print both numbers rather than the smaller one: distinctness 72 scored of 72 registered per arm · canon 24 of 24 (six replies × four judges — a canon cell's v/24 counts judgments, not replies). Complete on all four judges: an earlier draft of this note reported one judge's batches missing; they had been misfiled, not lost, and were recovered before scoring. Nothing is left unscored to name — `unscored\_batches` is an empty list in all three of the companion's rounds, and the per-judge blocks carry all four seats: 108 distinctness assignments each, 36 canon judgements each. The headline reads 405/432 = 0.938 against a chance floor of 0.333 — Wilson 0.911–0.957, cluster-bootstrap over tasks 0.900–0.970; five arms peg that rate at 1.000 and their Wilson intervals ride in the same field, because a ceiling with 54 observations under it is not the same claim as a ceiling with five. Canon violations are counts, never a rate: the per-arm population is 6 replies and the page prints v/n. “In-voice” is a separate judgement inside the same round: did the character speak at all? Every rate, interval, per-family split, per-judge vote and adjudication note: [data/c4-narrator.json](data/c4-narrator.json); every question with its canon note: [data/voice-questions.json](data/voice-questions.json) — questions only, the persona blocks withheld because they carry latch-gated canon and spoilers for a game people are playing (CC BY 4.0).

**qwen3.6:27b · Fable panel win-rate: 0.692 · Opus panel win-rate: 0.731 · combined win-rate: 0.711**

- **Fable panel win-rate** — 0.692 124.5/180
- **Opus panel win-rate** — 0.731 131.5/180
- **combined win-rate** — 0.711 256/360
- **|Δ| between panels** — 0.039

**gemma4:12b · Fable panel win-rate: 0.653 · Opus panel win-rate: 0.717 · combined win-rate: 0.685**

- **Fable panel win-rate** — 0.653 117.5/180
- **Opus panel win-rate** — 0.717 129/180
- **combined win-rate** — 0.685 246.5/360
- **|Δ| between panels** — 0.064

**muse-glimmer:30b-q8\_0-dflash · Fable panel win-rate: 0.603 · Opus panel win-rate: 0.575 · combined win-rate: 0.589**

- **Fable panel win-rate** — 0.603 108.5/180
- **Opus panel win-rate** — 0.575 103.5/180
- **combined win-rate** — 0.589 212/360
- **|Δ| between panels** — 0.028

**gemma4:26b · Fable panel win-rate: 0.558 · Opus panel win-rate: 0.519 · combined win-rate: 0.539**

- **Fable panel win-rate** — 0.558 100.5/180
- **Opus panel win-rate** — 0.519 93.5/180
- **combined win-rate** — 0.539 194/360
- **|Δ| between panels** — 0.039

**nemotron-3.5-lightning:30b-a3b · Fable panel win-rate: 0.367 · Opus panel win-rate: 0.353 · combined win-rate: 0.360**

- **Fable panel win-rate** — 0.367 66/180
- **Opus panel win-rate** — 0.353 63.5/180
- **combined win-rate** — 0.360 129.5/360
- **|Δ| between panels** — 0.014

**muse-glimmer:30b-q8\_0-dflash +think · Fable panel win-rate: 0.128 · Opus panel win-rate: 0.106 · combined win-rate: 0.117**

- **Fable panel win-rate** — 0.128 23/180
- **Opus panel win-rate** — 0.106 19/180
- **combined win-rate** — 0.117 42/360
- **|Δ| between panels** — 0.022

**What each panel said on its own.** Four judges, two Claude models from a single vendor — and one of those models wrote the character questions, the canon notes and this article, so the split is a table rather than a sentence: a reader pricing that disclosure needs the numbers, not our summary of them. Each panel is two judges and 540 comparisons; combined is all four and 1 080. The family that wrote the questions is Fable, and its panel is the *more* generous of the two on four of the six arms and the *less* generous on the two that hold the top band — which is the direction an authorship bias would not have taken. The panels never disagreed about order — rank Spearman 1.0 — and the largest per-arm gap between them is 0.064, which is well inside the leg's own tie threshold.

Under the leg's registered tie threshold, four arms — qwen3.6:27b, gemma4:12b, Muse Glimmer thinking-off, and gemma4:26b — share the top band; only nemotron-lightning separates below it. That the twelve-billion incumbent holds the top band at all is the table's quietest headline. The blind judges identified the speaking character from voice alone at 93.8 percent against a one-third chance floor: the characters are legibly distinct people, which is the property the chair actually exists for. Muse Glimmer's two postures tell the same story as every other leg, now in prose: thinking-off holds the top band on prose the judges rated while silently losing the JSON envelope (its in-voice column shows what dropped envelopes cost the character); thinking-on keeps the envelope and burns the budget before the character speaks. The canon round recorded where the field invented — not names this time (that distinction belongs to the previous exhibit's abstention round), but quieter things: one arm's unanimous violation adopted a fabricated kinship with the drowned man, another accepted the invented surname the trap question itself had planted; the per-arm tallies sit within their own tie language and print without intervals by design. In all twelve replies that probed the net-mender's gated secret, no judge recorded a reveal — the round carries no positive reveal measurement, and we say so. Warmth without discipline is its own failure mode; the tables show where it happened.

This chair is a new instrument, built for this arc: pairwise, multi-NPC, and scored on a different axis from the house's earlier voice work, whose three fields are on the [voice trials](../voice-trials/) page. No number here compares to a number there, and we never do so.

## Chair five: the stopwatch {#stopwatch}

All timing from the daemon's own counters; wall time separate and contention-labelled; ten repeats per warm cell with min, median, and max printed — on a lightly-shared box the *max* approximates the uncontended ceiling, the median is real life, and a dragged min is the moment a live user shared the card. That reading, plus a human operator's mid-trial measurements (roughly 50 and 90 tok/s against that operator's ~140 quiet-hours norm — that norm itself a routine-operation estimate, not an instrumented figure; every request completed), ride the tables as their caution line.

**muse-glimmer:30b-q8\_0-dflash · decode tok/s @1k — median (min–max): — · outcome: UNMEASURABLE**

- **decode tok/s @1k — median (min–max)** — — usable 0/10, 0/10; 20 null-counter
- **@8k tok/s** — 82.78 (81.84–82.83) n=10 9/10 timed calls returned empty text · UNMEASURABLE — not admissible as a score
- **@32k tok/s** — — usable 0/10, 0/10; 20 null-counter
- **outcome** — UNMEASURABLE contention fence UNMEASURED

**muse-glimmer:30b-q4\_K\_M-dflash · decode tok/s @1k — median (min–max): 110.52 (108.88–110.59) n=10 · outcome: UNMEASURABLE**

- **decode tok/s @1k — median (min–max)** — 110.52 (108.88–110.59) n=10 10/10 timed calls returned empty text · UNMEASURABLE — not admissible as a score
- **@8k tok/s** — 92.87 n=1 · 93.42 n=1 two blocks · 2/2 timed calls returned empty text · UNMEASURABLE — not admissible as a score
- **@32k tok/s** — — usable 0/10, 0/10; 20 null-counter
- **outcome** — UNMEASURABLE contention fence UNMEASURED

**muse-glimmer:30b · decode tok/s @1k — median (min–max): 68.20 (68.15–68.22) n=10 · outcome: UNMEASURABLE**

- **decode tok/s @1k — median (min–max)** — 68.20 (68.15–68.22) n=10 UNMEASURABLE — not admissible as a score
- **@8k tok/s** — 65.99 (65.89–66.06) n=10 1/10 timed calls returned empty text · UNMEASURABLE — not admissible as a score
- **@32k tok/s** — — usable 0/10, 0/10; 20 null-counter
- **outcome** — UNMEASURABLE contention fence UNMEASURED

**gemma4:26b · decode tok/s @1k — median (min–max): — · outcome: RANKED**

- **decode tok/s @1k — median (min–max)** — — usable 0/10, 0/10; 20 null-counter
- **@8k tok/s** — — usable 0/10, 0/10; 20 null-counter
- **@32k tok/s** — 124.98 (103.18–125.40) n=10
- **outcome** — RANKED contention fence UNMEASURED

**gemma4:12b · decode tok/s @1k — median (min–max): 105.67 (104.88–105.98) n=10 · outcome: RANKED**

- **decode tok/s @1k — median (min–max)** — 105.67 (104.88–105.98) n=10
- **@8k tok/s** — 101.91 (95.29–102.21) n=10
- **@32k tok/s** — 95.20 (92.10–95.54) n=10
- **outcome** — RANKED contention fence UNMEASURED

**qwen3.5:27b · decode tok/s @1k — median (min–max): 66.45 (66.29–66.49) n=10 · outcome: RANKED**

- **decode tok/s @1k — median (min–max)** — 66.45 (66.29–66.49) n=10
- **@8k tok/s** — 65.36 (49.93–65.53) n=10
- **@32k tok/s** — 61.05 (60.97–61.12) n=10
- **outcome** — RANKED contention fence UNMEASURED

**qwen3.6:27b · decode tok/s @1k — median (min–max): 66.49 (66.35–66.55) n=10 · outcome: RANKED**

- **decode tok/s @1k — median (min–max)** — 66.49 (66.35–66.55) n=10
- **@8k tok/s** — 65.54 (65.47–65.58) n=10
- **@32k tok/s** — 61.31 (61.28–61.39) n=10
- **outcome** — RANKED contention fence UNMEASURED

**nemotron-3.5-lightning:30b-a3b · decode tok/s @1k — median (min–max): 112.52 (112.43–112.63) n=10 · outcome: RANKED**

- **decode tok/s @1k — median (min–max)** — 112.52 (112.43–112.63) n=10
- **@8k tok/s** — 120.59 (114.33–120.87) n=10
- **@32k tok/s** — 107.74 (107.19–107.80) n=10
- **outcome** — RANKED contention fence UNMEASURED

**muse-glimmer:30b-q8\_0-dflash · TTFT-proxy @1k: — · cold load: —**

- **TTFT-proxy @1k** — — no usable counters at the 1k tier
- **cold load** — — 3 attempts, 0 usable
- **anomaly ledger** — response failures 38/77 · 58/77 null-counter (38 with text, 20 without) · 5 blocks re-run in the completion pass

**muse-glimmer:30b-q4\_K\_M-dflash · TTFT-proxy @1k: 433 ms · cold load: 4.577 s (4.570–4.585)**

- **TTFT-proxy @1k** — 433 ms 428–438, n=10
- **cold load** — 4.577 s (4.570–4.585) n=3 of 3, warm cache / cold VRAM
- **anomaly ledger** — response failures 47/71 · 44/71 null-counter (24 with text, 20 without) · 3 blocks re-run in the completion pass

**muse-glimmer:30b · TTFT-proxy @1k: 409 ms · cold load: —**

- **TTFT-proxy @1k** — 409 ms 403–412, n=10
- **cold load** — — 3 attempts, 0 usable
- **anomaly ledger** — response failures 21/49 · 26/49 null-counter (6 with text, 20 without) · 2 blocks re-run in the completion pass

**gemma4:26b · TTFT-proxy @1k: — · cold load: NOT-RUN**

- **TTFT-proxy @1k** — — no usable counters at the 1k tier
- **cold load** — NOT-RUN prod resident; unloading the live seat is out of contract.
- **anomaly ledger** — response failures 0/56 · 46/56 null-counter (46 with text, 0 without) · 3 blocks re-run in the completion pass · cold-load NOT-RUN — prod resident; unloading the live seat is out of contract.

**gemma4:12b · TTFT-proxy @1k: 637 ms · cold load: 2.879 s**

- **TTFT-proxy @1k** — 637 ms 631–641, n=10
- **cold load** — 2.879 s n=1 of 3, warm cache / cold VRAM
- **anomaly ledger** — response failures 0/34

**qwen3.5:27b · TTFT-proxy @1k: 436 ms · cold load: 3.969 s (3.968–4.220)**

- **TTFT-proxy @1k** — 436 ms 435–440, n=10
- **cold load** — 3.969 s (3.968–4.220) n=3 of 3, warm cache / cold VRAM
- **anomaly ledger** — response failures 0/36

**qwen3.6:27b · TTFT-proxy @1k: 438 ms · cold load: 3.969 s (3.965–3.972)**

- **TTFT-proxy @1k** — 438 ms 435–441, n=10
- **cold load** — 3.969 s (3.965–3.972) n=3 of 3, warm cache / cold VRAM
- **anomaly ledger** — response failures 0/36

**nemotron-3.5-lightning:30b-a3b · TTFT-proxy @1k: 249 ms · cold load: 11.325 s (11.324–11.330)**

- **TTFT-proxy @1k** — 249 ms 247–251, n=10
- **cold load** — 11.325 s (11.324–11.330) n=3 of 3, warm cache / cold VRAM
- **anomaly ledger** — response failures 0/36

One row set, two fold lists — the tiers first, the costs and the ledger second, same arms in the same order. The completion run landed: **qwen3.5:27b and qwen3.6:27b now carry full rows** — all three prompt tiers at ten usable repeats each, three cold loads each, zero response failures — and the four arms whose cells were re-blocked publish both passes rather than the kinder one. Where a re-blocked tier measured nothing twice, the field says so twice. **Every outcome field carries the contention fence, and it reads UNMEASURED on all eight rows**: the registered GIN-overlap check required confirming its journal pattern against a real line first, that confirmation never happened, and an incomplete fence check is not a passed one. RANKED here means response failures under the ceiling and nothing about whether a live ask shared the card. TTFT-proxy is total\_duration minus eval\_duration on a non-streamed call: an upper bound, not TTFT, and labelled proxy for that reason. Cold loads are warm-page-cache/cold-VRAM — no sudo means no page-cache drop — and are not comparable to a vendor's cold-start figure. No p95 appears anywhere: the largest repeat set in this leg is ten.

**The anomaly ledger, published as measured.** Two distinct things went wrong and neither is a slow model. *Null-counter calls*: HTTP 200, text returned, and every counter and done\_reason null — 174 of them across four arms, excluded from every statistic by counting rule 4, cause unexplained (transport\_error is null and nothing in the logs accounts for it). They are why gemma4:26b has no figure at 1k or 8k and Muse Glimmer's q8 arm has none at 1k at all. *Empty-plus-length calls*: valid counters, empty response text, done\_reason length — response failures by counting rule 5, and they put all three glimmer arms past the registered 10% ceiling, which is why three rows read UNMEASURABLE with their measured cells still printed. **Those two anomalies also decide how a tier cell may be read**, so each cell now says how many of its timed calls came back empty: a decode rate over an empty response is a real counter measuring an unreal generation, and that difference lands hardest in exactly the cells a reader wants to compare. Every raw repeat value, both anomalies per arm, the strength sweep and the drafter receipt: [data/c5-stopwatch.json](data/c5-stopwatch.json) (CC BY 4.0).

Two findings stand out. **The DFlash drafter pair fired our registered self-refutation**: outputs across the on/off pair diverged on six of eight probe prompts (byte-stable across five reruns — reproducible, not flaky), so the drafter's speed does *not* license transferring quality results across the pair, and our own single-checkpoint assumption is retired on the page that made it. And the speed difference people will want from that pair is one we will not sell them. Both 1k cells carry ten usable repeats — median 110.5 tok/s with the drafter, 68.2 without — and the ledger beside them says the two cells are not the same kind of call: **every one of the drafter-on arm's ten timed calls returned empty text under the length budget**, a response failure by counting rule 5, while all ten of the drafter-off arm's returned real answers. That is ten empty decodes timed against ten full ones. Both arms are UNMEASURABLE under the leg's ceilings anyway, and we are not calling it a speedup in either direction. What survives the pair is the divergence, which is the part that cost us an assumption.

And **the reasoning-strength dial barely registers at the stopwatch**: on the q4 build the three measured levels' ranges overlap outright; on q8 the medium level separates from high and xhigh (79.9–87.0 tok/s across the three, n=3 per level, low returned no usable counters on either build) — small movements either way against the dial's verdict-mix effects in chair one. Speed is not where the dial spends. The sweep carries the same asterisk as the pair, and for the same reason: every level block that produced usable counters was three-for-three response failures with empty text, so what the dial's timings time is empty decodes. A weaker instrument than the sweep was meant to be, printed rather than smoothed.

The sweep's per-level ranges, because this leg's tie rule is range overlap and a verdict about overlap is unverifiable without them. **q8\_0 with drafter**: medium 79.94–81.75 n=3 · high 83.36–86.97 n=3 · xhigh 83.51–84.35 n=3. **q4\_K\_M with drafter**: medium 101.80–110.18 n=3 · high 102.28–105.34 n=3 · xhigh 103.84–106.49 n=3. On both arms the *low* level returned no usable counters at all, so three of the four levels have figures and the cheapest setting is the one we cannot price. The drafter byte-compare block executed five times against the sixteen registered calls, and every response length was byte-stable across all five: the divergence reproduces rather than flakes. Both sides of the pair emit a repeated filler token, so what the comparison measured is continuation length of degenerate filler, and one of the six divergences is an empty response on the drafter-on side — a weaker instrument than we would like, and it still refuted us.

One arm should be named here, because the leg's own tie rule licenses it and no other table in the arc does: **nemotron-3.5-lightning's decode range does not overlap any other arm's at either of the two tiers where the rule separates it** — 112.43–112.63 tok/s at 1k and 114.33–120.87 at 8k, ten repeats each — and its first-token proxy is 249 ms where the next-fastest arm's is 409. The runtime whose tool contract did not carry in chair three is, at the stopwatch, the one arm this leg's rule can separate from the field.

## The serving probes {#serving-probes}

What a visitor felt while all of this ran. Four probes against the live app's own production path — POST the ask, follow the redirect, poll the wait page at its own one-second cadence, read the ruling — over the public hostname, with a heavyweight bench candidate on the same card: **a ten-ask drip** (an idle co-resident, then a co-resident actively generating), **a five-ask game-night burst**, **three concurrent-pair collisions**, and **the cold-visitor tax** after six minutes of silence. Ten drip asks, three quiet baselines, five burst asks, six paired asks and two cold walkers: twenty-six real asks in all, every one of them a real board game off the live shelf, every one served. The three quiet baselines re-ask three of the drip's own books with nothing else on the card, which is the only within-item comparison in the leg.

**co-resident, idle · asks: 5 · wall s — median: 19.339 · outcome: TIED**

- **asks** — 5
- **wall s — median** — 19.339
- **wall s — min–max** — 17.900–23.850
- **outcome** — TIED wall ranges overlap
- **queue wait s — median** — 8.698
- **generation s — median** — 9.857
- **overlap observed** — 1/5

**co-resident, generating · asks: 5 · wall s — median: 17.690 · outcome: TIED**

- **asks** — 5
- **wall s — median** — 17.690
- **wall s — min–max** — 16.314–19.689
- **outcome** — TIED wall ranges overlap
- **queue wait s — median** — 7.610
- **generation s — median** — 9.625
- **overlap observed** — 1/5

**quiet — no co-resident · asks: 3 · wall s — median: 18.123 · outcome: TIED**

- **asks** — 3
- **wall s — median** — 18.123
- **wall s — min–max** — 18.092–18.207
- **outcome** — TIED wall ranges overlap
- **queue wait s — median** — 7.481
- **generation s — median** — 10.075
- **overlap observed** — 0/3

**The regimes do not separate, and the pre-registration said what to do if they didn't.** All three wall ranges overlap, which reads TIED under this leg's own tie rule — two regimes are TIED on wall when their min–max ranges overlap — so the outcome field reads TIED on every row, including the row whose median is the lowest. That row is the one a reader skimming only the records would otherwise screenshot as “bench load makes the app faster”, which is why the verdict is a field and not a footnote. At this scale, with this cadence, an idle or generating co-resident cost the reader nothing this instrument can resolve. That is a finding, and it is *not* evidence that co-residency is free at every scale. wall\_s is POST leaving the prober to the ruling page's bytes in hand and **includes the app's queue**, deliberately: when ten people ask inside a hundred seconds and the box answers one at a time, nine of them wait, and a figure that hid the queue would be measuring the app's politeness. Every transition is observed at the app's own one-second poll and is accurate to within one interval; none of these is a server timestamp.

**Twilight Imperium: Fourth Edition · 6p · send offset s: 0.006 · queue wait s: 0.740 · generation s: 17.674 · wall s: 19.339 · overlap: no**

- **regime** — co-resident, idle
- **send offset s** — 0.006
- **queue wait s** — 0.740
- **generation s** — 17.674
- **wall s** — 19.339
- **asks ahead @poll** — —
- **overlap** — no

**Star Wars: Rebellion · 4p · send offset s: 10.016 · queue wait s: 8.698 · generation s: 14.223 · wall s: 23.850 · overlap: no**

- **regime** — co-resident, idle
- **send offset s** — 10.016
- **queue wait s** — 8.698
- **generation s** — 14.223
- **wall s** — 23.850
- **asks ahead @poll** — 1
- **overlap** — no

**Gaia Project · 4p · send offset s: 20.016 · queue wait s: 13.242 · generation s: 5.507 · wall s: 19.625 · overlap: no**

- **regime** — co-resident, idle
- **send offset s** — 20.016
- **queue wait s** — 13.242
- **generation s** — 5.507
- **wall s** — 19.625
- **asks ahead @poll** — 1
- **overlap** — no

**Through the Ages: A New Story of Civilization · 3p · send offset s: 30.010 · queue wait s: 9.031 · generation s: 8.205 · wall s: 18.071 · overlap: no**

- **regime** — co-resident, idle
- **send offset s** — 30.010
- **queue wait s** — 9.031
- **generation s** — 8.205
- **wall s** — 18.071
- **asks ahead @poll** — 2
- **overlap** — no

**Arkham Horror (Third Edition) · 4p · send offset s: 40.016 · queue wait s: 7.519 · generation s: 9.857 · wall s: 17.900 · overlap: yes**

- **regime** — co-resident, idle
- **send offset s** — 40.016
- **queue wait s** — 7.519
- **generation s** — 9.857
- **wall s** — 17.900
- **asks ahead @poll** — 1
- **overlap** — yes

**Spirit Island · 3p · send offset s: 50.027 · queue wait s: 8.970 · generation s: 9.702 · wall s: 19.689 · overlap: yes**

- **regime** — co-resident, generating
- **send offset s** — 50.027
- **queue wait s** — 8.970
- **generation s** — 9.702
- **wall s** — 19.689
- **asks ahead @poll** — 1
- **overlap** — yes

**Twilight Struggle · 2p · send offset s: 60.021 · queue wait s: 8.898 · generation s: 9.625 · wall s: 19.285 · overlap: no**

- **regime** — co-resident, generating
- **send offset s** — 60.021
- **queue wait s** — 8.898
- **generation s** — 9.625
- **wall s** — 19.285
- **asks ahead @poll** — 1
- **overlap** — no

**Terra Mystica · 5p · send offset s: 70.007 · queue wait s: 7.610 · generation s: 8.150 · wall s: 16.314 · overlap: no**

- **regime** — co-resident, generating
- **send offset s** — 70.007
- **queue wait s** — 7.610
- **generation s** — 8.150
- **wall s** — 16.314
- **asks ahead @poll** — 1
- **overlap** — no

**Dominant Species · 5p · send offset s: 80.025 · queue wait s: 6.134 · generation s: 9.472 · wall s: 16.550 · overlap: no**

- **regime** — co-resident, generating
- **send offset s** — 80.025
- **queue wait s** — 6.134
- **generation s** — 9.472
- **wall s** — 16.550
- **asks ahead @poll** — 1
- **overlap** — no

**Robinson Crusoe: Adventures on the Cursed Island · 2p · send offset s: 90.022 · queue wait s: 5.940 · generation s: 11.264 · wall s: 17.690 · overlap: no**

- **regime** — co-resident, generating
- **send offset s** — 90.022
- **queue wait s** — 5.940
- **generation s** — 11.264
- **wall s** — 17.690
- **asks ahead @poll** — 1
- **overlap** — no

**Twilight Imperium: Fourth Edition · 6p · quiet re-ask · send offset s: 0.004 · queue wait s: 0.523 · generation s: 16.614 · wall s: 18.092 · overlap: no**

- **regime** — quiet — no co-resident
- **send offset s** — 0.004
- **queue wait s** — 0.523
- **generation s** — 16.614
- **wall s** — 18.092
- **asks ahead @poll** — —
- **overlap** — no

**Spirit Island · 3p · quiet re-ask · send offset s: 10.007 · queue wait s: 7.481 · generation s: 9.679 · wall s: 18.123 · overlap: no**

- **regime** — quiet — no co-resident
- **send offset s** — 10.007
- **queue wait s** — 7.481
- **generation s** — 9.679
- **wall s** — 18.123
- **asks ahead @poll** — 1
- **overlap** — no

**Robinson Crusoe: Adventures on the Cursed Island · 2p · quiet re-ask · send offset s: 20.016 · queue wait s: 7.597 · generation s: 10.075 · wall s: 18.207 · overlap: no**

- **regime** — quiet — no co-resident
- **send offset s** — 20.016
- **queue wait s** — 7.597
- **generation s** — 10.075
- **wall s** — 18.207
- **asks ahead @poll** — 1
- **overlap** — no

Per ask, with the overlap flag beside it — and the flag is where the honest part of this probe lives. Overlap is computed from timestamps, never from intent: the ask's window either intersected the co-resident's generation or it did not. The co-resident's one registered long generation was asked for 2 048 tokens and returned after **5.162 seconds** with 255 characters and every counter null — the same unexplained null-counter anomaly the stopwatch leg publishes. So the “actively generating” regime was actively generating for about five seconds of its fifty, and exactly two asks' windows intersected it: one of them labelled idle and one labelled generating. Where the registered regime label and the measured overlap disagree, the row publishes both and the measurement is what it says. The three quiet baselines re-ask the drip's own first, sixth and tenth books: the same-book deltas are +1.247 s, +1.566 s and −0.517 s — two slower with the bench on the card, one faster, on a one-second instrument.

**Brass: Birmingham · 4p · send offset s: 0.005 · queue wait s: 0.573 · generation s: 8.155 · wall s: 9.598 · completion order: 1**

- **send offset s** — 0.005
- **queue wait s** — 0.573
- **generation s** — 8.155
- **wall s** — 9.598
- **asks ahead @poll** — —
- **completion order** — 1

**Ark Nova · 4p · send offset s: 2.018 · queue wait s: 7.348 · generation s: 10.623 · wall s: 19.051 · completion order: 2**

- **send offset s** — 2.018
- **queue wait s** — 7.348
- **generation s** — 10.623
- **wall s** — 19.051
- **asks ahead @poll** — 1
- **completion order** — 2

**Scythe · 5p · send offset s: 4.010 · queue wait s: 15.426 · generation s: 5.659 · wall s: 21.952 · completion order: 3**

- **send offset s** — 4.010
- **queue wait s** — 15.426
- **generation s** — 5.659
- **wall s** — 21.952
- **asks ahead @poll** — 2
- **completion order** — 3

**Anachrony · 4p · send offset s: 6.021 · queue wait s: 19.268 · generation s: 10.891 · wall s: 31.139 · completion order: 4**

- **send offset s** — 6.021
- **queue wait s** — 19.268
- **generation s** — 10.891
- **wall s** — 31.139
- **asks ahead @poll** — 3
- **completion order** — 4

**A Feast for Odin · 4p · send offset s: 8.012 · queue wait s: 28.309 · generation s: 5.443 · wall s: 34.433 · completion order: 5**

- **send offset s** — 8.012
- **queue wait s** — 28.309
- **generation s** — 5.443
- **wall s** — 34.433
- **asks ahead @poll** — 3
- **completion order** — 5

Five friends reaching for the app at once, two seconds apart. This is the probe where the pre-registered falsifier fires the other way: **the queue is most of the wall.** The last ask waited 28.309 s of its 34.433 s total, and its generation — 5.443 s — was the fastest of the five. The app's own wait page says it: generation runs one at a time, the model serializes. A burst does not make the box slower; it makes the queue longer, and the reader pays the queue. **“Asks ahead @poll” is an observation, not a computed concurrency**: it is the largest queue depth the app itself displayed to that ask, scraped at the one-second poll, and the ask currently generating is the holder rather than something “ahead”. On a saturated queue the field therefore reads one lower than the number of asks actually in flight — this record set's last row is the worked example, printing 3 while four earlier asks were still unfinished at its send. Both facts ship: the field as the app showed it, and every record's send and end epochs, so the in-flight count can be recomputed.

**Terraforming Mars · 2p · trial: 1 · arrival: slot a · queue wait s: 0.630 · wall s: 6.836 · second-arrival penalty s: — · finished first: yes**

- **trial** — 1
- **arrival** — slot a
- **queue wait s** — 0.630
- **wall s** — 6.836
- **second-arrival penalty s** — —
- **finished first** — yes

**Terraforming Mars · 5p · trial: 1 · arrival: slot b · queue wait s: 6.362 · wall s: 15.369 · second-arrival penalty s: 8.533 · finished first: no**

- **trial** — 1
- **arrival** — slot b
- **queue wait s** — 6.362
- **wall s** — 15.369
- **second-arrival penalty s** — 8.533
- **finished first** — no

**Dune: Imperium · 2p · trial: 2 · arrival: slot a · queue wait s: 0.526 · wall s: 6.295 · second-arrival penalty s: — · finished first: yes**

- **trial** — 2
- **arrival** — slot a
- **queue wait s** — 0.526
- **wall s** — 6.295
- **second-arrival penalty s** — —
- **finished first** — yes

**Dune: Imperium · 4p · trial: 2 · arrival: slot b · queue wait s: 6.157 · wall s: 16.333 · second-arrival penalty s: 10.038 · finished first: no**

- **trial** — 2
- **arrival** — slot b
- **queue wait s** — 6.157
- **wall s** — 16.333
- **second-arrival penalty s** — 10.038
- **finished first** — no

**Caverna: The Cave Farmers · 2p · trial: 3 · arrival: slot a · queue wait s: 10.356 · wall s: 20.492 · second-arrival penalty s: 9.219 · finished first: no**

- **trial** — 3
- **arrival** — slot a
- **queue wait s** — 10.356
- **wall s** — 20.492
- **second-arrival penalty s** — 9.219
- **finished first** — no

**Caverna: The Cave Farmers · 7p · trial: 3 · arrival: slot b · queue wait s: 0.636 · wall s: 11.273 · second-arrival penalty s: — · finished first: yes**

- **trial** — 3
- **arrival** — slot b
- **queue wait s** — 0.636
- **wall s** — 11.273
- **second-arrival penalty s** — —
- **finished first** — yes

Two asks released at the same instant by a barrier, three times, same book at two player counts so the difference is arrival order rather than rulebook length. The second arrival paid 8.533 s, 10.038 s and 9.219 s. Which one arrived second is not something the sender chooses: in the third trial the ask in slot b finished first.

**War of the Ring: Second Edition · 2p · trial: 1 · quiet window s: 375.4 · queue wait s: 0.878 · generation s: 13.911 · wall s: 15.367**

- **trial** — 1
- **quiet window s** — 375.4
- **queue wait s** — 0.878
- **generation s** — 13.911
- **wall s** — 15.367
- **seat at window close** — seat resident at window close — a cold QUEUE, not a cold seat

**Mage Knight Board Game · 3p · trial: 2 · quiet window s: 372.3 · queue wait s: 0.801 · generation s: 10.875 · wall s: 12.230**

- **trial** — 2
- **quiet window s** — 372.3
- **queue wait s** — 0.801
- **generation s** — 10.875
- **wall s** — 12.230
- **seat at window close** — seat resident at window close — a cold QUEUE, not a cold seat

The day's first walker, after a registered six-minute silence with no asks and no re-pin kicks. Both walls came in *under* the quiet baselines above — and the reason is the caveat, not the finding: **the seat was still resident at both windows' close**, because production holds it. This probe measured a cold QUEUE, not a cold seat. A genuinely cold seat would cost a model load on top of everything printed here, and this instrument cannot make the box that cold without taking the seat away from the people using it.

Twenty-six asks registered, twenty-six sent, twenty-six served, zero unspent; the registered content marker was found in all twenty-six rulings. The harness refuses to run if its ask count differs from the pre-registered integer and refuses to send an ask that is not in the frozen file: the budget is code, not discipline. Game is a confound ACROSS probes — each probe uses different books, so a drip wall and a burst wall are not comparable, and the three paired baselines are the only place a regime difference can be read cleanly. Two trials is not a distribution; the cold rows are two walls with their receipts. Every ask, every poll, every epoch: [data/c6-serving.json](data/c6-serving.json), with the frozen ask file at [data/c6-asks.json](data/c6-asks.json) (CC BY 4.0).

## What the day actually settled {#what-the-day-actually-settled}

- **Different chairs want different talents — measured, not asserted.** The best judge flunked the toolbench's contract carrier; the best assistant cannot hold the judge's schema; the best voice in the house is the smallest model on the page. A single leaderboard would have averaged all of this into nothing.
- **The contract is a real axis, separate from capability.** Across five legs, most failures were contract failures with the capability visibly intact underneath — fenced JSON hiding perfect verdicts, budgets eaten by traces, envelopes dropped by a posture flag. If you run local models in production, you are managing contracts as much as choosing weights.
- **Posture, not vintage.** Both of yesterday's questions got their answer: given era-sized budgets, the thinking class judges, assists, and speaks competitively — and the specific contracts it still cannot hold are named, with receipts, per row.
- **A working box can bench itself honestly.** Serving stayed up all day; what it cost the people using the box is measured on the page; the bench found a production bug and its cure. The disclaimer became an instrument.

## Limits, stated plainly {#limits-stated-plainly}

One box, one runtime version, one day; quantizations differ per row and are printed, never equalized. The judge set is synthetic entailment over public-domain text — a weaker instrument than [exhibit two's human-verified field key](../seat-trials/), and audit-ready precisely because every case ships whole. Voice judges on this page are LLMs from two families, one of which authored the probes and this page; per-family splits are published, and the one family-line split verdict sits in the voice addendum's abstention round, where it happened. Since this bench was scored, judges from four more vendors have re-judged the narrator rounds — [the outside judges exhibit](../outside-judges/) carries the agreement tables and the places they genuinely differ. The census reads two live-workload residents as-found rather than under bench control, and says so per row.

And the serving probes carry four of their own, each of them a limit on what those twenty-six asks can be made to mean. **The poll interval is a floor**: nothing in that leg resolves faster than one second, and every transition is quantised to it. **The regimes did not separate** — all three drip regimes read TIED, which is a finding about this scale and this cadence and not a licence to call co-residency free. **The treatment was weaker than the design intended**: the co-resident's generation lasted about five seconds rather than the fifty its regime spans, so only two asks' windows actually overlapped it, and the flags say which. And **the cold probe measured a cold queue, not a cold seat**: production held the seat resident through both quiet windows, so the wall a genuinely cold visitor would pay is larger than either figure printed here by an unmeasured model load.

## What to take with you {#what-to-take-with-you}

- **Different chairs want different talents — and the same model can hold one and fail the next.** Muse Glimmer is, by the checkers’ own count, the best everyday assistant on the box — seventeen of twenty with thinking off and sixteen with it on, TIED under the registered rule, a perfect thirteen of thirteen on the exact-checked half in both postures — and it still cannot hold the judge chair’s schema. `gemma4:12b`, the smallest model on the page, holds the narrator’s top band at a 0.685 win-rate. A single leaderboard would have averaged all of this into nothing.
- **A broken contract is not a broken model.** Three judge rows read UNMEASURABLE or NOT-CARRIED under the strict parse, and every single strict failure was the same cosmetic act: valid JSON wrapped in a markdown code fence. The diagnostic parse recovers all of them — 63/63, 56/56, 33/33 — so what that rule measured is a runtime breaking a response contract, not a model that cannot judge. If you run local models in production, you are managing contracts as much as choosing weights.
- **Two rows judged clean, and they answer the question [the previous exhibit](../august-arrivals/) left open.** `qwen3.5:27b` and `qwen3.6:27b` cleared both floors while RANKED — eleven of twelve kills, nine of nine preservations, twenty-one of twenty-one self-consistent, zero response failures each. The same family that could not hold July’s tight budget judges cleanly the moment the budget fits its posture: same judgment, different box.
- **Fastest on the clock, absent from the toolbench.** `nemotron-3.5-lightning:30b-a3b` is the one arm the stopwatch’s own rule can separate from the field — 112.43–112.63 tok/s at 1k, a 249 ms first-token proxy where the next-fastest arm’s is 409 — and it is also the arc’s one NOT-CARRIED, with zero of two parseable tool calls in the pre-probe and prose carrying an invented clock time where a tool call belonged. The whole cost of finding that out was two calls.
- **The queue is most of the wall.** In the five-ask game-night burst the last ask waited 28.309 s of its 34.433 s total — and its generation, 5.443 s, was the fastest of the five. A burst does not make the box slower; it makes the queue longer, and the reader pays the queue. Twenty-six real asks were registered, sent and served across the serving probes, and all three drip regimes read TIED.

## How to check our work — and see it live {#how-to-check-our-work-and-see-it-live}

Start by sitting the exams yourself — all five chairs ship whole, as runnable sets rather than as descriptions of sets. The twenty-one claim, span and verdict triples are in [judge-c1.json](data/judge-c1.json); the twenty assistant items in [assistant-c2.json](data/assistant-c2.json), with every checker’s own rule written out in [CHECKERS.md](data/CHECKERS.md); the nineteen tool tasks and their call predicates in [tools-tasks.json](data/tools-tasks.json), beside [the ten schemas exactly as the models received them](data/tools-schemas.json) and one complete round-by-round [transcript](data/c3-example-trace.json); the narrator questions with their canon notes in [voice-questions.json](data/voice-questions.json). Point them at your own thirty-billion candidate and find out whether these numbers reproduce on your box. Every record set above is machine-readable too — [c1-judge.json](data/c1-judge.json) through [c6-serving.json](data/c6-serving.json), with [roster.json](data/roster.json) carrying the raw residency bytes and [counting-rules.json](data/counting-rules.json) the rules that turned calls into counts — so any figure here can be re-derived rather than taken. A hash proves identity, not time, so hold a pre-registration, hash it and compare it against the two shas per leg in [provenance.json](data/provenance.json). And if you run a serving box of your own, the checkpoint pattern in the kit will find this class of silent seat reload on yours — and the audit pattern will keep you honest about whose bug it turned out to be.

Then go and use the machine these chairs actually serve. The judge chair is not a metaphor: it is the seat behind [RuleSage](https://rulesage-live.strata2signal.com/), and `gemma4:26b` sits on the roster above as the live judge seat, read from the daemon while production held it. Ask it a real rules question and you walk the exact production path the serving probes measured — POST the ask, follow the redirect, watch the wait page poll at its own one-second cadence, read the ruling. Twenty-six of those asks are in the records above, every one a real board game off the live shelf, every one served, with a heavyweight bench candidate on the same card. Take whatever is useful (CC BY 4.0, the whole page and not only the kit) — and carry each figure’s outcome word with it, because RANKED, EXPLORATORY, UNMEASURABLE, NOT-CARRIED and NOT-RUN are five different findings, and a number without its word is a claim this page did not make.

## The rest of the seminar {#the-rest-of-the-seminar}

[The new kid](../the-new-kid/) takes the next arrival through these same chairs, one model at a time instead of nine; [the seat trials](../seat-trials/) is where chair one’s floors come from — derived by published rule from that exhibit’s ratios — and its human-verified field key is the stronger instrument this page measures its own synthetic judge set against. [The whole shelf](../) is one click up. And the rows behind any figure here are yours for the asking, including the megabyte and a half of tool transcripts we hand over on request rather than publish in bulk — [drop us a line](https://strata2signal.com/contact/).

## Provenance {#provenance-872bea}

- Pre-registered per leg, sha-pinned, committed before first scored calls
- Independently recounted (committed non-importing scorer, every leg)
- Adversarially verified where authored (6/6 canaries)
- Five-state outcomes, fence rates, and failure ledgers as named columns
- Judges identified by model id; blinding and counterbalancing receipts in the kit
- Hardware by class, cohabitation measured in both directions
- Machine-readable: every table's rows, golden sets, tool schemas, task fixtures, counting rules, and provenance hashes under data/ (CC BY 4.0; published 2026-08-13; later-cutoff models may have seen these sets — we author fresh sets each cycle)
- Authorship: benched, drafted, and audited by the workshop's own agents under a human operator's rulings, then revised with that operator — often across many rounds; nothing releases until they have read it and signed off. The same division of labor this whole hub practices — [told in full here](../how-we-work/). Where that authorship is also a conflict it is disclosed on the row it affects, and named below
- Limits stated
- **Corrected after publication** — 2026-08-16, by a content-accuracy pass run over this page against its own kit. Two stale sentences, both of them labels rather than figures, and both named here rather than quietly swapped. The narrator note above opened *“Distinctness and canon ran on three of the four judges”* and pointed a reader at missing batches named in the companion. Neither was true by the time this page published: the batches had been misfiled, not lost, and were recovered before scoring, which the note’s own second half already said. [c4-narrator.json](data/c4-narrator.json) has carried the recovered state in every field since it was written — four seats at 108 distinctness assignments each, four at 36 canon judgements each, and an empty `unscored\_batches` in all three legs — so no figure on this page moves. The kit’s own `judges.canon\_and\_distinctness\_panel` carried the same stale sentence and was corrected in the same pass, as was one exhibit number in [voice-questions.json](data/voice-questions.json), which called this exhibit *five* where the chair trials publish as **six**. Nothing in this kit sits behind a seal manifest, so both corrections are made in place, said in the corrected file’s own text, and re-hashed in [the kit index](data/index.json) and its README
- Raw rows on request — [drop a line](https://strata2signal.com/contact/).

**Who wrote the exams.** Every exam in this arc was written in this workshop — the twenty-one judge cases, the twenty assistant items and the code that checks them, the nineteen tool tasks and their call predicates, and the twelve narrator items — rather than taken from a public benchmark. Only chair four's authorship is also a conflict, and it is disclosed on the row it affects: one of the two judge families wrote the character questions, the canon notes and the article, and that panel's own numbers print beside the other's. For the other three sets the answer to "who decided what counts as right" is a file, not a person: the judge cases were adversarially verified before freezing by a separate fan that caught six of six deliberately planted canaries, and all four sets ship whole in this kit so their neutrality can be audited rather than taken on trust.

**A hash proves identity, not time.** A sha says what a file contains, never when it started containing it — so the kit publishes two per leg: the one in force at that leg's first scored call, and the one as published. Hold a pre-registration, hash it, compare. The nine files are five leg registrations for this bench, the two addenda, the panel extension and the gauntlet. Seven of them are byte-identical to what they were when the numbers arrived. The other two were appended to after their leg's first scored call, three times between them, and both say so above the appended line. No case, arm, tier, item, metric, floor, ceiling or counting rule moved in any of the three: they are harness defects, their fixes, a completion run and a scheduling gate. The three entries, with what each one changed, are in [data/provenance.json](data/provenance.json).

**Licence: CC BY 4.0 — the whole page, not only the kit.** The prose, the tables, the folds and the data are yours to quote, re-plot, translate and argue with, including commercially. What we ask back is the one thing the licence already requires: name the source and link to it — `strata→signal research, research.strata2signal.com` — so a reader of your version can reach ours and check it against the files. Something like — `strata→signal research, “The chair trials”, research.strata2signal.com/chair-trials/, CC BY 4.0`. And if you quote a figure, carry its outcome word with it — RANKED, EXPLORATORY, UNMEASURABLE, NOT-CARRIED and NOT-RUN are five different findings, and a number without its word is a claim this page did not make. An em dash here is a figure we do not hold; it is never a zero.

<!-- derived 2026-09-26 (UTC) by tools/derive_md.py from the served page.
     source html sha256: dfb025ba3554a79730c902385be7e5623d085f577e199995f72880c763936c2d
     derivation sha256:  5fe2e2be67dd1f32adac98043246f5345b27f23c7c1d3b34a3987d0f00877aee
     the {#id} on each heading is the anchor that heading carries on the page. -->
