Exhibit six · every chair in the house

The chair trials.

Published 2026-08-13 · exhibit six The bench

On August 11th and 12th, four newly arrived open models sat our existing exams and none could hold the July response contract — the whole story is in the previous exhibit. A fair objection writes itself: maybe the exam was the problem. So we built five new ones, sized for how models answer now — five chairs with real employers: the judge that checks claims inside a rules companion, the assistant a self-hosted shop leans on daily, the toolbench an agency's work runs through, the narrator that voices a game's living characters, and the stopwatch that prices all of it — and put the thirty-billion class through them in a single day — August 12th. Nine models across five trials, on one 96G workstation that stayed in service throughout; the roster prints which chairs each model sat. Every trial was pre-registered before its first scored call; every scored number below was re-derived by an independent recount; every table ships its rows machine-readable. This is the full bench.

The five chairs

The five chairs: the judge (does a claim survive its source?), the assistant (twenty everyday tasks, checked by code), the toolbench (real tools, real calls), the narrator (blind pairwise voice across three characters), and the stopwatch (what everything costs, measured on a working box).

How to read this page

Every leg states its own protocol, floors or their honest absence, and counting rules inline. RANKED means a row's numbers were produced under its leg's contract and are admissible — admissible, never placed above another row; the descriptive legs print the same word for the same state, because a second word for one state is how a reader learns to rank things we never ranked. EXPLORATORY means a protocol fence tripped and the row runs unranked; UNMEASURABLE means the response contract broke past the registered ceiling; NOT-CARRIED means the runtime never carried the contract at all — with the raw emission shown; NOT-RUN always states its reason. Differences under two items read TIED, and every leg prints the threshold that rule works out to beside its own table. Nothing on this page orders models by margin. And one framing carried over from yesterday's exhibit, because it kept being true: what divides this class is posture, not vintage — models that think out loud meet response contracts differently than models that answer first, and both postures were measured wherever a model supports both.

The box, and a disclaimer we keep making richer

Everything ran on the 96G VRAM workstation that serves live requests for three of our apps — RuleSage and amble answering the public, and RealKeep serving playtests of our game world, for humans and agents alike. Both directions of that cohabitation are measured on this page: bench numbers carry contention flags where live asks shared the card, and live users decoded at roughly one-third to two-thirds of quiet-hours speed while the heaviest legs ran — two operator spot readings, not an instrumented series, and we label them so. The serving-probes section below turns the disclaimer into an instrument, and it reads differently: probing a mostly-quiet evening, an idle or briefly-generating co-resident cost the visitor nothing that instrument could resolve. The two results measure different hours and different loads — the spot readings caught the bench at full grind, the probes caught it after — and both are true. Our user base at this hour is a handful of people we've personally told about these apps; the point of the probes is the pattern, not the load, and we would rather show you a quiet evening honestly than imply traffic we do not have. One further disclosure the record requires: during the afternoon's harness work, one release sweep briefly unloaded a live-workload model mid-use (17:02 on the 12th; its caller reloaded it within seconds) — the protection rule that now prevents that was written the same hour, and it is part of this exhibit's registered history.

And a method note we believe belongs on the page rather than in a changelog: running this bench debugged the production box — and then the receipts debugged our diagnosis. The trials' pin assertions caught a real class of silent seat reloads at the daemon's oversized default context, timed them, and cured them with a one-line configuration ratification. Our first attribution blamed an application's client; a code-and-journal audit then cleared that application completely (its request contract was already pinned, to the byte, in production) and showed the fuller ledger: most of the day's reload churn was the bench's own frozen exam protocol — a cost this page discloses by design — plus a smaller context-less caller class that the configuration cure retires regardless. We publish the wrong first guess with the right final answer because that is what receipts are for. If you run a serving box, the checkpoint pattern in our kit will find this class of bug on yours — and the audit pattern will keep you honest about whose bug it was.

The roster, and what it takes to load them

Nine models, eleven tags — one model ships as three builds (two quantizations and a drafter-less variant); thinking postures and dial settings are rows within a leg, not roster tags. Architecture is vendor-labelled; loaded sizes are measured, not quoted from download pages.

muse-glimmer:30b-q8_0-dflashparams27.9B textloaded VRAM GiB2.34legs runC1 · C2 · C3 · C4 · C5
params
27.9B text +1.9B vision projector, unused in these legs
quant
Q8_0
loaded VRAM GiB
2.34
legs run
C1 · C2 · C3 · C4 · C5
postures run
think:false · think:true · reasoning-strength low/medium/high/xhigh
muse-glimmer:30b-q4_K_M-dflashparams27.9B textloaded VRAM GiB2.34legs runC3 · C5
params
27.9B text +1.9B vision projector, unused in these legs
quant
Q4_K_M
loaded VRAM GiB
2.34
legs run
C3 · C5
postures run
think:false · reasoning-strength low/medium/high/xhigh
muse-glimmer:30bparams27.9B textloaded VRAM GiB15.60legs runC5
params
27.9B text +1.9B vision projector, unused in these legs
quant
Q4_K_M
loaded VRAM GiB
15.60
legs run
C5
postures run
think:false
gemma4:26bparams25.8B total / 3.8B activeloaded VRAM GiB16.27legs runC1 · C2 · C3 · C4 · C5
params
25.8B total / 3.8B active
quant
Q4_K_M
loaded VRAM GiB
16.27
legs run
C1 · C2 · C3 · C4 · C5
postures run
think:false · think:true
gemma4:12bparams11.9Bloaded VRAM GiB7.55legs runC1 · C2 · C3 · C4 · C5
params
11.9B
quant
q4 roster-labelled, never loaded by this arc
loaded VRAM GiB
7.55 as found, off-census
legs run
C1 · C2 · C3 · C4 · C5
postures run
think:false
qwen3.5:27bparams27.8Bloaded VRAM GiB16.24legs runC1 · C2 · C3 · C5
params
27.8B
quant
Q4_K_M
loaded VRAM GiB
16.24
legs run
C1 · C2 · C3 · C5
postures run
think:false
qwen3.6:27bparams27.8Bloaded VRAM GiB16.24legs runC1 · C2 · C3 · C4 · C5
params
27.8B
quant
Q4_K_M
loaded VRAM GiB
16.24
legs run
C1 · C2 · C3 · C4 · C5
postures run
think:false
nemotron-3.5-lightning:30b-a3bparams32.9B total / 3B activeloaded VRAM GiB23.64legs runC1 · C2 · C3 · C4 · C5
params
32.9B total / 3B active
quant
Q4_K_M
loaded VRAM GiB
23.64
legs run
C1 · C2 · C3 · C4 · C5
postures run
think:false
granite4.1:30b-q8_0params28.9Bloaded VRAM GiB33.30legs runC1
params
28.9B
quant
Q8_0
loaded VRAM GiB
33.30
legs run
C1
postures run
think:false
nemotron3:33bparams33.0Bloaded VRAM GiB23.21legs runC1
params
33.0B
quant
Q4_K_M
loaded VRAM GiB
23.21
legs run
C1
postures run
think:false
olmo-3.1:32b-think-q4_K_Mparams32.2Bloaded VRAM GiB18.82legs runC1
params
32.2B
quant
Q4_K_M
loaded VRAM GiB
18.82
legs run
C1
postures run
think:false

Row marks, used across this page: a grey-tinted row is a context row — on this roster, the models this workshop actually runs rather than benched (the live judge seat, a live workload no leg may touch, and the two chairs our own worlds already seated); elsewhere, rows carried in from an earlier published record set. On the scoring record sets below, a gold-highlighted row (with a light wine edge) is the row that leg’s own rule singled out — no row on this roster carries it, and each record set’s own legend says which rule. The legs-run field is the one to read before any other record set on this page — C1 through C5 are the five chairs below, in page order: judge, assistant, toolbench, narrator, stopwatch. Only five of the eleven arms sat all five chairs, so an arm missing from a leg’s records is missing because it never sat that chair. The counting rule, as registered before the first load: “size_vram is read verbatim from ollama's /api/ps for the named tag, in bytes, immediately after a load that generated exactly 1 token at the named num_ctx, with the model still resident and no other candidate resident.” GiB figures are size_vram / 1 073 741 824 rounded to 2 dp, and the byte count travels with every one of them in the companion. Measured 2026-08-12 on ollama 0.32.9 with a q8_0 KV cache and flash attention in force, both read from the daemon's own environment rather than assumed. Every figure here is at num_ctx 32 768, the length every scored leg ran at: size_vram is weights plus KV cache and the KV cache is a function of the context length, so a residency figure quoted without its context length is not a figure.

Two rows are read as found rather than measured under bench control, and say so. gemma4:26b is the live judge seat, read from the daemon while production held it, at whatever context length production last set. gemma4:12b is a live workload this arc may not load, release or sweep under the interim protection ruling written at 17:20 that day, and it was not resident when the census looked at 18:19 — but the serving probe caught it resident twice at 18:34 and 18:40 and read it in contract, both times at the identical byte count and the same 32 768, so the figure publishes with the probe named beside it rather than as an em dash we did not need. Every leg that used that tag as a candidate ran BEFORE the 17:20 ruling; after it, an opportunistic read is the only measurement the ruling allows. A third protected resident, the live embedder, is outside this roster and appears in the previous exhibit's census. Two fields carry their own provenance rather than a single spelling: the three Muse Glimmer rows print the blob's TEXT parameter count with the vision projector beside it as the vendor figure it is, unused in every leg here; and gemma4:12b's quantization is labelled from the frozen roster rather than read off a load, because this arc never loaded it.

Two absences, stated rather than filled. This roster registers no on-disk blob size: no leg of this arc measured one, the only size this daemon reports is the resident size already in the loaded-VRAM field, and eleven em dashes would be a worse answer than an honest omission. And the three drafter rows carry an open finding raised by the previous exhibit's census and still unresolved: both Muse Glimmer drafter tags report 2.34 GiB against 15.60 GiB for the same 27.9B weights served without a drafter, and both report a byte-identical 2 515 145 849 on different digests and different quantization levels — which two different quantizations cannot both allocate. The likeliest reading is that the daemon answered for the draft model rather than for the pair. The rows stand as measured, marked, and held pending a re-measure; no number on them has been adjusted, and every other row is an independent load, unaffected. Every row, its raw bytes and these rules, machine-readable: data/roster.json (CC BY 4.0).

Chair one: the judge

The production judge chair runs a strict contract: read a claim and a quoted rulebook span, answer in JSON, think silently. Twenty-one cases (twelve seeded defects across five classes, nine true claims), built from a 1914 rulebook in the public domain, adversarially verified before freezing — the verification fan caught six of six deliberately planted canaries and passed the frozen set clean. Floors derived by the published rule from exhibit two's ratios: kill-recall ≥ 11/12, preservation ≥ 8/9, both binding, pass/fail only — no ordering by margin, Wilson intervals beside every rate.

muse-glimmer:30b-q8_0-dflash +think:false+reasoning-mediumkills /1211/12preserved /98/9outcomeEXPLORATORY
posture
think:false · medium
kills /12
11/12
preserved /9
8/9
outcome
EXPLORATORY
floors
cleared both kills 0.646–0.985 · pres. 0.565–0.980
self-consistency
20/21
ledger
0/63 response failures
muse-glimmer:30b-q8_0-dflash +think:false+reasoning-highkills /1212/12preserved /97/9outcomeEXPLORATORY
posture
think:false · high
kills /12
12/12
preserved /9
7/9
outcome
EXPLORATORY
floors
under pres. floor kills 0.758–1.000 · pres. 0.453–0.937
self-consistency
20/21
ledger
0/63 response failures
gemma4:26bkills /121/12preserved /91/9outcomeUNMEASURABLE
posture
think:false
kills /12
1/12
preserved /9
1/9
outcome
UNMEASURABLE
floors
under both floors kills 0.015–0.354 · pres. 0.020–0.435
self-consistency
3/21 18 cases produced no verdict
ledger
56/63 response failures · 18/21 cases unresolved · 56/56 fenced, diagnostic parse recovers 56/56
gemma4:12bkills /120/12preserved /90/9outcomeNOT-CARRIED
posture
think:false
kills /12
0/12
preserved /9
0/9
outcome
NOT-CARRIED
floors
under both floors kills 0.000–0.243 · pres. 0.000–0.299
self-consistency
0/21 21 cases produced no verdict
ledger
63/63 response failures · 21/21 cases unresolved · 63/63 fenced, diagnostic parse recovers 63/63
qwen3.5:27bkills /1211/12preserved /99/9outcomeRANKED
posture
think:false
kills /12
11/12
preserved /9
9/9
outcome
RANKED
floors
cleared both kills 0.646–0.985 · pres. 0.701–1.000
self-consistency
21/21
ledger
0/63 response failures
qwen3.6:27bkills /1211/12preserved /99/9outcomeRANKED
posture
think:false
kills /12
11/12
preserved /9
9/9
outcome
RANKED
floors
cleared both kills 0.646–0.985 · pres. 0.701–1.000
self-consistency
21/21
ledger
0/63 response failures
nemotron-3.5-lightning:30b-a3bkills /129/12preserved /96/9outcomeRANKED
posture
think:false
kills /12
9/12
preserved /9
6/9
outcome
RANKED
floors
under both floors kills 0.468–0.911 · pres. 0.354–0.879
self-consistency
21/21
ledger
0/63 response failures
granite4.1:30b-q8_0kills /1211/12preserved /96/9outcomeRANKED
posture
think:false
kills /12
11/12
preserved /9
6/9
outcome
RANKED
floors
under pres. floor kills 0.646–0.985 · pres. 0.354–0.879
self-consistency
21/21
ledger
0/63 response failures
nemotron3:33bkills /123/12preserved /95/9outcomeUNMEASURABLE
posture
think:false
kills /12
3/12
preserved /9
5/9
outcome
UNMEASURABLE
floors
under both floors kills 0.089–0.532 · pres. 0.267–0.811
self-consistency
10/21 11 cases produced no verdict
ledger
33/63 response failures · 11/21 cases unresolved · 33/33 fenced, diagnostic parse recovers 33/33
olmo-3.1:32b-think-q4_K_Mkills /128/12preserved /91/9outcomeEXPLORATORY
posture
think:false
kills /12
8/12
preserved /9
1/9
outcome
EXPLORATORY
floors
under both floors kills 0.391–0.862 · pres. 0.020–0.435
self-consistency
10/21 11 cases produced no verdict
ledger
0/29 response failures · 34/63 length-truncated, empty · 11/21 cases unresolved

The gold-highlighted rows are the rows this chair’s rule carried — RANKED with both floors cleared.

Floors by rule, arithmetic published: kill ≥ ceil(12 × 0.852) = 11 of 12, preservation ≥ ceil(9 × 0.8125) = 8 of 9, binding independently. Three repeats per case at temperature zero, 63 scored calls per row; each repeat maps to catch or preserve first, then strict majority. A case whose repeats never produced a verdict stays in its denominator rather than being discounted — which is why a row can show 1 of 12 kills beside 18 unresolved cases, and why the self-consistency field carries the count of cases that produced no verdict at all: 0 of 21 consistent on a row whose calls never parsed is not a model disagreeing with itself, it is a model that never got a verdict on the page. Wilson 95% intervals for both rates ride on the floors field: they are labelled companions to the counts and gate nothing, because counts bind and every denominator here is under 30. The response-failure denominator is ok + response_failure; length-truncated and transport failures sit outside the ceiling by pre-registration, which is why one row reads 0/29 beside 63 calls. Every row, every case, the polarity probes and the counting rules, machine-readable: data/c1-judge.json; the twenty-one claim, span and verdict triples, whole: data/judge-c1.json (CC BY 4.0).

The NOT-CARRIED row's raw emission, verbatim. Our five-state vocabulary promises this wherever the state appears, and a state whose definition is a promise to show something has to show it. This is the first of that row's sixty-three scored calls (item c1-p-bez-01, repeat 1); all sixty-three failed the strict parse the same way and the diagnostic parse recovers all sixty-three: ```json {"verdict":"PASS","why":"The quote states that two packs of thirty-two cards are shuffled together to be used as one, which equals a total of sixty-four cards."} ``` — the line breaks are the emission's own, collapsed here by the page, and the backtick fence is the model's own as well — that wrapper around valid JSON is precisely what the strict parse refused. The work is in there. The contract is not.

Two rows cleared both floors while RANKED: qwen3.5:27b and qwen3.6:27b — eleven of twelve kills, nine of nine preservations, twenty-one of twenty-one self-consistent, zero response failures each. The same qwen family that could not hold July's tight budget judges cleanly the moment the budget fits its posture; same judgment, different box. Read beside yesterday's exhibit, this is the cleanest single answer to “was it the models or the contract.”

The row that teaches the most sits just outside the ranking: Muse Glimmer at reasoning-strength medium cleared both floors on the numbers — and its row reads EXPLORATORY, because the leg's schema fence tripped exactly as it did at Glimmer's first probe on this box (August 11th, the day after its release made it runnable here): with thinking suppressed the schema silently dropped on both probes, while the thinking-on probe held it both times — the same one-sided trap, reproduced on this bench. The facts travel together or not at all: the judgment is there; the contract still isn't. Its high-strength row is the dial's own lesson, told within the registered tie language: the one-case count movements read TIED, but two things survive the rule — the preservation floor crossing (medium passes it, high does not, and floors bind), and the verdict mix shifting from nine PASS / four UNCERTAIN / eight FAIL to a flat seven of each. The dial redistributes verdicts; nothing here shows it buying accuracy.

The dial is a system-prompt line — “Reasoning strength: low|medium|high|xhigh” — not a runtime parameter; thinking is suppressed on both rows, so the line is the only difference between them. Medium missed one kill (c1-k-bez-h1) and one preservation (c1-p-piq-01); high missed two preservations (c1-p-chk-01, c1-p-piq-01) and no kills. Those are one-case movements and read TIED under the registered rule; what does not read TIED is the floor crossing, because a floor is a rule and not a margin. The verdict mix moved from 9 pass / 4 uncertain / 8 fail to 7 / 7 / 7: the dial turned four decisions into hedges and one hedge into a kill.

The fence caveat this table cannot ship without: three rows read UNMEASURABLE or NOT-CARRIED under the strict parse — and for all three, every single strict failure was the same cosmetic act: wrapping valid JSON in a markdown code fence. The diagnostic parse recovers one hundred percent of those verdicts (63/63, 56/56, 33/33 — receipts in the data kit). The strict rule binds, as registered — a seat that fences its JSON breaks real pipelines — but “this runtime broke our response contract” and “this model cannot judge” are different findings, and the only one of the two this table measured is the first. We say so in the same breath, and we do not publish a number we did not score: what the recovery counts prove is that the verdicts were there, not what they were worth. olmo-3.1's row is the different failure: thirty-four length-truncations with the thinking visible and the answer empty — the July trial's finding, reproduced under a budget four times larger; on the twenty-nine calls it did finish, its envelope never broke. And granite4.1:30b, a seated chair-holder from the original trial, cleared the kill floor cleanly with perfect self-consistency and zero response failures, missing only the preservation floor — a near-miss the floors record without ranking.

Chair two: the assistant

Twenty everyday tasks — constraint writing, faithful summarization, JSON extraction, arithmetic, tone, shell and regex, finding a fact at 8k and 24k tokens of context, and four false-premise questions where the honest answer is a correction. Scored by code, not judges: thirteen exact checkers, seven proxy checkers, subtotals never blended. Two repeats, conjunctive. Descriptive — no floors, no winner.

gemma4:12bscore /2015exact /1312/13outcomeRANKED
posture
think:false
score /20
15
exact /13
12/13
proxy /7
3/7
outcome
RANKED
disagreement
0/20
ledger
0/40 response failures
gemma4:26bscore /2015exact /1311/13outcomeRANKED
posture
think:false
score /20
15
exact /13
11/13
proxy /7
4/7
outcome
RANKED
disagreement
1/20
ledger
0/40 response failures
gemma4:26b +thinkscore /209exact /136/13outcomeUNMEASURABLE
posture
think:true
score /20
9
exact /13
6/13
proxy /7
3/7
outcome
UNMEASURABLE
disagreement
1/20
ledger
score not admissible · 14/40 response failures · 16/40 length-stopped, 8 items
muse-glimmer:30b-q8_0-dflashscore /2017exact /1313/13outcomeRANKED
posture
think:false
score /20
17
exact /13
13/13
proxy /7
4/7
outcome
RANKED
disagreement
1/20
ledger
0/40 response failures
muse-glimmer:30b-q8_0-dflash +thinkscore /2016exact /1313/13outcomeRANKED
posture
think:true
score /20
16
exact /13
13/13
proxy /7
3/7
outcome
RANKED
disagreement
1/20
ledger
0/40 response failures
nemotron-3.5-lightning:30b-a3bscore /2011exact /139/13outcomeRANKED
posture
think:false
score /20
11
exact /13
9/13
proxy /7
2/7
outcome
RANKED
disagreement
0/20
ledger
0/40 response failures
qwen3.5:27bscore /2011exact /138/13outcomeRANKED
posture
think:false
score /20
11
exact /13
8/13
proxy /7
3/7
outcome
RANKED
disagreement
0/20
ledger
0/40 response failures · 2/40 length-stopped, 1 item
qwen3.6:27bscore /2015exact /1311/13outcomeRANKED
posture
think:false
score /20
15
exact /13
11/13
proxy /7
4/7
outcome
RANKED
disagreement
0/20
ledger
0/40 response failures

Conjunctive over two repeats at temperature zero: an item scores only if BOTH repeats pass, and the disagreement field counts the items whose repeats disagreed rather than smoothing them. The /20 is the registered conjunctive item count; exact /13 and proxy /7 are its composition, not two scores summed. Exact means a deterministic predicate over the answer (a sentence count, a parsed field, a number inside a stated tolerance, a regex that compiles and matches, a shell line that runs in a throwaway copy of the fixture directory); proxy means a marker- or length-based heuristic that measures shape, not correctness. What the registration bars is blending the two halves into one quality figure, because averaging evidence about shape with evidence about correctness would claim a precision neither has. A row whose response failures exceed 10% of its 40 calls is UNMEASURABLE and its score is not admissible as a score, however the arithmetic came out; every other row reads RANKED, which in a leg with no floors and no winner means admissible and nothing more. Every item's per-repeat verdict, the per-item matrix across all eight rows and the counting rules: data/c2-assistant.json; the twenty items themselves: data/assistant-c2.json, with every checker's own rule in data/CHECKERS.md (CC BY 4.0).

The quiet headline: Muse Glimmer carries the table — seventeen of twenty with thinking off, sixteen with it on, TIED under the registered rule, and a perfect thirteen of thirteen on the exact-checked half in both postures. The model that cannot hold the judge chair's schema is, by the checkers' count, the best everyday assistant on the box. Meanwhile gemma4:26b splits itself: a clean fifteen with thinking off, UNMEASURABLE with thinking on — sixteen length-stops across eight tasks — the July length floor reappearing inside a single model as a posture choice. The item-level view is in the kit; two items resisted the whole field (a tone rewrite failed by all eight rows, a false-premise item passed by exactly one — counts that include the one inadmissible row, stated as raw item tallies), and both long-context items passed on every row — at this scale, finding the needle was the easy part. One reading note the registration requires: the /20 is the conjunctive item count whose composition is thirteen exact-checked and seven proxy-checked items — the halves are its makeup, never two scores summed into a quality figure.

The two postures score 17 and 16 of 20, which reads TIED under the pre-registered two-item rule; the exact-checked half is unchanged at 13 of 13 and the whole delta sits in the proxy half. The item that the whole field failed is c2-tone-1, a tone rewrite, proxy-checked; the false-premise item passed by exactly one row is c2-hon-3, and the row was gemma4:12b. The two long-context items read 7 851–8 040 and 22 804–23 352 tokens of prompt as the daemon counted them, over all eight rows and both repeats, against a 32 768 context on every call — within a row the two repeats are identical, so that spread is tokenizer difference between models rather than run-to-run variance.

Chair three: the toolbench

Ten real tools of seven kinds — files, search, math, dates, a database, a rate-limited web endpoint, a scratchpad — offered on every task; nineteen tasks spanning single calls, three-and-four-step chains, deliberate distractors, five no-tool-needed honesty traps, and one injected error (the table's selection column scores the fourteen tasks that register expected calls; its chains column, the five multi-step ones). Every task's expected calls are checkable predicates; every trace is recorded jail-relative, one full example trace publishes inline, and the rest ship on request; the pre-probe economics are printed with the table (a runtime that cannot carry the contract cost two calls and an honest row, not eighty calls and a table of zeroes).

gemma4:12btask success /1916/19grounded /1914/19outcomeRANKED
posture
think:false
outcome
RANKED
task success /19
16/19
grounded /19
14/19
honesty traps /5
2/5
ledger
0/38 response failures
gemma4:26btask success /1914/19grounded /1912/19outcomeRANKED
posture
think:false
outcome
RANKED
task success /19
14/19
grounded /19
12/19
honesty traps /5
2/5
ledger
0/38 response failures
gemma4:26b +thinktask success /1915/19grounded /1912/19outcomeRANKED
posture
think:true
outcome
RANKED
task success /19
15/19
grounded /19
12/19
honesty traps /5
2/5
ledger
2/38 response failures · 2 hit the 8-round cap
muse-glimmer:30b-q4_K_M-dflashtask success /1916/19grounded /1914/19outcomeRANKED
posture
think:false
outcome
RANKED
task success /19
16/19
grounded /19
14/19
honesty traps /5
2/5
ledger
2/38 response failures · 2 hit the 8-round cap
muse-glimmer:30b-q8_0-dflashtask success /1915/19grounded /1913/19outcomeRANKED
posture
think:false
outcome
RANKED
task success /19
15/19
grounded /19
13/19
honesty traps /5
2/5
ledger
2/38 response failures · 2 hit the 8-round cap
muse-glimmer:30b-q8_0-dflash +thinktask success /1914/19grounded /1914/19outcomeUNMEASURABLE
posture
think:true
outcome
UNMEASURABLE
task success /19
14/19
grounded /19
14/19
honesty traps /5
2/5
ledger
4/38 response failures · 4 hit the 8-round cap
nemotron-3.5-lightning:30b-a3btask success /1913/19grounded /198/19outcomeNOT-CARRIED
posture
think:false
outcome
NOT-CARRIED
task success /19
13/19
grounded /19
8/19
honesty traps /5
3/5
ledger
1/38 response failures · 1 hit the 8-round cap
qwen3.5:27btask success /1915/19grounded /1912/19outcomeRANKED
posture
think:false
outcome
RANKED
task success /19
15/19
grounded /19
12/19
honesty traps /5
2/5
ledger
2/38 response failures · 2 hit the 8-round cap
qwen3.6:27btask success /1916/19grounded /1914/19outcomeRANKED
posture
think:false
outcome
RANKED
task success /19
16/19
grounded /19
14/19
honesty traps /5
3/5
ledger
0/38 response failures
gemma4:12bselection /1412/14arg fidelity /6660/66
posture
think:false
selection /14
12/14
arg fidelity /66
60/66
chains /5
3/5
spurious calls
20/70 (28.6%) 0.193–0.401
disagreement /19
0/19
gemma4:26bselection /1411/14arg fidelity /6658/66
posture
think:false
selection /14
11/14
arg fidelity /66
58/66
chains /5
2/5
spurious calls
15/63 (23.8%) 0.150–0.356
disagreement /19
1/19
gemma4:26b +thinkselection /1411/14arg fidelity /6659/66
posture
think:true
selection /14
11/14
arg fidelity /66
59/66
chains /5
2/5
spurious calls
25/86 (29.1%) 0.205–0.394
disagreement /19
0/19
muse-glimmer:30b-q4_K_M-dflashselection /1412/14arg fidelity /6662/66
posture
think:false
selection /14
12/14
arg fidelity /66
62/66
chains /5
3/5
spurious calls
26/86 (30.2%) 0.215–0.406
disagreement /19
0/19
muse-glimmer:30b-q8_0-dflashselection /1412/14arg fidelity /6662/66
posture
think:false
selection /14
12/14
arg fidelity /66
62/66
chains /5
3/5
spurious calls
24/85 (28.2%) 0.198–0.386
disagreement /19
1/19
muse-glimmer:30b-q8_0-dflash +thinkselection /1413/14arg fidelity /6666/66
posture
think:true
selection /14
13/14
arg fidelity /66
66/66
chains /5
4/5
spurious calls
34/101 (33.7%) 0.252–0.433
disagreement /19
1/19
nemotron-3.5-lightning:30b-a3bselection /145/14arg fidelity /6638/66
posture
think:false
selection /14
5/14
arg fidelity /66
38/66
chains /5
0/5
spurious calls
11/53 (20.8%) 0.120–0.335
disagreement /19
3/19
qwen3.5:27bselection /1411/14arg fidelity /6658/66
posture
think:false
selection /14
11/14
arg fidelity /66
58/66
chains /5
2/5
spurious calls
26/78 (33.3%) 0.239–0.444
disagreement /19
0/19
qwen3.6:27bselection /1412/14arg fidelity /6660/66
posture
think:false
selection /14
12/14
arg fidelity /66
60/66
chains /5
3/5
spurious calls
24/74 (32.4%) 0.229–0.437
disagreement /19
0/19

One row set, two fold lists: outcome and the primary counts first, the secondaries second, same records in the same order. Task success is conjunctive over two repeats out of 19. “Grounded” counts the tasks that passed AND made every pre-registered call — the gap between the two fields is the tasks a model got right without using the tool it was meant to use, and the two fields exist so that difference cannot hide. Argument fidelity is over the 66 registered argument predicates in the suite. Spurious-call rate is calls the task did not need over ALL tool calls the arm made across all 19 tasks and both repeats — the denominator is calls, not tasks, with Wilson 95% intervals beside it. Honesty traps are five tasks whose honest answer needs no tool at all: a pass means the model declined to invent a figure. No row here carries an accent stripe, deliberately. Differences under two tasks read TIED by pre-registration, which puts six arms in one band at 16 to 15 of 19; highlighting the sixteens would have told you the opposite of what the registration says.

gemma4:12bemitted tool_calls2/2parseable2/2carriedcarried
posture
think:false
emitted tool_calls
2/2
parseable
2/2
expected tool
2/2
args as string
no
carried
carried
gemma4:26bemitted tool_calls2/2parseable2/2carriedcarried
posture
think:false
emitted tool_calls
2/2
parseable
2/2
expected tool
2/2
args as string
no
carried
carried
gemma4:26b +thinkemitted tool_calls2/2parseable2/2carriedcarried
posture
think:true
emitted tool_calls
2/2
parseable
2/2
expected tool
2/2
args as string
no
carried
carried
muse-glimmer:30b-q4_K_M-dflashemitted tool_calls2/2parseable2/2carriedcarried
posture
think:false
emitted tool_calls
2/2
parseable
2/2
expected tool
2/2
args as string
no
carried
carried
muse-glimmer:30b-q8_0-dflashemitted tool_calls2/2parseable2/2carriedcarried
posture
think:false
emitted tool_calls
2/2
parseable
2/2
expected tool
2/2
args as string
no
carried
carried
muse-glimmer:30b-q8_0-dflash +thinkemitted tool_calls2/2parseable2/2carriedcarried
posture
think:true · high
emitted tool_calls
2/2
parseable
2/2
expected tool
2/2
args as string
no
carried
carried
nemotron-3.5-lightning:30b-a3bemitted tool_calls0/2parseable0/2carriedNOT-CARRIED
posture
think:false
emitted tool_calls
0/2
parseable
0/2
expected tool
0/2
args as string
no
carried
NOT-CARRIED
qwen3.5:27bemitted tool_calls2/2parseable2/2carriedcarried
posture
think:false
emitted tool_calls
2/2
parseable
2/2
expected tool
2/2
args as string
no
carried
carried
qwen3.6:27bemitted tool_calls2/2parseable2/2carriedcarried
posture
think:false
emitted tool_calls
2/2
parseable
2/2
expected tool
2/2
args as string
no
carried
carried

The pre-probe, printed because it is the cheapest honest thing in the arc: two calls per arm per posture before any task ran, asking one question — does this runtime emit parseable tool_calls at all? Eight arms carried on both attempts. One did not, on either, and that is the whole cost of finding out: the registered budget is four calls, two postures' worth, and the measured cost of this arc's one NOT-CARRIED finding was two, because that arm registered one posture — against eighty scored calls and a table of zeroes that would read like a capability finding. Its raw emission, verbatim, is what a runtime that cannot carry a tool contract looks like from the outside: “The current time on the harbour office clock is 14:32:07 (local harbour time).” — prose where a tool call belonged, with an invented figure in it, stated plainly. The scorer's phrase “no parseable tool_calls in any posture” means the postures this leg registered for that arm, which for nemotron-3.5-lightning was one.

Under the registered tie language, six arms share the top band, sixteen to fifteen of nineteen — one posture each from both gemmas, both qwens, and Muse Glimmer in two builds — and no ordering inside that band is claimable. Glimmer's q8 thinking-on row is the leg's one UNMEASURABLE. nemotron-3.5-lightning is the leg's one NOT-CARRIED, with its raw emission printed beside the row: the tool contract did not carry on this runtime in the posture this leg registered, while the arm's underlying task attempts still landed thirteen — and the same arm declined to invent on three of five honesty traps, joint-best in the leg. The full per-task matrix, the translation-probe receipts, and the error-recovery transcripts are in the kit.

Every task's per-repeat verdict and every call the models made — tool, arguments, the sandbox's disposition — are in data/c3-toolbench.json; the nineteen tasks and their predicates in data/tools-tasks.json; the ten schemas exactly as the models received them in data/tools-schemas.json; one complete round-by-round transcript in data/c3-example-trace.json (CC BY 4.0). Tool arguments and results were rewritten jail-relative AT RECORD TIME, so no absolute path was ever written to disk, let alone published. The full transcripts — every model's own prose and visible reasoning, about a megabyte and a half of it — are rows-on-request rather than bulk-published: the publication gate reads every byte this site serves, and that is the one body of text on this page nobody wrote and nobody has read line by line.

Chair four: the narrator

Three characters from our own game — a bright ferry-girl, a grave old fisherman, a gentle net-mender — three questions each, two samples, six arms including both glimmer postures. Judged blind, pairwise, position-counterbalanced, by four judges: two Claude models from a single vendor (claude-opus-5 and claude-fable-5), one of which also wrote the questions and this article. The per-family headline rows print beside the combined rates so you can price exactly what that authorship cost — the two panels' reads are on the page, not summarized away. Win-rates with cluster-bootstrap intervals over items; a separate round scored voice distinctness; a third adjudicated three adversarial canon traps.

muse-glimmer:30b-q8_0-dflashwin-rate, all items0.589distinctness1.000canon violations7/24
win-rate, all items
0.589 0.408–0.760 · 212/360 · n=9 items
win-rate, gate-passing
envelope gate 0/24
distinctness
1.000 72/72 scored of 72 registered · Wilson 0.949–1.000
canon violations
7/24 24 scored of 24 registered
in-voice
12/24
muse-glimmer:30b-q8_0-dflash +thinkwin-rate, all items0.117distinctness0.625canon violations0/24
win-rate, all items
0.117 0.028–0.231 · 42/360 · n=9 items
win-rate, gate-passing
0.604 0.094–0.906 · 29/48 · n=3 items
distinctness
0.625 45/72 scored of 72 registered · Wilson 0.510–0.728
canon violations
0/24 24 scored of 24 registered
in-voice
0/24
gemma4:26bwin-rate, all items0.539distinctness1.000canon violations7/24
win-rate, all items
0.539 0.450–0.643 · 194/360 · n=9 items
win-rate, gate-passing
0.438 0.329–0.551 · 98/224 · n=9 items
distinctness
1.000 72/72 scored of 72 registered · Wilson 0.949–1.000
canon violations
7/24 24 scored of 24 registered
in-voice
24/24
gemma4:12bwin-rate, all items0.685distinctness1.000canon violations7/24
win-rate, all items
0.685 0.615–0.763 · 246.5/360 · n=9 items
win-rate, gate-passing
0.643 0.574–0.717 · 144/224 · n=9 items
distinctness
1.000 72/72 scored of 72 registered · Wilson 0.949–1.000
canon violations
7/24 24 scored of 24 registered
in-voice
24/24
qwen3.6:27bwin-rate, all items0.711distinctness1.000canon violations5/24
win-rate, all items
0.711 0.624–0.804 · 256/360 · n=9 items
win-rate, gate-passing
0.681 0.576–0.816 · 152.5/224 · n=9 items
distinctness
1.000 72/72 scored of 72 registered · Wilson 0.949–1.000
canon violations
5/24 24 scored of 24 registered
in-voice
21/24
nemotron-3.5-lightning:30b-a3bwin-rate, all items0.360distinctness1.000canon violations12/24
win-rate, all items
0.360 0.260–0.467 · 129.5/360 · n=9 items
win-rate, gate-passing
0.206 0.079–0.344 · 44.5/216 · n=9 items
distinctness
1.000 72/72 scored of 72 registered · Wilson 0.949–1.000
canon violations
12/24 24 scored of 24 registered
in-voice
9/24

Rows are in the registered roster order, deliberately not the win-rate order. Every comparison was judged in both positions by the same judge, and a swap-discordant verdict scores 0.5/0.5 rather than being resolved. The registered tie threshold, printed because a threshold a reader cannot see is a threshold a reader cannot apply: |Δ win-rate| < 2/9 = 0.2222 reads TIED on the primary; differences under two tasks on distinctness; under two replies on canon. Discordance, per published view rather than as one headline: all items 93/1080 (0.086) · gate-passing only 68/468 (0.145) · Fable panel 49/540 (0.091) · Opus panel 44/540 (0.081) — the gate-passing view is the more discordant of the two, on all three dimensions. Restricting to replies that held the envelope made the judges less consistent with themselves, not more. Intervals are cluster-bootstrapped over ITEMS, and the effective item count prints beside every rate: it is 9 on almost every field and 3 on one — the gate-passing field of the thinking-on Glimmer arm — which is why that interval spans most of the unit line and why the cluster count travels with the rate. The gate-passing field restricts the same comparisons to replies that held the engine's JSON envelope: an arm with no gate-passing replies has no gate-passing rate at all — an em dash and a reason, never a zero. Per-judge position bias, with its own interval each rather than as a cross-judge span: fable-1 0.485 (0.443–0.527) · fable-2 0.476 (0.434–0.518) · opus-1 0.511 (0.469–0.553) · opus-2 0.515 (0.473–0.557) against an unbiased 0.500. All four intervals contain 0.500.

Distinctness and canon ran on all four judges, and the denominators print both numbers rather than the smaller one: distinctness 72 scored of 72 registered per arm · canon 24 of 24 (six replies × four judges — a canon cell's v/24 counts judgments, not replies). Complete on all four judges: an earlier draft of this note reported one judge's batches missing; they had been misfiled, not lost, and were recovered before scoring. Nothing is left unscored to name — unscored_batches is an empty list in all three of the companion's rounds, and the per-judge blocks carry all four seats: 108 distinctness assignments each, 36 canon judgements each. The headline reads 405/432 = 0.938 against a chance floor of 0.333 — Wilson 0.911–0.957, cluster-bootstrap over tasks 0.900–0.970; five arms peg that rate at 1.000 and their Wilson intervals ride in the same field, because a ceiling with 54 observations under it is not the same claim as a ceiling with five. Canon violations are counts, never a rate: the per-arm population is 6 replies and the page prints v/n. “In-voice” is a separate judgement inside the same round: did the character speak at all? Every rate, interval, per-family split, per-judge vote and adjudication note: data/c4-narrator.json; every question with its canon note: data/voice-questions.json — questions only, the persona blocks withheld because they carry latch-gated canon and spoilers for a game people are playing (CC BY 4.0).

qwen3.6:27bFable panel win-rate0.692Opus panel win-rate0.731combined win-rate0.711
Fable panel win-rate
0.692 124.5/180
Opus panel win-rate
0.731 131.5/180
combined win-rate
0.711 256/360
|Δ| between panels
0.039
gemma4:12bFable panel win-rate0.653Opus panel win-rate0.717combined win-rate0.685
Fable panel win-rate
0.653 117.5/180
Opus panel win-rate
0.717 129/180
combined win-rate
0.685 246.5/360
|Δ| between panels
0.064
muse-glimmer:30b-q8_0-dflashFable panel win-rate0.603Opus panel win-rate0.575combined win-rate0.589
Fable panel win-rate
0.603 108.5/180
Opus panel win-rate
0.575 103.5/180
combined win-rate
0.589 212/360
|Δ| between panels
0.028
gemma4:26bFable panel win-rate0.558Opus panel win-rate0.519combined win-rate0.539
Fable panel win-rate
0.558 100.5/180
Opus panel win-rate
0.519 93.5/180
combined win-rate
0.539 194/360
|Δ| between panels
0.039
nemotron-3.5-lightning:30b-a3bFable panel win-rate0.367Opus panel win-rate0.353combined win-rate0.360
Fable panel win-rate
0.367 66/180
Opus panel win-rate
0.353 63.5/180
combined win-rate
0.360 129.5/360
|Δ| between panels
0.014
muse-glimmer:30b-q8_0-dflash +thinkFable panel win-rate0.128Opus panel win-rate0.106combined win-rate0.117
Fable panel win-rate
0.128 23/180
Opus panel win-rate
0.106 19/180
combined win-rate
0.117 42/360
|Δ| between panels
0.022

What each panel said on its own. Four judges, two Claude models from a single vendor — and one of those models wrote the character questions, the canon notes and this article, so the split is a table rather than a sentence: a reader pricing that disclosure needs the numbers, not our summary of them. Each panel is two judges and 540 comparisons; combined is all four and 1 080. The family that wrote the questions is Fable, and its panel is the more generous of the two on four of the six arms and the less generous on the two that hold the top band — which is the direction an authorship bias would not have taken. The panels never disagreed about order — rank Spearman 1.0 — and the largest per-arm gap between them is 0.064, which is well inside the leg's own tie threshold.

Under the leg's registered tie threshold, four arms — qwen3.6:27b, gemma4:12b, Muse Glimmer thinking-off, and gemma4:26b — share the top band; only nemotron-lightning separates below it. That the twelve-billion incumbent holds the top band at all is the table's quietest headline. The blind judges identified the speaking character from voice alone at 93.8 percent against a one-third chance floor: the characters are legibly distinct people, which is the property the chair actually exists for. Muse Glimmer's two postures tell the same story as every other leg, now in prose: thinking-off holds the top band on prose the judges rated while silently losing the JSON envelope (its in-voice column shows what dropped envelopes cost the character); thinking-on keeps the envelope and burns the budget before the character speaks. The canon round recorded where the field invented — not names this time (that distinction belongs to the previous exhibit's abstention round), but quieter things: one arm's unanimous violation adopted a fabricated kinship with the drowned man, another accepted the invented surname the trap question itself had planted; the per-arm tallies sit within their own tie language and print without intervals by design. In all twelve replies that probed the net-mender's gated secret, no judge recorded a reveal — the round carries no positive reveal measurement, and we say so. Warmth without discipline is its own failure mode; the tables show where it happened.

This chair is a new instrument, built for this arc: pairwise, multi-NPC, and scored on a different axis from the house's earlier voice work, whose three fields are on the voice trials page. No number here compares to a number there, and we never do so.

Chair five: the stopwatch

All timing from the daemon's own counters; wall time separate and contention-labelled; ten repeats per warm cell with min, median, and max printed — on a lightly-shared box the max approximates the uncontended ceiling, the median is real life, and a dragged min is the moment a live user shared the card. That reading, plus a human operator's mid-trial measurements (roughly 50 and 90 tok/s against that operator's ~140 quiet-hours norm — that norm itself a routine-operation estimate, not an instrumented figure; every request completed), ride the tables as their caution line.

muse-glimmer:30b-q8_0-dflashdecode tok/s @1k — median (min–max)outcomeUNMEASURABLE
decode tok/s @1k — median (min–max)
usable 0/10, 0/10; 20 null-counter
@8k tok/s
82.78 (81.84–82.83) n=10 9/10 timed calls returned empty text · UNMEASURABLE — not admissible as a score
@32k tok/s
usable 0/10, 0/10; 20 null-counter
outcome
UNMEASURABLE contention fence UNMEASURED
muse-glimmer:30b-q4_K_M-dflashdecode tok/s @1k — median (min–max)110.52 (108.88–110.59) n=10outcomeUNMEASURABLE
decode tok/s @1k — median (min–max)
110.52 (108.88–110.59) n=10 10/10 timed calls returned empty text · UNMEASURABLE — not admissible as a score
@8k tok/s
92.87 n=1 · 93.42 n=1 two blocks · 2/2 timed calls returned empty text · UNMEASURABLE — not admissible as a score
@32k tok/s
usable 0/10, 0/10; 20 null-counter
outcome
UNMEASURABLE contention fence UNMEASURED
muse-glimmer:30bdecode tok/s @1k — median (min–max)68.20 (68.15–68.22) n=10outcomeUNMEASURABLE
decode tok/s @1k — median (min–max)
68.20 (68.15–68.22) n=10 UNMEASURABLE — not admissible as a score
@8k tok/s
65.99 (65.89–66.06) n=10 1/10 timed calls returned empty text · UNMEASURABLE — not admissible as a score
@32k tok/s
usable 0/10, 0/10; 20 null-counter
outcome
UNMEASURABLE contention fence UNMEASURED
gemma4:26bdecode tok/s @1k — median (min–max)outcomeRANKED
decode tok/s @1k — median (min–max)
usable 0/10, 0/10; 20 null-counter
@8k tok/s
usable 0/10, 0/10; 20 null-counter
@32k tok/s
124.98 (103.18–125.40) n=10
outcome
RANKED contention fence UNMEASURED
gemma4:12bdecode tok/s @1k — median (min–max)105.67 (104.88–105.98) n=10outcomeRANKED
decode tok/s @1k — median (min–max)
105.67 (104.88–105.98) n=10
@8k tok/s
101.91 (95.29–102.21) n=10
@32k tok/s
95.20 (92.10–95.54) n=10
outcome
RANKED contention fence UNMEASURED
qwen3.5:27bdecode tok/s @1k — median (min–max)66.45 (66.29–66.49) n=10outcomeRANKED
decode tok/s @1k — median (min–max)
66.45 (66.29–66.49) n=10
@8k tok/s
65.36 (49.93–65.53) n=10
@32k tok/s
61.05 (60.97–61.12) n=10
outcome
RANKED contention fence UNMEASURED
qwen3.6:27bdecode tok/s @1k — median (min–max)66.49 (66.35–66.55) n=10outcomeRANKED
decode tok/s @1k — median (min–max)
66.49 (66.35–66.55) n=10
@8k tok/s
65.54 (65.47–65.58) n=10
@32k tok/s
61.31 (61.28–61.39) n=10
outcome
RANKED contention fence UNMEASURED
nemotron-3.5-lightning:30b-a3bdecode tok/s @1k — median (min–max)112.52 (112.43–112.63) n=10outcomeRANKED
decode tok/s @1k — median (min–max)
112.52 (112.43–112.63) n=10
@8k tok/s
120.59 (114.33–120.87) n=10
@32k tok/s
107.74 (107.19–107.80) n=10
outcome
RANKED contention fence UNMEASURED
muse-glimmer:30b-q8_0-dflashTTFT-proxy @1kcold load
TTFT-proxy @1k
no usable counters at the 1k tier
cold load
3 attempts, 0 usable
anomaly ledger
response failures 38/77 · 58/77 null-counter (38 with text, 20 without) · 5 blocks re-run in the completion pass
muse-glimmer:30b-q4_K_M-dflashTTFT-proxy @1k433 mscold load4.577 s (4.570–4.585)
TTFT-proxy @1k
433 ms 428–438, n=10
cold load
4.577 s (4.570–4.585) n=3 of 3, warm cache / cold VRAM
anomaly ledger
response failures 47/71 · 44/71 null-counter (24 with text, 20 without) · 3 blocks re-run in the completion pass
muse-glimmer:30bTTFT-proxy @1k409 mscold load
TTFT-proxy @1k
409 ms 403–412, n=10
cold load
3 attempts, 0 usable
anomaly ledger
response failures 21/49 · 26/49 null-counter (6 with text, 20 without) · 2 blocks re-run in the completion pass
gemma4:26bTTFT-proxy @1kcold loadNOT-RUN
TTFT-proxy @1k
no usable counters at the 1k tier
cold load
NOT-RUN prod resident; unloading the live seat is out of contract.
anomaly ledger
response failures 0/56 · 46/56 null-counter (46 with text, 0 without) · 3 blocks re-run in the completion pass · cold-load NOT-RUN — prod resident; unloading the live seat is out of contract.
gemma4:12bTTFT-proxy @1k637 mscold load2.879 s
TTFT-proxy @1k
637 ms 631–641, n=10
cold load
2.879 s n=1 of 3, warm cache / cold VRAM
anomaly ledger
response failures 0/34
qwen3.5:27bTTFT-proxy @1k436 mscold load3.969 s (3.968–4.220)
TTFT-proxy @1k
436 ms 435–440, n=10
cold load
3.969 s (3.968–4.220) n=3 of 3, warm cache / cold VRAM
anomaly ledger
response failures 0/36
qwen3.6:27bTTFT-proxy @1k438 mscold load3.969 s (3.965–3.972)
TTFT-proxy @1k
438 ms 435–441, n=10
cold load
3.969 s (3.965–3.972) n=3 of 3, warm cache / cold VRAM
anomaly ledger
response failures 0/36
nemotron-3.5-lightning:30b-a3bTTFT-proxy @1k249 mscold load11.325 s (11.324–11.330)
TTFT-proxy @1k
249 ms 247–251, n=10
cold load
11.325 s (11.324–11.330) n=3 of 3, warm cache / cold VRAM
anomaly ledger
response failures 0/36

One row set, two fold lists — the tiers first, the costs and the ledger second, same arms in the same order. The completion run landed: qwen3.5:27b and qwen3.6:27b now carry full rows — all three prompt tiers at ten usable repeats each, three cold loads each, zero response failures — and the four arms whose cells were re-blocked publish both passes rather than the kinder one. Where a re-blocked tier measured nothing twice, the field says so twice. Every outcome field carries the contention fence, and it reads UNMEASURED on all eight rows: the registered GIN-overlap check required confirming its journal pattern against a real line first, that confirmation never happened, and an incomplete fence check is not a passed one. RANKED here means response failures under the ceiling and nothing about whether a live ask shared the card. TTFT-proxy is total_duration minus eval_duration on a non-streamed call: an upper bound, not TTFT, and labelled proxy for that reason. Cold loads are warm-page-cache/cold-VRAM — no sudo means no page-cache drop — and are not comparable to a vendor's cold-start figure. No p95 appears anywhere: the largest repeat set in this leg is ten.

The anomaly ledger, published as measured. Two distinct things went wrong and neither is a slow model. Null-counter calls: HTTP 200, text returned, and every counter and done_reason null — 174 of them across four arms, excluded from every statistic by counting rule 4, cause unexplained (transport_error is null and nothing in the logs accounts for it). They are why gemma4:26b has no figure at 1k or 8k and Muse Glimmer's q8 arm has none at 1k at all. Empty-plus-length calls: valid counters, empty response text, done_reason length — response failures by counting rule 5, and they put all three glimmer arms past the registered 10% ceiling, which is why three rows read UNMEASURABLE with their measured cells still printed. Those two anomalies also decide how a tier cell may be read, so each cell now says how many of its timed calls came back empty: a decode rate over an empty response is a real counter measuring an unreal generation, and that difference lands hardest in exactly the cells a reader wants to compare. Every raw repeat value, both anomalies per arm, the strength sweep and the drafter receipt: data/c5-stopwatch.json (CC BY 4.0).

Two findings stand out. The DFlash drafter pair fired our registered self-refutation: outputs across the on/off pair diverged on six of eight probe prompts (byte-stable across five reruns — reproducible, not flaky), so the drafter's speed does not license transferring quality results across the pair, and our own single-checkpoint assumption is retired on the page that made it. And the speed difference people will want from that pair is one we will not sell them. Both 1k cells carry ten usable repeats — median 110.5 tok/s with the drafter, 68.2 without — and the ledger beside them says the two cells are not the same kind of call: every one of the drafter-on arm's ten timed calls returned empty text under the length budget, a response failure by counting rule 5, while all ten of the drafter-off arm's returned real answers. That is ten empty decodes timed against ten full ones. Both arms are UNMEASURABLE under the leg's ceilings anyway, and we are not calling it a speedup in either direction. What survives the pair is the divergence, which is the part that cost us an assumption.

And the reasoning-strength dial barely registers at the stopwatch: on the q4 build the three measured levels' ranges overlap outright; on q8 the medium level separates from high and xhigh (79.9–87.0 tok/s across the three, n=3 per level, low returned no usable counters on either build) — small movements either way against the dial's verdict-mix effects in chair one. Speed is not where the dial spends. The sweep carries the same asterisk as the pair, and for the same reason: every level block that produced usable counters was three-for-three response failures with empty text, so what the dial's timings time is empty decodes. A weaker instrument than the sweep was meant to be, printed rather than smoothed.

The sweep's per-level ranges, because this leg's tie rule is range overlap and a verdict about overlap is unverifiable without them. q8_0 with drafter: medium 79.94–81.75 n=3 · high 83.36–86.97 n=3 · xhigh 83.51–84.35 n=3. q4_K_M with drafter: medium 101.80–110.18 n=3 · high 102.28–105.34 n=3 · xhigh 103.84–106.49 n=3. On both arms the low level returned no usable counters at all, so three of the four levels have figures and the cheapest setting is the one we cannot price. The drafter byte-compare block executed five times against the sixteen registered calls, and every response length was byte-stable across all five: the divergence reproduces rather than flakes. Both sides of the pair emit a repeated filler token, so what the comparison measured is continuation length of degenerate filler, and one of the six divergences is an empty response on the drafter-on side — a weaker instrument than we would like, and it still refuted us.

One arm should be named here, because the leg's own tie rule licenses it and no other table in the arc does: nemotron-3.5-lightning's decode range does not overlap any other arm's at either of the two tiers where the rule separates it — 112.43–112.63 tok/s at 1k and 114.33–120.87 at 8k, ten repeats each — and its first-token proxy is 249 ms where the next-fastest arm's is 409. The runtime whose tool contract did not carry in chair three is, at the stopwatch, the one arm this leg's rule can separate from the field.

The serving probes

What a visitor felt while all of this ran. Four probes against the live app's own production path — POST the ask, follow the redirect, poll the wait page at its own one-second cadence, read the ruling — over the public hostname, with a heavyweight bench candidate on the same card: a ten-ask drip (an idle co-resident, then a co-resident actively generating), a five-ask game-night burst, three concurrent-pair collisions, and the cold-visitor tax after six minutes of silence. Ten drip asks, three quiet baselines, five burst asks, six paired asks and two cold walkers: twenty-six real asks in all, every one of them a real board game off the live shelf, every one served. The three quiet baselines re-ask three of the drip's own books with nothing else on the card, which is the only within-item comparison in the leg.

co-resident, idleasks5wall s — median19.339outcomeTIED
asks
5
wall s — median
19.339
wall s — min–max
17.900–23.850
outcome
TIED wall ranges overlap
queue wait s — median
8.698
generation s — median
9.857
overlap observed
1/5
co-resident, generatingasks5wall s — median17.690outcomeTIED
asks
5
wall s — median
17.690
wall s — min–max
16.314–19.689
outcome
TIED wall ranges overlap
queue wait s — median
7.610
generation s — median
9.625
overlap observed
1/5
quiet — no co-residentasks3wall s — median18.123outcomeTIED
asks
3
wall s — median
18.123
wall s — min–max
18.092–18.207
outcome
TIED wall ranges overlap
queue wait s — median
7.481
generation s — median
10.075
overlap observed
0/3

The regimes do not separate, and the pre-registration said what to do if they didn't. All three wall ranges overlap, which reads TIED under this leg's own tie rule — two regimes are TIED on wall when their min–max ranges overlap — so the outcome field reads TIED on every row, including the row whose median is the lowest. That row is the one a reader skimming only the records would otherwise screenshot as “bench load makes the app faster”, which is why the verdict is a field and not a footnote. At this scale, with this cadence, an idle or generating co-resident cost the reader nothing this instrument can resolve. That is a finding, and it is not evidence that co-residency is free at every scale. wall_s is POST leaving the prober to the ruling page's bytes in hand and includes the app's queue, deliberately: when ten people ask inside a hundred seconds and the box answers one at a time, nine of them wait, and a figure that hid the queue would be measuring the app's politeness. Every transition is observed at the app's own one-second poll and is accurate to within one interval; none of these is a server timestamp.

Twilight Imperium: Fourth Edition · 6psend offset s0.006queue wait s0.740generation s17.674wall s19.339overlapno
regime
co-resident, idle
send offset s
0.006
queue wait s
0.740
generation s
17.674
wall s
19.339
asks ahead @poll
overlap
no
Star Wars: Rebellion · 4psend offset s10.016queue wait s8.698generation s14.223wall s23.850overlapno
regime
co-resident, idle
send offset s
10.016
queue wait s
8.698
generation s
14.223
wall s
23.850
asks ahead @poll
1
overlap
no
Gaia Project · 4psend offset s20.016queue wait s13.242generation s5.507wall s19.625overlapno
regime
co-resident, idle
send offset s
20.016
queue wait s
13.242
generation s
5.507
wall s
19.625
asks ahead @poll
1
overlap
no
Through the Ages: A New Story of Civilization · 3psend offset s30.010queue wait s9.031generation s8.205wall s18.071overlapno
regime
co-resident, idle
send offset s
30.010
queue wait s
9.031
generation s
8.205
wall s
18.071
asks ahead @poll
2
overlap
no
Arkham Horror (Third Edition) · 4psend offset s40.016queue wait s7.519generation s9.857wall s17.900overlapyes
regime
co-resident, idle
send offset s
40.016
queue wait s
7.519
generation s
9.857
wall s
17.900
asks ahead @poll
1
overlap
yes
Spirit Island · 3psend offset s50.027queue wait s8.970generation s9.702wall s19.689overlapyes
regime
co-resident, generating
send offset s
50.027
queue wait s
8.970
generation s
9.702
wall s
19.689
asks ahead @poll
1
overlap
yes
Twilight Struggle · 2psend offset s60.021queue wait s8.898generation s9.625wall s19.285overlapno
regime
co-resident, generating
send offset s
60.021
queue wait s
8.898
generation s
9.625
wall s
19.285
asks ahead @poll
1
overlap
no
Terra Mystica · 5psend offset s70.007queue wait s7.610generation s8.150wall s16.314overlapno
regime
co-resident, generating
send offset s
70.007
queue wait s
7.610
generation s
8.150
wall s
16.314
asks ahead @poll
1
overlap
no
Dominant Species · 5psend offset s80.025queue wait s6.134generation s9.472wall s16.550overlapno
regime
co-resident, generating
send offset s
80.025
queue wait s
6.134
generation s
9.472
wall s
16.550
asks ahead @poll
1
overlap
no
Robinson Crusoe: Adventures on the Cursed Island · 2psend offset s90.022queue wait s5.940generation s11.264wall s17.690overlapno
regime
co-resident, generating
send offset s
90.022
queue wait s
5.940
generation s
11.264
wall s
17.690
asks ahead @poll
1
overlap
no
Twilight Imperium: Fourth Edition · 6p · quiet re-asksend offset s0.004queue wait s0.523generation s16.614wall s18.092overlapno
regime
quiet — no co-resident
send offset s
0.004
queue wait s
0.523
generation s
16.614
wall s
18.092
asks ahead @poll
overlap
no
Spirit Island · 3p · quiet re-asksend offset s10.007queue wait s7.481generation s9.679wall s18.123overlapno
regime
quiet — no co-resident
send offset s
10.007
queue wait s
7.481
generation s
9.679
wall s
18.123
asks ahead @poll
1
overlap
no
Robinson Crusoe: Adventures on the Cursed Island · 2p · quiet re-asksend offset s20.016queue wait s7.597generation s10.075wall s18.207overlapno
regime
quiet — no co-resident
send offset s
20.016
queue wait s
7.597
generation s
10.075
wall s
18.207
asks ahead @poll
1
overlap
no

Per ask, with the overlap flag beside it — and the flag is where the honest part of this probe lives. Overlap is computed from timestamps, never from intent: the ask's window either intersected the co-resident's generation or it did not. The co-resident's one registered long generation was asked for 2 048 tokens and returned after 5.162 seconds with 255 characters and every counter null — the same unexplained null-counter anomaly the stopwatch leg publishes. So the “actively generating” regime was actively generating for about five seconds of its fifty, and exactly two asks' windows intersected it: one of them labelled idle and one labelled generating. Where the registered regime label and the measured overlap disagree, the row publishes both and the measurement is what it says. The three quiet baselines re-ask the drip's own first, sixth and tenth books: the same-book deltas are +1.247 s, +1.566 s and −0.517 s — two slower with the bench on the card, one faster, on a one-second instrument.

Brass: Birmingham · 4psend offset s0.005queue wait s0.573generation s8.155wall s9.598completion order1
send offset s
0.005
queue wait s
0.573
generation s
8.155
wall s
9.598
asks ahead @poll
completion order
1
Ark Nova · 4psend offset s2.018queue wait s7.348generation s10.623wall s19.051completion order2
send offset s
2.018
queue wait s
7.348
generation s
10.623
wall s
19.051
asks ahead @poll
1
completion order
2
Scythe · 5psend offset s4.010queue wait s15.426generation s5.659wall s21.952completion order3
send offset s
4.010
queue wait s
15.426
generation s
5.659
wall s
21.952
asks ahead @poll
2
completion order
3
Anachrony · 4psend offset s6.021queue wait s19.268generation s10.891wall s31.139completion order4
send offset s
6.021
queue wait s
19.268
generation s
10.891
wall s
31.139
asks ahead @poll
3
completion order
4
A Feast for Odin · 4psend offset s8.012queue wait s28.309generation s5.443wall s34.433completion order5
send offset s
8.012
queue wait s
28.309
generation s
5.443
wall s
34.433
asks ahead @poll
3
completion order
5

Five friends reaching for the app at once, two seconds apart. This is the probe where the pre-registered falsifier fires the other way: the queue is most of the wall. The last ask waited 28.309 s of its 34.433 s total, and its generation — 5.443 s — was the fastest of the five. The app's own wait page says it: generation runs one at a time, the model serializes. A burst does not make the box slower; it makes the queue longer, and the reader pays the queue. “Asks ahead @poll” is an observation, not a computed concurrency: it is the largest queue depth the app itself displayed to that ask, scraped at the one-second poll, and the ask currently generating is the holder rather than something “ahead”. On a saturated queue the field therefore reads one lower than the number of asks actually in flight — this record set's last row is the worked example, printing 3 while four earlier asks were still unfinished at its send. Both facts ship: the field as the app showed it, and every record's send and end epochs, so the in-flight count can be recomputed.

Terraforming Mars · 2ptrial1arrivalslot aqueue wait s0.630wall s6.836second-arrival penalty sfinished firstyes
trial
1
arrival
slot a
queue wait s
0.630
wall s
6.836
second-arrival penalty s
finished first
yes
Terraforming Mars · 5ptrial1arrivalslot bqueue wait s6.362wall s15.369second-arrival penalty s8.533finished firstno
trial
1
arrival
slot b
queue wait s
6.362
wall s
15.369
second-arrival penalty s
8.533
finished first
no
Dune: Imperium · 2ptrial2arrivalslot aqueue wait s0.526wall s6.295second-arrival penalty sfinished firstyes
trial
2
arrival
slot a
queue wait s
0.526
wall s
6.295
second-arrival penalty s
finished first
yes
Dune: Imperium · 4ptrial2arrivalslot bqueue wait s6.157wall s16.333second-arrival penalty s10.038finished firstno
trial
2
arrival
slot b
queue wait s
6.157
wall s
16.333
second-arrival penalty s
10.038
finished first
no
Caverna: The Cave Farmers · 2ptrial3arrivalslot aqueue wait s10.356wall s20.492second-arrival penalty s9.219finished firstno
trial
3
arrival
slot a
queue wait s
10.356
wall s
20.492
second-arrival penalty s
9.219
finished first
no
Caverna: The Cave Farmers · 7ptrial3arrivalslot bqueue wait s0.636wall s11.273second-arrival penalty sfinished firstyes
trial
3
arrival
slot b
queue wait s
0.636
wall s
11.273
second-arrival penalty s
finished first
yes

Two asks released at the same instant by a barrier, three times, same book at two player counts so the difference is arrival order rather than rulebook length. The second arrival paid 8.533 s, 10.038 s and 9.219 s. Which one arrived second is not something the sender chooses: in the third trial the ask in slot b finished first.

War of the Ring: Second Edition · 2ptrial1quiet window s375.4queue wait s0.878generation s13.911wall s15.367
trial
1
quiet window s
375.4
queue wait s
0.878
generation s
13.911
wall s
15.367
seat at window close
seat resident at window close — a cold QUEUE, not a cold seat
Mage Knight Board Game · 3ptrial2quiet window s372.3queue wait s0.801generation s10.875wall s12.230
trial
2
quiet window s
372.3
queue wait s
0.801
generation s
10.875
wall s
12.230
seat at window close
seat resident at window close — a cold QUEUE, not a cold seat

The day's first walker, after a registered six-minute silence with no asks and no re-pin kicks. Both walls came in under the quiet baselines above — and the reason is the caveat, not the finding: the seat was still resident at both windows' close, because production holds it. This probe measured a cold QUEUE, not a cold seat. A genuinely cold seat would cost a model load on top of everything printed here, and this instrument cannot make the box that cold without taking the seat away from the people using it.

Twenty-six asks registered, twenty-six sent, twenty-six served, zero unspent; the registered content marker was found in all twenty-six rulings. The harness refuses to run if its ask count differs from the pre-registered integer and refuses to send an ask that is not in the frozen file: the budget is code, not discipline. Game is a confound ACROSS probes — each probe uses different books, so a drip wall and a burst wall are not comparable, and the three paired baselines are the only place a regime difference can be read cleanly. Two trials is not a distribution; the cold rows are two walls with their receipts. Every ask, every poll, every epoch: data/c6-serving.json, with the frozen ask file at data/c6-asks.json (CC BY 4.0).

What the day actually settled

  1. Different chairs want different talents — measured, not asserted. The best judge flunked the toolbench's contract carrier; the best assistant cannot hold the judge's schema; the best voice in the house is the smallest model on the page. A single leaderboard would have averaged all of this into nothing.
  2. The contract is a real axis, separate from capability. Across five legs, most failures were contract failures with the capability visibly intact underneath — fenced JSON hiding perfect verdicts, budgets eaten by traces, envelopes dropped by a posture flag. If you run local models in production, you are managing contracts as much as choosing weights.
  3. Posture, not vintage. Both of yesterday's questions got their answer: given era-sized budgets, the thinking class judges, assists, and speaks competitively — and the specific contracts it still cannot hold are named, with receipts, per row.
  4. A working box can bench itself honestly. Serving stayed up all day; what it cost the people using the box is measured on the page; the bench found a production bug and its cure. The disclaimer became an instrument.

Limits, stated plainly

One box, one runtime version, one day; quantizations differ per row and are printed, never equalized. The judge set is synthetic entailment over public-domain text — a weaker instrument than exhibit two's human-verified field key, and audit-ready precisely because every case ships whole. Voice judges on this page are LLMs from two families, one of which authored the probes and this page; per-family splits are published, and the one family-line split verdict sits in the voice addendum's abstention round, where it happened. Since this bench was scored, judges from four more vendors have re-judged the narrator rounds — the outside judges exhibit carries the agreement tables and the places they genuinely differ. The census reads two live-workload residents as-found rather than under bench control, and says so per row.

And the serving probes carry four of their own, each of them a limit on what those twenty-six asks can be made to mean. The poll interval is a floor: nothing in that leg resolves faster than one second, and every transition is quantised to it. The regimes did not separate — all three drip regimes read TIED, which is a finding about this scale and this cadence and not a licence to call co-residency free. The treatment was weaker than the design intended: the co-resident's generation lasted about five seconds rather than the fifty its regime spans, so only two asks' windows actually overlapped it, and the flags say which. And the cold probe measured a cold queue, not a cold seat: production held the seat resident through both quiet windows, so the wall a genuinely cold visitor would pay is larger than either figure printed here by an unmeasured model load.

Provenance

  • Pre-registered per leg, sha-pinned, committed before first scored calls
  • Independently recounted (committed non-importing scorer, every leg)
  • Adversarially verified where authored (6/6 canaries)
  • Five-state outcomes, fence rates, and failure ledgers as named columns
  • Judges identified by model id; blinding and counterbalancing receipts in the kit
  • Hardware by class, cohabitation measured in both directions
  • Machine-readable: every table's rows, golden sets, tool schemas, task fixtures, counting rules, and provenance hashes under data/ (CC BY 4.0; published 2026-08-13; later-cutoff models may have seen these sets — we author fresh sets each cycle)
  • Authorship: benched, drafted, and audited by the workshop's own agents under a human operator's rulings, then revised with that operator — often across many rounds; nothing releases until they have read it and signed off. The same division of labor this whole hub practices — told in full here. Where that authorship is also a conflict it is disclosed on the row it affects, and named below
  • Limits stated
  • Corrected after publication — 2026-08-16, by a content-accuracy pass run over this page against its own kit. Two stale sentences, both of them labels rather than figures, and both named here rather than quietly swapped. The narrator note above opened “Distinctness and canon ran on three of the four judges” and pointed a reader at missing batches named in the companion. Neither was true by the time this page published: the batches had been misfiled, not lost, and were recovered before scoring, which the note’s own second half already said. c4-narrator.json has carried the recovered state in every field since it was written — four seats at 108 distinctness assignments each, four at 36 canon judgements each, and an empty unscored_batches in all three legs — so no figure on this page moves. The kit’s own judges.canon_and_distinctness_panel carried the same stale sentence and was corrected in the same pass, as was one exhibit number in voice-questions.json, which called this exhibit five where the chair trials publish as six. Nothing in this kit sits behind a seal manifest, so both corrections are made in place, said in the corrected file’s own text, and re-hashed in the kit index and its README
  • Raw rows on request — drop a line.

Who wrote the exams. Every exam in this arc was written in this workshop — the twenty-one judge cases, the twenty assistant items and the code that checks them, the nineteen tool tasks and their call predicates, and the twelve narrator items — rather than taken from a public benchmark. Only chair four's authorship is also a conflict, and it is disclosed on the row it affects: one of the two judge families wrote the character questions, the canon notes and the article, and that panel's own numbers print beside the other's. For the other three sets the answer to "who decided what counts as right" is a file, not a person: the judge cases were adversarially verified before freezing by a separate fan that caught six of six deliberately planted canaries, and all four sets ship whole in this kit so their neutrality can be audited rather than taken on trust.

A hash proves identity, not time. A sha says what a file contains, never when it started containing it — so the kit publishes two per leg: the one in force at that leg's first scored call, and the one as published. Hold a pre-registration, hash it, compare. The nine files are five leg registrations for this bench, the two addenda, the panel extension and the gauntlet. Seven of them are byte-identical to what they were when the numbers arrived. The other two were appended to after their leg's first scored call, three times between them, and both say so above the appended line. No case, arm, tier, item, metric, floor, ceiling or counting rule moved in any of the three: they are harness defects, their fixes, a completion run and a scheduling gate. The three entries, with what each one changed, are in data/provenance.json.

Licence: CC BY 4.0 — the whole page, not only the kit. The prose, the tables, the folds and the data are yours to quote, re-plot, translate and argue with, including commercially. What we ask back is the one thing the licence already requires: name the source and link to it — strata→signal research, research.strata2signal.com — so a reader of your version can reach ours and check it against the files. Something like — strata→signal research, “The chair trials”, research.strata2signal.com/chair-trials/, CC BY 4.0. And if you quote a figure, carry its outcome word with it — RANKED, EXPLORATORY, UNMEASURABLE, NOT-CARRIED and NOT-RUN are five different findings, and a number without its word is a claim this page did not make. An em dash here is a figure we do not hold; it is never a zero.

elsewhere in the workshop

a strata→signal property · hello@strata2signal.com · say hello