Exhibit two · the judge seat
The seat trials.
First published 2026-08-11 · addendum 2026-08-12 · extension 2026-08-13 · exhibit two The bench
Inside our products sits a judge: the model that decides whether a claim is backed by its source or invented. Before any model gets that chair, it sits this trial — and the trial is allowed to say no. Twenty-three models sat it, across twenty-five scored runs. Five local seats earned chairs — and the four newest models sat it and did not.
The instrument
A judge that lets one bad sentence through doesn't cause a scandal — it puts a fact nobody can source into a walking tour somebody reads aloud to their family, believing it. That is who the floors below are written down for, and why they are written down before any model runs.
The trial is 43 claim-cases drawn from a 113-row human-verified answer key over historical walking-tour dossiers: 27 cases where the right answer is "kill this claim" (it is not in the source) and 16 where the right answer is "leave it alone" (it is). Three repeats per case at temperature zero, one byte-frozen prompt, and two floors written down before any model ran: kill-recall ≥ 23/27 and preservation ≥ 13/16. Each floor binds on its own — a judge that catches every fabrication but shreds true claims still fails.
The records below hold the whole roster — 25 scored runs: sixteen local seats and nine cloud runs — five ranked cloud seats, three cloud seats in the thinking fence, and one nine-repeat re-test. One model, gpt-oss:120b, sat both a local and a cloud seat and appears in both rosters; one further candidate was blocked at the billing probe and never ran. Twenty-three distinct models in all — nothing sampled, nothing dropped: the four models added on 2026-08-12 are all on the roster below, and all four of them broke the response contract often enough to be unmeasurable. A row we could not score is still a row.
Every report header pins the fixture, prompt, and thresholds by sha256; every score is a count of rows in an append-only file. For this write-up we re-derived every roster with an independent scorer: 25 of 25 model-runs matched the published numbers exactly, and all three hashes recomputed from disk. Those hashes, so you can hold us to them:
fixture sha256 6a226b62dceb005dbd658bc13cf229e38f6f0a2f07ae51091f444ec3e96b56d1
thresholds sha256 37af75ffa18484a969528e894ef16f7fdcd10eb364e0669ac4f9333ce5279ba0
prompt sha256 92af6030dddface2c863fe9d672bec6a6bb1017122c0028e435214acd9fb298d
The local roster · 96G VRAM workstation, one seat on the 24G rig
mistral-medium-3.5:128b kills /2727keeps /1615verdictPASS
- size
- 128B q4
- kills /27
- 27
- keeps /16
- 15
- verdict
- PASS
- warm p50 ms
- 3 280
- note
- —
command-a:111b kills /2727keeps /1614verdictPASS
- size
- 111B q4
- kills /27
- 27
- keeps /16
- 14
- verdict
- PASS
- warm p50 ms
- 3 077
- note
- —
granite4.1:30b kills /2727keeps /1613verdictPASS
- size
- 30B q8
- kills /27
- 27
- keeps /16
- 13
- verdict
- PASS
- warm p50 ms
- 951
- note
- —
llama3.3:70b kills /2727keeps /1613verdictPASS
- size
- 70.6B q4
- kills /27
- 27
- keeps /16
- 13
- verdict
- PASS
- warm p50 ms
- 1 386
- note
- —
nemotron3:33b kills /2727keeps /1613verdictPASS
- size
- 33.0B q4
- kills /27
- 27
- keeps /16
- 13
- verdict
- PASS
- warm p50 ms
- 1 281
- note
- —
gpt-oss:120b kills /2727keeps /1611verdictFAIL
- size
- 116.8B MXFP4
- kills /27
- 27
- keeps /16
- 11
- verdict
- FAIL
- warm p50 ms
- 1 774
- note
- the 120B story, below
mistral-smallkills /2727keeps /1610verdictFAIL
- size
- 23.6B q4
- kills /27
- 27
- keeps /16
- 10
- verdict
- FAIL
- warm p50 ms
- 530
- note
- —
nemotron-cascade-2:30b kills /2726keeps /169verdictFAIL
- size
- 31.6B q4
- kills /27
- 26
- keeps /16
- 9
- verdict
- FAIL
- warm p50 ms
- 2 325
- note
- 12/129 responses unmeasurable
gemma4:12b (24G rig) kills /2727keeps /167verdictUNMEASURABLE
- size
- 11.9B q4
- kills /27
- 27
- keeps /16
- 7
- verdict
- UNMEASURABLE
- warm p50 ms
- 5 193
- note
- 14% of calls broke the contract
muse-glimmer:30b-q8_0-dflash kills /2727keeps /166verdictUNMEASURABLE
- size
- 27.9B q8
- kills /27
- 27
- keeps /16
- 6
- verdict
- UNMEASURABLE
- warm p50 ms
- 5 099
- note
- 16.3% truncated
gemma4:26b kills /2726keeps /168verdictUNMEASURABLE
- size
- 25.8B q4
- kills /27
- 26
- keeps /16
- 8
- verdict
- UNMEASURABLE
- warm p50 ms
- 2 310
- note
- 18.6% broke the contract
nemotron-3.5-lightning:30b-a3b kills /2718keeps /162verdictUNMEASURABLE
- size
- 30B/3B q4
- kills /27
- 18
- keeps /16
- 2
- verdict
- UNMEASURABLE
- warm p50 ms
- 5 731
- note
- 51.2% truncated
olmo-3.1:32b-think-q4_K_M kills /2715keeps /164verdictUNMEASURABLE
- size
- 32B q4
- kills /27
- 15
- keeps /16
- 4
- verdict
- UNMEASURABLE
- warm p50 ms
- 12 193
- note
- 55.8% truncated
qwen3.6:27b kills /2711keeps /165verdictUNMEASURABLE
- size
- 27B
- kills /27
- 11
- keeps /16
- 5
- verdict
- UNMEASURABLE
- warm p50 ms
- 13 314
- note
- 62.8% truncated
qwen3.5:27b kills /277keeps /160verdictUNMEASURABLE
- size
- 27.8B q4
- kills /27
- 7
- keeps /16
- 0
- verdict
- UNMEASURABLE
- warm p50 ms
- 13 494
- note
- 83.7% truncated
granite4.1-guardian:8b kills /270keeps /160verdictUNMEASURABLE
- size
- 8B bf16
- kills /27
- 0
- keeps /16
- 0
- verdict
- UNMEASURABLE
- warm p50 ms
- —
- note
- speaks its own schema, see below
Gold marks the rows this trial’s rule carried — PASS on both floors; it never asserts a margin between them.
UNMEASURABLE = the runtime broke the response contract too often for the judge to be measured at all — a verdict about the serving stack, not the model's judgment. Mistral-family seats needed a code-fence stripper on 100% of responses; that is declared as a named deviation in the reports, because a candidate rescued on every row is visibly not the same measurement as one that never fenced. The four models added on 2026-08-12 needed no fence at all — zero rescued rows between them — and every one of their failures was the same kind: a thinking trace that ran out of answer budget before it closed its JSON. They failed the length floor, not the format floor. Their runs also disabled the runner's courtesy step, which re-warms a co-resident model without pinning its context; the step fires only between batches, never inside one, and is declared in the reports as a named deviation. The four rows, their failure ledgers, and the counting rules, machine-readable: data/addendum-2026-08-12.json (CC BY 4.0).
The cloud roster · latency deliberately unranked
Why cloud models are on this page at all: they are a reference ceiling, not candidates. Nothing in our products ever calls a hosted model — the chair is only ever given to a local one, running on your own hardware. And what was sent to those endpoints is stated, not implied: the fixture is our own authored walking-tour dossiers and their public historical source material. No customer or user data was in any prompt that left our machines.
mistral-large-3:675b kills /2727keeps /1615verdictPASS
- kills /27
- 27
- keeps /16
- 15
- verdict
- PASS
- wall p50 ms
- 1 444
- self-consistency
- 42/43
- note
- —
nemotron-3-ultrakills /2727keeps /1615verdictPASS
- kills /27
- 27
- keeps /16
- 15
- verdict
- PASS
- wall p50 ms
- 2 253
- self-consistency
- 36/43
- note
- —
glm-5.2kills /2727keeps /1613verdictPASS
- kills /27
- 27
- keeps /16
- 13
- verdict
- PASS
- wall p50 ms
- 1 893
- self-consistency
- 40/43
- note
- —
qwen3.5:397b kills /2727keeps /1611verdictFAIL
- kills /27
- 27
- keeps /16
- 11
- verdict
- FAIL
- wall p50 ms
- 1 903
- self-consistency
- 42/43
- note
- —
deepseek-v4-prokills /2727keeps /169verdictFAIL
- kills /27
- 27
- keeps /16
- 9
- verdict
- FAIL
- wall p50 ms
- 1 040
- self-consistency
- 40/43
- note
- —
Gold marks the rows this exam’s rule carried — PASS on both floors.
self-consistency = cases where all three repeats agreed, of cases that produced a verdict; percentages in row notes are of the seat’s 129 scored calls (43 cases × 3 repeats).
"Cloud responses carry no load_duration, so warm-equivalent is uncomputable… the latency column reads — rather than a verdict it cannot make."
One further seat was blocked at the billing probe and sits as a NOT-RUN record with no score fields at all — in the instrument's own words, a blank where a score would go invites a reader to infer one.
Extension, 2026-08-13: that NOT-RUN row resolved — the blocked candidate returned funded, sat the fully probed protocol, and PASSED both floors. The same night, both of the models this hub judges with — claude-opus-5 and claude-fable-5 — sat this exam as candidates, and the instrument withheld their verdicts under its own unprobed-transport law. The runs, the bill to the cent, and the cross-vendor agreement tables are in the outside judges exhibit.
The 120B story · and the thinking fence
The most interesting seat is the one that thought out loud. On the cloud endpoint, gpt-oss:120b was probed on a real case and returned 630 characters of thinking after being told think: false — so the runner recorded the dishonor, kept its 27/10 score out of the ranked roster, and printed why: a thinking judge is a different seat. Three cloud seats landed in that fence, and the fence is on protocol conformance, not on score — note minimax-m3, the last row, which cleared both floors and is still not ranked:
gpt-oss:120b kills /2727keeps /1610floorsunder pres. floor
- kills /27
- 27
- keeps /16
- 10
- floors
- under pres. floor
- wall p50 ms
- 3 064
- self-consistency
- 38/43
- ranked?
- no — thinking fence
gpt-oss:20b kills /2727keeps /1612floorsone short, one floor
- kills /27
- 27
- keeps /16
- 12
- floors
- one short, one floor
- wall p50 ms
- 2 297
- self-consistency
- 40/42
- ranked?
- no — thinking fence
gpt-oss:20b (9-repeat re-test) kills /2727keeps /1612floorsheld: still one short
- kills /27
- 27
- keeps /16
- 12
- floors
- held: still one short
- wall p50 ms
- 2 290
- self-consistency
- 41/43
- ranked?
- no — thinking fence
minimax-m3kills /2727keeps /1613floorsclears both
- kills /27
- 27
- keeps /16
- 13
- floors
- clears both
- wall p50 ms
- 2 281
- self-consistency
- 33/43
- ranked?
- no — thinking fence
gpt-oss:20b missed exactly one floor by exactly one case — the pre-registered near-miss condition — so it was re-run at nine repeats, 387 calls, and confirmed its 27/12: the re-test buys a better estimate, not a second chance. Inside it sits the clearest single artifact on the bench: one preservation case that came back four PASS, five FAIL across nine calls at temperature zero — a genuine coin flip, and the reason three repeats is a floor on evidence, not a luxury.
On our own hardware, where the protocol never asks about thinking, the 120B ranked: a perfect 27/27 on kills — and 11/16 on preservation, failing the floor. That 117B model caught every fabrication and then shredded five true claims — while granite4.1:30b, a seat two size-tiers below it, cleared both floors. That is the refuser failure mode the second floor exists to catch, in its purest form: a perfect kill score is not a seat.
The finding we did not expect
Across all 25 model-runs, no seat ever served a fabrication. Not one kill-case was majority-passed by any model, local or frontier, 8B to 675B. Every sub-27 kill score in the rosters is explained entirely by rows the runtime failed to measure — truncations and broken JSON, not bad judgment. The kill floor eliminated nobody. All of the discrimination happened on the other two axes: whether a judge preserves true claims, and whether its runtime can hold the response contract.
There are two honest readings of a floor nobody touches, and we owe you both. One: modern instruction-followers, 8B up, have genuinely learned this job — even the hard tier, the eight cases where the source neither contradicts nor supports and the only right move is to refuse, was cleared by every measurable seat. Two: our 27 kill-cases are too easy, and a floor no candidate touches has measured nothing. A bench cannot tell you which from the inside. So the next cut of the kill set is being built harder, with the stated goal of breaking at least one seat — and until it runs, read this result as no seat failed this fixture, not local judges don't fabricate.
What this run does say: on this fixture, the fashionable fear — a judge that waves fabrications through — never showed up. The real discriminators were a judge too eager to kill, and a serving stack that silently drops the format you asked for.
The format floor
Two silent failure modes of the same field, both measured, both turned into refusals: one cloud endpoint accepts the format parameter and silently ignores it — HTTP 200, bare prose back — and one local runtime quirk where think: false beside a format schema silently disables the schema. This is why the protocol pins think per endpoint instead of trusting defaults.
Failures come in exactly two kinds — truncation (the response ran out of budget mid-thought) and not-JSON-at-all — and neither is ever mapped to a verdict. The instrument's words: that is not a measurement, it is a laundry.
The sharpest artifact of the whole trial: an 8B guardrail model that answered every one of its 129 calls with <score> no </score> — nineteen characters, HTTP 200, and very possibly the right answer, in a schema nobody asked for. The stack was healthy; the model simply does not speak the contract. It sits in the rosters as UNMEASURABLE, which is the honest verdict.
Where the chair sits
This is not an abstract leaderboard — the chair is a working seat. The judge this trial fills grades claims inside Kiln, the engine that turns documents into verified artifacts: a claim that cannot cite its source does not ship. And the same abstain-honesty bar — say "not in the book" when the book is silent — is the manner RuleSage answers rules questions in, cited to the page. Different products, one standard: the model in the chair is allowed to be wrong, and never allowed to bluff.
Provenance
- Pre-registered — both floors were written down before any model ran; the thresholds file is sha256-pinned in every report.
- Recounted — every roster above re-derived by an independent scorer from the raw row files; 25/25 matched.
- Authorship — benched, drafted, and audited by the workshop's own agents under a human operator's rulings, then revised with that operator — often across many rounds; nothing releases until they have read it and signed off. The same division of labor this whole hub practices — told in full here.
- Hardware by class — the 96G VRAM workstation and the 24G VRAM rig; within-rig comparisons are exact. Specific specs and row-level data are available on request — drop a line.
Licence: CC BY 4.0 — the whole page, not only the kit. The prose, the tables, the folds and the data are yours to quote, re-plot, translate and argue with, including commercially. What we ask back is the one thing the licence already requires: name the source and link to it — strata→signal research, research.strata2signal.com — so a reader of your version can reach ours and check it against the files. Something like — strata→signal research, “The seat trials”, research.strata2signal.com/seat-trials/, CC BY 4.0. And if you quote a figure, quote the floor it was measured against: clearing a floor is a measurement here, and seating is a separate decision this page never makes on anyone’s behalf.