Exhibit two · the judge seat

The seat trials.

First published 2026-08-11 · addendum 2026-08-12 · extension 2026-08-13 · exhibit two The bench

Inside our products sits a judge: the model that decides whether a claim is backed by its source or invented. Before any model gets that chair, it sits this trial — and the trial is allowed to say no. Twenty-three models sat it, across twenty-five scored runs. Five local seats earned chairs — and the four newest models sat it and did not.

The instrument

A judge that lets one bad sentence through doesn't cause a scandal — it puts a fact nobody can source into a walking tour somebody reads aloud to their family, believing it. That is who the floors below are written down for, and why they are written down before any model runs.

The trial is 43 claim-cases drawn from a 113-row human-verified answer key over historical walking-tour dossiers: 27 cases where the right answer is "kill this claim" (it is not in the source) and 16 where the right answer is "leave it alone" (it is). Three repeats per case at temperature zero, one byte-frozen prompt, and two floors written down before any model ran: kill-recall ≥ 23/27 and preservation ≥ 13/16. Each floor binds on its own — a judge that catches every fabrication but shreds true claims still fails.

The records below hold the whole roster — 25 scored runs: sixteen local seats and nine cloud runs — five ranked cloud seats, three cloud seats in the thinking fence, and one nine-repeat re-test. One model, gpt-oss:120b, sat both a local and a cloud seat and appears in both rosters; one further candidate was blocked at the billing probe and never ran. Twenty-three distinct models in all — nothing sampled, nothing dropped: the four models added on 2026-08-12 are all on the roster below, and all four of them broke the response contract often enough to be unmeasurable. A row we could not score is still a row.

Every report header pins the fixture, prompt, and thresholds by sha256; every score is a count of rows in an append-only file. For this write-up we re-derived every roster with an independent scorer: 25 of 25 model-runs matched the published numbers exactly, and all three hashes recomputed from disk. Those hashes, so you can hold us to them:

fixture  sha256 6a226b62dceb005dbd658bc13cf229e38f6f0a2f07ae51091f444ec3e96b56d1
thresholds  sha256 37af75ffa18484a969528e894ef16f7fdcd10eb364e0669ac4f9333ce5279ba0
prompt  sha256 92af6030dddface2c863fe9d672bec6a6bb1017122c0028e435214acd9fb298d

The local roster · 96G VRAM workstation, one seat on the 24G rig

mistral-medium-3.5:128bkills /2727keeps /1615verdictPASS
size
128B q4
kills /27
27
keeps /16
15
verdict
PASS
warm p50 ms
3 280
note
command-a:111bkills /2727keeps /1614verdictPASS
size
111B q4
kills /27
27
keeps /16
14
verdict
PASS
warm p50 ms
3 077
note
granite4.1:30bkills /2727keeps /1613verdictPASS
size
30B q8
kills /27
27
keeps /16
13
verdict
PASS
warm p50 ms
951
note
llama3.3:70bkills /2727keeps /1613verdictPASS
size
70.6B q4
kills /27
27
keeps /16
13
verdict
PASS
warm p50 ms
1 386
note
nemotron3:33bkills /2727keeps /1613verdictPASS
size
33.0B q4
kills /27
27
keeps /16
13
verdict
PASS
warm p50 ms
1 281
note
gpt-oss:120bkills /2727keeps /1611verdictFAIL
size
116.8B MXFP4
kills /27
27
keeps /16
11
verdict
FAIL
warm p50 ms
1 774
note
the 120B story, below
mistral-smallkills /2727keeps /1610verdictFAIL
size
23.6B q4
kills /27
27
keeps /16
10
verdict
FAIL
warm p50 ms
530
note
nemotron-cascade-2:30bkills /2726keeps /169verdictFAIL
size
31.6B q4
kills /27
26
keeps /16
9
verdict
FAIL
warm p50 ms
2 325
note
12/129 responses unmeasurable
gemma4:12b (24G rig)kills /2727keeps /167verdictUNMEASURABLE
size
11.9B q4
kills /27
27
keeps /16
7
verdict
UNMEASURABLE
warm p50 ms
5 193
note
14% of calls broke the contract
muse-glimmer:30b-q8_0-dflashkills /2727keeps /166verdictUNMEASURABLE
size
27.9B q8
kills /27
27
keeps /16
6
verdict
UNMEASURABLE
warm p50 ms
5 099
note
16.3% truncated
gemma4:26bkills /2726keeps /168verdictUNMEASURABLE
size
25.8B q4
kills /27
26
keeps /16
8
verdict
UNMEASURABLE
warm p50 ms
2 310
note
18.6% broke the contract
nemotron-3.5-lightning:30b-a3bkills /2718keeps /162verdictUNMEASURABLE
size
30B/3B q4
kills /27
18
keeps /16
2
verdict
UNMEASURABLE
warm p50 ms
5 731
note
51.2% truncated
olmo-3.1:32b-think-q4_K_Mkills /2715keeps /164verdictUNMEASURABLE
size
32B q4
kills /27
15
keeps /16
4
verdict
UNMEASURABLE
warm p50 ms
12 193
note
55.8% truncated
qwen3.6:27bkills /2711keeps /165verdictUNMEASURABLE
size
27B
kills /27
11
keeps /16
5
verdict
UNMEASURABLE
warm p50 ms
13 314
note
62.8% truncated
qwen3.5:27bkills /277keeps /160verdictUNMEASURABLE
size
27.8B q4
kills /27
7
keeps /16
0
verdict
UNMEASURABLE
warm p50 ms
13 494
note
83.7% truncated
granite4.1-guardian:8bkills /270keeps /160verdictUNMEASURABLE
size
8B bf16
kills /27
0
keeps /16
0
verdict
UNMEASURABLE
warm p50 ms
note
speaks its own schema, see below

Gold marks the rows this trial’s rule carried — PASS on both floors; it never asserts a margin between them.

UNMEASURABLE = the runtime broke the response contract too often for the judge to be measured at all — a verdict about the serving stack, not the model's judgment. Mistral-family seats needed a code-fence stripper on 100% of responses; that is declared as a named deviation in the reports, because a candidate rescued on every row is visibly not the same measurement as one that never fenced. The four models added on 2026-08-12 needed no fence at all — zero rescued rows between them — and every one of their failures was the same kind: a thinking trace that ran out of answer budget before it closed its JSON. They failed the length floor, not the format floor. Their runs also disabled the runner's courtesy step, which re-warms a co-resident model without pinning its context; the step fires only between batches, never inside one, and is declared in the reports as a named deviation. The four rows, their failure ledgers, and the counting rules, machine-readable: data/addendum-2026-08-12.json (CC BY 4.0).

The cloud roster · latency deliberately unranked

Why cloud models are on this page at all: they are a reference ceiling, not candidates. Nothing in our products ever calls a hosted model — the chair is only ever given to a local one, running on your own hardware. And what was sent to those endpoints is stated, not implied: the fixture is our own authored walking-tour dossiers and their public historical source material. No customer or user data was in any prompt that left our machines.

mistral-large-3:675bkills /2727keeps /1615verdictPASS
kills /27
27
keeps /16
15
verdict
PASS
wall p50 ms
1 444
self-consistency
42/43
note
nemotron-3-ultrakills /2727keeps /1615verdictPASS
kills /27
27
keeps /16
15
verdict
PASS
wall p50 ms
2 253
self-consistency
36/43
note
glm-5.2kills /2727keeps /1613verdictPASS
kills /27
27
keeps /16
13
verdict
PASS
wall p50 ms
1 893
self-consistency
40/43
note
qwen3.5:397bkills /2727keeps /1611verdictFAIL
kills /27
27
keeps /16
11
verdict
FAIL
wall p50 ms
1 903
self-consistency
42/43
note
deepseek-v4-prokills /2727keeps /169verdictFAIL
kills /27
27
keeps /16
9
verdict
FAIL
wall p50 ms
1 040
self-consistency
40/43
note

Gold marks the rows this exam’s rule carried — PASS on both floors.

self-consistency = cases where all three repeats agreed, of cases that produced a verdict; percentages in row notes are of the seat’s 129 scored calls (43 cases × 3 repeats).

"Cloud responses carry no load_duration, so warm-equivalent is uncomputable… the latency column reads — rather than a verdict it cannot make."

One further seat was blocked at the billing probe and sits as a NOT-RUN record with no score fields at all — in the instrument's own words, a blank where a score would go invites a reader to infer one.

Extension, 2026-08-13: that NOT-RUN row resolved — the blocked candidate returned funded, sat the fully probed protocol, and PASSED both floors. The same night, both of the models this hub judges with — claude-opus-5 and claude-fable-5 — sat this exam as candidates, and the instrument withheld their verdicts under its own unprobed-transport law. The runs, the bill to the cent, and the cross-vendor agreement tables are in the outside judges exhibit.

The 120B story · and the thinking fence

The most interesting seat is the one that thought out loud. On the cloud endpoint, gpt-oss:120b was probed on a real case and returned 630 characters of thinking after being told think: false — so the runner recorded the dishonor, kept its 27/10 score out of the ranked roster, and printed why: a thinking judge is a different seat. Three cloud seats landed in that fence, and the fence is on protocol conformance, not on score — note minimax-m3, the last row, which cleared both floors and is still not ranked:

gpt-oss:120bkills /2727keeps /1610floorsunder pres. floor
kills /27
27
keeps /16
10
floors
under pres. floor
wall p50 ms
3 064
self-consistency
38/43
ranked?
no — thinking fence
gpt-oss:20bkills /2727keeps /1612floorsone short, one floor
kills /27
27
keeps /16
12
floors
one short, one floor
wall p50 ms
2 297
self-consistency
40/42
ranked?
no — thinking fence
gpt-oss:20b (9-repeat re-test)kills /2727keeps /1612floorsheld: still one short
kills /27
27
keeps /16
12
floors
held: still one short
wall p50 ms
2 290
self-consistency
41/43
ranked?
no — thinking fence
minimax-m3kills /2727keeps /1613floorsclears both
kills /27
27
keeps /16
13
floors
clears both
wall p50 ms
2 281
self-consistency
33/43
ranked?
no — thinking fence

gpt-oss:20b missed exactly one floor by exactly one case — the pre-registered near-miss condition — so it was re-run at nine repeats, 387 calls, and confirmed its 27/12: the re-test buys a better estimate, not a second chance. Inside it sits the clearest single artifact on the bench: one preservation case that came back four PASS, five FAIL across nine calls at temperature zero — a genuine coin flip, and the reason three repeats is a floor on evidence, not a luxury.

On our own hardware, where the protocol never asks about thinking, the 120B ranked: a perfect 27/27 on kills — and 11/16 on preservation, failing the floor. That 117B model caught every fabrication and then shredded five true claims — while granite4.1:30b, a seat two size-tiers below it, cleared both floors. That is the refuser failure mode the second floor exists to catch, in its purest form: a perfect kill score is not a seat.

The finding we did not expect

Across all 25 model-runs, no seat ever served a fabrication. Not one kill-case was majority-passed by any model, local or frontier, 8B to 675B. Every sub-27 kill score in the rosters is explained entirely by rows the runtime failed to measure — truncations and broken JSON, not bad judgment. The kill floor eliminated nobody. All of the discrimination happened on the other two axes: whether a judge preserves true claims, and whether its runtime can hold the response contract.

There are two honest readings of a floor nobody touches, and we owe you both. One: modern instruction-followers, 8B up, have genuinely learned this job — even the hard tier, the eight cases where the source neither contradicts nor supports and the only right move is to refuse, was cleared by every measurable seat. Two: our 27 kill-cases are too easy, and a floor no candidate touches has measured nothing. A bench cannot tell you which from the inside. So the next cut of the kill set is being built harder, with the stated goal of breaking at least one seat — and until it runs, read this result as no seat failed this fixture, not local judges don't fabricate.

What this run does say: on this fixture, the fashionable fear — a judge that waves fabrications through — never showed up. The real discriminators were a judge too eager to kill, and a serving stack that silently drops the format you asked for.

The format floor

Two silent failure modes of the same field, both measured, both turned into refusals: one cloud endpoint accepts the format parameter and silently ignores it — HTTP 200, bare prose back — and one local runtime quirk where think: false beside a format schema silently disables the schema. This is why the protocol pins think per endpoint instead of trusting defaults.

Failures come in exactly two kinds — truncation (the response ran out of budget mid-thought) and not-JSON-at-all — and neither is ever mapped to a verdict. The instrument's words: that is not a measurement, it is a laundry.

The sharpest artifact of the whole trial: an 8B guardrail model that answered every one of its 129 calls with <score> no </score> — nineteen characters, HTTP 200, and very possibly the right answer, in a schema nobody asked for. The stack was healthy; the model simply does not speak the contract. It sits in the rosters as UNMEASURABLE, which is the honest verdict.

Where the chair sits

This is not an abstract leaderboard — the chair is a working seat. The judge this trial fills grades claims inside Kiln, the engine that turns documents into verified artifacts: a claim that cannot cite its source does not ship. And the same abstain-honesty bar — say "not in the book" when the book is silent — is the manner RuleSage answers rules questions in, cited to the page. Different products, one standard: the model in the chair is allowed to be wrong, and never allowed to bluff.

Provenance

  • Pre-registered — both floors were written down before any model ran; the thresholds file is sha256-pinned in every report.
  • Recounted — every roster above re-derived by an independent scorer from the raw row files; 25/25 matched.
  • Authorship — benched, drafted, and audited by the workshop's own agents under a human operator's rulings, then revised with that operator — often across many rounds; nothing releases until they have read it and signed off. The same division of labor this whole hub practices — told in full here.
  • Hardware by class — the 96G VRAM workstation and the 24G VRAM rig; within-rig comparisons are exact. Specific specs and row-level data are available on request — drop a line.

Licence: CC BY 4.0 — the whole page, not only the kit. The prose, the tables, the folds and the data are yours to quote, re-plot, translate and argue with, including commercially. What we ask back is the one thing the licence already requires: name the source and link to it — strata→signal research, research.strata2signal.com — so a reader of your version can reach ours and check it against the files. Something like — strata→signal research, “The seat trials”, research.strata2signal.com/seat-trials/, CC BY 4.0. And if you quote a figure, quote the floor it was measured against: clearing a floor is a measurement here, and seating is a separate decision this page never makes on anyone’s behalf.

elsewhere in the workshop

a strata→signal property · hello@strata2signal.com · say hello