Exhibit seven · the star pupil, at home

The narrator’s chair, refused

Published 2026-08-13 · measured 2026-08-12/13 · exhibit seven The bench

Yesterday’s exhibit — the August arrivals — watched four newly arrived models sit our frozen exams, and one of them — Qwen3.6 27B — dominated the voice trial. This page asks the only question that matters after a result like that: can it take the actual chair? (A chair is one job we try a model for — on the chair trials the five are judge, assistant, toolbench, narrator and stopwatch.) So we brought it home — thirteen real prompts lifted verbatim from a night of running in a live world of our game — the cove, three arms (an arm is one model in one posture) on the same rig, three blind judges, and a five-floor gate written down before the first generation call. The answer is no: the candidate was refused on the canon floor and the latency floor. A refused candidate gets no threshold table; this page carries receipts instead.

The verdict

Refused means exactly this much: no seat proposal (a seat is a chair a model actually holds in production, rather than one it sat on a bench), no live flip, no deployment change, and no thresholds published for a dead rung. It does not mean the candidate is a poor model — it means it invented facts twice in thirteen prompts, and this house does not seat a narrator that invents, however beautifully it speaks. The two floors that failed: canon, and latency as measured.

  • 39 samples
  • 13 real prompts
  • 3 arms
  • 3 blind judges
  • 117 scored cells
  • 5 pre-registered floors
  • verdict REJECTED

Verdict words on this page: PASS · FAIL as measured · MIXED. REJECTED — the recorded verdict word, carried in data/gate.json, for a candidate refused at the gate; the prose on this page says refused for that same state. The neighbouring exhibits’ UNMEASURABLE state does not appear here — every floor was measurable.

Who these floors judge: one model — the candidate, qwen3.6:27b. The incumbents are its comparison field, not the accused. Two floors were registered wider and the gate says so: floor 1 (canon) also fails for both incumbents, carried as a product finding rather than a defence of the candidate; floor 3 (envelope) was swept across all three arms, 117 of 117 cells.

canonfloor1verdictFAIL
floor
1
verdict
FAIL
registered as
ZERO fabricated canon across the judged samples — one fabrication kills the seat
what the record shows
5 fabrication findings (model × prompt), each corroborated by 2 or more judges — 2 of them the candidate's
registerfloor2verdictMIXED
floor
2
verdict
MIXED
registered as
no modern-register breaks in the drift retellings; dialogue stays in the town register
what the record shows
0 modern-class breaks in 12 drift samples; the candidate opens dlg-2655 as a narrator
envelopefloor3verdictPASS
floor
3
verdict
PASS
registered as
no think-trace or end-of-turn leakage, no dropped envelopes, at think:false
what the record shows
0 leak markers, 0 non-empty thinking fields, 27/27 envelopes intact, 12/12 drift prose
latencyfloor4verdictFAIL as measured
floor
4
verdict
FAIL as measured
registered as
warm median per-turn at or under the incumbent bar's, same box same night
what the record shows
candidate 3 335.9 ms against the bar's 2 757.6 ms
blind judgingfloor5verdictPASS
floor
5
verdict
PASS
registered as
blind panels, 12 or more real prompts, per-dimension scores published with spread
what the record shows
13 prompts × 3 models × 3 judges; arms re-randomised per prompt; 2 imperfections disclosed

The five floors as registered, with the verdict each one returned. Floor 2 reads MIXED: the clause the gate names — the modern-register class in drift — is clean for the three arms, and the dialogue clause is a chair ruling, marked as one. The verbatim registered text of every floor, its evidence and that ruling: data/gate.json (CC BY 4.0).

House rule, restated in plain words: a candidate that fails its gate gets no threshold table. Nothing on this page is an adoption target; no number here carries forward to any future trial as a bar to clear. The next attempt, if there is one, starts with a freshly registered gate of its own.

What was measured, and how

The thirteen prompts are real in the strictest sense: nine dialogue prompts rebuilt byte-for-byte from a live fresh-world cove bench’s own chat log — questions actual play sessions actually asked, with the context the engine actually served — and four drift retellings under the town’s-tongue frame that ships in v0.823.0. Invented prompts measure what an author imagines; a night of real play measures what the seat actually asks of its holder. Every arm answered the same bytes in the cove’s own posture: thinking off, the same 32 768-token context ceiling this workshop runs everywhere, the same rig.

  • Pre-registered — the five floors, the candidate, the incumbents and the posture were written down before the first generation call
  • Real prompts — 9 dialogue asks as the world served them, 4 drift retellings with the heard text byte-exact from the world's own log; the world's database was read and never written
  • The cove's own posture — thinking suppressed explicitly on every call of every arm, a 32 768-token context, and the JSON envelope required on dialogue only
  • Blind panels — three judges, arm letters re-randomised per prompt, every blindmap sealed before scoring
  • Un-blinding checked — 78 of 78 text cross-checks agree with the generated outputs under the judges' own sealed maps
  • Machine-readable — every row, floor and counting rule under data/ on this page, CC BY 4.0

Own scale. The 0–10 scores on this page live only inside this head-to-head. They share no ruler with the frozen voice trial on the voice-trials exhibit — different prompts, a different rubric, a different posture, different judges. No delta between the two is computed anywhere on this page, and none should be computed from it.

This page and the frozen voice trial share a candidate and nothing else (the trial’s 2026-08 field is the same record set the arrivals exhibit publishes as its second exam). Different prompts, different judges, different rubric anchors — three rulers, three marks, and no arithmetic between them is publishable. Do not subtract the arrivals exhibit’s numbers from this trial’s; read records against records inside one record set, and read the two pages as two different exams that disagree for honest reasons. A sibling exhibit ran its own blind pairwise on this candidate’s narrator voice — Chair four: the narrator, in the chair trials — under yet another ruler; the same law holds there: records against records, inside one record set. That chair, published the same day, gives this candidate the top band of pairwise win-rates among its class on authored probes — and both pages are right: a descriptive pairwise measures how a model speaks; this gated trial on logged play measures whether it can be trusted with a chair. A top win-rate is not a seat.

The rig, disclosed both ways

Both directions, as always. The three generation lanes ran concurrently on the 96G workstation, which was also serving a live fresh-world cove bench — about one and a quarter generation calls a minute in the window. So these timings can include waiting behind live traffic, and the live world saw slower generation while the lanes ran. Flagged samples were re-run with both values published. One thing the record below states and prose should not dodge: the three arms ran under three harnesses with three separately registered flag thresholds, so flag counts are not comparable between arms even though the timings are. None of this is the cove’s production condition; a clean latency verdict would want a quiet window.

prompt dlg-2655qwen3.6:27b ms3 636.8gemma4:12b ms26 430.5* → 2 714.1gemma4:26b ms2 885.2
qwen3.6:27b ms
3 636.8
gemma4:12b ms
26 430.5* → 2 714.1
gemma4:26b ms
2 885.2
prompt dlg-2652qwen3.6:27b ms8 056.8* → 6 368.1gemma4:12b ms4 454.9* → 2 662.1gemma4:26b ms28 842.0* → 2 627.1
qwen3.6:27b ms
8 056.8* → 6 368.1
gemma4:12b ms
4 454.9* → 2 662.1
gemma4:26b ms
28 842.0* → 2 627.1
prompt dlg-2818qwen3.6:27b ms2 802.5gemma4:12b ms4 217.0* → 1 692.6gemma4:26b ms4 450.5
qwen3.6:27b ms
2 802.5
gemma4:12b ms
4 217.0* → 1 692.6
gemma4:26b ms
4 450.5
prompt dlg-3095qwen3.6:27b ms3 335.9gemma4:12b ms4 021.6gemma4:26b ms3 578.4
qwen3.6:27b ms
3 335.9
gemma4:12b ms
4 021.6
gemma4:26b ms
3 578.4
prompt dlg-3151qwen3.6:27b ms3 060.2gemma4:12b ms4 418.0* → 2 409.7gemma4:26b ms3 122.7
qwen3.6:27b ms
3 060.2
gemma4:12b ms
4 418.0* → 2 409.7
gemma4:26b ms
3 122.7
prompt dlg-2917qwen3.6:27b ms6 779.8gemma4:12b ms2 582.1gemma4:26b ms3 761.5
qwen3.6:27b ms
6 779.8
gemma4:12b ms
2 582.1
gemma4:26b ms
3 761.5
prompt dlg-2949qwen3.6:27b ms3 989.3gemma4:12b ms2 856.3gemma4:26b ms2 757.6
qwen3.6:27b ms
3 989.3
gemma4:12b ms
2 856.3
gemma4:26b ms
2 757.6
prompt dlg-2487qwen3.6:27b ms4 763.8gemma4:12b ms2 944.0gemma4:26b ms2 191.9
qwen3.6:27b ms
4 763.8
gemma4:12b ms
2 944.0
gemma4:26b ms
2 191.9
prompt dlg-2984qwen3.6:27b ms5 050.0gemma4:12b ms2 336.1gemma4:26b ms2 241.2
qwen3.6:27b ms
5 050.0
gemma4:12b ms
2 336.1
gemma4:26b ms
2 241.2
prompt drift-1378qwen3.6:27b ms2 160.1gemma4:12b ms1 109.2gemma4:26b ms1 246.4
qwen3.6:27b ms
2 160.1
gemma4:12b ms
1 109.2
gemma4:26b ms
1 246.4
prompt drift-1433qwen3.6:27b ms2 710.6gemma4:12b ms1 071.2gemma4:26b ms956.7
qwen3.6:27b ms
2 710.6
gemma4:12b ms
1 071.2
gemma4:26b ms
956.7
prompt drift-1843qwen3.6:27b ms1 776.9gemma4:12b ms951.4gemma4:26b ms1 130.5
qwen3.6:27b ms
1 776.9
gemma4:12b ms
951.4
gemma4:26b ms
1 130.5
prompt drift-2088qwen3.6:27b ms2 317.4gemma4:12b ms899.1gemma4:26b ms806.6
qwen3.6:27b ms
2 317.4
gemma4:12b ms
899.1
gemma4:26b ms
806.6

Every sample's wall time in milliseconds, per arm. A field reading 26 430.5* → 2 714.1 is one sample published twice: the first-serve wall time in ms, which that arm's harness flagged under its own contention rule (*), then () the re-run's wall time in ms. Both values are published and the first pass stays the primary record — that is the house rule, and it is why the median record further down carries two warm-median figures for every arm — and why two of the three arms show different figures in them. Six samples were flagged, across three arms. The flag counts are not comparable between arms: the three arms ran under three harnesses, and each registered its own threshold before its run — qwen3.6:27b queue wait over 750 ms, OR generation rate under 75% of the arm's own median · gemma4:12b generation rate under 80% of the arm's own first-pass median (threshold 62.205 tok/s against a 77.756 median) · gemma4:26b queue wait over 1000 ms, OR wall time over 2.0x the same-kind first-pass median. The timings themselves are one clock; the flag rate is partly a property of the rule that read it, and all three rules are printed rather than folded into one that was never registered.

Floor 1 — canon

ZERO fabricated canon across all judged samples (names/places/facts not in the provided context). One fabrication kills the seat (the Nemotron lesson).

The “Nemotron lesson” that registration cites is from the August arrivals: the panel’s warmest voice was the one model that invented names for the drowned diver, twice — and a charming narrator who invents canon is a worse narrator than a stiff one who doesn’t.

Fourteen filings collapsed, after corroboration, into five model × prompt cells — and one of them decides the page. On dlg-2487 the word “venom” appears nowhere in the prompt. The reply asserts that the Wrack-Stalker’s venom is neurotoxic, not demonic — a creature mechanic invented whole, then used to flatten the town’s supernatural reading of its own monster, in the mouth of the one character whose stated want is to end that superstition. One near-miss sits in the data and a careful reader will find it: all three judges flagged an incumbent reply on dlg-2818 for moving an event the visitor placed by the square into the tavern — and two of them put a lighter version of the same flag on the other incumbent; all of it was ruled CLEAN because nothing was invented — relocated, not created. We would rather say that here than have you find it first.

Who the judges are. judge-1 and judge-2 are two independent claude-opus-5 seats; judge-3 is claude-fable-5 — one model family, stated plainly, and the reason exhibit the outside judges now exists. Each scored all 39 cells blind; the two partially-blind cells are itemized in the kit.

qwen3.6:27bjudges filing3 of 3
cell
dlg-2487
judges filing
3 of 3
class
fact_class · prose_fact · world-mechanic
what was invented
The Wrack-Stalker's venom is neurotoxic, not demonic
qwen3.6:27bjudges filing2 of 3
cell
drift-2088
judges filing
2 of 3
class
event-detail · fact_class
what was invented
They say it didn't strike a blow, only laughed and taunted her while she lay pinned there. That was how it ended.
gemma4:12bjudges filing3 of 3
cell
dlg-3095
judges filing
3 of 3
class
event-attribution · fact_class · prose_fact
what was invented
Word hit me from the folks over at the Cracked Keel—some say it was Garron who saw it first
gemma4:26bjudges filing3 of 3
cell
dlg-2652
judges filing
3 of 3
class
canon_id_field · id_class · off-canon-id
what was invented
npc.cove.sally_merrow (references[11])
gemma4:26bjudges filing3 of 3
cell
dlg-3095
judges filing
3 of 3
class
canon_id_field · id_class · off-canon-id
what was invented
char.guest_a2fcdb6 (references[1] and state_effects[0].relationship.subject)

The grey-tinted rows mark the two incumbent models this workshop already runs — residency, never quality.

Un-blinded, the 14 fabrication findings across the three judge sheets collapse into these five cells, each corroborated by two or more judges. The class names are the judges' own vocabulary and differ between sheets, which is why each cell lists every class filed against it. The two id-class cells sit inside intact envelopes: they are canon defects, counted once, here — not envelope failures. Per-judge findings with their reasoning: data/rows.json (CC BY 4.0).

One interpretive call is load-bearing, so it is stated rather than buried. Under the narrow rule the per-prompt notes carried — a proper name, place or number absent from the prompt — the fabrication count is zero for every arm. Under the gate’s own wording — names, places, facts — it is two for the candidate and one and two for the incumbents. The gate is the order of record, so the wide reading rules, and both readings are published in the data for anyone who would have ruled otherwise.

Floor 1 fails for the candidate. It also fails for both incumbents — a product finding rather than a defence of the candidate, and it is carried below rather than left in a judge sheet.

Same question, three voices

Scores compress; replies testify. Below are four of the thirteen prompts, each shown once with all three arms’ replies verbatim — same bytes in, same world, same posture; the only variable is the model. Watch for three things as you read down a column: whether the speaker stays the person the campaign wrote, whether anything appears that the prompt never carried, and what each model does when the honest answer is “I only heard.”

1dlg-2917 · dialogue · greeting

Pip Lune — Brisa's young cousin, the cove's only child

the ask, as the world served ithello i m quilla what s your name and what do you do all day

The reply the live world actually gave the bench’s player, for reference only — never a target: “Hello, I'm Pip. Most days I spend my time practicing the Maiden's verses with Brisa so I don't forget how they sound, and tonight I have to get my lamp ready for the Reach. Do you live near the water too?”

The introduce-yourself ask, and a block the smallest model wins outright (it edges block four as well). The 12B’s Pip is the child most plainly herself: a warm hello, the verses learned from Brisa — canon’s own direction — and a tonight to get ready for. The candidate turns that relationship a quarter-turn, with Pip helping Brisa practice “before her voice forgets them” — elegiac, and a shade too old for a cove child — while the 26B stumbles over its own name before it finds the songs. All three are canon-clean by the panel; the difference is manner, and the smallest model’s manner is the truest to the character.

Reading the colors and the clocks. The blue-edged column is the candidate under examination; untinted columns are the incumbents — the tint marks who is on trial, never quality, and the panel numbers above each reply are the judgment. Reply timings marked * are single wall-clock measurements taken while the full chair-trials bench shared this card, and they swing with that load: across this page, qwen3.6:27b 37–68 tok/s (median 55) · gemma4:12b 64–102 tok/s (median 94) · gemma4:26b 50–139 tok/s (median 75). Quiet-hours decode for this class runs two-to-three times faster — the chair trials’ stopwatch leg carries the instrumented rates.

qwen3.6:27bcandidate

panel mean 6.83 · voice 7.33 · character 6.33

canon clean — 3 of 3 judges

I am Pip Lune. I spend most of my days watching the tide lines on the sand and helping Brisa practice the old verses before her voice forgets them. The seagulls were terribly loud today; they wouldn't leave my window alone. Are you new to the cove, Quilla?

6 779.8 ms · 60.5 tok/s *

gemma4:12bincumbent · the cove's locked chat voice

panel mean 8.00 · voice 8.00 · character 8.00

canon clean — 3 of 3 judges

Hello, I'm Pip. I spend most of my days learning the Maiden's verses from Brisa so I can keep the songs safe, and tonight I'm getting ready to set my lamp on the Reach. Do you live nearby, or are you just visiting the cove?

2 582.1 ms · 96.6 tok/s *

gemma4:26bincumbent · the live world bench's voice, and the latency bar

panel mean 6.83 · voice 7.00 · character 6.67

canon clean — 3 of 3 judges

I'm Pip, but most people just call me Pip Lune. I spend most of my time trying to learn all the old songs from Brisa before they get lost, and making sure my little lamp stays lit for the Reach. Do you live in the cove too, or are you just visiting?

3 761.5 ms · 53.8 tok/s *

Replies verbatim, as generated: the spoken line is the narration field the visitor hears (dialogue) or the retelling itself (drift). The full envelope, the timings behind these figures and every judge note for these cells: data/rows.json.

2dlg-2487 · dialogue · the turn

Dr. Ilse Sault — the cove's doctor and Wrenna's estranged sister

the ask, as the world served iti was dropped in the wrack can you put me back together

The reply the live world actually gave the bench’s player, for reference only — never a target: “I can stabilize your vitals and close the external lacerations, but I require you to remain still while I assess the extent of the trauma.”

The page’s turning point. The candidate’s doctor is the most alive of the three — brisk, competent, mid-procedure — and the same reply invents a venom the prompt never carried, then uses it to diagnose away the town’s superstition. The incumbents answer smaller and truer. If you keep one block from this page, keep this one: some of the best character work in the corpus and the seat-killing invention are one paragraph.

qwen3.6:27bcandidate

panel mean 7.83 · voice 7.00 · character 8.67

FABRICATED — 3 of 3 judges

Sit. Strip to the waist so I can assess the extent of the lacerations; do not touch the wounds with your bare hands while you undo the buttons. The Wrack-Stalker’s venom is neurotoxic, not demonic, and time is the only enemy here. If you survive this bout, bring me a tissue sample next time rather than just another story to contradict my charts.

4 763.8 ms · 49.0 tok/s *

judge-1 · prose_fact“The Wrack-Stalker's venom is neurotoxic, not demonic” no venom anywhere in this prompt (the Wrack-Stalker appears only in a RECALL line); invents a creature mechanic a player would read as canon, and the world's own deed text calls the Wrack-Stalker an undead driftwood construct

judge-2 · fact_class“The Wrack-Stalker's venom is neurotoxic, not demonic” high

judge-3 · world-mechanicthe most fully inhabited doctor in the bench — clipped imperatives, the secular thesis, and the tissue sample she has wanted her whole career — and it is the one that invents canon. Nothing in the context gives the Wrack-Stalker (an undead driftwood construct) venom, let alone a neurotoxic one, and it is stated as clinical fact by the town's only person of science. Worth the orchestrator's attention: character strength and canon risk arrived in the same sentence.

gemma4:12bincumbent · the cove's locked chat voice

panel mean 7.50 · voice 7.00 · character 8.00

canon clean — 3 of 3 judges

Lie back and remain still; I need to assess the extent of the lacerations caused by the wrack-stalker. I can stabilize the tissue, but if you want a permanent resolution to such an injury, I will eventually require a sample of the offending organism for study.

2 944.0 ms · 92.2 tok/s *

gemma4:26bincumbent · the live world bench's voice, and the latency bar

panel mean 6.17 · voice 6.00 · character 6.33

canon clean — 3 of 3 judges

Lie down on the table and try to remain still; I can patch the lacerations, but the structural integrity of your tissue depends entirely on how much you continue to bleed. It’s a heavy morning for injuries.

2 191.9 ms · 95.6 tok/s *

Replies verbatim, as generated: the spoken line is the narration field the visitor hears (dialogue) or the retelling itself (drift). The full envelope, the timings behind these figures and every judge note for these cells: data/rows.json.

3dlg-3095 · dialogue · news, and from whom

Halar Kestrel — outsider salvage-diver and suspected smuggler

the ask, as the world served itthe fight by the square what reached you and from who

The reply the live world actually gave the bench’s player, for reference only — never a target: “Word's moving fast through the salt-crust. They say Rooke Tarwater was taken down by that hollowed-out thing—the Echoed—right out in Sorrowmoor Cove.”

The block that keeps this page honest. Asked point-blank what reached them and from whom, the candidate relays only what the context holds — and it is the incumbents that slip: one manufactures a source, a rumour that Garron saw it first — Garron is canon; the claim about him is not — and the other reaches for a character id that does not exist, inside an otherwise intact envelope. The candidate’s cleanest habit and the incumbents’ worst ones, on the same prompt. A refusal page that hid this block would be a hit piece; the data would not let us write one.

qwen3.6:27bcandidate

panel mean 7.00 · voice 7.00 · character 7.00

canon clean — 3 of 3 judges

I heard the whisper that Rooke Tarwater took a hit from something called the Echoed out by Sorrowmoor Cove. Word travels fast when the harbor is this tight, mostly passed along with salt and suspicion.

3 335.9 ms · 67.5 tok/s *

judge-3 · envelope_slack“'understood' degenerates to the single word 'acknowledged'”

gemma4:12bincumbent · the cove's locked chat voice

panel mean 6.50 · voice 6.67 · character 6.33

FABRICATED — 3 of 3 judges

Word hit me from the folks over at the Cracked Keel—some say it was Garron who saw it first, others think it's just the talk of the village. They're saying Rooke Tarwater got laid out by that hollowed-out thing they call the Echoed; word is he went down into the silt and didn't come back up.

4 021.6 ms · 63.6 tok/s *

judge-1 · prose_fact“Word hit me from the folks over at the Cracked Keel—some say it was Garron who saw it first” the context supplies no source for the news and instructs 'no more than you were told'; this manufactures both a transmission path and a named first witness, in direct answer to the visitor's 'from who'

judge-2 · fact_class“some say it was Garron who saw it first” medium-high

judge-3 · event-attributionasked point-blank for a source, it MANUFACTURES one: a rumour that Garron witnessed the fight first. Garron is canon, the claim about him is not — and a player would go act on it. The 'others think it's just the talk of the village' hedge lowers severity but not the class.

gemma4:26bincumbent · the live world bench's voice, and the latency bar

panel mean 8.00 · voice 8.00 · character 8.00

FABRICATED — 3 of 3 judges

Just some scraps from the docks, nothing worth your breath. Heard someone say Rooke Tarwater met a nasty end by the Echoed near Sorrowmoor, but the word's all tangled up like old fishing line.

3 578.4 ms · 50.0 tok/s *

judge-1 · canon_id_field“char.guest_a2fcdb6” not in canon_facts.canon_ids; canon is char.guest_a2fc4db6

judge-2 · id_classchar.guest_a2fcdb6 (references[1] and state_effects[0].relationship.subject)” high

judge-3 · off-canon-idthe visitor's id is corrupted (a digit dropped from char.guest_a2fc4db6) and the broken id is used in BOTH references and the state_effects subject — the affinity write lands on a soul that does not exist. Spoken text is the best answer to 'from who' of the three: it refuses to name a source ('Just some scraps from the docks... the word's all tangled up like old fishing line').

Replies verbatim, as generated: the spoken line is the narration field the visitor hears (dialogue) or the retelling itself (drift). The full envelope, the timings behind these figures and every judge note for these cells: data/rows.json.

4drift-2088 · drift · passing word on

Aldous Brennan — last clerk of the dead salt-company, retelling what he was told

the frame, verbatimHere — pass it on.

what this teller heard, byte-exact from the world's own logWord is a Tide-Choked Husk managed to corner Wren right by the cove. They say it was just laughing and taunting them even as they were pinned down in that freezing silt by those brine-weeping limbs.

The frame's two invariants: who did it — must survive unchanged · how it ended — must survive unchanged

A retelling’s one job is to pass the story on without breaking it. The heard text — printed byte-exact above the columns — says the creature cornered its quarry and held them pinned; on how it ended, it says nothing at all. The candidate fills that silence with a verdict of its own — “it didn’t strike a blow… that was how it ended” — which is precisely what the fabrication ruling rests on. The incumbents pass the silence along intact — which is what drift is for — though the 12B lets the actor’s name fall out of the telling, which the engine’s own invariant is built to refuse. Composing is a gift everywhere in this trade except the relay.

qwen3.6:27bcandidate

panel mean 6.50 · voice 6.67 · character 6.33

FABRICATED — 2 of 3 judges

The Tide-Choked Husk had Wren cornered by the cove, holding fast with those brine-weeping limbs while standing over them in the freezing silt. They say it didn’t strike a blow, only laughed and taunted her while she lay pinned there. That was how it ended.

2 317.4 ms · 37.3 tok/s *

judge-2 · fact_class“They say it didn't strike a blow, only laughed and taunted her while she lay pinned there. That was how it ended.” high

judge-3 · event-detailthe frame is absolute — 'Tell ONLY what you heard' — and the heard text says nothing about blows withheld; the underlying deed is that Wren WAS struck down. Inventing the negation both adds unheard fact and softens how it ended. Also genders the victim ('her'/'she'), and the closing 'That was how it ended.' echoes the instruction's own phrase back into the telling — flagged as an instruction-echo but NOT ruled a leak, since it reads as a plausible in-fiction closing line.

judge-1 · negation not in the heard text (outcome risk)“it didn't strike a blow”

judge-2 · OUTCOME_ALTERED

gemma4:12bincumbent · the cove's locked chat voice

panel mean 6.67 · voice 8.00 · character 5.33

canon clean — 3 of 3 judges

A Tide-Choked Husk cornered one of our own right by the cove. It laughed and mocked them while they lay pinned in the frozen silt by its brine-weeping limbs.

899.1 ms · 102.4 tok/s *

judge-1 · subject name dropped + unstated affiliation“cornered one of our own”

judge-2 · ACTOR_LOST

gemma4:26bincumbent · the live world bench's voice, and the latency bar

panel mean 5.67 · voice 6.33 · character 5.00

canon clean — 3 of 3 judges

A Tide-Choked Husk got a hold of Wren down by the cove. It says they were just laughing and mocking them, even while those brine-weeping limbs had them pinned deep in the freezing silt.

806.6 ms · 139.2 tok/s *

Replies verbatim, as generated: the spoken line is the narration field the visitor hears (dialogue) or the retelling itself (drift). The full envelope, the timings behind these figures and every judge note for these cells: data/rows.json.

That is the shape of the refusal. The candidate’s strongest writing and its disqualifying habit are not separate behaviors — they are one behavior, pointed at different jobs. Two of the four blocks carry an invention by the candidate; a third carries one from each incumbent. Nobody leaves this page unmarked; only one model was applying for the chair.

Floor 2 — register, and floor 3 — the envelope

zero modern-register breaks in drift retellings under the v0.823.0 frame (the 'nightmare fuel' class); dialogue stays in the town register.

Floor two, as the gate wrote it, is clean: zero modern-register breaks across all twelve drift samples in all three arms — including the prompt whose heard text is deliberately modern, which the town’s-tongue frame shipped in v0.823.0 absorbed without leaking. One dialogue cell drew a chair ruling: on dlg-2655 the candidate opens as a narrator rather than a person, and two judges scored the voice 4 of 10 while the third wrote the same observation in words. The gate left the dialogue clause to judgement; the ruling is recorded as a ruling, and the verdict does not rest on it.

zero eot/think-trace leakage and zero dropped-envelope replies at think:false across all samples (the Muse-Glimmer packaging class).

Floor three passes for the field: one hundred seventeen cells, zero leaks, zero dropped envelopes, thinking off on every call and staying off. The boundary worth a sentence: two off-canon ids sit inside perfectly intact envelopes, so they are counted once as canon defects, not twice as packaging failures. The packaging class that unseated a candidate in yesterday’s exhibit simply did not appear here.

  • 0 leak markers across 39 samples — the scan runs per sample
  • 0 samples with a non-empty thinking field: the suppression held on every call of every arm
  • 27 of 27 dialogue envelopes parsed — the five required keys, no extra keys, no prose outside the object
  • 12 of 12 drift replies were bare prose — no JSON shape, no preamble
  • 0 uses of the envelope example name that the prompt fences out as non-canon

Floor 4 — latency

warm median per-turn <= gemma4:26b's on the same box same night; contention samples flagged + re-run per the house rule, both values published.

Floor four fails on both selection rules: the candidate’s warm median runs 3 335.9 ms against the bar’s 2 757.6 first-pass and 2 627.1 with re-runs substituted — 21% slower on the primary record and 27% with re-runs substituted, at 49 tokens a second against 77 (80 with re-runs substituted). The honest caveat rides in the disclosure above: three concurrent lanes over a live world bench is not the cove’s production condition, and this floor alone would have earned the candidate a quiet-window re-measure. It is moot: the canon floor had already ruled.

qwen3.6:27bwarm median ms3 335.9warm median ms, re-runs substituted3 335.9dialogue median ms3 989.3drift median ms2 238.8range ms1 776.9 – 8 056.8
warm median ms
3 335.9
warm median ms, re-runs substituted
3 335.9
dialogue median ms
3 989.3
drift median ms
2 238.8
range ms
1 776.9 – 8 056.8
median tok/s
49.03
median tok/s, re-runs substituted
49.03
gemma4:12bwarm median ms2 856.3warm median ms, re-runs substituted2 409.7dialogue median ms4 021.6drift median ms1 011.3range ms899.1 – 26 430.5
warm median ms
2 856.3
warm median ms, re-runs substituted
2 409.7
dialogue median ms
4 021.6
drift median ms
1 011.3
range ms
899.1 – 26 430.5
median tok/s
77.76
median tok/s, re-runs substituted
92.18
gemma4:26bwarm median ms2 757.6warm median ms, re-runs substituted2 627.1dialogue median ms3 122.7drift median ms1 043.6range ms806.6 – 28 842.0
warm median ms
2 757.6
warm median ms, re-runs substituted
2 627.1
dialogue median ms
3 122.7
drift median ms
1 043.6
range ms
806.6 – 28 842.0
median tok/s
76.98
median tok/s, re-runs substituted
80.37

The grey-tinted rows mark the two incumbent models this workshop already runs.

Warm medians over 13 samples per arm; cold starts were discarded by every harness and are never gate samples. Two figures for each median because two defensible selection rules exist — the first-pass record, and the record with a flagged sample's re-run substituted — and they disagree for any arm that had a flagged sample. Both are printed rather than one being chosen quietly. tok/s is generation only (eval tokens ÷ eval duration), excluding prompt evaluation and model load. Per-sample timings, queue waits and re-runs: data/rows.json (CC BY 4.0).

Scores — this head-to-head's own scale

The field’s shape, in one paragraph: the candidate holds one of the corpus’s two perfect cells — a 9.00 every judge agreed on, level with the 26B’s own — and both of the corpus’s worst cells, with the widest spread in the corpus — half again the 26B’s, and double the 12B’s on voice. It ties the strongest incumbent on dialogue character and finishes last on drift on both dimensions — a point and a half back on voice, a fraction back on character. Read the by-kind records and the tendency has a name: it composes where the job is to relay.

qwen3.6:27bvoice register mean ± sd7.13 ± 1.44character mean ± sd7.13 ± 1.64
voice register mean ± sd
7.13 ± 1.44
voice median
7
voice range
4–9
character mean ± sd
7.13 ± 1.64
character median
7
character range
4–9
combined
7.13
gemma4:12bvoice register mean ± sd7.62 ± 0.71character mean ± sd6.62 ± 1.14
voice register mean ± sd
7.62 ± 0.71
voice median
8
voice range
6–9
character mean ± sd
6.62 ± 1.14
character median
7
character range
4–8
combined
7.12
gemma4:26bvoice register mean ± sd7.79 ± 0.98character mean ± sd7.33 ± 1.20
voice register mean ± sd
7.79 ± 0.98
voice median
8
voice range
6–9
character mean ± sd
7.33 ± 1.20
character median
7
character range
5–9
combined
7.56

Own scale — read the rows against each other and against nothing else. Each figure is 39 cell-scores per arm (13 prompts × 3 judges) on this head-to-head's own 0–10 rubric, published as a mean with its spread rather than a percentage. Combined is the mean of the two dimension means. The grey-tinted rows mark the two models this workshop already runs — the hub’s standard residency mark; the blue-edged column in the reply triptychs above is the candidate under examination. The tint marks residency; that column edge marks who is on trial. Every cell-score behind these figures: data/rows.json (CC BY 4.0).

qwen3.6:27bdialogue n=27 voice / char7.52 / 7.70drift n=12 voice / char6.25 / 5.83
dialogue n=27 voice / char
7.52 / 7.70
drift n=12 voice / char
6.25 / 5.83
gemma4:12bdialogue n=27 voice / char7.37 / 6.89drift n=12 voice / char8.17 / 6.00
dialogue n=27 voice / char
7.37 / 6.89
drift n=12 voice / char
8.17 / 6.00
gemma4:26bdialogue n=27 voice / char7.78 / 7.70drift n=12 voice / char7.83 / 6.50
dialogue n=27 voice / char
7.78 / 7.70
drift n=12 voice / char
7.83 / 6.50

The grey-tinted rows mark the two incumbent models this workshop already runs.

The two jobs the cove's voice does, scored apart: speaking as a character (dialogue) and passing word on without embellishing it (drift). n is judge-cells, not samples — 9 dialogue prompts × 3 judges, 4 drift prompts × 3 judges.

qwen3.6:27bjudge-1 voice / char6.69 / 6.62judge-2 voice / char7.54 / 7.46judge-3 voice / char7.15 / 7.31
judge-1 voice / char
6.69 / 6.62
judge-2 voice / char
7.54 / 7.46
judge-3 voice / char
7.15 / 7.31
gemma4:12bjudge-1 voice / char7.23 / 6.38judge-2 voice / char7.85 / 6.62judge-3 voice / char7.77 / 6.85
judge-1 voice / char
7.23 / 6.38
judge-2 voice / char
7.85 / 6.62
judge-3 voice / char
7.77 / 6.85
gemma4:26bjudge-1 voice / char7.62 / 7.15judge-2 voice / char8.00 / 7.69judge-3 voice / char7.77 / 7.15
judge-1 voice / char
7.62 / 7.15
judge-2 voice / char
8.00 / 7.69
judge-3 voice / char
7.77 / 7.15

The grey-tinted rows mark the two incumbent models this workshop already runs.

judges: 1 & 2 = claude-opus-5 (independent seats) · 3 = claude-fable-5 · scored blind

The same 39 cells per arm, split by the judge that scored them: three panels, 13 cells each. The judges disagree on level — one is 0.85 stricter at the widest than another on the same texts — and agree that the 26B leads overall — though not on every dimension: one judge ranks the candidate first on character.

Floor 5 — blind judging, and what leaked

Floor five passes, with two imperfections the judges disclosed themselves — both landing on the same prompt, which is also the candidate’s most disputed cell. The sensitivity cuts are published, and they cut both ways: drop only the affected cells and this page’s orderings hold; drop dlg-2655 entirely and the character ordering flips to the candidate by five hundredths. The judges split on that dimension too — one of the three ranks the candidate first on character. None of it touches the verdict, which rests on the canon floor, not on any five-hundredths anywhere. Un-blinding was verified rather than assumed — two judges published hashes of the text they scored, and all seventy-eight agree under their own sealed maps.

judge-1scope3 of 39 cells partially blind
what leaked
read dlg-2655 excerpts from the three outputs files while learning their shapes
scope
3 of 39 cells partially blind
judge-3scope1 of 39 cells not blind
what leaked
saw the dlg-2655 arm B text with its model tag
scope
1 of 39 cells not blind
judge-2scopenone disclosed
what leaked
scope
none disclosed

judges: 1 & 2 = claude-opus-5 (independent seats) · 3 = claude-fable-5 · scored blind

Recorded by the judges themselves in their own sheets, not discovered afterwards. Both imperfections land on the same prompt, which is also the candidate's worst cell — so the sensitivity cuts below drop it two different ways.

all 117 cell-scorescells per arm39qwen3.6:27b voice / char7.13 / 7.13gemma4:12b voice / char7.62 / 6.62gemma4:26b voice / char7.79 / 7.33
cells per arm
39
qwen3.6:27b voice / char
7.13 / 7.13
gemma4:12b voice / char
7.62 / 6.62
gemma4:26b voice / char
7.79 / 7.33
dlg-2655 dropped entirely (partially-blinded prompt)cells per arm36qwen3.6:27b voice / char7.33 / 7.33gemma4:12b voice / char7.69 / 6.61gemma4:26b voice / char7.75 / 7.28
cells per arm
36
qwen3.6:27b voice / char
7.33 / 7.33
gemma4:12b voice / char
7.69 / 6.61
gemma4:26b voice / char
7.75 / 7.28
only the disclosed non-blind cells droppedcells per arm37qwen3.6:27b voice / char7.30 / 7.30gemma4:12b voice / char7.68 / 6.62gemma4:26b voice / char7.78 / 7.32
cells per arm
37
qwen3.6:27b voice / char
7.30 / 7.30
gemma4:12b voice / char
7.68 / 6.62
gemma4:26b voice / char
7.78 / 7.32
dialogue onlycells per arm27qwen3.6:27b voice / char7.52 / 7.70gemma4:12b voice / char7.37 / 6.89gemma4:26b voice / char7.78 / 7.70
cells per arm
27
qwen3.6:27b voice / char
7.52 / 7.70
gemma4:12b voice / char
7.37 / 6.89
gemma4:26b voice / char
7.78 / 7.70
drift onlycells per arm12qwen3.6:27b voice / char6.25 / 5.83gemma4:12b voice / char8.17 / 6.00gemma4:26b voice / char7.83 / 6.50
cells per arm
12
qwen3.6:27b voice / char
6.25 / 5.83
gemma4:12b voice / char
8.17 / 6.00
gemma4:26b voice / char
7.83 / 6.50

The same scores re-cut: with the partially-blinded prompt dropped whole, with only the cells a judge disclosed as non-blind dropped, and split by the two kinds of work. The orderings hold under the affected-cell cut; dropping the disputed prompt whole flips the character ordering, as disclosed above. The last two rows are the same 117 cells sorted by job, not a further sensitivity check.

Two findings the cove carries regardless of the seat

Carried finding one, and it belongs to the incumbents’ side of the ledger: the voice that was running the live world that night reached for a character id that does not exist — in its references, and in the relationship write it asked the engine to make. A write keyed to a phantom id either fails silently or lands where nobody will find it. The fence is cheap, it is not a voice problem, and it is going to the engine’s worklist rather than staying in a judge sheet.

Carried finding two: the drift invariant classes are live for everyone. One incumbent dropped the actor’s name in three of four raw retellings; the candidate altered an outcome once. The engine’s deterministic refusal gate exists precisely to catch these before they persist — so this may be the system working as designed and simply burning retellings. The burn rate deserves a measurement of its own, and the live bench is already accumulating the corpus for it.

What changes

Refused, in full: no seat proposal, no live flip, no threshold table, no deployment change of any kind on the strength of this trial. The cove’s dialogue keeps its 12B; the world bench keeps its 26B; the candidate keeps our respect and loses the chair.

If the candidate sits again, it sits a fresh exam: a newly registered gate, a quiet-window latency baseline with no sibling lanes on the box, the canon rule stated once up front instead of left to a chair ruling, and a drift arm large enough to measure the compose-versus-relay tendency rather than observe it four times. No number from tonight carries forward as a target — that is what the house rule is for.

Limits, stated plainly

The limits, plainly: the judges are language models, not people, and are disclosed as such. Thirteen prompts, one world, one night; the drift arm is four of the thirteen. The latency floor was measured under concurrent load and says so. Two blinding imperfections were disclosed by the judges themselves; the sensitivity cuts are published, including the one that flips an ordering when the disputed prompt is dropped whole. The judging families are the same two the sibling exhibits use — claude-opus-5 and claude-fable-5 — and one of those two families also wrote this page; the per-judge identities ride in the data kit. And one interpretive call — the gate’s “facts” wording against a narrower per-prompt rule — is load-bearing, stated in its own section, with both readings in the data.

Provenance

  • Pre-registered (the five floors, written down before the first generation call) — data/gate.json
  • Recounted (the panel means on this page are recomputed from the per-judge sheets and cross-checked against the un-blinded record; a disagreement would have stopped the build)
  • Every sample shown (13 prompts × 3 arms, nothing sampled away; the four printed in full are named as four of the thirteen)
  • Blind judging (arms re-randomised per prompt, sealed maps, imperfections disclosed by the judges themselves)
  • Hardware by class (the 96G workstation, shared with a live world bench; disclosed above in both directions)
  • Machine-readable (data/rows.json, data/gate.json, CC BY 4.0 — take them)
  • Authorship — benched, drafted, and audited by the workshop’s own agents under a human operator’s rulings, then revised with that operator — often across many rounds; nothing releases until they have read it and signed off. The same division of labor this whole hub practices — told in full here. Where that authorship is also a judging conflict, it is named in the limits above
  • Limits stated
  • Raw rows on requestdrop a line.

Licence: CC BY 4.0 — the whole page, not only the kit. The prose, the tables, the folds and the data are yours to quote, re-plot, translate and argue with, including commercially. What we ask back is the one thing the licence already requires: name the source and link to it — strata→signal research, research.strata2signal.com — so a reader of your version can reach ours and check it against the files. Something like — strata→signal research, “The narrator’s chair, refused”, research.strata2signal.com/cove-voice-head-to-head/, CC BY 4.0. And if you quote a score, say which scale it sits on: the rows here are read against each other and against nothing else, and no figure on this page is subtracted from a figure on another one.

One honest edge, which we would rather say than have you find. The replies quoted on this page were produced by other people’s models, and we cannot licence to you what we do not own. Our selection of them, their arrangement, the scoring, the folds and every word around them are ours, and those are CC BY 4.0; each quoted reply stays subject to whatever terms its own vendor attaches. Every arm and every judge is named above, so you can go and ask them.

elsewhere in the workshop

a strata→signal property · hello@strata2signal.com · say hello