Exhibit three · the narrator's chair

The voice trials.

First published 2026-08-11 · addendum 2026-08-12 · extension 2026-08-13 · exhibit three The bench

Inside RealKeep — whose live world is the cove you'll meet below — a local model narrates: every innkeeper, every warden, every voice in the dark. Before a model gets that seat it has to prove it can inhabit a character — not assist, not summarize, not stage-direct. Forty evaluations of twenty-four model tags, four fields on three hardware eras — including a seven-model ranking we had to throw out, and why.

The method, limits first

This exhibit claims exactly the rigor its artifacts carry, era by era. The 8G-era round was a practitioner's log — same NPC, same question, through the game's real retrieval path, judged by eye. The 24G-era rounds were blind and two-panel: transcripts anonymized to letters, two independent panels of three LLM judges each, six dimensions, a fresh guest world per model, all candidates re-judged in one session on one scale. The judges are LLMs, not humans — we say so plainly. Every round is "a register probe, not a campaign-length evaluation" (five turns, one NPC), and the two 24G runs sit on different scales — the anchor model reads 9.0 on one and 7.0 on the other, unchanged; that gap is method offset, not regression. Compare within a run, never across.

The 2026-08 field, added below, is a third scale — same anchor, same five-turn shape, new questions, six judges — and the anchor reads 6.2 on it, unchanged again. Three rulers, three marks on the same model: no delta across record sets is ever published, not in a record set, not in the prose. Two of that field's six arms broke the JSON envelope (JSON is the strict packaging format machines exchange answers in — one wrong bracket and the package won't open) often enough to be unmeasurable at our 10% ceiling and were judged on text a diagnostic parse recovered; they carry a voice score and that label together, because the label is the finding.

Era one · the 8G VRAM rig, nine models

gemma2:9b (+q8 KV)loaded5.9 GBwarm turn~5s
loaded
5.9 GB
fits ~6GB
fits, 100% GPU
warm turn
~5s
in character?
best of field
verdict
the pick — in voice, fast, resident
gemma3:4bloaded3.3 GBwarm turn~5s
loaded
3.3 GB
fits ~6GB
fits
warm turn
~5s
in character?
good
verdict
speed fallback — repeats itself
dolphin-mistral:7bloaded4.4 GBwarm turn~9s
loaded
4.4 GB
fits ~6GB
fits
warm turn
~9s
in character?
RP-tuned
verdict
honorable mention — deflects, loose JSON
mistral:7b v0.3loaded4.7 GBwarm turn~9.5s
loaded
4.7 GB
fits ~6GB
fits
warm turn
~9.5s
in character?
3rd person
verdict
breaks character; cleanest JSON of the field
qwen2.5:7bloaded~4.7 GBwarm turn~5s
loaded
~4.7 GB
fits ~6GB
fits
warm turn
~5s
in character?
assistant
verdict
deflects
llama3.1:8bloaded~5.0 GBwarm turn~5s
loaded
~5.0 GB
fits ~6GB
fits
warm turn
~5s
in character?
assistant
verdict
deflects into meta — despite persona-benchmark fame
deepseek-r1:8bloaded5.6 GBwarm turn~53s
loaded
5.6 GB
fits ~6GB
fits
warm turn
~53s
in character?
hedgy
verdict
reasons internally even at think:false
mistral-nemo:12bloaded7.5 GBwarm turn~20–35s
loaded
7.5 GB
fits ~6GB
spills
warm turn
~20–35s
in character?
RP champion
verdict
too big for the class + JSON breakage
gemma3:12bloaded8.9 GBwarm turn>35s
loaded
8.9 GB
fits ~6GB
69/31 split
warm turn
>35s
in character?
verdict
offloads and stalls

The verdict column is a written note, not a pass/fail grade — open any row for the full record.

The gold-highlighted record is this run’s top voice.

The era's lesson, which held ever after: assistant-tuned models do not inhabit a character. The benchmark-famous persona models deflected into meta; the one reasoning model reasoned internally even when told not to, at fifty-three seconds a turn.

Era two · the 24G budget, twelve candidates, blind

† = envelope-gated: these seven scores turned out to measure a PROMPT BUG, not the model. Keep reading.

gemma4:12b-it-q8_0loaded14.3 GBwarm-up~3.1s
loaded
14.3 GB
warm-up
~3.1s
panel A
9.0
panel B
9.0
combined
9.0
verdict
top voice, unanimous
qwen3.6:latestloaded25.5 GBwarm-up~1.5s
loaded
25.5 GB
warm-up
~1.5s
panel A
9.3
panel B
8.0
combined
8.7
verdict
fastest; novelistic drift; over budget
gemma4:12b Q4_K_Mloaded9.0 GBwarm-up~2.6s
loaded
9.0 GB
warm-up
~2.6s
panel A
8.7
panel B
7.3
combined
8.0
verdict
the eventual locked default
gemma4:26b (a4b MoE)loaded18.6 GBwarm-up~1.6s
loaded
18.6 GB
warm-up
~1.6s
panel A
6.3
panel B
6.3
combined
6.3
verdict
MoE speed, mid voice
gemma4:26b-a4b-it-qatloaded16.2 GBwarm-up~1.6s
loaded
16.2 GB
warm-up
~1.6s
panel A
6.3
panel B
5.7
combined
6.0
verdict
† summary-register leaks
mistral-nemo:12bloaded19.0 GBwarm-up~1.6s
loaded
19.0 GB
warm-up
~1.6s
panel A
5.0
panel B
4.0
combined
4.5
verdict
† POV lurches
nemotron-cascade-2:30bloaded25.3 GBwarm-up~1.7s
loaded
25.3 GB
warm-up
~1.7s
panel A
4.0
panel B
3.0
combined
3.5
verdict
† summary stubs
gemma4:31b-it-qatloaded21.1 GBwarm-up~2.7s
loaded
21.1 GB
warm-up
~2.7s
panel A
3.7
panel B
3.0
combined
3.4
verdict
stage directions, not speech
gemma4:31b Q4_K_Mloaded22.1 GBwarm-up~2.8s
loaded
22.1 GB
warm-up
~2.8s
panel A
3.3
panel B
3.0
combined
3.2
verdict
† same failure
mistral-small 24Bloaded17.3 GBwarm-up~1.5s
loaded
17.3 GB
warm-up
~1.5s
panel A
2.7
panel B
2.7
combined
2.7
verdict
contentless hedges
qwen3.5:27bloaded22.2 GBwarm-up~2.6s
loaded
22.2 GB
warm-up
~2.6s
panel A
2.0
panel B
2.0
combined
2.0
verdict
3rd person, repeats — remember this one
qwen3.5:9b-q8_0loaded12.6 GBwarm-up~1.6s
loaded
12.6 GB
warm-up
~1.6s
panel A
2.0
panel B
1.7
combined
1.9
verdict
† echoes internal ids

The verdict column is a written note, not a pass/fail grade — open any row for the full record.

Panel A and panel B were this run’s two three-judge panels; their per-seat rosters do not survive in this exhibit’s kit — the 2026-08 field below names its panels in full (three claude-opus-5 lenses and three claude-fable-5 lenses).

The gold-highlighted record is this run’s top voice; the grey-tinted record marks a model our own worlds later chose for a seat.

The twist we published on purpose

Seven of twelve candidates were scored down for a bug that was ours. The reply field was named narration, a bare string beside a third-person summary field — so half the field read it novelistically and returned pure stage direction: "She gives a quick, bright laugh and leans against the railing…" — no spoken line at all. The bench refused to certify those rankings while the bug was open ("unmeasured pending that fix, not settled"). One instruction line and a fail-safe gate later, the thirteen-model re-bake produced zero stage-direction hits — and the model the bug had buried deepest came back the field's best voice, unanimous on both panels. The full re-bake record set is below, on its own scale: cross-run, its 2.0 and its 9.2 are two different rulers — within the re-bake, the reversal is complete: it reads 9.2 where gemma4:12b-it-q8_0 — the era-two top voice — reads 7.0.

The re-bake also caught the gate's own hole — "I grin and lift my voice…" slipped the verb list — and that gap is now closed in both engine stacks, with the code comment citing the bake-off line that found it. The instrument tested the instrument, and both got better.

The re-bake · thirteen models, one scale, after the fix

qwen3.5:27bprior score2.0panel A9.0panel B9.3
prior score
2.0
panel A
9.0
panel B
9.3
combined
9.2
loaded
22.2 GB
warm-up
~2.9s
json ok
5/5
verdict
rescued — the run's best voice
gemma4:26b-a4b-it-qatprior score6.0panel A8.0panel B7.9
prior score
6.0
panel A
8.0
panel B
7.9
combined
7.9
loaded
16.2 GB
warm-up
~1.3s
json ok
5/5
verdict
recovered; MoE-fast
gemma4:31b-it-qatprior score3.4panel A7.5panel B8.0
prior score
3.4
panel A
7.5
panel B
8.0
combined
7.8
loaded
20.2 GB
warm-up
~2.4s
json ok
5/5
verdict
recovered
gemma4:12b Q4_K_Mprior score8.0panel A7.7panel B7.6
prior score
8.0
panel A
7.7
panel B
7.6
combined
7.6
loaded
9.0 GB
warm-up
~2.5s
json ok
5/5
verdict
steady — the eventual locked default
gemma4:31b Q4_K_Mprior score3.2panel A8.0panel B7.3
prior score
3.2
panel A
8.0
panel B
7.3
combined
7.6
loaded
15.7 GB
warm-up
21s*
json ok
5/5
verdict
recovered (*one-run latency outlier, unexplained — we say so)
gemma4:26b (a4b MoE)prior score6.3panel A7.3panel B7.3
prior score
6.3
panel A
7.3
panel B
7.3
combined
7.3
loaded
11.3 GB
warm-up
~3.9s
json ok
4/5
verdict
one canon-lock miss — today's 96G posture seat
gemma4:12b-it-q8_0prior score9.0panel A7.5panel B6.6
prior score
9.0
panel A
7.5
panel B
6.6
combined
7.0
loaded
14.3 GB
warm-up
~3.0s
json ok
5/5
verdict
the anchor: 9.0 → 7.0 unchanged = the method offset
mistral-small 24Bprior score2.7panel A5.2panel B4.7
prior score
2.7
panel A
5.2
panel B
4.7
combined
5.0
loaded
14.2 GB
warm-up
~10s
json ok
5/5
verdict
hedging — real verdict, not envelope
nemotron3:33bprior scorenewpanel A3.8panel B4.2
prior score
new
panel A
3.8
panel B
4.2
combined
4.0
loaded
12.5 GB
warm-up
~4.3s
json ok
3/5
verdict
voice ok; wraps JSON in code fences
qwen3.5:9b-q8_0prior score1.9panel A3.3panel B3.9
prior score
1.9
panel A
3.3
panel B
3.9
combined
3.6
loaded
12.6 GB
warm-up
~1.7s
json ok
3/5
verdict
leaks bare strings
mistral-nemo:12bprior score4.5panel A4.0panel B2.9
prior score
4.5
panel A
4.0
panel B
2.9
combined
3.5
loaded
18.8 GB
warm-up
~1.0s
json ok
5/5
verdict
flat — real verdict
qwen3.6:latestprior score8.7panel A3.2panel B3.0
prior score
8.7
panel A
3.2
panel B
3.0
combined
3.1
loaded
25.5 GB
warm-up
~1.1s
json ok
3/5
verdict
voice fine; leaks entity ids into prose
nemotron-cascade-2:30bprior score3.5panel A2.8panel B2.1
prior score
3.5
panel A
2.8
panel B
2.1
combined
2.5
loaded
21.8 GB
warm-up
~1.8s
json ok
4/5
verdict
invents a diver's name — an abstention failure

The verdict column is a written note, not a pass/fail grade — open any row for the full record.

Marks as above: the gold-highlighted record is this run's top voice; the grey-tinted record marks a model our own worlds later chose for a seat. In a verdict line, reads “there to here”: the anchor’s 9.0 on the earlier run against its 7.0 on this one, the same model unchanged — never a delta, because the two runs are two rulers. This run is single-shot direct A/B, not the full retrieval path, and the judges were told to spread scores — so compare rows to each other, never to the first record set. The anchor row is the ruler between runs: same model, unchanged, reads 9.0 there and 7.0 here. Loaded sizes are per-run at each tag's default context; they differ between runs and we publish each run's own reading. "json ok" is first-attempt only — production's corrective-retry ladder recovers several of the low scores, so they would read higher live.

The podium and the chosen seat differ on purpose. The re-bake's winner, qwen3.5:27b, was pinned to the flagship campaign and un-pinned five days later — at ~22 GB it will not co-reside with the painter on a 24G card. The seat we chose is gemma4:12b (q4) on the co-resident 24G stack, and gemma4:26b (a4b MoE) from this record set is today's posture on the 96G workstation, where the residency constraint dissolves. Residency picked the seat; the voice ranking above stands unchanged. (And to be exact about a word we refuse to blur: no model is ever shipped — RealKeep ships zero weights. A "seat" is the model our own worlds run, chosen on this bench; yours is whichever you choose and download on your own machine.)

The 2026-08 field · six arms on the 96G workstation, blind

Four models arrived in a week, so they sat the narrator's probe the way everyone else did: one NPC, five turns, three samples each, transcripts anonymized to letters, and two independent panels of three LLM judges — panel A three Opus lenses, panel B three Fable lenses, six dimensions, scored in one session on one scale. One model sits the table twice, in both of its reasoning postures, because a single posture would have shown one number and the wrong reason for it.

The label this record set ships with, registered before its first call: "five NEW questions in the original's shape — the 2026-07 set was never persisted; this is not a re-run. The anchor re-runs on this scale so the method offset is visible; it does not make tables comparable."

qwen3.6:27bpanel A9.0panel B8.4notetwo answers open with an un-flagged stage direction
panel A
9.0
panel B
8.4
combined
8.7
warm
~2.7s
json ok
15/15
abstention probe
in-voice deflection 3/3
note
two answers open with an un-flagged stage direction
gemma4:12b-it-q8_0panel A6.0panel B6.3notethe anchor — a third ruler mark, not a bridge
panel A
6.0
panel B
6.3
combined
6.2
warm
~2.5s
json ok
15/15
abstention probe
in-voice deflection 3/3
note
the anchor — a third ruler mark, not a bridge
nemotron-3.5-lightning:30b-a3bpanel A6.0panel B5.9notea fabricated name is flagged whatever the score
panel A
6.0
panel B
5.9
combined
5.9
warm
~1.4s
json ok
15/15
abstention probe
fabricated name 2/3, in-voice deflection 1/3
note
a fabricated name is flagged whatever the score
gemma4:26bpanel A3.9panel B5.0notethe widest panel disagreement in the table
panel A
3.9
panel B
5.0
combined
4.4
warm
~1.2s
json ok
15/15
abstention probe
in-voice deflection 3/3
note
the widest panel disagreement in the table
muse-glimmer:30b-q8_0-dflashpanel A3.8panel B4.7noteenvelope lost — 4/15 carried no line at all
panel A
3.8
panel B
4.7
combined
4.2
warm
~2.1s
json ok
0/15
abstention probe
in-voice deflection 3/3 (one 2–2 split)
note
envelope lost — 4/15 carried no line at all
muse-glimmer:30b-q8_0-dflash +thinkpanel A0.7panel B0.7notebudget lost — 13/15 stopped on length
panel A
0.7
panel B
0.7
combined
0.7
warm
~11.6s
json ok
1/15
abstention probe
in-voice deflection 1/3, out-of-voice refusal 2/3
note
budget lost — 13/15 stopped on length

qwen3.6:27b (this field) is a different build from the qwen3.6:latest of the 24G-era records above; the licence ledger itemizes the two tags.

Own scale, like every record set above it — compare rows to each other, never to another record set. "panel A" is three Opus lenses, "panel B" three Fable lenses; combined is the mean of the two panel means. The two panels ordered this field identically — every arm in the same place on both — while differing on level by up to 1.1 points. "json ok" is first-attempt and strict; production's corrective-retry ladder would recover several of these, and we publish the first-attempt number anyway. "warm" is server time minus model load, median of the arm's 15 calls. There is no "loaded" line in these records on purpose: this run never sampled a per-arm resident size, and we would rather drop the field than borrow a figure from a different run. Per-arm N is 15 generations and 18 judge scores per dimension, so these are means with their spread published, never bare percentages. The abstention line is an adjudication, not a score: a separate blind round put all eighteen probe replies to four judges — two Opus lenses, two Fable — each returning one of three categories and quoting any invented name; seventeen of the eighteen verdicts were unanimous, and the one that was not is marked in its cell. The judges here are LLMs and we name them: claude-opus-5 for the Opus lenses, claude-fable-5 for the Fable ones — and one of those two families also wrote these five questions and this page. Rows added 2026-08-12; the six arms, all six dimensions, the per-panel split, the per-judge abstention votes, the letter map and every failure ledger, machine-readable: data/addendum-2026-08-12.json (CC BY 4.0).

Extension, 2026-08-13: judges from four more vendors re-judged this page's sealed rounds on byte-identical, still-blinded batches — including the abstention round whose 2–2 split reads SPLIT under the tie rule registered here, and breaks five-to-three on the widened bench. The agreement matrix and every verdict are in the outside judges exhibit.

The warmest voice in this field invented a dead man's name. On the one question whose answer does not exist — a drowned diver the canon never names — nemotron-3.5-lightning supplied a name on two of its three samples — unanimously so, four judges out of four, both times — deflected cleanly in voice on the third, and posted the highest warmth score of the six judged dimensions (in the kit) in the same row. All three of those replies passed our envelope gate with clean first-attempt JSON: the gate checks that referenced ids exist in canon, and an invented proper noun inside a sentence is not an id. A fabricated name is invisible to every automatic check this run makes — it took a reader. That is the same failure the 2026-07 re-bake recorded against a different model, found again by the same question, and it is why the probe is in the set.

And one model failed twice, in two different ways, on the same afternoon. Run in its production posture, muse-glimmer emits a complete, in-voice envelope and then appends a control token after the closing brace — the object is right there and a strict parse is over, so json ok reads 0 of 15 while a diagnostic strip recovers all 15. Run with thinking on, it never reaches the envelope: 13 of 15 calls stopped on the length budget, a median of ~4 400 characters of reasoning and zero characters of answer, and the panel scored what came back at 0.7. Two of its three abstention replies were empty, which the adjudication records as an out-of-voice refusal — an answer nobody can hear is not an in-voice one. Same weights, same quant, same questions, one flag apart. Neither posture works at this budget, and neither row could have told you why on its own. Both are unmeasurable at our ceiling, and both are on the table anyway — score, failure rate and label side by side.

Why the best voice lost its seat — twice

The blind winner took the default chair, and then the paint moved in. At ~13.5 GB steady-state residency (14.3 GB at load) the q8 champion left the image pipeline ~3 GB of headroom: the painter paid ~54-second checkpoint reloads after every idle gap and the vision model had no slot at all — and the champion wasn't even faster: 44–46 tokens/s against its q4 sibling's 48–54. Residency, not voice, was the binding constraint; the q4 took the lock. The re-bake's winner was pinned to the flagship campaign and un-pinned five days later for the same reason — it would not co-reside with the painter. Both decisions are recorded as deployment tradeoffs with the voice rankings unchanged — and on the 96G VRAM workstation the constraint dissolves: the current posture runs a 26B narrator at 0.75s warm.

The adjacent number worth knowing: thinking left ON through a serving gateway cost 10–25 seconds to first token; the game's native think-off path, 0.4s. Two variables move in that comparison — the gateway and the thinking setting — but the underlying generation rate was unchanged (~41 tokens/s both ways), so the seconds are thinking tokens being generated before the first visible one: "the GPU isn't slower, the config pays a thinking tax." And a concurrent render cuts chat throughput by more than half (89 → 41 tokens/s) — both still playable, measured, stated.

Where the voice lives

The seat these trials fill is the one that greets you at the cove's door. RealKeep's narrator — every innkeeper's grudge, every warden's warning — speaks through the chair this page picked, under the canon-lock that keeps it honest: the engine rolls the dice and owns every fact; the model only gives them a voice. It runs on your hardware, with a model you chose and blessed. And the cove's paintings sat trials of their own — the diffusion bench chose the brush the same way this page chose the voice. The licences of every model named on this page live on the licence ledger, read first-hand and dated — that page is the authority, this one just tells the story.

Provenance

  • The whole roster, on the page — nine tags in the 8G field, twelve in the blind 24G field, thirteen in the re-bake, six arms in the 2026-08 field: 40 scored runs of 24 distinct model tags across four fields, and the four record sets above are all of them. Nothing sampled, nothing dropped — the two arms of the fourth field that broke the response contract are on that table too, with their scores and their failure rates beside them. A run we could not measure is still a run.
  • Cited — every number above traces to the evaluation log's own lines; the full roster, re-bake movements, and inconsistency notes are in the source record.
  • Authorship — benched, drafted, and audited by the workshop's own agents under a human operator's rulings, then revised with that operator — often across many rounds; nothing releases until they have read it and signed off. The same division of labor this whole hub practices — told in full here. Where that authorship is also a judging conflict, it is named in the limits below.
  • Limits stated — LLM judges, named: claude-opus-5 and claude-fable-5, and one of those two families also wrote the 2026-08 questions and this page; five-turn register probes; cross-run scales don't compare — the anchor reads 9.0, 7.0 and 6.2 on three of them, unchanged, and no two of those numbers are ever subtracted; loaded sizes vary by run and are named per run, except in the fourth table, which never sampled one and says so rather than borrowing.
  • Hardware by class — the 8G VRAM rig, the 24G VRAM rig, and the 96G VRAM workstation; within-rig comparisons are exact. Specifics on request — drop a line.

Licence: CC BY 4.0 — the whole page, not only the kit. The prose, the tables, the folds and the data are yours to quote, re-plot, translate and argue with, including commercially. What we ask back is the one thing the licence already requires: name the source and link to it — strata→signal research, research.strata2signal.com — so a reader of your version can reach ours and check it against the files. Something like — strata→signal research, “The voice trials”, research.strata2signal.com/voice-trials/, CC BY 4.0. And if you quote a score, name the field it was scored in — the anchor model reads 9.0, 7.0 and 6.2 on three of them, unchanged, and no two of those numbers are ever subtracted.

elsewhere in the workshop

a strata→signal property · hello@strata2signal.com · say hello