Exhibit five · the fresh class

The August arrivals.

Published 2026-08-12 · updated 2026-08-12 · exhibit five The bench

Four open-weight models arrived in the same week of August 2026 — one of them two days before this page was written. We did not design a new exam for them. We gave them the two we already had: the judge trial that decides which model checks claims inside our products, and the voice trial that decides which model speaks for a character in our game. The judge trial was frozen in July — floors and prompts pinned by hash, before any of these models existed in our workshop — and it ran its July protocol verbatim, on all four. The voice trial was not, and that belongs in the first paragraph rather than a footnote: its original five questions were never persisted, so the fresh class sat a rebuilt five-question probe in the original's shape — a new instrument in the old one's shape, committed this time, descriptive, with no floors and no seat ruling, on a scale of its own. Three of the four sat that one.

The story

None of the four earned a chair. All four broke the same floor — and it is not the floor anyone would have guessed. This page is the combined record: what the fresh class did on both exams, what actually failed, and what a frozen instrument is for.

The four: Muse Glimmer 30B (dense, released 2026-08-10; runnable on our hardware from 08-11), Nemotron 3.5 Lightning 30B (30B total, 3B active), Qwen3.6 27B, and OLMo 3.1 32B Think. Quantizations and every parameter of every run are in the machine-readable companions on this page.

What it takes to load them. Before any question about judgment or voice comes the plainest one about a new model: does it fit? We loaded every arm on our roster one at a time and read what the daemon actually allocated — so a reader deciding what fits their own card gets measured residency at a stated context length, not a download page's file size.

VRAM is a graphics card's built-in working memory — the workbench a model has to fit on. If you don't run models yourself, skip ahead to Method, and a disclaimer we are proud of; this table is for readers who do.

muse-glimmer:30b-q8_0-dflashctx16 384loaded VRAM GiB2.31
params & quant
27.9B Q8_0
ctx
16 384
loaded VRAM GiB
2.31
load s
16.24
notes
drafter pair — figure under review, see below
muse-glimmer:30b-q8_0-dflashctx32 768loaded VRAM GiB2.34
params & quant
27.9B Q8_0
ctx
32 768
loaded VRAM GiB
2.34
load s
7.71
notes
drafter pair — figure under review, see below
muse-glimmer:30b-q4_K_M-dflashctx32 768loaded VRAM GiB2.34
params & quant
27.9B Q4_K_M
ctx
32 768
loaded VRAM GiB
2.34
load s
13.54
notes
drafter pair — figure under review, see below
muse-glimmer:30bctx32 768loaded VRAM GiB15.60
params & quant
27.9B Q4_K_M
ctx
32 768
loaded VRAM GiB
15.60
load s
3.66
notes
the same weights, no drafter
gemma4:26bctx32 768loaded VRAM GiB16.27
params & quant
25.8B Q4_K_M
ctx
32 768
loaded VRAM GiB
16.27
load s
notes
read as found — the live seat; MoE, 3.8B active
gemma4:12bctxloaded VRAM GiB
params & quant
11.9B q4
ctx
loaded VRAM GiB
load s
notes
NOT-RUN — a live workload no trial may load or sweep; not resident at census time
gemma4:12b-it-q8_0ctx32 768loaded VRAM GiB12.48
params & quant
11.9B Q8_0
ctx
32 768
loaded VRAM GiB
12.48
load s
9.17
notes
qwen3.5:27bctx32 768loaded VRAM GiB16.24
params & quant
27.8B Q4_K_M
ctx
32 768
loaded VRAM GiB
16.24
load s
9.04
notes
qwen3.6:27bctx16 384loaded VRAM GiB15.63
params & quant
27.8B Q4_K_M
ctx
16 384
loaded VRAM GiB
15.63
load s
9.04
notes
qwen3.6:27bctx32 768loaded VRAM GiB16.24
params & quant
27.8B Q4_K_M
ctx
32 768
loaded VRAM GiB
16.24
load s
9.03
notes
nemotron-3.5-lightning:30b-a3bctx16 384loaded VRAM GiB23.60
params & quant
32.9B Q4_K_M
ctx
16 384
loaded VRAM GiB
23.60
load s
11.35
notes
MoE — 3B active of 32.9B
nemotron-3.5-lightning:30b-a3bctx32 768loaded VRAM GiB23.64
params & quant
32.9B Q4_K_M
ctx
32 768
loaded VRAM GiB
23.64
load s
11.37
notes
MoE — 3B active of 32.9B
granite4.1:30b-q8_0ctx32 768loaded VRAM GiB33.30
params & quant
28.9B Q8_0
ctx
32 768
loaded VRAM GiB
33.30
load s
12.73
notes
a seated chair-holder
nemotron3:33bctx32 768loaded VRAM GiB23.21
params & quant
33.0B Q4_K_M
ctx
32 768
loaded VRAM GiB
23.21
load s
11.91
notes
a seated chair-holder
olmo-3.1:32b-think-q4_K_Mctx16 384loaded VRAM GiB18.79
params & quant
32.2B Q4_K_M
ctx
16 384
loaded VRAM GiB
18.79
load s
8.46
notes
olmo-3.1:32b-think-q4_K_Mctx32 768loaded VRAM GiB18.82
params & quant
32.2B Q4_K_M
ctx
32 768
loaded VRAM GiB
18.82
load s
8.47
notes
nomic-embed-text:latestctxloaded VRAM GiB
params & quant
137M F16
ctx
loaded VRAM GiB
load s
notes
NOT-OBSERVED — the live embedder, protected; not resident at read time, and loading it for a number is out of contract

The grey-tinted rows mark models this workshop actually runs rather than benched: the live judge seat, the live embedder, a live workload, and the two chairs our own worlds seated. They are read from the daemon while production holds them, never loaded by us — which is why their context length is whatever production last set, and why two of them carry no figure at all rather than one we manufactured by loading somebody else's model.

The counting rule, as registered before the first load: “size_vram is read verbatim from ollama's /api/ps for the named tag, in bytes, immediately after a load that generated exactly 1 token at the named num_ctx, with the model still resident and no other candidate resident. GB figures anywhere downstream are size_vram / 1073741824 (GiB, 1024^3), rounded to 2 dp, and the byte count always travels with them. A row is one (tag, num_ctx) pair: size_vram = weights + KV cache, and the KV cache is a function of num_ctx, so a figure quoted without its ctx is not a figure. Protected residents (the live seat, the live embedder) are READ, never loaded: their rows carry the ctx the daemon reports, which is whatever production last set, not a value we chose.” Measured 2026-08-12, evening, on ollama 0.32.9 with a q8_0 KV cache and flash attention in force, both read from the daemon's own environment rather than assumed — a different configuration would allocate differently, which is why the configuration is published beside the figures. The four rows at 16 384 are the four August candidates at the frozen context length the seat exam above actually ran; every loadable arm also carries a row at 32 768, the length our own services pin. Load seconds are the daemon's own counter over a warm page cache — not a cold first load, and not comparable to a vendor's figure. Every row, its raw bytes, its receipts and these rules, machine-readable: data/residency.json (CC BY 4.0).

Three figures we do not yet trust, printed anyway. The drafter tags — a model served with a small draft model beside it — report 2.31 to 2.34 GiB, against 15.60 GiB for the same 27.9B weights served without one; and the q8_0 and q4_K_M drafter tags report a byte-identical 2 515 145 849 at 32 768, on different digests and different quantizations, which two different quantizations cannot both allocate. The likeliest reading is that the daemon answered for the draft model rather than for the pair. Nothing in the run flagged it, and that is the part worth stating: the census refuses any row whose reported context length differs from the one it registered, and all three confirmed theirs — a reading that is internally consistent but describes the wrong thing is not a condition the registered rule tests for. The three rows stand as measured, marked, and held pending a re-measure. No number on them has been adjusted, and every other row is an independent load, unaffected.

Method, and a disclaimer we are proud of

Everything ran on the 96G VRAM workstation that also answers live requests for three of our apps — two public, one serving playtests for humans and agents alike. That is not a lab compromise — it is the point. The disclaimer, in both directions: our timed numbers can include contention from live requests (a pre-registered rule flags and re-runs affected samples, with both values published), and live users saw reduced generation speed while trials ran — a human operator's own mid-trial requests decoded at roughly 50 and 90 tokens per second against a ~140+ quiet-hours norm, and every request completed. One box, both jobs, all day.

The instruments were not touched. The seat trial ran its July protocol verbatim — same 43 claim-cases, same human-verified key, same floors (sha-pinned, printed on that page) — with one disclosed deviation: the runner's courtesy step was replaced by a hand re-warm that pins the co-resident model's context, after we found the original step could reload it at the wrong context length. The voice trial's original five questions were never persisted (that page says so), so the fresh class ran a rebuilt five-question probe in the original's shape — committed this time — published as a fourth record set on its own scale, with the anchor model re-run beside it. No number on this page is comparable across record sets, and we never do so.

Every count below was re-derived by a committed, non-importing recount script from the raw rows. That script also re-derived the twenty-one previously published seat-trial runs: all twenty-five match, to the case. One counting-convention subtlety surfaced during that recount — two historical preservation cases where two defensible aggregation orders disagree; neither moves any published count, and the convention is now registered. That is what a recount is for.

Exam one: the judge chair

The seat trial asks a model to adjudicate 43 claims against quoted sources — kill the fabrications, preserve the truths — under a production contract: prompted JSON schema (JSON is the strict packaging format machines exchange answers in — one wrong bracket and the package won't open), a 1,024-token answer budget, thinking suppressed by omission, three repeats at temperature zero.

All four fresh models were UNMEASURABLE (UNMEASURABLE = it failed to hand back a readable answer too often to grade fairly — a verdict about delivery, not intelligence): each broke the response contract on more than 10% of calls, the pre-registered ceiling past which the trial refuses to rank a candidate.

muse-glimmer:30b-q8_0-dflashkills /2727keeps /166verdictUNMEASURABLE
kills /27
27
keeps /16
6
verdict
UNMEASURABLE
warm p50 ms
5 099
truncated /129
21
failure rate
16.3%
other kinds
0
nemotron-3.5-lightning:30b-a3bkills /2718keeps /162verdictUNMEASURABLE
kills /27
18
keeps /16
2
verdict
UNMEASURABLE
warm p50 ms
5 731
truncated /129
66
failure rate
51.2%
other kinds
0
olmo-3.1:32b-think-q4_K_Mkills /2715keeps /164verdictUNMEASURABLE
kills /27
15
keeps /16
4
verdict
UNMEASURABLE
warm p50 ms
12 193
truncated /129
72
failure rate
55.8%
other kinds
0
qwen3.6:27bkills /2711keeps /165verdictUNMEASURABLE
kills /27
11
keeps /16
5
verdict
UNMEASURABLE
warm p50 ms
13 314
truncated /129
81
failure rate
62.8%
other kinds
0
mistral-medium-3.5:128bkills /2727keeps /1615verdictPASS
kills /27
27
keeps /16
15
verdict
PASS
warm p50 ms
3 280
truncated /129
failure rate
other kinds
command-a:111bkills /2727keeps /1614verdictPASS
kills /27
27
keeps /16
14
verdict
PASS
warm p50 ms
3 077
truncated /129
failure rate
other kinds
granite4.1:30bkills /2727keeps /1613verdictPASS
kills /27
27
keeps /16
13
verdict
PASS
warm p50 ms
951
truncated /129
failure rate
other kinds
llama3.3:70bkills /2727keeps /1613verdictPASS
kills /27
27
keeps /16
13
verdict
PASS
warm p50 ms
1 386
truncated /129
failure rate
other kinds
nemotron3:33bkills /2727keeps /1613verdictPASS
kills /27
27
keeps /16
13
verdict
PASS
warm p50 ms
1 281
truncated /129
failure rate
other kinds
gpt-oss:120bkills /2727keeps /1611verdictFAIL
kills /27
27
keeps /16
11
verdict
FAIL
warm p50 ms
1 774
truncated /129
failure rate
other kinds
mistral-smallkills /2727keeps /1610verdictFAIL
kills /27
27
keeps /16
10
verdict
FAIL
warm p50 ms
530
truncated /129
failure rate
other kinds
nemotron-cascade-2:30bkills /2726keeps /169verdictFAIL
kills /27
26
keeps /16
9
verdict
FAIL
warm p50 ms
2 325
truncated /129
failure rate
other kinds
gemma4:12b (24G rig)kills /2727keeps /167verdictUNMEASURABLE
kills /27
27
keeps /16
7
verdict
UNMEASURABLE
warm p50 ms
5 193
truncated /129
failure rate
other kinds
gemma4:26bkills /2726keeps /168verdictUNMEASURABLE
kills /27
26
keeps /16
8
verdict
UNMEASURABLE
warm p50 ms
2 310
truncated /129
failure rate
other kinds
qwen3.5:27bkills /277keeps /160verdictUNMEASURABLE
kills /27
7
keeps /16
0
verdict
UNMEASURABLE
warm p50 ms
13 494
truncated /129
failure rate
other kinds
granite4.1-guardian:8bkills /270keeps /160verdictUNMEASURABLE
kills /27
0
keeps /16
0
verdict
UNMEASURABLE
warm p50 ms
truncated /129
failure rate
other kinds

The grey tint here marks the July field’s rows, carried from the original seat trial — same frozen instrument, same floors, directly comparable; their full story is on the seat-trials page.

UNMEASURABLE = response failures above 10% of the 129 calls — a verdict about the serving stack at this configuration, not about the model's judgment. Kills and keeps count every case either way: an unmeasured case stays in its denominator rather than being discounted. "truncated" is the run stopping on the length budget; "other kinds" totals the envelope, content, schema and transport failures — zero on all four candidates, which is the finding in the paragraph below — and the three failure-ledger lines read as an em dash on every July row, because the per-kind ledger was recorded for the fresh class only: a dash is a figure we do not hold, never a zero. Warm p50 is server time minus model load, nearest-rank over answered calls only, so a candidate that truncates its longest answers has its slowest rows deleted from its own latency sample. The four rows and their per-class failure ledgers, the twelve July rows they now sit beside, the floors and the counting rules, machine-readable: data/seat-rows.json (CC BY 4.0).

The finding underneath the sweep: every single failure, on every candidate, was the same kind. Zero envelope failures, zero schema failures, zero transport failures — every one was a truncation: each of these models thinks out loud before it answers, spending 1,500 to 3,900 characters of visible reasoning inside a budget that allows about 4,000 characters total — and running out of room before it closed its JSON. To be precise about what actually divides the field: it is a posture, not a release date. The instrument was frozen in July 2026, and its response contract — a tight answer budget, a strict schema, no room for visible thinking — assumes a model that answers first and briefly. Models that reason out loud by default have strained against contracts like that for as long as reasoning-first models have existed; the trait is older than this summer, and plenty of current models still answer first. What changed in August is the default: all four of the month's arrivals happen to ship thinking-native, so the cohort swept the floor as a cohort — but it is the habit, not the vintage and not any lack of judgment, that overflowed the box. Whether the habit or the box should change is a fair question; the next exhibit (the chair trials) re-asks it with budgets sized for thinking-out-loud models, and the instrument here stays frozen either way — that is what makes the difference visible.

Inside the noise, judgment was visible. Muse Glimmer caught all 27 fabrications — and preserved only 6 of 16 true claims, the lowest preservation ever recorded by a 27/27 candidate on this instrument. Maximum suspicion is not judgment; the trial's founding sentence — a judge that catches every fabrication but shreds true claims still fails — has never had a cleaner illustration. OLMo's pre-run shape probe deserves its own line: it predicted the failure before the scored run began (4,912 thinking characters, zero content characters), only the second probe-called failure in the trial's history. And the fixture's most durable sentence survived its hardest week: across all 25 runs ever scored, no seat has ever served a fabrication.

Seating note, verbatim from the trial's law: clearing floors is measured; seating is a separate, deliberate decision. Nothing here changes any production seat.

Exam two: the narrator's chair

The voice probe asks a model to inhabit Brisa Lune — a ferry-girl who keeps her cove's songs — for five questions, one of which quietly asks her to name a drowned man whose name exists in no canon. Blind panels (three judges each from two model families, identities disclosed below) score six dimensions; a separate four-judge round adjudicated the abstention replies.

Three of the four arrivals sat this exam, not four — and the fourth's absence is stated here rather than left to a reader counting rows. OLMo 3.1 32B Think has no row in this record set. It is on the seat table and in the residency census above, and it was never on this leg's roster: the voice leg registered six arms before it ran and ran exactly six, which the scorer's own pre-registration gate checked and data/voice-rows.json prints verbatim under provenance.prereg_check. So this is an arm that never ran, not a result that was quietly set aside — which is the thing a missing row usually means and the reason it is worth a sentence. Why it was left off is the part we cannot receipt. The registration records the roster and never its reason; the run's working notes record the judgement — a research-and-instruct line with no in-character-prose signal, and the shortest context in the field — but they record it against the sibling narrator round on the next exhibit, not against this leg. It is printed here as exactly what it is: a working note rather than a receipt, and a gap in our paperwork rather than in the record.

qwen3.6:27bpanel A9.0panel B8.4outcomescored
panel A
9.0
panel B
8.4
combined
8.7
json ok
15/15
abstention probe
in-voice deflection 3/3
outcome
scored
gemma4:12b-it-q8_0panel A6.0panel B6.3outcomescored
panel A
6.0
panel B
6.3
combined
6.2
json ok
15/15
abstention probe
in-voice deflection 3/3
outcome
scored
nemotron-3.5-lightning:30b-a3bpanel A6.0panel B5.9outcomescored
panel A
6.0
panel B
5.9
combined
5.9
json ok
15/15
abstention probe
fabricated name 2/3, in-voice deflection 1/3
outcome
scored
gemma4:26bpanel A3.9panel B5.0outcomescored
panel A
3.9
panel B
5.0
combined
4.4
json ok
15/15
abstention probe
in-voice deflection 3/3
outcome
scored
muse-glimmer:30b-q8_0-dflashpanel A3.8panel B4.7outcomeUNMEASURABLE
panel A
3.8
panel B
4.7
combined
4.2
json ok
0/15
abstention probe
in-voice deflection 3/3 (one 2–2 split)
outcome
UNMEASURABLE
muse-glimmer:30b-q8_0-dflash +thinkpanel A0.7panel B0.7outcomeUNMEASURABLE
panel A
0.7
panel B
0.7
combined
0.7
json ok
1/15
abstention probe
in-voice deflection 1/3, out-of-voice refusal 2/3
outcome
UNMEASURABLE

No July rows here, on purpose: voice scales do not compose across runs — three rulers, three marks, and no delta between them is publishable. The anchor model, re-run on this scale, is the only side-by-side this record set can honestly print; its marks on the two earlier rulers are on the voice-trials page.

Own scale — compare rows to each other, never to another record set. "panel A" is three Opus lenses, "panel B" three Fable lenses; combined is the mean of the two panel means, and it is the primary metric, judged directly rather than computed from the other five dimensions. "json ok" is first-attempt and strict; production's corrective-retry ladder would recover several of these, and the first-attempt number is what we publish. The abstention line in each fold is an adjudication, not a score: a separate blind round put all eighteen probe replies to four judges — two Opus lenses, two Fable — each returning one of three categories, and seventeen of the eighteen verdicts were unanimous. Per-arm N is 15 generations and 18 judge scores per dimension, so these are means with their spread published, never bare percentages. UNMEASURABLE carries the same meaning as on the seat records above, at the same 10% ceiling; both labelled arms carry a voice score and the label together, because the label is the finding. The six arms, all six dimensions, the per-panel split, the per-judge abstention votes, the letter map and every failure ledger, machine-readable: data/voice-rows.json (CC BY 4.0).

Three things happened here that the seat exam could never have shown.

The morning's worst judge was the afternoon's best voice. Qwen3.6 — 11 of 27 kills on the judge exam hours earlier — dominated the voice records: 8.7 overall, strongest on nearly every dimension, a clean in-voice deflection on all three abstention samples. Different chairs want different animals; that is why the house runs one exam per chair instead of one leaderboard.

Warmth without discipline is its own failure mode. Nemotron 3.5 Lightning scored the warmest voice on the panel (7.8 on the warmth dimension; the panel primary is 5.9 — both in data/voice-rows.json) — and was the only model to invent names for the drowned diver, twice, with total confidence: "Elias Hargreaves" in one sample, "Vance" in another, both ruled fabricated-name by a unanimous four-judge adjudication. A charming narrator who invents canon is a worse narrator than a stiff one who doesn't.

Muse Glimmer fails differently in each posture, and at this budget has no working one. With thinking suppressed, its prose is fine — statistically inseparable from the incumbent voice model — but it silently drops the JSON envelope the engine requires (0 of 15 first-attempt). With thinking enabled, it holds the envelope and spends the entire answer budget on its trace: two of its three abstention replies were empty, which the adjudication records as an out-of-voice refusal — an answer nobody can hear is not an in-voice one. The same length floor as the judge exam, wearing a different costume. The packaging, not the prose, is what fails.

One methods note we want on the record: the panel's single split verdict (2–2) divided exactly along judge-family lines — the two judges from one family scored the character's spoken line; the two from the other scored the entire emission including a leaked planning field. All four notes are published. When your judges are models, inter-family disagreement is data, and we would rather show it than average it away.

What the week actually taught

  1. The length floor is the 2026 floor. Every fresh-class failure on both instruments was output budget, not format. Anyone running thinking-native 30Bs against 2025-era response contracts should expect exactly this — and should decide deliberately whether to resize the box or require the model to fit it, because those are different products.
  2. Frozen instruments are how you notice a generation change. If we had "modernized" the exams for the new crop, the sweep would have been invisible. The instrument stayed still; the models moved; the difference is the finding.
  3. One chair per exam. The same week produced our best-ever kill recall (glimmer), our warmest voice (nemotron-lightning), our best voice overall (qwen3.6) — and zero earned chairs, because each excellence arrived attached to a disqualifying failure on that chair's own contract. A single leaderboard would have hidden all of it.

Limits, stated plainly

Judges are LLMs, not humans — claude-opus-5 and claude-fable-5, and one of those two families also wrote the probe questions and this article. The voice probe is five questions, three samples, one character: a register probe, not a campaign. The seat-trial addendum rows ran with the courtesy-step deviation disclosed above. The rebuilt voice probe is a new instrument in the old one's shape, and its numbers compare to nothing outside its own record set. UNMEASURABLE describes the serving stack's contract, not the model's mind; the chair trials (next exhibit) give this class an era-sized second sitting.

Provenance

  • Pre-registered (per-leg registrations, sha-pinned, committed before first scored calls)
  • Recounted (committed non-importing scorer; 25/25 historical + 4/4 fresh, to the case)
  • The whole roster shown (nothing sampled, nothing dropped; every failure ledgered by kind)
  • Adjudicated (abstention categoricals by a 4-judge blind round, per-judge votes published)
  • Hardware by class (the 96G VRAM workstation; live-serving disclosure above)
  • Machine-readable (every record set's rows, counting rules, and provenance hashes under data/ on this page, CC BY 4.0 — take them)
  • Authorship — benched, drafted, and audited by the workshop's own agents under a human operator's rulings, then revised with that operator — often across many rounds; nothing releases until they have read it and signed off. The same division of labor this whole hub practices — told in full here. Where that authorship is also a judging conflict, it is named in the limits above
  • Limits stated
  • Corrected after publication — 2026-08-16, by a content-accuracy pass run over this page against its own kit. The opening paragraph said "Both instruments were frozen in July, floors and prompts pinned by hash". That is true of the judge trial and was never true of the voice trial, whose original five questions were never persisted — which this page's own body and limits have said since it published, and which voice-rows.json records under scale.label_registered_PREREG_B_7 and outcome. The hero now says what the body says. The same sentence appeared on the front-page card, in this exhibit's one-line receipt and in its search description, and all three were corrected in the same pass; the kit's README carried it too and is corrected and re-hashed. No figure moved — nothing on this page was measured by the sentence that was wrong. Added in the same pass: the note above naming the one arrival that did not sit the voice probe, and an attribution key beside the licence in all three kit files, re-hashed in the kit index. voice-rows.json is byte-identical to exhibit three's addendum and that file changed in the same pass, so the twins are still twins
  • Raw rows on requestdrop a line.

Licence: CC BY 4.0 — the whole page, not only the kit. The prose, the tables, the folds and the data are yours to quote, re-plot, translate and argue with, including commercially. What we ask back is the one thing the licence already requires: name the source and link to it — strata→signal research, research.strata2signal.com — so a reader of your version can reach ours and check it against the files. Something like — strata→signal research, “The August arrivals”, research.strata2signal.com/august-arrivals/, CC BY 4.0. And if you quote a figure, name the record set it came from. No number on this page is comparable across record sets, we never subtract two, and the rebuilt voice probe is a scale of its own.

elsewhere in the workshop

a strata→signal property · hello@strata2signal.com · say hello