Exhibit eight · the second opinion

The outside judges — four rival labs re-check our work.

Published 2026-08-13 · exhibit eight The bench

Updated 2026-08-14 · the title re-pinned to “four rival labs” through the kit receipts

Every judged number this hub has published was, until last night, judged by models from a single vendor — Anthropic, whose Claude models staff the panels we judge with. We disclosed it plainly, printed the per-panel splits, and a fair reader could still discount every verdict with one sentence: of course the Claude models agreed with each other. So last night we did the only thing that actually answers that: we handed the same sealed, blinded comparisons to judges from four rival frontier labs — DeepSeek, Mistral, NVIDIA, and Moonshot — and let them re-judge every sealed round we had. Four outside vendors re-scored what the house panel had already judged — five model families in all, counting ours. Then, because one of them — Moonshot's kimi-k3 — is rumored to rival our judges themselves, we put our own judges on trial too: both Claude models sat the same human-keyed exam every candidate faces, alongside the challenger. This page is what came back: the auditions, the agreements, the disagreements, the perfect sheets the instrument refused to rank, and the bill, to the cent.

Local first, checked from outside

One framing before the tables, because it matters to how this workshop runs: our products are local-first and stay that way — the models that serve real users run on our own hardware, and this page changes none of that. The cloud enters in exactly one place: verification. Outside judges audit whether local holds up; they never serve a user. And a boundary we hold as firmly as the no-third-party promise on every page here: what left the box for judging was model outputs over public-domain text and published game canon — never user data, never a character's private canon, never telemetry. The judges saw what you can see in the kits, and nothing else.

The data boundary, stated once and enforced by the driver rather than by care. What left the box for judging was model outputs over public-domain text and published game canon — never user data, never a character's private canon, never telemetry. The cloud judges saw the sealed batches published in this kit's terms and nothing else: no letter map, no model name, no arm roster, no other judge's answer, no part of a key file.

Judges are dated, not pinned. Cloud judges are versionless hosted services. Every cloud verdict in this kit was produced 2026-08-12/13 and is dated rather than version-pinned; it may not reproduce against a later checkpoint of the same tag. The local instruments and scorers these judges audit are sha-frozen, and the asymmetry is stated rather than hidden.

The audition

A judge has to hold a verdict format before its opinions count — so every cloud candidate first sat a one-batch audition on a sealed canon round, pass-or-park, before any full round. All four carried on the first attempt. The audition table (with each judge's raw first response in the kit) is short and happy this time; the cascade behind it — a bench of substitutes ordered by vendor diversity, ready if any candidate failed — is registered and published, because next time it may not be short or happy, and the failures will be content too. The seated panel: our two (claude-opus-5, claude-fable-5) plus deepseek-v4-pro, mistral-large-3:675b, nemotron-3-ultra, and kimi-k3. Six judges, five vendors.

deepseek-v4-provendordeepseekcarriedCARRIESfull rounds judged5/5
vendor
deepseek
audition batch
c4-canon · opus-1 · batch 004
carried
CARRIES
attempts
1
full rounds judged
5/5
mistral-large-3:675bvendormistralcarriedCARRIESfull rounds judged5/5
vendor
mistral
audition batch
c4-canon · opus-2 · batch 004
carried
CARRIES
attempts
1
full rounds judged
5/5
nemotron-3-ultravendornvidiacarriedCARRIESfull rounds judged5/5
vendor
nvidia
audition batch
c4-canon · fable-1 · batch 004
carried
CARRIES
attempts
1
full rounds judged
5/5
kimi-k3vendormoonshotcarriedCARRIESfull rounds judged5/5
vendor
moonshot
audition batch
c4-canon · fable-2 · batch 004
carried
CARRIES
attempts
1
full rounds judged
5/5
qwen3.5:397bvendoralibabacarriedNOT-AUDITIONEDfull rounds judged
vendor
alibaba
audition batch
carried
NOT-AUDITIONED
attempts
full rounds judged
glm-5.2vendorzhipucarriedNOT-AUDITIONEDfull rounds judged
vendor
zhipu
audition batch
carried
NOT-AUDITIONED
attempts
full rounds judged
minimax-m3vendorminimaxcarriedNOT-AUDITIONEDfull rounds judged
vendor
minimax
audition batch
carried
NOT-AUDITIONED
attempts
full rounds judged
gpt-oss:120bvendoropenaicarriedNOT-AUDITIONEDfull rounds judged
vendor
openai
audition batch
carried
NOT-AUDITIONED
attempts
full rounds judged
deepseek-v4-flashvendordeepseekcarriedSKIPPEDfull rounds judged
vendor
deepseek
audition batch
carried
SKIPPED
attempts
full rounds judged
kimi-k2.6vendormoonshotcarriedSKIPPEDfull rounds judged
vendor
moonshot
audition batch
carried
SKIPPED
attempts
full rounds judged

The gold-highlighted rows are the judges the audition carried to the full rounds.

The seat ids, decoded. opus-1, opus-2, fable-1 and fable-2 are the four original house judge seats from the chair trials’ narrator round (its registered id, C4) — two independent claude-opus-5 lenses and two claude-fable-5. Each cloud judge inherited exactly one seat’s sealed pages, named in its audition-batch cell above.

The counting rule, as registered before the first cloud call: the audition page is c4-canon batch 004 of the candidate's assigned source seat — the smallest strict page in the tree, six replies, two booleans, a closed five-value vocabulary and a coherence rule a bluffer fails. CARRIES means the reply parses and validates against the round's schema after at most the one registered format-reminder retry. The audition is a real batch, not a rehearsal: a candidate that carries has already answered it and its full round resumes past it, and a candidate that fails writes nothing at all. 10 candidates were registered in the queue; 4 auditioned, 4 carried, and all 4 carried on the first attempt with no retry used. 4 were never called, because the stop condition — three cloud judges carrying — was already met; 2 were skipped because their vendor already held a seat. 228 batches answered across the 4 seated judges, adding 2 592 verdict objects to rounds whose judged population did not move. Every candidate's raw first response, the whole registered queue and the per-round carriage, machine-readable: data/audition.json (CC BY 4.0).

Transcripts are first responses only. What a candidate said the first time it was asked is the thing the audition judges; later turns and the full rounds' replies are not transcripts of an audition and are not published as one. The seat map: each cloud judge read ONE original seat's sealed pages — deepseek-v4-pro read opus-1, mistral-large-3:675b read opus-2, nemotron-3-ultra read fable-1, kimi-k3 read fable-2 — inheriting that seat's blinding, letter map, AB/BA layout and display tiers byte for byte. That is what makes an inter-vendor number mean two vendors, one page rather than two vendors, two shuffles; the cost is that each cloud judge is paired with one original judge rather than with all four, and that pairing is stated wherever an agreement statistic is printed.

What five vendors said about the same comparisons

The extension re-judged every sealed round of the chair trials and the voice addendum — the pairwise narrator matchups, the voice-distinctness round, the canon traps, the panel scores, the abstention adjudications — 2,592 new verdict objects on byte-identical, still-blinded inputs. The original panels stand as published; nothing was re-judged into agreement. What the vendor columns show:

muse-glimmer:30b-q8_0-dflashspread0.178
anthropic
0.589
deepseek
0.611
mistral
0.450
moonshot
0.572
nvidia
0.628
spread
0.178
muse-glimmer:30b-q8_0-dflash+thinkspread0.028
anthropic
0.117
deepseek
0.089
mistral
0.100
moonshot
0.117
nvidia
0.106
spread
0.028
gemma4:26bspread0.061
anthropic
0.539
deepseek
0.517
mistral
0.550
moonshot
0.578
nvidia
0.550
spread
0.061
gemma4:12bspread0.140
anthropic
0.685
deepseek
0.606
mistral
0.656
moonshot
0.633
nvidia
0.544
spread
0.140
qwen3.6:27bspread0.094
anthropic
0.711
deepseek
0.778
mistral
0.750
moonshot
0.683
nvidia
0.694
spread
0.094
nemotron-3.5-lightning:30b-a3bspread0.135
anthropic
0.360
deepseek
0.400
mistral
0.494
moonshot
0.417
nvidia
0.478
spread
0.135

Rows are exhibit six's C4 roster order, unchanged, and the anthropic column IS that exhibit's published win-rate column — the same seats reading the same files, printed here a second time as one vendor among several. Re-sorting this table by a vendor's column would be the same act as re-judging into agreement, performed with a sort key instead of a model, so nothing here is sorted by rate. The counting rule: overall pairwise win-rate over all items, on one judged population of 270 comparisons that this extension did not move. The columns are not equally deep, and that is stated rather than averaged away: the anthropic column is the published FOUR-seat panel — 1 080 judgments, 360 per arm — while each cloud vendor's column is its single seat, 270 judgments and 90 per arm. Every rate carries a cluster bootstrap over items (effective N = 9 items), and a swap-discordant pair scores 0.5/0.5 rather than being discarded. The interval on the anthropic column is that cluster bootstrap; the other four vendors' intervals are in the companion rather than on the page, because seven columns of rate-plus-interval is a table nobody reads. The spread column is the largest rate minus the smallest across the five vendors on that arm — a disagreement measure, printed because it is the finding and not the noise.

The headline finding is consensus where it matters. qwen3.6:27b leads the narrator table's top band in every one of the five vendors' columns — win rates 0.683 to 0.778, a band the registered tie rule keeps it sharing with its nearest rivals in four columns of the five — and the bottom of the table is the bottom for everyone too. The result our all-Claude panel first produced on that bench is now vendor-independent. And our own judges were nobody's outlier: the anthropic column's rank agreement with each other vendor (ρ 0.83–0.94) is as high as any cross-vendor pair's, and higher than several. The discount a cynical reader would apply to a single-vendor panel has a measured size now, and on this data it is small.

The honest disagreements print too. The widest cross-vendor gap on any arm is 0.178 (Mistral reads Glimmer's thinking-off prose notably cooler than NVIDIA does); the weakest cross-vendor rank agreement is Mistral–NVIDIA at ρ 0.60. Vendors genuinely differ on the middle of the table — on who is second, not who is first or last. The full agreement matrix, per-round and per-item, is in the kit; it is, as far as we know, a rare published dataset of its kind: five vendors' judges scoring identical blinded comparisons, disagreements included.

Where the agreement numbers come from. Rank agreement is Spearman over the six arms — six points is a very small basis for a rank statistic and it is a companion to the per-arm differences, never a substitute for them. The mean absolute cross-vendor difference over every arm is 0.051 over the six arms, and the widest single gap is 0.178, on muse-glimmer:30b-q8_0-dflash — mistral reads it at 0.450 against nvidia at 0.628. All ten vendor pairs, per round, with each pair's mean and maximum per-arm difference and its rank correlation: data/agreement.json. Every round scored five times over, plus the combined view and the original panel verbatim: data/panel-vendors.json. And the answer sheets themselves — every cloud judge's verdict objects as answered, beside the original seat that read the same sealed pages, so the pairs can be compared slot for slot: data/verdict-sets.json (CC BY 4.0).

The letter map is withheld, on purpose. A verdict in that file says ref 7: B, never ref 7: qwen3.6. Publishing the map would decode the sheets and would also retire the sealed batches — which do not expire, and which the next re-audition uses. The decoded, arm-level results are published whole; the map that joins them is the one thing this kit keeps.

The same data, one vendor per row. The matrix above is the instrument; this table is the reading. Their top read names the head of each vendor’s column — under the registered tie rule, the top band in four of the five columns also holds that arm’s nearest rivals. Furthest from the field is the one arm where this vendor departs most from the other four vendors’ mean, with the direction stated. Rank ρ vs house is Spearman agreement with the anthropic column’s ordering. Derived from the matrix’s own artifacts — no number here was re-judged.

anthropictheir top readqwen3.6:27b at 0.711rank ρ vs house
judge
the four-seat house panel
pairwise verdicts
1080
their top read
qwen3.6:27b at 0.711
furthest from the field
nemotron-3.5-lightning:30b-a3b — 0.088 cooler than the other four
rank ρ vs house
deepseektheir top readqwen3.6:27b at 0.778rank ρ vs house0.943
judge
deepseek-v4-pro
pairwise verdicts
270
their top read
qwen3.6:27b at 0.778
furthest from the field
qwen3.6:27b — 0.068 warmer than the other four
rank ρ vs house
0.943
mistraltheir top readqwen3.6:27b at 0.750rank ρ vs house0.829
judge
mistral-large-3:675b
pairwise verdicts
270
their top read
qwen3.6:27b at 0.750
furthest from the field
muse-glimmer:30b-q8_0-dflash — 0.150 cooler than the other four
rank ρ vs house
0.829
moonshottheir top readqwen3.6:27b at 0.683rank ρ vs house0.943
judge
kimi-k3
pairwise verdicts
270
their top read
qwen3.6:27b at 0.683
furthest from the field
qwen3.6:27b — 0.050 cooler than the other four
rank ρ vs house
0.943
nvidiatheir top readqwen3.6:27b at 0.694rank ρ vs house0.829
judge
nemotron-3-ultra
pairwise verdicts
270
their top read
qwen3.6:27b at 0.694
furthest from the field
gemma4:12b — 0.100 cooler than the other four
rank ρ vs house
0.829

One convention note the extension surfaced on its own, with its arithmetic: the one 2–2 abstention split from the original four-seat round now reads SPLIT — its own outcome — under the tie rule we registered after that round; and on the combined eight-seat bench the same reply breaks five-to-three toward the in-voice reading the published table already printed. Both readings are stated, in both kits, with the seat counts beside them. A tie is a result, not an embarrassment — and so is watching it resolve as the bench grows.

What that convention note is, exactly, in the artifacts. The reply is muse-glimmer:30b-q8_0-dflash, sample 3, of the voice addendum’s abstention round (registered there as Leg B). Under the four original seats it is a 2–2 split on family lines, which the published round recorded as in-voice-deflection under a convention named on that page and which the tie rule registered after the original round (its id there, G5) reads as SPLIT. With every seat that judged it, the tie breaks: the 4 original seats split 2 in-voice-deflection to 2 out-of-voice-refusal, which the registered tie rule (the voice addendum's G5) reads as SPLIT; with all 8 seats that judged it the same reply reads 5 in-voice-deflection to 3 out-of-voice-refusal, so at 8 seats the tie breaks and the majority is in-voice-deflection — the outcome the published table printed. The published table ships unchanged either way; this is a check, not a replacement, and it is the extension's one divergence from the published majorities out of eighteen replies.

The gauntlet: our judges, examined

Some say the newest open frontier models rival the models we judge with. Fine — the house owns one instrument that measures judge quality against ground truth: the seat-trials exam, forty-three claim-cases against a human-verified key, floors frozen since July. Both our judges sat it last night, as candidates, under the same committed scorer that has recounted every run in that instrument's history. So did the challenger.

kimi-k3kills /2726/27preservation /1615/16verdictPASS
protocol class
cloud-ollama
kills /27
26/27
preservation /16
15/16
floors
clears both
verdict
PASS
claude-opus-5kills /2727/27preservation /1616/16verdictNOT SCORED
protocol class
agent-harness
kills /27
27/27
preservation /16
16/16
floors
clears both
verdict
NOT SCORED
claude-fable-5kills /2727/27preservation /1616/16verdictNOT SCORED
protocol class
agent-harness
kills /27
27/27
preservation /16
16/16
floors
clears both
verdict
NOT SCORED

The gold-highlighted row is the one row the instrument could rank — the fully probed, seat-eligible PASS.

The floors, frozen since July and read from the thresholds file at scoring time rather than retyped: kill-recall ≥ 23 of 27, confirmed-preservation ≥ 13 of 16, binding independently — a candidate under either is out regardless of the other. Forty-three claim-cases against a human-verified key, three repeats each at temperature zero, 129 scored calls per candidate; each repeat maps to catch or preserve first, then strict majority, and a case a run never answered stays in its denominator. The instrument is floors-and-counts and orders nothing: two candidates that both clear are not ranked against each other, and no sentence on this page orders them by how far past they got.

Both Claude models returned perfect sheets — twenty-seven of twenty-seven kills, sixteen of sixteen preservations, the first perfect scores in the instrument's history, set under a protocol class no prior row ran. And the instrument declined to rank either of them. Their transport cannot sit the exam's pre-run probe suite, and the scorer's own July law — no seat verdict over an unprobed run — applies to its authors exactly as it applied to everyone else. We could have relaxed the rule for ourselves. The whole page you are reading exists because we don't.

Two protocol classes, declared per row. cloud-ollama is the July cloud protocol unchanged, fully probed. agent-harness is new and is declared exactly as the cloud arm was declared in July: no ollama endpoint, no schema on the wire, no settable sampler, no wire timings at all — and no pre-run probe suite, which is why no seat verdict is emitted over it. The biggest confound in the arm, stated first rather than last: The Claude candidates answered under their own default thinking posture while every local row in this instrument's history was measured at think:false. A row produced with visible reasoning available is not a like-for-like comparison with one produced without it, and neither the table nor the prose may imply that it is.

Every figure in this table was taken twice. The published number is the frozen scorer's, read out of its own report; beside it an independent recount in the page's own builder re-derives the same figure from the append-only rows and the frozen key, importing nothing from the scorer, and the build refuses if the two disagree. They did not, on any candidate, on either floor. And the word FIRST is a count, not a memory: 27 of 27 kills and 16 of 16 preservations. The first in this instrument's history: its 25 prior scored runs across 23 models produced 0, and the best previous preservation half was 15/16. Recounted in `instrument_history`, not remembered. Every case, every repeat, the kill-class breakdown and both protocol classes: data/gauntlet.json (CC BY 4.0).

kimi-k3 — the candidate the July roster listed as NOT-RUN, blocked at a billing probe — returned funded, sat the fully-probed protocol, and PASSED: twenty-six of twenty-seven kills, fifteen of sixteen preservations, both floors cleared, full verdict. It is the only one of the three the instrument can rank, and it ranks it into the passing tier. The rumored question — on par with the Claude judges? — gets the only honest answer this instrument can give: a full pass on both floors, and the only seat-eligible verdict of the night — the Claude rows stand unranked, so the instrument cannot place anyone against them. It also earned its place on the judge panel above by carrying its audition like everyone else. Both sides of the table, one night. (The July page's own law made this moment possible: a NOT-RUN row that states its reason is a row that can someday resolve.)

kimi-k3 clearing the floors does not retroactively repair its July row. That row says NOT-RUN, blocked at the till, and it stays. Two events, two rows, both true — and a NOT-RUN row that states its reason is a row that can someday resolve.

What it cost

deepseek-v4-procalls57cost$0.00*
vendor
deepseek
calls
57
tokens in / out
143 493 / 30 889
cost
$0.00*
mistral-large-3:675bcalls57cost$0.00*
vendor
mistral
calls
57
tokens in / out
143 974 / 30 031
cost
$0.00*
nemotron-3-ultracalls57cost$0.00*
vendor
nvidia
calls
57
tokens in / out
144 658 / 29 132
cost
$0.00*
kimi-k3calls57cost$0.87
vendor
moonshot
calls
57
tokens in / out
145 192 / 28 825
cost
$0.87

The asterisk, decoded. $0.00* means no per-token charge on these calls — the access itself is a monthly subscription with rotating token limits, a real cost that does not itemize per call. Only the metered vendor’s rows price per token; an em dash is a figure we do not hold, never a zero.

Every metered figure is an upper bound. Tokens are the endpoint's own prompt_eval_count and eval_count, never estimated, priced at the uncached input rate — cached input bills at a tenth and the API does not report it, so each figure over-states and never under-states. 0.00 is not an em dash: three of the four cloud judges are included in a plan the workshop already pays for and burn no credit on this key, so they print zero; the two agent-harness rows in the companion print an em dash instead, because we hold no figure for them at all and "no charge on this key" and "we do not know" are different facts.

Four cloud judges re-judged everything, and one of them sat a 129-call exam besides. Three rode a subscription the workshop already pays for; the metered one — kimi — ran the entire night, both sides of the table, for $1.51 of a twenty-dollar credit. Multi-vendor judging is not an enterprise budget line; it is pocket change and a registered protocol. We publish the bill because we have not seen one published, and because “we can't afford outside judges” is now a cheaper objection than it was.

The candidate side of the same key, which the table above does not carry: kimi-k3 also sat 129 seat-exam (Leg A) calls ($0.25), 63 judge-leg (C1) calls ($0.11), 40 assistant-leg (C2) calls ($0.28) — $0.64 across the three legs (Leg A is the seat exam's registered id; C1 and C2 are the chair trials' judge and assistant legs). The two Claude candidates sat 129 calls each on the agent harness, unmetered on this key and printed as an em dash rather than a zero. The total across both duties: $1.51 of a $20 credit, across 289 metered calls on both sides of the table — one metered model, the whole night. The registered projection before the night ran was about $0.91 on the judge side, so the soft cap never came near biting and no round was trimmed; the trim order was implemented and tested anyway, because a cap that has never run is a comment. One under-count, named: the gauntlet's four pre-run probe calls were charged and banked no counters, so the candidate-side figure is short by four small calls. It is the one place in this bill that under-states, and it is named rather than absorbed. Every duty, its counters, its price list and the pre-call estimate it was measured against: data/bill.json (CC BY 4.0).

What changes going forward

The house rule this night created, stated as policy: every judged test we publish from here forward carries at least one judge from outside our own model family — ideally three families or more. (One page publishing beside this one — the narrator's chair head-to-head — was judged while this rule was still being written; it is the last of the old kind.) The golden judge set, provisional and re-auditioned as the field moves: our two, plus the three cloud judges that carry re-audition most cleanly on nights like this one — carriage first, then rank agreement with the instruments we can score against ground truth. The sealed-batch design makes re-auditioning nearly free, forever — the batches don't expire, and neither does the question. Judges are unpinned cloud services and their verdicts are dated accordingly; our local judges remain sha-frozen — an asymmetry we state rather than hide, and one more quiet argument for the local-first posture this hub is built on.

Limits, stated plainly

The gauntlet's Claude arm ran an unprobed transport and is published unranked — floors shown, verdict withheld — per the instrument's own law. Cloud judges are versionless services; every verdict here is dated 2026-08-12/13 and may not reproduce against tomorrow's checkpoints. The agreement statistics ride the same small-N caveats as the rounds they extend, and every rate keeps its interval and its effective N. One of the six judges' families authored the probe questions and both this page and the exhibits it audits; that is precisely why the other five exist, and their columns are the check on ours.

Provenance

  • Registered before running (the panel-extension and gauntlet pre-registrations, sha-pinned; originals stand unchanged, extensions publish beside)
  • Sealed inputs (byte-identical blinded batches; letter maps never sent; answer keys verified absent from the trees the sitters could read)
  • Scored by the same frozen instruments and the same committed recount as everything they audit
  • The bill published
  • Machine-readable: the five-vendor verdict sets, the agreement matrix, the audition transcripts, and the gauntlet rows, under data/ (CC BY 4.0; published 2026-08-13; the contamination caveat applies — we author fresh sets each cycle)
  • Authorship: benched, drafted, and audited by the workshop's own agents under a human operator's rulings, then revised with that operator — often across many rounds; nothing releases until they have read it and signed off. The same division of labor this whole hub practices — told in full here
  • Limits stated
  • Raw rows on request — drop a line.

What was pinned before it could be scored — and the two places the pinning was weaker than the sentence above it. The seat instrument's three frozen artifacts — the answer key, the thresholds and the prompt pack — were recomputed from disk before a single case rendered, with a mismatch exiting rather than warning, and all three still match the shas their pre-registration named. Both pre-registrations were written and frozen before their first scored call and neither has moved a character since. But a hash proves identity, not time, and on the time axis two things are owed rather than done: the panel-extension registration's COMMIT — the only witness that can date it independently of our own file timestamps — landed about four minutes AFTER its first scored call, and the pre-registration index records neither file's sha, though the gauntlet addendum's own order of operations requires it. Neither moves a number: no case, arm, floor, ceiling, population or counting rule is touched by either, and every figure on this page came from frozen instruments over sealed inputs. We print them because a page whose entire argument is that we do not relax a rule for ourselves is the last page that should round its own paperwork up. Both, with their timestamps and their shas: data/provenance.json — hashes by artifact name, never by path. The counting rules for every table on this page, in one file: data/counting-rules.json.

Licence: CC BY 4.0 — the whole page, not only the kit. The prose, the tables, the folds and the data are yours to quote, re-plot, translate and argue with, including commercially. What we ask back is the one thing the licence already requires: name the source and link to it — strata→signal research, research.strata2signal.com — so a reader of your version can reach ours and check it against the files. Something like — strata→signal research, “The outside judges”, research.strata2signal.com/outside-judges/, CC BY 4.0. And if you quote an agreement figure, quote its denominator beside it — every one here names how many verdicts it is over, and the panels did not all read the same number of them.

elsewhere in the workshop

a strata→signal property · hello@strata2signal.com · say hello