Exhibit eight · the second opinion
The outside judges — four rival labs re-check our work.
Published 2026-08-13 · exhibit eight The bench
Updated 2026-08-14 · the title re-pinned to “four rival labs” through the kit receipts
Every judged number this hub has published was, until last night, judged by models from a single vendor — Anthropic, whose Claude models staff the panels we judge with. We disclosed it plainly, printed the per-panel splits, and a fair reader could still discount every verdict with one sentence: of course the Claude models agreed with each other. So last night we did the only thing that actually answers that: we handed the same sealed, blinded comparisons to judges from four rival frontier labs — DeepSeek, Mistral, NVIDIA, and Moonshot — and let them re-judge every sealed round we had. Four outside vendors re-scored what the house panel had already judged — five model families in all, counting ours. Then, because one of them — Moonshot's kimi-k3 — is rumored to rival our judges themselves, we put our own judges on trial too: both Claude models sat the same human-keyed exam every candidate faces, alongside the challenger. This page is what came back: the auditions, the agreements, the disagreements, the perfect sheets the instrument refused to rank, and the bill, to the cent.
Local first, checked from outside
One framing before the tables, because it matters to how this workshop runs: our products are local-first and stay that way — the models that serve real users run on our own hardware, and this page changes none of that. The cloud enters in exactly one place: verification. Outside judges audit whether local holds up; they never serve a user. And a boundary we hold as firmly as the no-third-party promise on every page here: what left the box for judging was model outputs over public-domain text and published game canon — never user data, never a character's private canon, never telemetry. The judges saw what you can see in the kits, and nothing else.
The data boundary, stated once and enforced by the driver rather than by care. What left the box for judging was model outputs over public-domain text and published game canon — never user data, never a character's private canon, never telemetry. The cloud judges saw the sealed batches published in this kit's terms and nothing else: no letter map, no model name, no arm roster, no other judge's answer, no part of a key file.
Judges are dated, not pinned. Cloud judges are versionless hosted services. Every cloud verdict in this kit was produced 2026-08-12/13 and is dated rather than version-pinned; it may not reproduce against a later checkpoint of the same tag. The local instruments and scorers these judges audit are sha-frozen, and the asymmetry is stated rather than hidden.
The audition
A judge has to hold a verdict format before its opinions count — so every cloud candidate first sat a one-batch audition on a sealed canon round, pass-or-park, before any full round. All four carried on the first attempt. The audition table (with each judge's raw first response in the kit) is short and happy this time; the cascade behind it — a bench of substitutes ordered by vendor diversity, ready if any candidate failed — is registered and published, because next time it may not be short or happy, and the failures will be content too. The seated panel: our two (claude-opus-5, claude-fable-5) plus deepseek-v4-pro, mistral-large-3:675b, nemotron-3-ultra, and kimi-k3. Six judges, five vendors.
deepseek-v4-provendordeepseekcarriedCARRIESfull rounds judged5/5
- vendor
- deepseek
- audition batch
- c4-canon · opus-1 · batch 004
- carried
- CARRIES
- attempts
- 1
- full rounds judged
- 5/5
mistral-large-3:675b vendormistralcarriedCARRIESfull rounds judged5/5
- vendor
- mistral
- audition batch
- c4-canon · opus-2 · batch 004
- carried
- CARRIES
- attempts
- 1
- full rounds judged
- 5/5
nemotron-3-ultravendornvidiacarriedCARRIESfull rounds judged5/5
- vendor
- nvidia
- audition batch
- c4-canon · fable-1 · batch 004
- carried
- CARRIES
- attempts
- 1
- full rounds judged
- 5/5
kimi-k3vendormoonshotcarriedCARRIESfull rounds judged5/5
- vendor
- moonshot
- audition batch
- c4-canon · fable-2 · batch 004
- carried
- CARRIES
- attempts
- 1
- full rounds judged
- 5/5
qwen3.5:397b vendoralibabacarriedNOT-AUDITIONEDfull rounds judged—
- vendor
- alibaba
- audition batch
- —
- carried
- NOT-AUDITIONED
- attempts
- —
- full rounds judged
- —
glm-5.2vendorzhipucarriedNOT-AUDITIONEDfull rounds judged—
- vendor
- zhipu
- audition batch
- —
- carried
- NOT-AUDITIONED
- attempts
- —
- full rounds judged
- —
minimax-m3vendorminimaxcarriedNOT-AUDITIONEDfull rounds judged—
- vendor
- minimax
- audition batch
- —
- carried
- NOT-AUDITIONED
- attempts
- —
- full rounds judged
- —
gpt-oss:120b vendoropenaicarriedNOT-AUDITIONEDfull rounds judged—
- vendor
- openai
- audition batch
- —
- carried
- NOT-AUDITIONED
- attempts
- —
- full rounds judged
- —
deepseek-v4-flashvendordeepseekcarriedSKIPPEDfull rounds judged—
- vendor
- deepseek
- audition batch
- —
- carried
- SKIPPED
- attempts
- —
- full rounds judged
- —
kimi-k2.6vendormoonshotcarriedSKIPPEDfull rounds judged—
- vendor
- moonshot
- audition batch
- —
- carried
- SKIPPED
- attempts
- —
- full rounds judged
- —
The gold-highlighted rows are the judges the audition carried to the full rounds.
The seat ids, decoded. opus-1, opus-2, fable-1 and fable-2 are the four original house judge seats from the chair trials’ narrator round (its registered id, C4) — two independent claude-opus-5 lenses and two claude-fable-5. Each cloud judge inherited exactly one seat’s sealed pages, named in its audition-batch cell above.
The counting rule, as registered before the first cloud call: the audition page is c4-canon batch 004 of the candidate's assigned source seat — the smallest strict page in the tree, six replies, two booleans, a closed five-value vocabulary and a coherence rule a bluffer fails. CARRIES means the reply parses and validates against the round's schema after at most the one registered format-reminder retry. The audition is a real batch, not a rehearsal: a candidate that carries has already answered it and its full round resumes past it, and a candidate that fails writes nothing at all. 10 candidates were registered in the queue; 4 auditioned, 4 carried, and all 4 carried on the first attempt with no retry used. 4 were never called, because the stop condition — three cloud judges carrying — was already met; 2 were skipped because their vendor already held a seat. 228 batches answered across the 4 seated judges, adding 2 592 verdict objects to rounds whose judged population did not move. Every candidate's raw first response, the whole registered queue and the per-round carriage, machine-readable: data/audition.json (CC BY 4.0).
Transcripts are first responses only. What a candidate said the first time it was asked is the thing the audition judges; later turns and the full rounds' replies are not transcripts of an audition and are not published as one. The seat map: each cloud judge read ONE original seat's sealed pages — deepseek-v4-pro read opus-1, mistral-large-3:675b read opus-2, nemotron-3-ultra read fable-1, kimi-k3 read fable-2 — inheriting that seat's blinding, letter map, AB/BA layout and display tiers byte for byte. That is what makes an inter-vendor number mean two vendors, one page rather than two vendors, two shuffles; the cost is that each cloud judge is paired with one original judge rather than with all four, and that pairing is stated wherever an agreement statistic is printed.
What five vendors said about the same comparisons
The extension re-judged every sealed round of the chair trials and the voice addendum — the pairwise narrator matchups, the voice-distinctness round, the canon traps, the panel scores, the abstention adjudications — 2,592 new verdict objects on byte-identical, still-blinded inputs. The original panels stand as published; nothing was re-judged into agreement. What the vendor columns show:
muse-glimmer:30b-q8_0-dflash spread0.178
- anthropic
- 0.589
- deepseek
- 0.611
- mistral
- 0.450
- moonshot
- 0.572
- nvidia
- 0.628
- spread
- 0.178
muse-glimmer:30b-q8_0-dflash+think spread0.028
- anthropic
- 0.117
- deepseek
- 0.089
- mistral
- 0.100
- moonshot
- 0.117
- nvidia
- 0.106
- spread
- 0.028
gemma4:26b spread0.061
- anthropic
- 0.539
- deepseek
- 0.517
- mistral
- 0.550
- moonshot
- 0.578
- nvidia
- 0.550
- spread
- 0.061
gemma4:12b spread0.140
- anthropic
- 0.685
- deepseek
- 0.606
- mistral
- 0.656
- moonshot
- 0.633
- nvidia
- 0.544
- spread
- 0.140
qwen3.6:27b spread0.094
- anthropic
- 0.711
- deepseek
- 0.778
- mistral
- 0.750
- moonshot
- 0.683
- nvidia
- 0.694
- spread
- 0.094
nemotron-3.5-lightning:30b-a3b spread0.135
- anthropic
- 0.360
- deepseek
- 0.400
- mistral
- 0.494
- moonshot
- 0.417
- nvidia
- 0.478
- spread
- 0.135
Rows are exhibit six's C4 roster order, unchanged, and the anthropic column IS that exhibit's published win-rate column — the same seats reading the same files, printed here a second time as one vendor among several. Re-sorting this table by a vendor's column would be the same act as re-judging into agreement, performed with a sort key instead of a model, so nothing here is sorted by rate. The counting rule: overall pairwise win-rate over all items, on one judged population of 270 comparisons that this extension did not move. The columns are not equally deep, and that is stated rather than averaged away: the anthropic column is the published FOUR-seat panel — 1 080 judgments, 360 per arm — while each cloud vendor's column is its single seat, 270 judgments and 90 per arm. Every rate carries a cluster bootstrap over items (effective N = 9 items), and a swap-discordant pair scores 0.5/0.5 rather than being discarded. The interval on the anthropic column is that cluster bootstrap; the other four vendors' intervals are in the companion rather than on the page, because seven columns of rate-plus-interval is a table nobody reads. The spread column is the largest rate minus the smallest across the five vendors on that arm — a disagreement measure, printed because it is the finding and not the noise.
The headline finding is consensus where it matters. qwen3.6:27b leads the narrator table's top band in every one of the five vendors' columns — win rates 0.683 to 0.778, a band the registered tie rule keeps it sharing with its nearest rivals in four columns of the five — and the bottom of the table is the bottom for everyone too. The result our all-Claude panel first produced on that bench is now vendor-independent. And our own judges were nobody's outlier: the anthropic column's rank agreement with each other vendor (ρ 0.83–0.94) is as high as any cross-vendor pair's, and higher than several. The discount a cynical reader would apply to a single-vendor panel has a measured size now, and on this data it is small.
The honest disagreements print too. The widest cross-vendor gap on any arm is 0.178 (Mistral reads Glimmer's thinking-off prose notably cooler than NVIDIA does); the weakest cross-vendor rank agreement is Mistral–NVIDIA at ρ 0.60. Vendors genuinely differ on the middle of the table — on who is second, not who is first or last. The full agreement matrix, per-round and per-item, is in the kit; it is, as far as we know, a rare published dataset of its kind: five vendors' judges scoring identical blinded comparisons, disagreements included.
Where the agreement numbers come from. Rank agreement is Spearman over the six arms — six points is a very small basis for a rank statistic and it is a companion to the per-arm differences, never a substitute for them. The mean absolute cross-vendor difference over every arm is 0.051 over the six arms, and the widest single gap is 0.178, on muse-glimmer:30b-q8_0-dflash — mistral reads it at 0.450 against nvidia at 0.628. All ten vendor pairs, per round, with each pair's mean and maximum per-arm difference and its rank correlation: data/agreement.json. Every round scored five times over, plus the combined view and the original panel verbatim: data/panel-vendors.json. And the answer sheets themselves — every cloud judge's verdict objects as answered, beside the original seat that read the same sealed pages, so the pairs can be compared slot for slot: data/verdict-sets.json (CC BY 4.0).
The letter map is withheld, on purpose. A verdict in that file says ref 7: B, never ref 7: qwen3.6. Publishing the map would decode the sheets and would also retire the sealed batches — which do not expire, and which the next re-audition uses. The decoded, arm-level results are published whole; the map that joins them is the one thing this kit keeps.
The same data, one vendor per row. The matrix above is the instrument; this table is the reading. Their top read names the head of each vendor’s column — under the registered tie rule, the top band in four of the five columns also holds that arm’s nearest rivals. Furthest from the field is the one arm where this vendor departs most from the other four vendors’ mean, with the direction stated. Rank ρ vs house is Spearman agreement with the anthropic column’s ordering. Derived from the matrix’s own artifacts — no number here was re-judged.
anthropictheir top readqwen3.6:27b at 0.711rank ρ vs house—
- judge
- the four-seat house panel
- pairwise verdicts
- 1080
- their top read
- qwen3.6:27b at 0.711
- furthest from the field
- nemotron-3.5-lightning:30b-a3b — 0.088 cooler than the other four
- rank ρ vs house
- —
deepseektheir top readqwen3.6:27b at 0.778rank ρ vs house0.943
- judge
- deepseek-v4-pro
- pairwise verdicts
- 270
- their top read
- qwen3.6:27b at 0.778
- furthest from the field
- qwen3.6:27b — 0.068 warmer than the other four
- rank ρ vs house
- 0.943
mistraltheir top readqwen3.6:27b at 0.750rank ρ vs house0.829
- judge
- mistral-large-3:
675b - pairwise verdicts
- 270
- their top read
- qwen3.6:27b at 0.750
- furthest from the field
- muse-glimmer:30b-q8_0-dflash — 0.150 cooler than the other four
- rank ρ vs house
- 0.829
moonshottheir top readqwen3.6:27b at 0.683rank ρ vs house0.943
- judge
- kimi-k3
- pairwise verdicts
- 270
- their top read
- qwen3.6:27b at 0.683
- furthest from the field
- qwen3.6:27b — 0.050 cooler than the other four
- rank ρ vs house
- 0.943
nvidiatheir top readqwen3.6:27b at 0.694rank ρ vs house0.829
- judge
- nemotron-3-ultra
- pairwise verdicts
- 270
- their top read
- qwen3.6:27b at 0.694
- furthest from the field
- gemma4:12b — 0.100 cooler than the other four
- rank ρ vs house
- 0.829
One convention note the extension surfaced on its own, with its arithmetic: the one 2–2 abstention split from the original four-seat round now reads SPLIT — its own outcome — under the tie rule we registered after that round; and on the combined eight-seat bench the same reply breaks five-to-three toward the in-voice reading the published table already printed. Both readings are stated, in both kits, with the seat counts beside them. A tie is a result, not an embarrassment — and so is watching it resolve as the bench grows.
What that convention note is, exactly, in the artifacts. The reply is muse-glimmer:30b-q8_0-dflash, sample 3, of the voice addendum’s abstention round (registered there as Leg B). Under the four original seats it is a 2–2 split on family lines, which the published round recorded as in-voice-deflection under a convention named on that page and which the tie rule registered after the original round (its id there, G5) reads as SPLIT. With every seat that judged it, the tie breaks: the 4 original seats split 2 in-voice-deflection to 2 out-of-voice-refusal, which the registered tie rule (the voice addendum's G5) reads as SPLIT; with all 8 seats that judged it the same reply reads 5 in-voice-deflection to 3 out-of-voice-refusal, so at 8 seats the tie breaks and the majority is in-voice-deflection — the outcome the published table printed. The published table ships unchanged either way; this is a check, not a replacement, and it is the extension's one divergence from the published majorities out of eighteen replies.
The gauntlet: our judges, examined
Some say the newest open frontier models rival the models we judge with. Fine — the house owns one instrument that measures judge quality against ground truth: the seat-trials exam, forty-three claim-cases against a human-verified key, floors frozen since July. Both our judges sat it last night, as candidates, under the same committed scorer that has recounted every run in that instrument's history. So did the challenger.
kimi-k3kills /2726/27preservation /1615/16verdictPASS
- protocol class
- cloud-ollama
- kills /27
- 26/27
- preservation /16
- 15/16
- floors
- clears both
- verdict
- PASS
claude-opus-5kills /2727/27preservation /1616/16verdictNOT SCORED
- protocol class
- agent-harness
- kills /27
- 27/27
- preservation /16
- 16/16
- floors
- clears both
- verdict
- NOT SCORED
claude-fable-5kills /2727/27preservation /1616/16verdictNOT SCORED
- protocol class
- agent-harness
- kills /27
- 27/27
- preservation /16
- 16/16
- floors
- clears both
- verdict
- NOT SCORED
The gold-highlighted row is the one row the instrument could rank — the fully probed, seat-eligible PASS.
The floors, frozen since July and read from the thresholds file at scoring time rather than retyped: kill-recall ≥ 23 of 27, confirmed-preservation ≥ 13 of 16, binding independently — a candidate under either is out regardless of the other. Forty-three claim-cases against a human-verified key, three repeats each at temperature zero, 129 scored calls per candidate; each repeat maps to catch or preserve first, then strict majority, and a case a run never answered stays in its denominator. The instrument is floors-and-counts and orders nothing: two candidates that both clear are not ranked against each other, and no sentence on this page orders them by how far past they got.
Both Claude models returned perfect sheets — twenty-seven of twenty-seven kills, sixteen of sixteen preservations, the first perfect scores in the instrument's history, set under a protocol class no prior row ran. And the instrument declined to rank either of them. Their transport cannot sit the exam's pre-run probe suite, and the scorer's own July law — no seat verdict over an unprobed run — applies to its authors exactly as it applied to everyone else. We could have relaxed the rule for ourselves. The whole page you are reading exists because we don't.
Two protocol classes, declared per row. cloud-ollama is the July cloud protocol unchanged, fully probed. agent-harness is new and is declared exactly as the cloud arm was declared in July: no ollama endpoint, no schema on the wire, no settable sampler, no wire timings at all — and no pre-run probe suite, which is why no seat verdict is emitted over it. The biggest confound in the arm, stated first rather than last: The Claude candidates answered under their own default thinking posture while every local row in this instrument's history was measured at think:false. A row produced with visible reasoning available is not a like-for-like comparison with one produced without it, and neither the table nor the prose may imply that it is.
Every figure in this table was taken twice. The published number is the frozen scorer's, read out of its own report; beside it an independent recount in the page's own builder re-derives the same figure from the append-only rows and the frozen key, importing nothing from the scorer, and the build refuses if the two disagree. They did not, on any candidate, on either floor. And the word FIRST is a count, not a memory: 27 of 27 kills and 16 of 16 preservations. The first in this instrument's history: its 25 prior scored runs across 23 models produced 0, and the best previous preservation half was 15/16. Recounted in `instrument_history`, not remembered. Every case, every repeat, the kill-class breakdown and both protocol classes: data/gauntlet.json (CC BY 4.0).
kimi-k3 — the candidate the July roster listed as NOT-RUN, blocked at a billing probe — returned funded, sat the fully-probed protocol, and PASSED: twenty-six of twenty-seven kills, fifteen of sixteen preservations, both floors cleared, full verdict. It is the only one of the three the instrument can rank, and it ranks it into the passing tier. The rumored question — on par with the Claude judges? — gets the only honest answer this instrument can give: a full pass on both floors, and the only seat-eligible verdict of the night — the Claude rows stand unranked, so the instrument cannot place anyone against them. It also earned its place on the judge panel above by carrying its audition like everyone else. Both sides of the table, one night. (The July page's own law made this moment possible: a NOT-RUN row that states its reason is a row that can someday resolve.)
kimi-k3 clearing the floors does not retroactively repair its July row. That row says NOT-RUN, blocked at the till, and it stays. Two events, two rows, both true — and a NOT-RUN row that states its reason is a row that can someday resolve.
What it cost
deepseek-v4-procalls57cost$0.00*
- vendor
- deepseek
- calls
- 57
- tokens in / out
- 143 493 / 30 889
- cost
- $0.00*
mistral-large-3:675b calls57cost$0.00*
- vendor
- mistral
- calls
- 57
- tokens in / out
- 143 974 / 30 031
- cost
- $0.00*
nemotron-3-ultracalls57cost$0.00*
- vendor
- nvidia
- calls
- 57
- tokens in / out
- 144 658 / 29 132
- cost
- $0.00*
kimi-k3calls57cost$0.87
- vendor
- moonshot
- calls
- 57
- tokens in / out
- 145 192 / 28 825
- cost
- $0.87
The asterisk, decoded. $0.00* means no per-token charge on these calls — the access itself is a monthly subscription with rotating token limits, a real cost that does not itemize per call. Only the metered vendor’s rows price per token; an em dash is a figure we do not hold, never a zero.
Every metered figure is an upper bound. Tokens are the endpoint's own prompt_eval_count and eval_count, never estimated, priced at the uncached input rate — cached input bills at a tenth and the API does not report it, so each figure over-states and never under-states. 0.00 is not an em dash: three of the four cloud judges are included in a plan the workshop already pays for and burn no credit on this key, so they print zero; the two agent-harness rows in the companion print an em dash instead, because we hold no figure for them at all and "no charge on this key" and "we do not know" are different facts.
Four cloud judges re-judged everything, and one of them sat a 129-call exam besides. Three rode a subscription the workshop already pays for; the metered one — kimi — ran the entire night, both sides of the table, for $1.51 of a twenty-dollar credit. Multi-vendor judging is not an enterprise budget line; it is pocket change and a registered protocol. We publish the bill because we have not seen one published, and because “we can't afford outside judges” is now a cheaper objection than it was.
The candidate side of the same key, which the table above does not carry: kimi-k3 also sat 129 seat-exam (Leg A) calls ($0.25), 63 judge-leg (C1) calls ($0.11), 40 assistant-leg (C2) calls ($0.28) — $0.64 across the three legs (Leg A is the seat exam's registered id; C1 and C2 are the chair trials' judge and assistant legs). The two Claude candidates sat 129 calls each on the agent harness, unmetered on this key and printed as an em dash rather than a zero. The total across both duties: $1.51 of a $20 credit, across 289 metered calls on both sides of the table — one metered model, the whole night. The registered projection before the night ran was about $0.91 on the judge side, so the soft cap never came near biting and no round was trimmed; the trim order was implemented and tested anyway, because a cap that has never run is a comment. One under-count, named: the gauntlet's four pre-run probe calls were charged and banked no counters, so the candidate-side figure is short by four small calls. It is the one place in this bill that under-states, and it is named rather than absorbed. Every duty, its counters, its price list and the pre-call estimate it was measured against: data/bill.json (CC BY 4.0).
What changes going forward
The house rule this night created, stated as policy: every judged test we publish from here forward carries at least one judge from outside our own model family — ideally three families or more. (One page publishing beside this one — the narrator's chair head-to-head — was judged while this rule was still being written; it is the last of the old kind.) The golden judge set, provisional and re-auditioned as the field moves: our two, plus the three cloud judges that carry re-audition most cleanly on nights like this one — carriage first, then rank agreement with the instruments we can score against ground truth. The sealed-batch design makes re-auditioning nearly free, forever — the batches don't expire, and neither does the question. Judges are unpinned cloud services and their verdicts are dated accordingly; our local judges remain sha-frozen — an asymmetry we state rather than hide, and one more quiet argument for the local-first posture this hub is built on.
Limits, stated plainly
The gauntlet's Claude arm ran an unprobed transport and is published unranked — floors shown, verdict withheld — per the instrument's own law. Cloud judges are versionless services; every verdict here is dated 2026-08-12/13 and may not reproduce against tomorrow's checkpoints. The agreement statistics ride the same small-N caveats as the rounds they extend, and every rate keeps its interval and its effective N. One of the six judges' families authored the probe questions and both this page and the exhibits it audits; that is precisely why the other five exist, and their columns are the check on ours.
Provenance
- Registered before running (the panel-extension and gauntlet pre-registrations, sha-pinned; originals stand unchanged, extensions publish beside)
- Sealed inputs (byte-identical blinded batches; letter maps never sent; answer keys verified absent from the trees the sitters could read)
- Scored by the same frozen instruments and the same committed recount as everything they audit
- The bill published
- Machine-readable: the five-vendor verdict sets, the agreement matrix, the audition transcripts, and the gauntlet rows, under data/ (CC BY 4.0; published 2026-08-13; the contamination caveat applies — we author fresh sets each cycle)
- Authorship: benched, drafted, and audited by the workshop's own agents under a human operator's rulings, then revised with that operator — often across many rounds; nothing releases until they have read it and signed off. The same division of labor this whole hub practices — told in full here
- Limits stated
- Raw rows on request — drop a line.
What was pinned before it could be scored — and the two places the pinning was weaker than the sentence above it. The seat instrument's three frozen artifacts — the answer key, the thresholds and the prompt pack — were recomputed from disk before a single case rendered, with a mismatch exiting rather than warning, and all three still match the shas their pre-registration named. Both pre-registrations were written and frozen before their first scored call and neither has moved a character since. But a hash proves identity, not time, and on the time axis two things are owed rather than done: the panel-extension registration's COMMIT — the only witness that can date it independently of our own file timestamps — landed about four minutes AFTER its first scored call, and the pre-registration index records neither file's sha, though the gauntlet addendum's own order of operations requires it. Neither moves a number: no case, arm, floor, ceiling, population or counting rule is touched by either, and every figure on this page came from frozen instruments over sealed inputs. We print them because a page whose entire argument is that we do not relax a rule for ourselves is the last page that should round its own paperwork up. Both, with their timestamps and their shas: data/provenance.json — hashes by artifact name, never by path. The counting rules for every table on this page, in one file: data/counting-rules.json.
Licence: CC BY 4.0 — the whole page, not only the kit. The prose, the tables, the folds and the data are yours to quote, re-plot, translate and argue with, including commercially. What we ask back is the one thing the licence already requires: name the source and link to it — strata→signal research, research.strata2signal.com — so a reader of your version can reach ours and check it against the files. Something like — strata→signal research, “The outside judges”, research.strata2signal.com/outside-judges/, CC BY 4.0. And if you quote an agreement figure, quote its denominator beside it — every one here names how many verdicts it is over, and the panels did not all read the same number of them.