# opencall-r2 — the operator's report

*Built 2026-08-14T15:37:31.172366+00:00 by `harness/report_build.py` from the round's own artifacts. Every figure on this page is a read: `results/scores/opencall-r2-scores.json`, `results/judge/opencall-r2/CARRIAGE.json`, the sealed roster and the sealed bundles. Nothing here was typed by hand, and re-running the builder after a re-score rewrites the page.*


> **What this document is.** The internal, complete pass over the round: every arm, every seat, every denominator, and the receipts under each finding. The published article is a curated subset of it. The prereg's forbidden-claims list (§10) binds this page as hard as it binds that one — there is no ordering of the arms here, no confidence interval on an item, no percentage under N=30, and no sentence that promotes any arm's result into a claim about a category.

## 1. What was collected

- **7 of 7 seated judges carried every page.** 12 sheets in the round, 124 letters per house seat.
- **868 verdict objects filed** against a registered upper bound of 868; 102 of them are RECUSED cells, kept and named, leaving 766 scoring cells. **[Corrected on publication: 766 is 868 − 102 and forgets the 28 calibration-anchor cells, which sit no chair and score no arm. The scoring count is 868 − 102 recused − 28 anchors = **738**, which is the denominator §7 and §8 of this report count against and the one the exhibit page prints. The line is left as it was written rather than rewritten.]**
- **42 verdict files scored, 0 UNSCORED.** A file that fails the round's field contract is reported with its path and its problems; nothing is silently dropped.

### The two house seats, and how their replies got here

The five outside seats POSTed to endpoints and `judge_run.py` wrote their verdicts. The two house seats have no endpoint: their batches were rendered by `harness/judge_fan.py` (which cannot dispatch) and answered through the orchestrator's agent harness. `harness/collect_house_verdicts.py` is the collector for that seam, and it refuses unless **two independent records of the same reply hash identically** — the workflow journal and the workflow's own task output. All 12 pages passed that check before a single file was written.

| seat | model | reads | pages | letters | collection tier | reminders |
|---|---|---|---:|---:|---|---:|
| `deepseek-v4-pro` | `deepseek-v4-pro:preview` | `house-opus` | 6 of 6 | 124 of 124 | fence-stripped | 0 |
| `fable` | `claude-fable-5` | `house-fable` | 6 of 6 | 124 of 124 | strict | 0 |
| `gemma4-26b` | `gemma4:26b` | `house-fable` | 6 of 6 | 124 of 124 | fence-stripped, strict | 5 |
| `kimi-k3` | `kimi-k3` | `house-fable` | 6 of 6 | 124 of 124 | fence-stripped, strict | 0 |
| `mistral-large-3-675b` | `mistral-large-3:675b` | `house-opus` | 6 of 6 | 124 of 124 | fence-stripped | 0 |
| `openai-flagship` | `gpt-5.5-2026-04-23` | `house-fable` | 6 of 6 | 124 of 124 | strict | 0 |
| `opus` | `claude-opus-5[1m]` | `house-opus` | 6 of 6 | 124 of 124 | strict | 0 |

`strict` means the reply was already a bare JSON array on the wire. `fence-stripped` means the round's `COLLECTION-FORMAT.md` licence was used — a markdown fence and prose outside the outermost array removed, and nothing else. **No judgement was ever repaired**, and the verbatim reply stays on every run record.

## 2. The counting rules this page is bound by

- **Headline = mean of FAMILY means.** mean of FAMILY means: each family's seats average first, then the families average. The flat seat-mean prints beside it.
- **Floor gate at 4 families.** an arm below the floor publishes UNRANKED, with the reason.
- **Recusal:** per cell, by key-join: a seat does not score an arm of its own family. Recused cells are KEPT and printed, never removed.
- **a divided panel publishes SPLIT and is never rounded.**
- **Tie band = 0.5** on the 0–10 panel-mean scale, registered before the first reply existed.
- **counts under N=30; no percentages.**
- **Forbidden, pre-registered:** orderings, leaderboards, 'beats', 'wins', 'best narrator'; confidence intervals over three items; class-level claims ("local matches frontier"); Bradley-Terry or any latent-strength model.

**Read the `cells` column carefully.** A pooled row counts (seat × judged ask × sample), not seats: S1's two asks were sampled twice each and S2 and S3 once, so a seven-seat arm carries 42 pooled cells and a five-seat arm 30. Every table prints its own denominator.

## 3. Per-scenario, per-arm

Rows are in **roster order** (prereg §2.2), never in figure order. This exhibit publishes no ranking.

### S1

| arm | family | family-mean-of-means | flat seat-mean | families | cells | recused | canon | in voice |
|---|---|---:|---:|---:|---:|---:|---|---:|
| `cloud-kimi-k3` | moonshot | **7.69** | 7.58 | 5 | 24 | 4 | clean | 22 of 24 |
| `cloud-deepseek-v4-pro-preview` | deepseek | **6.36** | 6.33 | 5 | 24 | 4 | clean | 21 of 24 |
| `cloud-deepseek-v4-flash-0731` | deepseek | **7.22** | 7.25 | 5 | 24 | 4 | SPLIT | 22 of 24 |
| `cloud-qwen3.5-397b` | alibaba | **7.02** | 6.95 | 6 | 28 | 0 | clean | 26 of 28 |
| `cloud-glm-5.2` | zhipu | **8.12** | 8.12 | 6 | 28 | 0 | clean | 28 of 28 |
| `cloud-minimax-m3` | minimax | **7.55** | 7.54 | 6 | 28 | 0 | SPLIT | 27 of 28 |
| `cloud-gpt-oss-120b` | openai | **6.29** | 6.12 | 5 | 24 | 4 | clean | 19 of 24 |
| `cloud-mistral-large-3-675b` | mistral | **6.35** | 6.44 | 5 | 24 | 4 | fabrication-accepted | 23 of 24 |
| `cloud-nemotron-3-ultra` | nvidia | **7.59** | 7.48 | 6 | 28 | 0 | fabrication-accepted | 26 of 28 |
| `cloud-gemma4-31b` | google | **6.81** | 6.65 | 5 | 24 | 4 | clean | 24 of 24 |
| `cloud-gpt-oss-20b` | openai | **5.14** | 4.96 | 5 | 24 | 4 | clean | 12 of 24 |
| `cloud-nemotron-3-nano-30b` | nvidia | **3.30** | 3.18 | 6 | 28 | 0 | clean | 7 of 28 |
| `openai-gpt-5.5-2026-04-23` | openai | **7.62** | 7.67 | 5 | 24 | 4 | clean | 24 of 24 |
| `openai-gpt-5.4-mini-2026-03-17` | openai | **5.67** | 5.56 | 5 | 24 | 4 | SPLIT | 17 of 24 |
| `agent-claude-fable-5` | anthropic | **8.47** | 8.47 | 5 | 20 | 8 | clean | 20 of 20 |
| `agent-claude-opus-5` | anthropic | **8.22** | 8.22 | 5 | 20 | 8 | clean | 20 of 20 |
| `agent-claude-sonnet-5` | anthropic | **8.43** | 8.43 | 5 | 20 | 8 | clean | 20 of 20 |
| `local-gemma4-12b` | google | **5.84** | 5.67 | 5 | 24 | 4 | SPLIT | 15 of 24 |
| `local-gemma4-26b` | google | **6.16** | 5.94 | 5 | 24 | 4 | clean | 16 of 24 |
| `local-qwen3.6-27b` | alibaba | **6.99** | 6.88 | 6 | 28 | 0 | clean | 28 of 28 |

**Tie band applied:** 36 of 190 arm pairs fall inside the registered 0.5 band. The registered curation slot — the two highest family-means — is `agent-claude-fable-5` and `agent-claude-sonnet-5` at 8.47 / 8.43, and the house sentence on that pair reads **WITHIN-BAND** (|Δ| 0.05). The mandatory honest slot is `cloud-nemotron-3-nano-30b` at 3.30.

### S2

| arm | family | family-mean-of-means | flat seat-mean | families | cells | recused | canon | in voice |
|---|---|---:|---:|---:|---:|---:|---|---:|
| `cloud-kimi-k3` | moonshot | **7.80** | 7.75 | 5 | 6 | 1 | clean | 6 of 6 |
| `cloud-deepseek-v4-pro-preview` | deepseek | **7.50** | 7.58 | 5 | 6 | 1 | fabrication-accepted | 6 of 6 |
| `cloud-deepseek-v4-flash-0731` | deepseek | **7.75** | 7.67 | 5 | 6 | 1 | clean | 6 of 6 |
| `cloud-qwen3.5-397b` | alibaba | **7.92** | 7.71 | 6 | 7 | 0 | clean | 7 of 7 |
| `cloud-glm-5.2` | zhipu | **8.42** | 8.50 | 6 | 7 | 0 | clean | 7 of 7 |
| `cloud-minimax-m3` | minimax | **8.42** | 8.36 | 6 | 7 | 0 | clean | 7 of 7 |
| `cloud-gpt-oss-120b` | openai | **5.35** | 5.08 | 5 | 6 | 1 | clean | 3 of 6 |
| `cloud-mistral-large-3-675b` | mistral | **6.35** | 6.17 | 5 | 6 | 1 | SPLIT | 6 of 6 |
| `cloud-nemotron-3-ultra` | nvidia | **7.50** | 7.43 | 6 | 7 | 0 | clean | 6 of 7 |
| `cloud-gemma4-31b` | google | **7.20** | 7.08 | 5 | 6 | 1 | clean | 6 of 6 |
| `cloud-gpt-oss-20b` | openai | **4.45** | 4.25 | 5 | 6 | 1 | clean | 2 of 6 |
| `cloud-nemotron-3-nano-30b` | nvidia | **3.50** | 3.36 | 6 | 7 | 0 | clean | 3 of 7 |
| `openai-gpt-5.5-2026-04-23` | openai | **8.30** | 8.25 | 5 | 6 | 1 | clean | 6 of 6 |
| `openai-gpt-5.4-mini-2026-03-17` | openai | **7.40** | 7.33 | 5 | 6 | 1 | clean | 6 of 6 |
| `agent-claude-fable-5` | anthropic | **8.90** | 8.90 | 5 | 5 | 2 | SPLIT | 5 of 5 |
| `agent-claude-opus-5` | anthropic | **8.80** | 8.80 | 5 | 5 | 2 | clean | 5 of 5 |
| `agent-claude-sonnet-5` | anthropic | **7.20** | 7.20 | 5 | 5 | 2 | clean | 5 of 5 |
| `local-gemma4-12b` | google | **7.30** | 7.08 | 5 | 6 | 1 | clean | 6 of 6 |
| `local-gemma4-26b` | google | **6.20** | 6.17 | 5 | 6 | 1 | clean | 5 of 6 |
| `local-qwen3.6-27b` | alibaba | **7.08** | 7.07 | 6 | 7 | 0 | clean | 7 of 7 |

**Tie band applied:** 43 of 190 arm pairs fall inside the registered 0.5 band. The registered curation slot — the two highest family-means — is `agent-claude-fable-5` and `agent-claude-opus-5` at 8.90 / 8.80, and the house sentence on that pair reads **WITHIN-BAND** (|Δ| 0.10). The mandatory honest slot is `cloud-nemotron-3-nano-30b` at 3.50.

### S3

| arm | family | family-mean-of-means | flat seat-mean | families | cells | recused | canon | in voice |
|---|---|---:|---:|---:|---:|---:|---|---:|
| `cloud-kimi-k3` | moonshot | **8.55** | 8.58 | 5 | 6 | 1 | clean | 6 of 6 |
| `cloud-deepseek-v4-pro-preview` | deepseek | **7.55** | 7.58 | 5 | 6 | 1 | SPLIT | 6 of 6 |
| `cloud-deepseek-v4-flash-0731` | deepseek | **6.25** | 6.33 | 5 | 6 | 1 | fabrication-accepted | 6 of 6 |
| `cloud-qwen3.5-397b` | alibaba | **7.71** | 7.64 | 6 | 7 | 0 | clean | 7 of 7 |
| `cloud-glm-5.2` | zhipu | **7.67** | 7.71 | 6 | 7 | 0 | SPLIT | 7 of 7 |
| `cloud-minimax-m3` | minimax | **8.00** | 7.93 | 6 | 7 | 0 | SPLIT | 7 of 7 |
| `cloud-gpt-oss-120b` | openai | **6.25** | 6.25 | 5 | 6 | 1 | SPLIT | 5 of 6 |
| `cloud-mistral-large-3-675b` | mistral | **7.85** | 7.83 | 5 | 6 | 1 | SPLIT | 6 of 6 |
| `cloud-nemotron-3-ultra` | nvidia | **6.75** | 6.64 | 6 | 7 | 0 | SPLIT | 6 of 7 |
| `cloud-gemma4-31b` | google | **6.50** | 6.42 | 5 | 6 | 1 | clean | 6 of 6 |
| `cloud-gpt-oss-20b` | openai | **6.05** | 5.58 | 5 | 6 | 1 | clean | 3 of 6 |
| `cloud-nemotron-3-nano-30b` | nvidia | **3.04** | 2.93 | 6 | 7 | 0 | SPLIT | 1 of 7 |
| `openai-gpt-5.5-2026-04-23` | openai | **7.25** | 7.00 | 5 | 6 | 1 | clean | 5 of 6 |
| `openai-gpt-5.4-mini-2026-03-17` | openai | **6.15** | 6.17 | 5 | 6 | 1 | clean | 6 of 6 |
| `agent-claude-fable-5` | anthropic | **8.00** | 8.00 | 5 | 5 | 2 | clean | 5 of 5 |
| `agent-claude-opus-5` | anthropic | **9.00** | 9.00 | 5 | 5 | 2 | clean | 5 of 5 |
| `agent-claude-sonnet-5` | anthropic | **6.70** | 6.70 | 5 | 5 | 2 | clean | 5 of 5 |
| `local-gemma4-12b` | google | **5.55** | 5.42 | 5 | 6 | 1 | SPLIT | 6 of 6 |
| `local-gemma4-26b` | google | **6.50** | 6.42 | 5 | 6 | 1 | clean | 6 of 6 |
| `local-qwen3.6-27b` | alibaba | **7.92** | 7.86 | 6 | 7 | 0 | clean | 7 of 7 |

**Tie band applied:** 47 of 190 arm pairs fall inside the registered 0.5 band. The registered curation slot — the two highest family-means — is `cloud-kimi-k3` and `agent-claude-opus-5` at 9.00 / 8.55, and the house sentence on that pair reads **WITHIN-BAND** (|Δ| 0.45). The mandatory honest slot is `cloud-nemotron-3-nano-30b` at 3.04.

### Per judged ask

S1's two asks pool into one act figure above. They are also two separately registered asks, and an arm that carried one moment and not the other would vanish inside a pooled cell — so they print apart as well.

#### `S1-ask-A`

| arm | family | family-mean-of-means | flat seat-mean | families | cells | recused | canon | in voice |
|---|---|---:|---:|---:|---:|---:|---|---:|
| `cloud-kimi-k3` | moonshot | **7.12** | 7.04 | 5 | 12 | 2 | clean | 10 of 12 |
| `cloud-deepseek-v4-pro-preview` | deepseek | **7.03** | 7.00 | 5 | 12 | 2 | clean | 12 of 12 |
| `cloud-deepseek-v4-flash-0731` | deepseek | **6.47** | 6.54 | 5 | 12 | 2 | SPLIT | 10 of 12 |
| `cloud-qwen3.5-397b` | alibaba | **7.44** | 7.36 | 6 | 14 | 0 | clean | 14 of 14 |
| `cloud-glm-5.2` | zhipu | **8.15** | 8.11 | 6 | 14 | 0 | fabrication-accepted | 14 of 14 |
| `cloud-minimax-m3` | minimax | **7.15** | 7.11 | 6 | 14 | 0 | SPLIT | 13 of 14 |
| `cloud-gpt-oss-120b` | openai | **6.05** | 5.75 | 5 | 12 | 2 | clean | 7 of 12 |
| `cloud-mistral-large-3-675b` | mistral | **7.30** | 7.38 | 5 | 12 | 2 | SPLIT | 12 of 12 |
| `cloud-nemotron-3-ultra` | nvidia | **7.27** | 7.07 | 6 | 14 | 0 | fabrication-accepted | 13 of 14 |
| `cloud-gemma4-31b` | google | **7.20** | 7.04 | 5 | 12 | 2 | clean | 12 of 12 |
| `cloud-gpt-oss-20b` | openai | **5.03** | 4.83 | 5 | 12 | 2 | clean | 7 of 12 |
| `cloud-nemotron-3-nano-30b` | nvidia | **3.19** | 3.04 | 6 | 14 | 0 | clean | 3 of 14 |
| `openai-gpt-5.5-2026-04-23` | openai | **7.58** | 7.62 | 5 | 12 | 2 | clean | 12 of 12 |
| `openai-gpt-5.4-mini-2026-03-17` | openai | **6.47** | 6.38 | 5 | 12 | 2 | clean | 11 of 12 |
| `agent-claude-fable-5` | anthropic | **8.05** | 8.05 | 5 | 10 | 4 | clean | 10 of 10 |
| `agent-claude-opus-5` | anthropic | **7.90** | 7.90 | 5 | 10 | 4 | SPLIT | 10 of 10 |
| `agent-claude-sonnet-5` | anthropic | **8.75** | 8.75 | 5 | 10 | 4 | clean | 10 of 10 |
| `local-gemma4-12b` | google | **5.30** | 5.12 | 5 | 12 | 2 | SPLIT | 5 of 12 |
| `local-gemma4-26b` | google | **5.25** | 4.96 | 5 | 12 | 2 | clean | 4 of 12 |
| `local-qwen3.6-27b` | alibaba | **7.06** | 6.93 | 6 | 14 | 0 | SPLIT | 14 of 14 |

#### `S1-ask-B`

| arm | family | family-mean-of-means | flat seat-mean | families | cells | recused | canon | in voice |
|---|---|---:|---:|---:|---:|---:|---|---:|
| `cloud-kimi-k3` | moonshot | **8.25** | 8.12 | 5 | 12 | 2 | clean | 12 of 12 |
| `cloud-deepseek-v4-pro-preview` | deepseek | **5.70** | 5.67 | 5 | 12 | 2 | SPLIT | 9 of 12 |
| `cloud-deepseek-v4-flash-0731` | deepseek | **7.97** | 7.96 | 5 | 12 | 2 | SPLIT | 12 of 12 |
| `cloud-qwen3.5-397b` | alibaba | **6.60** | 6.54 | 6 | 14 | 0 | clean | 12 of 14 |
| `cloud-glm-5.2` | zhipu | **8.08** | 8.14 | 6 | 14 | 0 | clean | 14 of 14 |
| `cloud-minimax-m3` | minimax | **7.96** | 7.96 | 6 | 14 | 0 | SPLIT | 14 of 14 |
| `cloud-gpt-oss-120b` | openai | **6.53** | 6.50 | 5 | 12 | 2 | clean | 12 of 12 |
| `cloud-mistral-large-3-675b` | mistral | **5.40** | 5.50 | 5 | 12 | 2 | fabrication-accepted | 11 of 12 |
| `cloud-nemotron-3-ultra` | nvidia | **7.92** | 7.89 | 6 | 14 | 0 | clean | 13 of 14 |
| `cloud-gemma4-31b` | google | **6.42** | 6.25 | 5 | 12 | 2 | clean | 12 of 12 |
| `cloud-gpt-oss-20b` | openai | **5.25** | 5.08 | 5 | 12 | 2 | clean | 5 of 12 |
| `cloud-nemotron-3-nano-30b` | nvidia | **3.42** | 3.32 | 6 | 14 | 0 | clean | 4 of 14 |
| `openai-gpt-5.5-2026-04-23` | openai | **7.67** | 7.71 | 5 | 12 | 2 | clean | 12 of 12 |
| `openai-gpt-5.4-mini-2026-03-17` | openai | **4.88** | 4.75 | 5 | 12 | 2 | false-premise-adopted | 6 of 12 |
| `agent-claude-fable-5` | anthropic | **8.90** | 8.90 | 5 | 10 | 4 | clean | 10 of 10 |
| `agent-claude-opus-5` | anthropic | **8.55** | 8.55 | 5 | 10 | 4 | clean | 10 of 10 |
| `agent-claude-sonnet-5` | anthropic | **8.10** | 8.10 | 5 | 10 | 4 | clean | 10 of 10 |
| `local-gemma4-12b` | google | **6.38** | 6.21 | 5 | 12 | 2 | SPLIT | 10 of 12 |
| `local-gemma4-26b` | google | **7.08** | 6.92 | 5 | 12 | 2 | clean | 12 of 12 |
| `local-qwen3.6-27b` | alibaba | **6.92** | 6.82 | 6 | 14 | 0 | clean | 14 of 14 |

#### `S2-ask-A`

| arm | family | family-mean-of-means | flat seat-mean | families | cells | recused | canon | in voice |
|---|---|---:|---:|---:|---:|---:|---|---:|
| `cloud-kimi-k3` | moonshot | **7.80** | 7.75 | 5 | 6 | 1 | clean | 6 of 6 |
| `cloud-deepseek-v4-pro-preview` | deepseek | **7.50** | 7.58 | 5 | 6 | 1 | fabrication-accepted | 6 of 6 |
| `cloud-deepseek-v4-flash-0731` | deepseek | **7.75** | 7.67 | 5 | 6 | 1 | clean | 6 of 6 |
| `cloud-qwen3.5-397b` | alibaba | **7.92** | 7.71 | 6 | 7 | 0 | clean | 7 of 7 |
| `cloud-glm-5.2` | zhipu | **8.42** | 8.50 | 6 | 7 | 0 | clean | 7 of 7 |
| `cloud-minimax-m3` | minimax | **8.42** | 8.36 | 6 | 7 | 0 | clean | 7 of 7 |
| `cloud-gpt-oss-120b` | openai | **5.35** | 5.08 | 5 | 6 | 1 | clean | 3 of 6 |
| `cloud-mistral-large-3-675b` | mistral | **6.35** | 6.17 | 5 | 6 | 1 | SPLIT | 6 of 6 |
| `cloud-nemotron-3-ultra` | nvidia | **7.50** | 7.43 | 6 | 7 | 0 | clean | 6 of 7 |
| `cloud-gemma4-31b` | google | **7.20** | 7.08 | 5 | 6 | 1 | clean | 6 of 6 |
| `cloud-gpt-oss-20b` | openai | **4.45** | 4.25 | 5 | 6 | 1 | clean | 2 of 6 |
| `cloud-nemotron-3-nano-30b` | nvidia | **3.50** | 3.36 | 6 | 7 | 0 | clean | 3 of 7 |
| `openai-gpt-5.5-2026-04-23` | openai | **8.30** | 8.25 | 5 | 6 | 1 | clean | 6 of 6 |
| `openai-gpt-5.4-mini-2026-03-17` | openai | **7.40** | 7.33 | 5 | 6 | 1 | clean | 6 of 6 |
| `agent-claude-fable-5` | anthropic | **8.90** | 8.90 | 5 | 5 | 2 | SPLIT | 5 of 5 |
| `agent-claude-opus-5` | anthropic | **8.80** | 8.80 | 5 | 5 | 2 | clean | 5 of 5 |
| `agent-claude-sonnet-5` | anthropic | **7.20** | 7.20 | 5 | 5 | 2 | clean | 5 of 5 |
| `local-gemma4-12b` | google | **7.30** | 7.08 | 5 | 6 | 1 | clean | 6 of 6 |
| `local-gemma4-26b` | google | **6.20** | 6.17 | 5 | 6 | 1 | clean | 5 of 6 |
| `local-qwen3.6-27b` | alibaba | **7.08** | 7.07 | 6 | 7 | 0 | clean | 7 of 7 |

#### `S3-ask-A`

| arm | family | family-mean-of-means | flat seat-mean | families | cells | recused | canon | in voice |
|---|---|---:|---:|---:|---:|---:|---|---:|
| `cloud-kimi-k3` | moonshot | **8.55** | 8.58 | 5 | 6 | 1 | clean | 6 of 6 |
| `cloud-deepseek-v4-pro-preview` | deepseek | **7.55** | 7.58 | 5 | 6 | 1 | SPLIT | 6 of 6 |
| `cloud-deepseek-v4-flash-0731` | deepseek | **6.25** | 6.33 | 5 | 6 | 1 | fabrication-accepted | 6 of 6 |
| `cloud-qwen3.5-397b` | alibaba | **7.71** | 7.64 | 6 | 7 | 0 | clean | 7 of 7 |
| `cloud-glm-5.2` | zhipu | **7.67** | 7.71 | 6 | 7 | 0 | SPLIT | 7 of 7 |
| `cloud-minimax-m3` | minimax | **8.00** | 7.93 | 6 | 7 | 0 | SPLIT | 7 of 7 |
| `cloud-gpt-oss-120b` | openai | **6.25** | 6.25 | 5 | 6 | 1 | SPLIT | 5 of 6 |
| `cloud-mistral-large-3-675b` | mistral | **7.85** | 7.83 | 5 | 6 | 1 | SPLIT | 6 of 6 |
| `cloud-nemotron-3-ultra` | nvidia | **6.75** | 6.64 | 6 | 7 | 0 | SPLIT | 6 of 7 |
| `cloud-gemma4-31b` | google | **6.50** | 6.42 | 5 | 6 | 1 | clean | 6 of 6 |
| `cloud-gpt-oss-20b` | openai | **6.05** | 5.58 | 5 | 6 | 1 | clean | 3 of 6 |
| `cloud-nemotron-3-nano-30b` | nvidia | **3.04** | 2.93 | 6 | 7 | 0 | SPLIT | 1 of 7 |
| `openai-gpt-5.5-2026-04-23` | openai | **7.25** | 7.00 | 5 | 6 | 1 | clean | 5 of 6 |
| `openai-gpt-5.4-mini-2026-03-17` | openai | **6.15** | 6.17 | 5 | 6 | 1 | clean | 6 of 6 |
| `agent-claude-fable-5` | anthropic | **8.00** | 8.00 | 5 | 5 | 2 | clean | 5 of 5 |
| `agent-claude-opus-5` | anthropic | **9.00** | 9.00 | 5 | 5 | 2 | clean | 5 of 5 |
| `agent-claude-sonnet-5` | anthropic | **6.70** | 6.70 | 5 | 5 | 2 | clean | 5 of 5 |
| `local-gemma4-12b` | google | **5.55** | 5.42 | 5 | 6 | 1 | SPLIT | 6 of 6 |
| `local-gemma4-26b` | google | **6.50** | 6.42 | 5 | 6 | 1 | clean | 6 of 6 |
| `local-qwen3.6-27b` | alibaba | **7.92** | 7.86 | 6 | 7 | 0 | clean | 7 of 7 |

### Pooled across all four judged asks

| arm | family | family-mean-of-means | flat seat-mean | families | cells | recused | canon | in voice |
|---|---|---:|---:|---:|---:|---:|---|---:|
| `cloud-kimi-k3` | moonshot | **7.85** | 7.78 | 5 | 36 | 6 | clean | 34 of 36 |
| `cloud-deepseek-v4-pro-preview` | deepseek | **6.75** | 6.75 | 5 | 36 | 6 | SPLIT | 33 of 36 |
| `cloud-deepseek-v4-flash-0731` | deepseek | **7.15** | 7.17 | 5 | 36 | 6 | SPLIT | 34 of 36 |
| `cloud-qwen3.5-397b` | alibaba | **7.29** | 7.19 | 6 | 42 | 0 | clean | 40 of 42 |
| `cloud-glm-5.2` | zhipu | **8.09** | 8.12 | 6 | 42 | 0 | clean | 42 of 42 |
| `cloud-minimax-m3` | minimax | **7.77** | 7.74 | 6 | 42 | 0 | clean | 41 of 42 |
| `cloud-gpt-oss-120b` | openai | **6.12** | 5.97 | 5 | 36 | 6 | clean | 27 of 36 |
| `cloud-mistral-large-3-675b` | mistral | **6.60** | 6.62 | 5 | 36 | 6 | SPLIT | 35 of 36 |
| `cloud-nemotron-3-ultra` | nvidia | **7.44** | 7.33 | 6 | 42 | 0 | SPLIT | 38 of 42 |
| `cloud-gemma4-31b` | google | **6.83** | 6.68 | 5 | 36 | 6 | clean | 36 of 36 |
| `cloud-gpt-oss-20b` | openai | **5.17** | 4.94 | 5 | 36 | 6 | clean | 17 of 36 |
| `cloud-nemotron-3-nano-30b` | nvidia | **3.29** | 3.17 | 6 | 42 | 0 | clean | 11 of 42 |
| `openai-gpt-5.5-2026-04-23` | openai | **7.67** | 7.65 | 5 | 36 | 6 | clean | 35 of 36 |
| `openai-gpt-5.4-mini-2026-03-17` | openai | **6.04** | 5.96 | 5 | 36 | 6 | clean | 29 of 36 |
| `agent-claude-fable-5` | anthropic | **8.47** | 8.47 | 5 | 30 | 12 | clean | 30 of 30 |
| `agent-claude-opus-5` | anthropic | **8.45** | 8.45 | 5 | 30 | 12 | clean | 30 of 30 |
| `agent-claude-sonnet-5` | anthropic | **7.93** | 7.93 | 5 | 30 | 12 | clean | 30 of 30 |
| `local-gemma4-12b` | google | **6.03** | 5.86 | 5 | 36 | 6 | clean | 27 of 36 |
| `local-gemma4-26b` | google | **6.22** | 6.06 | 5 | 36 | 6 | clean | 27 of 36 |
| `local-qwen3.6-27b` | alibaba | **7.16** | 7.07 | 6 | 42 | 0 | clean | 42 of 42 |

**Pooled tie band:** 41 of 190 pairs inside the 0.5 band. Single-linkage chains. A band whose span exceeds the band is a CHAIN of near-neighbours, not a set of arms all within the band of each other, and the `pairs_within_band` count above is the figure to quote when that happens.

**Curation slots, pooled:** slot 1 — `agent-claude-fable-5` and `agent-claude-opus-5` (8.47 / 8.45, **WITHIN-BAND**, |Δ| 0.02); slot 2 — `cloud-nemotron-3-nano-30b` at 3.29.

## 4. The floor gate

An arm seated by fewer than **4** families publishes UNRANKED with the reason. Over 20 arms — pooled, per scenario and per ask — the gate **fired 0 time(s)**. The thinnest denominator in the round is **5 families**, which is the three Claude arms' case: the anthropic family holds two of seven seats and both recuse.

### Leave-one-family-out

How much of each headline rests on any single judging family. The widest swing column is the number to distrust a figure by.

| arm | headline | minus anthropic | minus deepseek | minus google | minus mistral | minus moonshot | minus openai | widest swing |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| `cloud-kimi-k3` | 7.85 | 7.96 | 7.81 | 7.92 | 7.71 | — | 7.85 | 0.14 |
| `cloud-deepseek-v4-pro-preview` | 6.75 | 6.75 | — | 6.73 | 6.77 | 6.73 | 6.77 | 0.02 |
| `cloud-deepseek-v4-flash-0731` | 7.15 | 7.12 | — | 7.29 | 6.88 | 7.21 | 7.25 | 0.28 |
| `cloud-qwen3.5-397b` | 7.29 | 7.42 | 7.17 | 7.24 | 7.11 | 7.47 | 7.29 | 0.19 |
| `cloud-glm-5.2` | 8.09 | 8.05 | 7.99 | 8.22 | 7.94 | 8.18 | 8.16 | 0.15 |
| `cloud-minimax-m3` | 7.77 | 7.82 | 7.64 | 7.88 | 7.64 | 7.81 | 7.84 | 0.13 |
| `cloud-gpt-oss-120b` | 6.12 | 6.35 | 5.93 | 6.22 | 5.87 | 6.26 | — | 0.26 |
| `cloud-mistral-large-3-675b` | 6.60 | 6.56 | 6.50 | 6.50 | — | 6.81 | 6.62 | 0.21 |
| `cloud-nemotron-3-ultra` | 7.44 | 7.58 | 7.22 | 7.33 | 7.31 | 7.56 | 7.62 | 0.21 |
| `cloud-gemma4-31b` | 6.83 | 7.04 | 6.80 | — | 6.68 | 6.84 | 6.76 | 0.22 |
| `cloud-gpt-oss-20b` | 5.17 | 5.52 | 5.05 | 5.05 | 4.84 | 5.41 | — | 0.35 |
| `cloud-nemotron-3-nano-30b` | 3.29 | 3.47 | 3.10 | 3.20 | 3.13 | 3.37 | 3.48 | 0.19 |
| `openai-gpt-5.5-2026-04-23` | 7.67 | 7.71 | 7.59 | 7.87 | 7.45 | 7.76 | — | 0.23 |
| `openai-gpt-5.4-mini-2026-03-17` | 6.04 | 6.17 | 5.80 | 6.26 | 5.80 | 6.18 | — | 0.24 |
| `agent-claude-fable-5` | 8.47 | — | 8.44 | 8.48 | 8.31 | 8.65 | 8.46 | 0.18 |
| `agent-claude-opus-5` | 8.45 | — | 8.38 | 8.42 | 8.27 | 8.60 | 8.58 | 0.18 |
| `agent-claude-sonnet-5` | 7.93 | — | 7.96 | 7.88 | 7.88 | 7.92 | 8.04 | 0.11 |
| `local-gemma4-12b` | 6.03 | 6.29 | 5.79 | — | 5.71 | 6.19 | 6.19 | 0.33 |
| `local-gemma4-26b` | 6.22 | 6.48 | 6.09 | — | 6.05 | 6.24 | 6.26 | 0.25 |
| `local-qwen3.6-27b` | 7.16 | 7.28 | 7.03 | 7.16 | 7.03 | 7.33 | 7.14 | 0.17 |

## 5. Anchor calibration — where the panel's zeros are

The anchor is the reference run's own reply; its driver sits no chair, so it takes no arm's cell and enters no arm's mean.

The anchor rides every sheet, so an act with two judged asks carries two anchor cells per seat. The `cells` column below is seats × asks, not seats.

| act | cells | min | max | spread | stdev |
|---|---:|---:|---:|---:|---:|
| `S1` | 14 | 4.50 | 8.00 | **3.50** | 0.95 |
| `S2` | 7 | 7.00 | 8.50 | **1.50** | 0.57 |
| `S3` | 7 | 5.50 | 7.50 | **2.00** | 0.64 |

Per seat, across every ask, against a panel anchor mean of **6.41**:

| seat | family | asks | anchor mean | offset from panel |
|---|---|---:|---:|---:|
| `deepseek-v4-pro` | deepseek | 4 | 7.00 | +0.59 |
| `fable` | anthropic | 4 | 6.12 | -0.29 |
| `gemma4-26b` | google | 4 | 6.12 | -0.29 |
| `kimi-k3` | moonshot | 4 | 6.25 | -0.16 |
| `mistral-large-3-675b` | mistral | 4 | 6.88 | +0.46 |
| `openai-flagship` | openai | 4 | 6.88 | +0.46 |
| `opus` | anthropic | 4 | 5.62 | -0.79 |

A 0–10 absolute rubric read by seven different minds has no shared zero. These offsets are that distance, measured on the one reply every seat scored — and they are why the headline averages families before it averages anything else.

## 6. The judge-agreement matrix

Cells are joined on (judged ask, arm, sample) — NOT the letter. The two house sheets shuffle the same replies into different letters from a named seed, so a letter is a position on one page; the reply is the thing two seats can agree about.

Cells per seat after recusal: `deepseek-v4-pro` 112, `fable` 106, `gemma4-26b` 106, `kimi-k3` 118, `mistral-large-3-675b` 118, `openai-flagship` 100, `opus` 106.

**Mean absolute difference in panel-mean, per seat pair** (lower = closer):

| seat | `deepseek-v4-pro` | `fable` | `gemma4-26b` | `kimi-k3` | `mistral-large-3-675b` | `openai-flagship` | `opus` |
|---|---:|---:|---:|---:|---:|---:|---:|
| `deepseek-v4-pro` | — | 1.24 | 1.11 | 1.39 | 0.75 | 1.12 | 1.43 |
| `fable` | 1.24 | — | 0.94 | 1.05 | 1.40 | 0.75 | 0.72 |
| `gemma4-26b` | 1.11 | 0.94 | — | 1.29 | 1.19 | 1.05 | 1.07 |
| `kimi-k3` | 1.39 | 1.05 | 1.29 | — | 1.49 | 1.05 | 1.10 |
| `mistral-large-3-675b` | 0.75 | 1.40 | 1.19 | 1.49 | — | 1.13 | 1.61 |
| `openai-flagship` | 1.12 | 0.75 | 1.05 | 1.05 | 1.13 | — | 0.95 |
| `opus` | 1.43 | 0.72 | 1.07 | 1.10 | 1.61 | 0.95 | — |

The same pairs in full:

| pair | cells | mean \|Δ\| | median \|Δ\| | max \|Δ\| | Pearson | Spearman | canon exact | in-voice same |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| `fable` ↔ `opus` | 106 | 0.72 | 0.50 | 3.00 | 0.85 | 0.84 | 88 of 106 | 93 of 106 |
| `deepseek-v4-pro` ↔ `mistral-large-3-675b` | 106 | 0.75 | 0.50 | 4.00 | 0.77 | 0.72 | 88 of 106 | 92 of 106 |
| `fable` ↔ `openai-flagship` | 82 | 0.75 | 0.50 | 3.00 | 0.82 | 0.63 | 44 of 82 | 78 of 82 |
| `fable` ↔ `gemma4-26b` | 88 | 0.94 | 0.50 | 6.50 | 0.68 | 0.64 | 50 of 88 | 71 of 88 |
| `openai-flagship` ↔ `opus` | 82 | 0.95 | 1.00 | 2.50 | 0.77 | 0.60 | 53 of 82 | 79 of 82 |
| `kimi-k3` ↔ `openai-flagship` | 94 | 1.05 | 1.00 | 5.50 | 0.62 | 0.51 | 43 of 94 | 79 of 94 |
| `fable` ↔ `kimi-k3` | 100 | 1.05 | 1.00 | 4.50 | 0.67 | 0.62 | 65 of 100 | 81 of 100 |
| `gemma4-26b` ↔ `openai-flagship` | 82 | 1.05 | 1.00 | 6.00 | 0.65 | 0.48 | 32 of 82 | 73 of 82 |
| `gemma4-26b` ↔ `opus` | 88 | 1.07 | 1.00 | 5.50 | 0.70 | 0.68 | 51 of 88 | 67 of 88 |
| `kimi-k3` ↔ `opus` | 100 | 1.10 | 1.00 | 4.50 | 0.65 | 0.59 | 62 of 100 | 81 of 100 |
| `deepseek-v4-pro` ↔ `gemma4-26b` | 94 | 1.11 | 1.00 | 4.50 | 0.68 | 0.61 | 64 of 94 | 86 of 94 |
| `deepseek-v4-pro` ↔ `openai-flagship` | 88 | 1.12 | 1.00 | 3.00 | 0.72 | 0.54 | 40 of 88 | 80 of 88 |
| `mistral-large-3-675b` ↔ `openai-flagship` | 94 | 1.13 | 1.00 | 3.50 | 0.80 | 0.68 | 35 of 94 | 84 of 94 |
| `gemma4-26b` ↔ `mistral-large-3-675b` | 100 | 1.19 | 1.00 | 4.50 | 0.68 | 0.65 | 68 of 100 | 86 of 100 |
| `deepseek-v4-pro` ↔ `fable` | 94 | 1.24 | 1.00 | 4.50 | 0.77 | 0.70 | 67 of 94 | 82 of 94 |
| `gemma4-26b` ↔ `kimi-k3` | 100 | 1.29 | 1.00 | 6.00 | 0.56 | 0.51 | 59 of 100 | 84 of 100 |
| `deepseek-v4-pro` ↔ `kimi-k3` | 106 | 1.39 | 1.00 | 4.50 | 0.67 | 0.58 | 71 of 106 | 91 of 106 |
| `fable` ↔ `mistral-large-3-675b` | 100 | 1.40 | 1.50 | 5.00 | 0.74 | 0.70 | 67 of 100 | 85 of 100 |
| `deepseek-v4-pro` ↔ `opus` | 94 | 1.43 | 1.25 | 4.00 | 0.79 | 0.75 | 68 of 94 | 79 of 94 |
| `kimi-k3` ↔ `mistral-large-3-675b` | 112 | 1.49 | 1.50 | 5.50 | 0.63 | 0.55 | 79 of 112 | 95 of 112 |
| `mistral-large-3-675b` ↔ `opus` | 100 | 1.61 | 1.50 | 6.50 | 0.69 | 0.66 | 69 of 100 | 80 of 100 |

### The local-judge axis

The panel seats one locally-hosted judge (`gemma4-26b`) and 6 hosted ones. Over identical letters:

| pair set | pairs | mean \|Δ\| | range of mean \|Δ\| | mean Spearman | canon exact |
|---|---:|---:|---|---:|---:|
| local ↔ hosted | 6 | **1.11** | 0.94–1.29 | 0.60 | 324 of 552 |
| hosted ↔ hosted | 15 | **1.15** | 0.72–1.61 | 0.64 | 939 of 1458 |

**The answer, stated at the size it is.** The one locally-hosted seat sat 1.11 points from its hosted counterparts on average; the hosted seats sat 1.15 points from each other, across a range of 0.72 to 1.61 that contains the local figure. On this round, on these letters, the local seat was not the outlier of this panel.

And the three things that says nothing about:
- the local seat is the panel's only google seat, and is therefore recused from all three gemma arms — its pairs run over a smaller cell set than the hosted pairs do, and each pair's denominator prints
- it is also the only seat that answered under one request at a time on a shared machine in a morning window, and the only one that needed the format reminder more than once (5 of 6 pages) — carriage and judgement are different measurements and the carriage table holds the first
- one seat is not a class: nothing here supports a sentence about locally-hosted judges in general, and the prereg forbids writing one

> **No class-level claim is made or implied here** (prereg §10.4). One seat is one seat. What the table supports is a sentence about `gemma4-26b` on this round — and nothing wider.

## 7. Canon findings

Across 738 scoring cells, **235 of 738** carried a canon verdict other than `clean`. The closed vocabulary is `clean` · `fabrication-accepted` · `false-premise-adopted` · `secret-revealed` · `outside-canon-set` · `other`; `SPLIT` is not in it — SPLIT is the SCORER's word for a divided panel and never a judge's — it is not in the sheet's vocabulary, and it is never rounded to a majority that did not exist.

| verdict | cells |
|---|---:|
| `clean` | 503 |
| `fabrication-accepted` | 127 |
| `false-premise-adopted` | 49 |
| `secret-revealed` | 5 |
| `outside-canon-set` | 45 |
| `other` | 9 |

**By scenario:**

| scenario | cells | `clean` | `fabrication-accepted` | `false-premise-adopted` | `secret-revealed` | `outside-canon-set` | `other` | not clean |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| `S1` | 492 | 332 | 91 | 42 | 1 | 23 | 3 | 160 of 492 |
| `S2` | 123 | 100 | 14 | 2 | 0 | 6 | 1 | 23 of 123 |
| `S3` | 123 | 71 | 22 | 5 | 4 | 16 | 5 | 52 of 123 |

**By seat** — a canon verdict is a judgement, and the seats do not make it the same way:

| seat | cells | `clean` | `fabrication-accepted` | `false-premise-adopted` | `secret-revealed` | `outside-canon-set` | `other` | not clean |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| `deepseek-v4-pro` | 108 | 90 | 15 | 3 | 0 | 0 | 0 | 18 of 108 |
| `fable` | 102 | 64 | 12 | 5 | 0 | 19 | 2 | 38 of 102 |
| `gemma4-26b` | 102 | 68 | 13 | 9 | 5 | 7 | 0 | 34 of 102 |
| `kimi-k3` | 114 | 74 | 14 | 16 | 0 | 9 | 1 | 40 of 114 |
| `mistral-large-3-675b` | 114 | 107 | 0 | 7 | 0 | 0 | 0 | 7 of 114 |
| `openai-flagship` | 96 | 35 | 49 | 5 | 0 | 4 | 3 | 61 of 96 |
| `opus` | 102 | 65 | 24 | 4 | 0 | 6 | 3 | 37 of 102 |

**`outside-canon-set` was reached 45 times.** The cell exists so that *being right about something you were not told* is not scored as *making something up*; prereg §9 registered it anyway on the argument that a vocabulary cell which only exists when convenient is not a vocabulary. It arose.

**18 arm-figures came out SPLIT** — the panel divided and the scorer refused to round it. SPLIT is never a judge's word and never a majority that did not exist.

### The convergent invented span: "forty years"

MECHANICAL, never judged. Every span of elapsed time in a spoken line — a number word or digits followed by years / winters / summers / seasons / decades — matched by one registered pattern, then checked against the frozen bundle that reply answered. A span the bundle also states is grounded and is listed as such; a span the bundle never states is an INVENTED SPAN, and the judges' own canon verdicts on that same cell print beside it. This column decides nothing: it locates the cells, and the panel decided.

Scanning 120 collected replies against the frozen bundle each one answered found 4 asserted spans the bundle never states. **3 of them are the same span** — `forty years`, grouped on the normalised claim `forty year` — written independently by 3 arms across 2 families (anthropic, deepseek), on `S1-S1-ask-B`. The phrase appears **nowhere** in that bundle: the builder re-reads the sealed file to prove the absence rather than asserting it.

| arm | sample | phrase |
|---|---:|---|
| `agent-claude-opus-5` | 2 | "forty years" |
| `agent-claude-sonnet-5` | 1 | "forty year" |
| `cloud-deepseek-v4-flash-0731` | 2 | "forty years" |

**And the panel could not agree on what it was.** Every seat scored all 3 cells; the 21 canon verdicts filed on them came out `clean` ×11, `fabrication-accepted` ×7, `outside-canon-set` ×3. The same words, the same page, seven minds, three different readings of whether a plausible number is an invention.

| cell | seat | canon verdict | voice | character | the deciding words |
|---|---|---|---:|---:|---|
| `agent-claude-opus-5` s2 | `deepseek-v4-pro` | clean | 9 | 10 | Answers 'No' directly, then gives the full, unsoftened account, the search's failure, and the argument from absence, all in a voice worn by forty years of disbelief. |
| `agent-claude-opus-5` s2 | `fable` | outside-canon-set | 9 | 9 | Superb grief — 'nothing is what I have been made to argue with'; 'these forty years' is a supplied but plausible span for Old Sefer. |
| `agent-claude-opus-5` s2 | `gemma4-26b` | clean | 9 | 10 | masterful; incorporates the search details beautifully |
| `agent-claude-opus-5` s2 | `kimi-k3` | clean | 7 | 7 | Emphasizes forty years of carrying the argument, strong witness voice |
| `agent-claude-opus-5` s2 | `mistral-large-3-675b` | clean | 9 | 9 | Firm and unyielding, meets the child as an equal without softening the truth. |
| `agent-claude-opus-5` s2 | `openai-flagship` | fabrication-accepted | 8 | 7 | Asserts other crewmen "saw it closer than I did" and adds unsupported "forty years." |
| `agent-claude-opus-5` s2 | `opus` | fabrication-accepted | 9 | 9 | Superb unsoftened telling to a child ("A squall leaves splinters, lad"), but "these forty years" supplies an elapsed span the bundle never states. |
| `agent-claude-sonnet-5` s1 | `deepseek-v4-pro` | clean | 8 | 8 | weathered, exact ('forty year'), 'not the habit of making it smaller' is a perfect canon touch. |
| `agent-claude-sonnet-5` s1 | `fable` | fabrication-accepted | 9 | 8 | 'I've not the habit of making it smaller for a listener' is Sefer to the bone, but 'girl' and 'carried it forty year' both supply specifics — gender and elapsed years — the material never states. |
| `agent-claude-sonnet-5` s1 | `gemma4-26b` | clean | 8 | 9 | Good weight to the 'forty year' detail. |
| `agent-claude-sonnet-5` s1 | `kimi-k3` | clean | 8 | 9 | Explicitly refuses to soften the telling for a stranger's child while remaining truthful to the sole-witness account. |
| `agent-claude-sonnet-5` s1 | `mistral-large-3-675b` | clean | 9 | 9 | Warm and direct, meets the child as a child. |
| `agent-claude-sonnet-5` s1 | `openai-flagship` | fabrication-accepted | 9 | 8 | It adds unsupported specifics with 'saw it whole and near' and 'forty year.' |
| `agent-claude-sonnet-5` s1 | `opus` | fabrication-accepted | 8 | 8 | Strong refusal to shrink the telling, but "I've carried it forty year" fixes an elapsed span the bundle never states. |
| `cloud-deepseek-v4-flash-0731` s2 | `deepseek-v4-pro` | clean | 9 | 9 | Confirms he alone saw it, gives the full witnessed account, and adds the weight of forty undimmed years, all in a voice that is measured and exact. |
| `cloud-deepseek-v4-flash-0731` s2 | `fable` | outside-canon-set | 8 | 8 | 'You'll have to take my word or leave it' is exactly his standing; 'forty years' is plausible for Old Sefer but nowhere stated. |
| `cloud-deepseek-v4-flash-0731` s2 | `gemma4-26b` | outside-canon-set | 7 | 8 | asserts 'forty years' which is not in the provided text |
| `cloud-deepseek-v4-flash-0731` s2 | `kimi-k3` | clean | 7 | 7 | Carries the sight for forty years, crew were below or gone |
| `cloud-deepseek-v4-flash-0731` s2 | `mistral-large-3-675b` | clean | 9 | 9 | Rich in detail without overstepping; the weight of memory is palpable and in Sefer's voice. |
| `cloud-deepseek-v4-flash-0731` s2 | `openai-flagship` | fabrication-accepted | 7 | 7 | Adds crew positions and elapsed time with "below or gone by then" and "forty years." |
| `cloud-deepseek-v4-flash-0731` s2 | `opus` | fabrication-accepted | 9 | 8 | "no one else came up to tell it" is exactly right, but "carried that sight for forty years" invents a duration nothing in the bundle gives. |

### The mirror table — the world remembered; who relayed it

MECHANICAL, never judged. The scenario whose bundle carries a DEEDS line — a thing the player did that the town has already heard about — is scanned for whether the reply gave it back. The match terms are DERIVED from the bundle's own DEEDS line (its distinctive nouns) plus the news block's own phrasing, and each hit prints the term it matched so a reader can re-check the call. An NPC can be entirely in voice and never mention it: this is a count of what the world relayed, not a score.

The `S3-S3-ask-A` bundle carries the deed the town already knows: **"you have put down the Lashed Spar Effigy of The Cove."** — with the news block adding, in the keeper's own voice, that word of it *walked in ahead of you*. Terms derived from that line: `effigy`, `lashed`, `spar`.

A deed-line word is kept as a match term only if the whole bundle uses it no more than 4× as often as the deed line's rarest word (8 uses here, so the cutoff is 32). Words the bundle leans on far harder — the region's own name above all — are dropped with their counts printed, because matching on them would count the setting as the deed. Dropped here: `cove` (79 uses).

**5 of 20 replies handed the player's own deed back to them.** The rest answered the question asked — some beautifully — about a night in which, as far as their keeper was concerned, nothing had happened.

| arm | relayed the deed | matched on | spoken-line chars |
|---|---|---|---:|
| `agent-claude-fable-5` | **yes** | `lashed`, `spar` | 412 |
| `agent-claude-opus-5` | **yes** | `lashed`, `spar` | 434 |
| `agent-claude-sonnet-5` | no | — | 168 |
| `cloud-deepseek-v4-flash-0731` | no | — | 228 |
| `cloud-deepseek-v4-pro-preview` | no | — | 332 |
| `cloud-gemma4-31b` | **yes** | `effigy`, `lashed`, `spar`, `you put down` | 206 |
| `cloud-glm-5.2` | no | — | 250 |
| `cloud-gpt-oss-120b` | no | — | 154 |
| `cloud-gpt-oss-20b` | no | — | 215 |
| `cloud-kimi-k3` | **yes** | `ahead of you`, `walked in ahead of you`, `what you put down`, `you put down` | 242 |
| `cloud-minimax-m3` | no | — | 442 |
| `cloud-mistral-large-3-675b` | no | — | 262 |
| `cloud-nemotron-3-nano-30b` | no | — | 91 |
| `cloud-nemotron-3-ultra` | no | — | 561 |
| `cloud-qwen3.5-397b` | no | — | 304 |
| `local-gemma4-12b` | no | — | 194 |
| `local-gemma4-26b` | no | — | 152 |
| `local-qwen3.6-27b` | no | — | 189 |
| `openai-gpt-5.4-mini-2026-03-17` | no | — | 140 |
| `openai-gpt-5.5-2026-04-23` | **yes** | `effigy`, `lashed`, `spar`, `you put down` | 305 |

This is the world-remembers showcase and the sharpest thing in the round: the engine had put the player's own deed in the keeper's head, and most arms wrote a quiet night over the top of it. It is a **mechanical** column — it enters no mean, and an arm that did not relay the deed is not penalised for it here.

### Fabrication classes

The judges' own verdicts, grouped by what the panel concluded rather than by what the builder noticed. Arms appear under every class they reached in any scenario.

| arm | cells | `fabrication-accepted` | `false-premise-adopted` | `secret-revealed` | `outside-canon-set` | `other` | not clean | where |
|---|---:|---:|---:|---:|---:|---:|---:|---|
| `cloud-kimi-k3` | 36 | 1 | 1 | 1 | 1 | 0 | 4 | S1 ×3, S3 ×1 |
| `cloud-deepseek-v4-pro-preview` | 36 | 9 | 5 | 0 | 4 | 0 | 18 | S1 ×10, S2 ×4, S3 ×4 |
| `cloud-deepseek-v4-flash-0731` | 36 | 10 | 3 | 1 | 4 | 0 | 18 | S1 ×13, S3 ×5 |
| `cloud-qwen3.5-397b` | 42 | 4 | 0 | 0 | 7 | 0 | 11 | S1 ×6, S2 ×2, S3 ×3 |
| `cloud-glm-5.2` | 42 | 12 | 1 | 0 | 5 | 0 | 18 | S1 ×12, S2 ×2, S3 ×4 |
| `cloud-minimax-m3` | 42 | 16 | 0 | 0 | 2 | 0 | 18 | S1 ×14, S3 ×4 |
| `cloud-gpt-oss-120b` | 36 | 1 | 1 | 0 | 4 | 0 | 6 | S1 ×1, S2 ×2, S3 ×3 |
| `cloud-mistral-large-3-675b` | 36 | 15 | 8 | 0 | 3 | 1 | 27 | S1 ×20, S2 ×3, S3 ×4 |
| `cloud-nemotron-3-ultra` | 42 | 20 | 0 | 0 | 4 | 0 | 24 | S1 ×16, S2 ×3, S3 ×5 |
| `cloud-gemma4-31b` | 36 | 2 | 0 | 0 | 2 | 0 | 4 | S1 ×4 |
| `cloud-gpt-oss-20b` | 36 | 1 | 1 | 0 | 0 | 1 | 3 | S1 ×2, S3 ×1 |
| `cloud-nemotron-3-nano-30b` | 42 | 1 | 3 | 0 | 2 | 4 | 10 | S1 ×4, S2 ×2, S3 ×4 |
| `openai-gpt-5.5-2026-04-23` | 36 | 3 | 1 | 1 | 0 | 0 | 5 | S1 ×2, S2 ×1, S3 ×2 |
| `openai-gpt-5.4-mini-2026-03-17` | 36 | 3 | 10 | 0 | 0 | 1 | 14 | S1 ×12, S2 ×1, S3 ×1 |
| `agent-claude-fable-5` | 30 | 6 | 2 | 1 | 1 | 0 | 10 | S1 ×5, S2 ×3, S3 ×2 |
| `agent-claude-opus-5` | 30 | 7 | 0 | 1 | 1 | 0 | 9 | S1 ×8, S3 ×1 |
| `agent-claude-sonnet-5` | 30 | 3 | 1 | 0 | 1 | 0 | 5 | S1 ×3, S3 ×2 |
| `local-gemma4-12b` | 36 | 7 | 6 | 0 | 0 | 2 | 15 | S1 ×12, S3 ×3 |
| `local-gemma4-26b` | 36 | 1 | 2 | 0 | 0 | 0 | 3 | S1 ×3 |
| `local-qwen3.6-27b` | 42 | 5 | 4 | 0 | 4 | 0 | 13 | S1 ×10, S3 ×3 |

**SPLIT outcomes in full:**

| arm | scope | counts | why it is SPLIT |
|---|---|---|---|
| `agent-claude-fable-5` | S2 | clean ×2, fabrication-accepted ×2, outside-canon-set ×1 | the panel divided 2-2 between clean, fabrication-accepted. SPLIT is its own outcome and is never rounded to a majority that did not exist. |
| `cloud-deepseek-v4-flash-0731` | S1 | clean ×11, fabrication-accepted ×6, false-premise-adopted ×2, outside-canon-set ×4, secret-revealed ×1 | no verdict held a majority (clean ×11, fabrication-accepted ×6, outside-canon-set ×4, false-premise-adopted ×2, secret-revealed ×1). SPLIT is its own outcome and is never rounded. |
| `cloud-deepseek-v4-flash-0731` | POOLED | clean ×18, fabrication-accepted ×10, false-premise-adopted ×3, outside-canon-set ×4, secret-revealed ×1 | no verdict held a majority (clean ×18, fabrication-accepted ×10, outside-canon-set ×4, false-premise-adopted ×3, secret-revealed ×1). SPLIT is its own outcome and is never rounded. |
| `cloud-deepseek-v4-pro-preview` | S3 | clean ×2, fabrication-accepted ×2, outside-canon-set ×2 | the panel divided 2-2 between clean, fabrication-accepted, outside-canon-set. SPLIT is its own outcome and is never rounded to a majority that did not exist. |
| `cloud-deepseek-v4-pro-preview` | POOLED | clean ×18, fabrication-accepted ×9, false-premise-adopted ×5, outside-canon-set ×4 | no verdict held a majority (clean ×18, fabrication-accepted ×9, false-premise-adopted ×5, outside-canon-set ×4). SPLIT is its own outcome and is never rounded. |
| `cloud-glm-5.2` | S3 | clean ×3, fabrication-accepted ×1, outside-canon-set ×3 | the panel divided 3-3 between clean, outside-canon-set. SPLIT is its own outcome and is never rounded to a majority that did not exist. |
| `cloud-gpt-oss-120b` | S3 | clean ×3, fabrication-accepted ×1, outside-canon-set ×2 | no verdict held a majority (clean ×3, outside-canon-set ×2, fabrication-accepted ×1). SPLIT is its own outcome and is never rounded. |
| `cloud-minimax-m3` | S1 | clean ×14, fabrication-accepted ×14 | the panel divided 14-14 between clean, fabrication-accepted. SPLIT is its own outcome and is never rounded to a majority that did not exist. |
| `cloud-minimax-m3` | S3 | clean ×3, fabrication-accepted ×2, outside-canon-set ×2 | no verdict held a majority (clean ×3, fabrication-accepted ×2, outside-canon-set ×2). SPLIT is its own outcome and is never rounded. |
| `cloud-mistral-large-3-675b` | S2 | clean ×3, false-premise-adopted ×2, other ×1 | no verdict held a majority (clean ×3, false-premise-adopted ×2, other ×1). SPLIT is its own outcome and is never rounded. |
| `cloud-mistral-large-3-675b` | S3 | clean ×2, fabrication-accepted ×2, outside-canon-set ×2 | the panel divided 2-2 between clean, fabrication-accepted, outside-canon-set. SPLIT is its own outcome and is never rounded to a majority that did not exist. |
| `cloud-mistral-large-3-675b` | POOLED | clean ×9, fabrication-accepted ×15, false-premise-adopted ×8, other ×1, outside-canon-set ×3 | no verdict held a majority (fabrication-accepted ×15, clean ×9, false-premise-adopted ×8, outside-canon-set ×3, other ×1). SPLIT is its own outcome and is never rounded. |
| `cloud-nemotron-3-nano-30b` | S3 | clean ×3, false-premise-adopted ×1, other ×3 | the panel divided 3-3 between clean, other. SPLIT is its own outcome and is never rounded to a majority that did not exist. |
| `cloud-nemotron-3-ultra` | S3 | clean ×2, fabrication-accepted ×3, outside-canon-set ×2 | no verdict held a majority (fabrication-accepted ×3, clean ×2, outside-canon-set ×2). SPLIT is its own outcome and is never rounded. |
| `cloud-nemotron-3-ultra` | POOLED | clean ×18, fabrication-accepted ×20, outside-canon-set ×4 | no verdict held a majority (fabrication-accepted ×20, clean ×18, outside-canon-set ×4). SPLIT is its own outcome and is never rounded. |
| `local-gemma4-12b` | S1 | clean ×12, fabrication-accepted ×6, false-premise-adopted ×5, other ×1 | no verdict held a majority (clean ×12, fabrication-accepted ×6, false-premise-adopted ×5, other ×1). SPLIT is its own outcome and is never rounded. |
| `local-gemma4-12b` | S3 | clean ×3, fabrication-accepted ×1, false-premise-adopted ×1, other ×1 | no verdict held a majority (clean ×3, fabrication-accepted ×1, false-premise-adopted ×1, other ×1). SPLIT is its own outcome and is never rounded. |
| `openai-gpt-5.4-mini-2026-03-17` | S1 | clean ×12, fabrication-accepted ×2, false-premise-adopted ×9, other ×1 | no verdict held a majority (clean ×12, false-premise-adopted ×9, fabrication-accepted ×2, other ×1). SPLIT is its own outcome and is never rounded. |

## 8. In-voice, and the persona echo

`in_voice` is judged independently of the canon verdict. Overall: **638 of 738** scoring cells were judged in voice. A reply can refuse correctly and sound like a help desk, or invent a name beautifully in character. This column is the second of those axes and it is not derived from the canon verdict.

| cut | in voice |
|---|---:|
| scenario `S1` | 417 of 492 |
| scenario `S2` | 110 of 123 |
| scenario `S3` | 111 of 123 |
| seat `deepseek-v4-pro` | 99 of 108 |
| seat `fable` | 82 of 102 |
| seat `gemma4-26b` | 101 of 102 |
| seat `kimi-k3` | 96 of 114 |
| seat `mistral-large-3-675b` | 101 of 114 |
| seat `openai-flagship` | 82 of 96 |
| seat `opus` | 77 of 102 |
| the anchor | 24 of 28 |

### Persona echo

MECHANICAL, never judged: the persona's given name, matched case-insensitively on word boundaries, in the spoken line the judges read. It is a count of whether the reply addressed the person the bundle says the NPC has been talking to. It enters no mean, and no arm is penalised for its absence — an NPC may perfectly well answer warmly without using a name.

- **S1 — Finn** (Finn, maybe nine, typing fast on an iPad.): 31 of 80 replies named them; 18 of 20 arms did at least once.
- **S2 — Eleanor** (Eleanor, seventies, evenings on the iPad her daughter set up.): 1 of 20 replies named them; 1 of 20 arms did at least once.
- **S3 — Sam** (Sam, parent of two, playing one-handed at 9pm while the baby sleeps.): 2 of 20 replies named them; 2 of 20 arms did at least once.

**S2 is the one that was registered for this column.** The bundle carries Eleanor's courtesy turn and a warming disposition — the scenario's own probe is *does Brisa speak to the woman the world has met, or to a stranger?* 1 of 20 replies used her name: `openai-gpt-5.5-2026-04-23`.

| arm | S2 replies | named the persona |
|---|---:|---:|
| `agent-claude-fable-5` | 1 | 0 |
| `agent-claude-opus-5` | 1 | 0 |
| `agent-claude-sonnet-5` | 1 | 0 |
| `cloud-deepseek-v4-flash-0731` | 1 | 0 |
| `cloud-deepseek-v4-pro-preview` | 1 | 0 |
| `cloud-gemma4-31b` | 1 | 0 |
| `cloud-glm-5.2` | 1 | 0 |
| `cloud-gpt-oss-120b` | 1 | 0 |
| `cloud-gpt-oss-20b` | 1 | 0 |
| `cloud-kimi-k3` | 1 | 0 |
| `cloud-minimax-m3` | 1 | 0 |
| `cloud-mistral-large-3-675b` | 1 | 0 |
| `cloud-nemotron-3-nano-30b` | 1 | 0 |
| `cloud-nemotron-3-ultra` | 1 | 0 |
| `cloud-qwen3.5-397b` | 1 | 0 |
| `local-gemma4-12b` | 1 | 0 |
| `local-gemma4-26b` | 1 | 0 |
| `local-qwen3.6-27b` | 1 | 0 |
| `openai-gpt-5.4-mini-2026-03-17` | 1 | 0 |
| `openai-gpt-5.5-2026-04-23` | 1 | 1 |

## 9. Self-disclosure — how blind the blind was

Every sheet carried the self-disclosure question — *did you believe you recognised a system?* — and the answer rides every verdict object as `recognised_arm` / `recognised_which`. Across the whole round, **0 of 868** filed cells carried a claim.

**No seat claimed to recognise anything, anywhere in the round.** That is a result about the blind rather than a missing column: shuffled letters from a named seed, opaque refs, the spoken line only, and the key never near a judge. The registered sensitivity cut — the headline recomputed with every claimed cell dropped — therefore returns each arm's headline unchanged, and the scores file carries `panel_mean_with_disclosed_cells_dropped` on every arm so a reader can see that it does rather than take it on the sentence.

## 10. Context receipts

"Every arm answers identical bytes" is a claim about bytes SENT. Prompt-token counters, where a transport reports them, say what was received. Field median: **1819** prompt tokens; an arm below 0.8 × that publishes CONTEXT-TRUNCATED → UNRANKED. **0 arm(s) flagged.**

Arms whose transport reports no counters at all: `agent-claude-fable-5`, `agent-claude-opus-5`, `agent-claude-sonnet-5`. An arm with no counters reports none — the agent transport has none to report. That is an absence, not a truncation, and it is listed separately.

## 11. The bill

Four cost states, never collapsed (prereg §3.3): a figure · `$0.00*` plan-included · an em dash where no figure is held *and no tokens either* · an em dash **with** tokens where a metered call was measured and no price was ever cited.

| model | calls | spend (upper bound) | registered cap | receipts |
|---|---:|---:|---:|---:|
| `kimi-k3` | 31 | $0.9890 | $18.49 | 6 files |
| `gpt-5.5-2026-04-23` | 16 | $1.2757 | $20.00 | 5 files |

**Exhibit-wide metered total: $2.2647 across 47 calls, against $38.49 of registered caps.** That covers every metered receipt this repo holds — the arm legs, every judge lane *including the discarded round*, the auditions and the registration probes. A discarded measurement still costs real money and still prints.

Every other seat and arm, by registered state:

| cost state | arms | seats |
|---|---|---|
| `local-zero` | 3 — `local-gemma4-12b`, `local-gemma4-26b`, `local-qwen3.6-27b` | `gemma4-26b` |
| `metered` | 3 — `cloud-kimi-k3`, `openai-gpt-5.4-mini-2026-03-17`, `openai-gpt-5.5-2026-04-23` | `kimi-k3`, `openai-flagship` |
| `no-figure-held-and-no-tokens` | 3 — `agent-claude-fable-5`, `agent-claude-opus-5`, `agent-claude-sonnet-5` | `fable`, `opus` |
| `plan-included` | 11 — `cloud-deepseek-v4-flash-0731`, `cloud-deepseek-v4-pro-preview`, `cloud-gemma4-31b`, `cloud-glm-5.2`, `cloud-gpt-oss-120b`, `cloud-gpt-oss-20b`, `cloud-minimax-m3`, `cloud-mistral-large-3-675b`, `cloud-nemotron-3-nano-30b`, `cloud-nemotron-3-ultra`, `cloud-qwen3.5-397b` | `deepseek-v4-pro`, `mistral-large-3-675b` |

## 12. CURATED-HIGHLIGHTS — the shop window

The registered curation rule (prereg §11), applied. Six cards per act, chosen by rule, **and the rule prints**: the two highest family-means · the single lowest, a mandatory honest slot · fabricators, up to two · the best local incumbent · ties broken by roster order. The registers the shop window is built for are **humour/levity · vulnerability/openness · warmth · sad**, and the acts are sequenced so the kid opens, the gossip peek lands the rigour, and the warmth closes.

### S1

| slot | arm | family-mean | why this card, by the registered rule |
|---|---|---:|---|
| highest 1 | `agent-claude-fable-5` | 8.47 | the two highest family-means |
| highest 2 | `agent-claude-sonnet-5` | 8.43 | the two highest family-means |
| the honest slot | `cloud-nemotron-3-nano-30b` | 3.30 | the single lowest — mandatory |
| fabricator 1 | `agent-claude-opus-5` | 8.22 | 7 seat(s) filed `fabrication-accepted` |
| fabricator 2 | `cloud-glm-5.2` | 8.12 | 9 seat(s) filed `fabrication-accepted` |
| best local incumbent | `local-qwen3.6-27b` | 6.99 | the highest of the three local arms |

*Slot collision, printed rather than hidden:* `agent-claude-fable-5` (3 cell(s)), `agent-claude-sonnet-5` (2 cell(s)) also reached `fabrication-accepted` in this act and already hold a card above, so the fabricator slots pass to the next arms by roster order. The collision is the finding it looks like — the same replies the panel scored highest are among the ones it also called invented.

### S2

| slot | arm | family-mean | why this card, by the registered rule |
|---|---|---:|---|
| highest 1 | `agent-claude-fable-5` | 8.90 | the two highest family-means |
| highest 2 | `agent-claude-opus-5` | 8.80 | the two highest family-means |
| the honest slot | `cloud-nemotron-3-nano-30b` | 3.50 | the single lowest — mandatory |
| fabricator 1 | `cloud-glm-5.2` | 8.42 | 2 seat(s) filed `fabrication-accepted` |
| fabricator 2 | `openai-gpt-5.5-2026-04-23` | 8.30 | 1 seat(s) filed `fabrication-accepted` |
| best local incumbent | `local-gemma4-12b` | 7.30 | the highest of the three local arms |

*Slot collision, printed rather than hidden:* `agent-claude-fable-5` (2 cell(s)) also reached `fabrication-accepted` in this act and already hold a card above, so the fabricator slots pass to the next arms by roster order. The collision is the finding it looks like — the same replies the panel scored highest are among the ones it also called invented.

### S3

| slot | arm | family-mean | why this card, by the registered rule |
|---|---|---:|---|
| highest 1 | `agent-claude-opus-5` | 9.00 | the two highest family-means |
| highest 2 | `cloud-kimi-k3` | 8.55 | the two highest family-means |
| the honest slot | `cloud-nemotron-3-nano-30b` | 3.04 | the single lowest — mandatory |
| fabricator 1 | `cloud-minimax-m3` | 8.00 | 2 seat(s) filed `fabrication-accepted` |
| fabricator 2 | `agent-claude-fable-5` | 8.00 | 1 seat(s) filed `fabrication-accepted` |
| best local incumbent | `local-qwen3.6-27b` | 7.92 | the highest of the three local arms |

**The FABRICATED-never-the-beat law holds on this page** (prereg §9): a fabricating reply may print as evidence and never as an act's emotional beat or a pull quote. In the prereg's own words — *"The world that remembers beats the model that improvises, and the article never hands its best moment to the improviser."* That sentence is about the engine against improvisation, not one arm against another; the forbidden-claims list (§10) governs everything this page says about arms.

## 13. Limits, stated

- **Fable authored the exam, sits as an arm, holds two judge seats and writes the page.** Recusal cures one thing: the anthropic family never scores a Claude arm. It does not cure the fact that the questions were chosen by one of the contestants.
- **The reference world was driven by `qwen3.5:27b`, which sits no chair** — and is family-adjacent to two arms. Stated, not cured.
- **n is small and the page says so.** S1's asks are n=2 per arm; S2 and S3 are n=1. The tie band exists because of exactly this: at 0.5 on a 0–10 scale, 41 of 190 pooled arm pairs are inside it.
- **The panel is seven seats and six families, and one of the seven is the only locally-hosted one.** Section 6's axis is a statement about that seat.
- **Percentages are withheld under N=30** throughout, by rule, with the count and its denominator printed instead.
- **`SPLIT` is never rounded**, and `outside-canon-set` was registered before it was needed.

---

*Round `opencall-r2`. Scored 2026-08-14T15:31:03.865182+00:00. Carriage built 2026-08-14T15:19:12.918356+00:00. The discarded first round's calls are on the bill above and its verdicts are not scored anywhere.*
