{
 "round": "two-frontiers-cove-r1",
 "gate": "G-OUTSIDE-READ",
 "verdict": "PASS",
 "detail": "one full-size call to mistral-large-3:675b on the pre-registration's own 154,566 bytes with the registered prompt; the raw reply is filed (14,693 characters)",
 "started_utc": "2026-09-06T18:31:22Z",
 "finished_utc": "2026-09-06T18:32:15Z",
 "evidence": {
  "seat": "mistral-large-3-675b",
  "model": "mistral-large-3:675b",
  "transport": "ollama-cloud (prompted, think:true per the cloud law; no options)",
  "registered_prompt": "find every choice in this document that favours one arm",
  "prompt_file": "prereg/outside-read-prompt-cove.md",
  "read_document": "prereg/PREREG-COVE-LEG.md",
  "sent_utc": "2026-09-06T18:31:22Z",
  "http_status": 200,
  "ok": true,
  "latency_ms": 53264,
  "request_sha256": {
   "system": "c10e0691044b4c3d573b18eeb36e9cb0748396dfb52d1fe35af2ae0cd9793485",
   "user": "2103146f99661c4f6f52aa2b314e24556ffa0c70f5fba1b14136f7a57de0830a",
   "frame": "d16eb28cdad99a49d6a708358ca648689e4e045d4d206b93e8cf31114a5cce42",
   "prereg_bytes": "f9d9669f7812e2141aa37f8e3974fee7e4b19e0ac2d13e30294cba5364db4908"
  },
  "request_chars": {
   "prereg": 154566,
   "user_total": 155743
  },
  "counters": {
   "prompt_eval_count": 41570,
   "eval_count": 3735,
   "done_reason": "stop",
   "total_duration_ns": 52974416813,
   "eval_duration_ns": null,
   "prompt_eval_duration_ns": null,
   "load_duration_ns": null
  },
  "reply_sha256": "9198598cc5df6fa6ee3291f61a4b1f96d5d49bbe3e97745cebb3945c2afa74fc",
  "reply_chars": 14693,
  "thinking_chars": 0,
  "failure": null,
  "ran_after_the_round": {
   "registered_moment": "before the seal (§8)",
   "actual_moment": "after the round and before the page",
   "why_it_is_disclosed_rather_than_absorbed": "the probe exists to find the choices in this document that favour an arm, and every choice it can name was frozen before the first call — so a late read cannot change what it is reading. What a late read CAN change is the page, and that is the only thing this round lets it change: the instrument, the denominators, the panel and the band are closed. Each finding carries a disposition in §12."
  },
  "publication_screen": {
   "findings": 0,
   "rule": "the raw reply publishes in the kit, so it is screened here with the same allowlist-aware rule list the kit is screened with. A reply with a finding is filed as a sha and a length with the reason, never as text nobody screened."
  },
  "reply": "Here are the findings that favour one arm, grouped by section and rated by impact:\n\n---\n\n### §2 · The instrument, frozen\n**Finding 1**\n*Passage:*\n> **S1's two judged asks are n=2 per ask, and the within-arm spread does NOT publish as a figure.** Two samples buy two more independent moments in the cluster that carries the canon trap. But **NOT-RUN — the pinned scorer's `per_seat` projection carries no `sample` field**, so no per-sample figure exists in `scores.json` and none is typed by hand: `by_ask` pools both samples into one figure, there is no `by_sample`, and the only rows that carry `sample` are the top-level `recused_cells` dicts — which the two frontier arms do not have, because they have zero recused cells.\n\n*Arm favoured:* `local-gemma4-26b`\n*Why:* The local arm is the only one with recused cells (8), so its `recused_cells` dicts are the only ones that include `sample` information. This means the within-arm spread for the local arm is *implicitly* available in the published data (via the `recused_cells` dicts), while the frontier arms' spreads are explicitly suppressed. This creates an asymmetry in what a reader can compute from the published bytes, favouring the local arm by making its variability more accessible.\n*Rating:* **MATERIAL** (would change a table cell if the spread were published for all arms).\n\n---\n\n### §3 · The arms\n**Finding 2**\n*Passage:*\n> **The local seat's enforced JSON — the asymmetry named, and its sign refused.** Only the local seat's runtime actually enforces the reply schema; the three hosted arms were asked for it in words.\n\n*Arm favoured:* `local-gemma4-26b`\n*Why:* The local arm is the only one whose replies are guaranteed to parse as JSON, removing one failure mode (`NOT-COLLECTED — TRUNCATED`) from its row. The document refuses to speculate on whether this helps or hurts the model's actual output, but the asymmetry is named and the local arm is the only one with this advantage. This could subtly bias a reader toward the local arm by reducing its visible failure rate.\n*Rating:* **MATERIAL** (would change a reader's perception of reliability).\n\n---\n\n**Finding 3**\n*Passage:*\n> **(2) `production_port_amendment: \"A2\"` — an id this leg ADOPTS rather than mints, disclosed because it is another round's number.** [...] **the arm's latency cells are EMPTY** — a timing taken off a contended production instance is not a serving figure, and this leg publishes no latency it cannot stand behind — and **no `keep_alive` is sent**, so nothing this round does evicts or re-pins what that box is serving.\n\n*Arm favoured:* `local-gemma4-26b`\n*Why:* The local arm's latency cells are left empty, while the other three arms' latencies are published. This could imply the local arm is \"faster\" by omission, even though the document states the latency is not a serving figure. The asymmetry is disclosed, but the absence of data could still subtly favour the local arm in a reader's mind.\n*Rating:* **COSMETIC** (latency is not a scored metric, but the omission could imply an advantage).\n\n---\n\n**Finding 4**\n*Passage:*\n> **The sealed CLI arm runs on a DIFFERENT subscription account than exhibit forty's**, on the **same plan tier** (D-20260906-28). Same tier, different budget and different daily window; the account itself is never named or printed.\n\n*Arm favoured:* `cli-claude-fable-5-1`\n*Why:* The CLI arm's subscription account has a different daily window and budget, which could affect its performance (e.g., rate limits, quota availability). The document discloses this but does not quantify the impact, leaving open the possibility that the CLI arm's performance is influenced by its account's specific constraints. The other arms' accounts are not discussed, so this asymmetry could subtly favour the CLI arm by implying it has \"fresh\" resources.\n*Rating:* **MATERIAL** (could affect performance, but the impact is unmeasured).\n\n---\n\n**Finding 5**\n*Passage:*\n> **A Claude wrote the exam.** The open call's asks, rubric wording and scoring code were authored by a Claude-family model, one of whose descendants is an arm here. Under this panel **no Claude sits a chair**, so the reply-scoring half of August's disclosure retires and the half that survives is stated plainly and high: *a Claude no longer judges, but a Claude still wrote the questions, and recusal never cured that.*\n\n*Arm favoured:* `cli-claude-fable-5-1`\n*Why:* The exam was authored by a Claude-family model, which could subtly favour the CLI arm by giving it an advantage in understanding the \"style\" or \"intent\" of the questions. The document discloses this but does not quantify the impact, leaving open the possibility that the CLI arm benefits from this authorship asymmetry. The other arms have no such potential advantage.\n*Rating:* **MATERIAL** (could affect performance, but the impact is unmeasured).\n\n---\n\n### §4 · The panel\n**Finding 6**\n*Passage:*\n> **Recusal is by FAMILY, joined at scoring**, over sheets that never differ between judges. Every recused cell is kept and **prints with the family named**.\n\n| arm | family | recusing chair | seats scoring | families scoring | recused cells |\n|---|---|---|---|---|---|\n| `cli-claude-fable-5-1` | anthropic | none — no Claude chair | **7** | **7** | **0** |\n| `openai-gpt-6-astra` | openai | none — no OpenAI chair | **7** | **7** | **0** |\n| `cloud-glm-5-3` | zhipu | `glm-5.3` (its own family) | 6 | 6 | **8** |\n| `local-gemma4-26b` | google | `gemma4-31b` | 6 | 6 | **8** |\n\n*Arm favoured:* `cli-claude-fable-5-1` and `openai-gpt-6-astra`\n*Why:* The two frontier arms are the only ones scored by all 7 families, while the recused arms are scored by 6. This gives the frontier arms a larger denominator (56 cells vs. 48) and a more \"complete\" panel reading. The document discloses this, but the asymmetry could subtly favour the frontier arms by making their scores appear more robust.\n*Rating:* **MATERIAL** (would change a reader's perception of score reliability).\n\n---\n\n**Finding 7**\n*Passage:*\n> **THE ANCHOR — the calibration reply, from each bundle's own `provenance.reference_reply_verbatim`.** [...] **Disclosed adjacency, corrected:** the anchor's author is of the **alibaba** family and one seated seat is alibaba's, so on **six letters of every sheet-set — six of 38, not one** — the alibaba seat scores a reply written by its own family.\n\n*Arm favoured:* `cloud-glm-5-3`\n*Why:* The alibaba seat (which is the same family as the `cloud-glm-5-3` arm) scores the anchor on 6 of 38 letters, creating a potential adjacency bias. The document discloses this, but the adjacency could subtly favour the `cloud-glm-5-3` arm by making the alibaba seat's scores for that arm more consistent with its scores for the anchor.\n*Rating:* **COSMETIC** (the adjacency is disclosed, and the anchor does not enter the arm's mean).\n\n---\n\n### §6 · The scorer\n**Finding 8**\n*Passage:*\n> **THE COMPARABLE NUMBER, registered before the data.** The two frontier arms stand on 7 families; the glm and gemma arms on 6 — and on two **different** sixes. Arithmetic makes those denominators commensurable (that is what the family-mean-of-means ruling is *for*) but it does not make the family **sets** identical. So:\n>\n> - **Frontier arms vs the glm arm** → the **6-family common panel**: `leave_one_family_out[\"zhipu\"]` for the frontier arms, beside the glm arm's own headline.\n> - **Frontier arms vs the local seat** → the **6-family common panel**: `leave_one_family_out[\"google\"]` for the frontier arms, beside the gemma arm's own headline.\n> - **The glm arm vs the local seat** → the **5-family common panel**: `leave_one_family_out[\"google\"]` for the glm arm beside `leave_one_family_out[\"zhipu\"]` for the gemma arm.\n\n*Arm favoured:* `cli-claude-fable-5-1` and `openai-gpt-6-astra`\n*Why:* The comparable figures for the frontier arms are computed on a 6-family panel (by dropping one family), while the comparable figure for the `cloud-glm-5-3` vs. `local-gemma4-26b` pair is computed on a 5-family panel. This gives the frontier arms a larger and more consistent denominator in their comparisons, subtly favouring them by making their scores appear more robust.\n*Rating:* **MATERIAL** (would change a reader's perception of comparability).\n\n---\n\n**Finding 9**\n*Passage:*\n> **The 0.80 context rule, stated as the pinned code actually computes it.** [...] **The CLI row is excluded from the rule's field median as well as from its threshold, and the exclusion is a property of the transport rather than a setting:** the sealed CLI reports no `prompt_eval_count`, so `context_receipts()` files it under `arms_without_counters` and it enters neither `medians` nor `flags` — the pinned scorer has **no per-arm exemption and none is added.** **The CLI row therefore has no ratio and no 0.80 verdict**, and prints `NOT-APPLICABLE — transport` where the other three print a ratio.\n\n*Arm favoured:* `cli-claude-fable-5-1`\n*Why:* The CLI arm is excluded from the context rule's field median, which means its prompt token count does not influence the threshold for the other arms. If the CLI arm's prompt tokens were unusually high or low, this could affect the threshold for the other arms, but the CLI arm is shielded from this effect. The document discloses this, but the asymmetry could subtly favour the CLI arm by making its scores appear more stable.\n*Rating:* **MATERIAL** (could affect the threshold for other arms, but the impact is unmeasured).\n\n---\n\n### §9 · Spend, caps, and the four cost states\n**Finding 10**\n*Passage:*\n> **The four cost states never collapse into one column**: `metered` · `plan-included` (`$0.00*`, the asterisk load-bearing) · `own-silicon` · `no-figure-held`. A subscription figure is an estimate wearing a label, and it never prints as a receipt.\n\n*Arm favoured:* `local-gemma4-26b`\n*Why:* The local arm is the only one with `own-silicon` cost state, which could subtly imply it is \"cheaper\" or \"more efficient\" than the other arms. The document discloses this, but the asymmetry could favour the local arm in a reader's mind by making it appear more cost-effective.\n*Rating:* **COSMETIC** (cost is not a scored metric, but the framing could imply an advantage).\n\n---\n\n### §12 · Amendments (A2)\n**Finding 11**\n*Passage:*\n> **What §6 and §9 registered.** §6: *\"the sealed CLI reports no `prompt_eval_count`, so `context_receipts()` files it under `arms_without_counters` and it enters neither `medians` nor `flags` … The CLI row therefore has no ratio and no 0.80 verdict, and prints `NOT-APPLICABLE — transport` where the other three print a ratio.\"* §9's row for that arm: *\"no counter exists on this transport.\"*\n>\n> **What the round measured.** The transport reports **`input_tokens`, `cache_creation_input_tokens` and `cache_read_input_tokens` on every call**, and the runner sums those three under the one key the pinned scorer reads — the A8 prompt shape §6 itself registers, *\"summed the way the transport's OWN accounting sums it\"*. So **all four arms report a prompt figure**, and the CLI row has a ratio after all.\n\n*Arm favoured:* `cli-claude-fable-5-1`\n*Why:* The amendment reveals that the CLI arm *does* report a prompt token count, contradicting the earlier claim that it does not. This means the CLI arm's prompt tokens are now included in the context rule's field median, which could affect the threshold for all arms. The earlier exclusion could have subtly favoured the CLI arm by shielding it from the rule's effects, and the amendment removes this asymmetry.\n*Rating:* **MATERIAL** (changes the context rule's field median, affecting all arms).\n\n---\n\n### §12 · Amendments (A5)\n**Finding 12**\n*Passage:*\n> **THE DEFECT, named.** The hosted shelf serving `cloud-glm-5-3` returned the model's **chain of thought inline, in front of the envelope**, on every one of its eight cells, despite the `think: false` this round's shelf law registers [...] So that arm's `response_text` values are *reasoning, then the reply*: **6,780 to 23,899 characters** where the other three arms' emissions are **287 to 1,056**.\n\n*Arm favoured:* `cloud-glm-5-3`\n*Why:* The `cloud-glm-5-3` arm's replies included leaked reasoning (6,780–23,899 characters), which was initially stored as part of the reply. This could have subtly favoured the arm by making its replies appear longer or more detailed, even though the reasoning was later stripped. The document discloses this, but the initial inclusion of reasoning could have biased the panel's perception of the arm's output.\n*Rating:* **MATERIAL** (could have affected the panel's perception of reply length/quality).\n\n---\n\n### §12 · Amendments (A6)\n**Finding 13**\n*Passage:*\n> **THE DENOMINATORS AS THEY RAN, derived from the G-PANEL receipt rather than typed.** 6 sheets × 6 seats = **36 verdict files**; 6 seats × 38 letters = **228 cells filed**; **8 recused** cells (one chair × eight replies, once — not sixteen, because the *other* recusing chair is the seat that went dark and a chair that does not sit recuses nothing); **48 / 48 / 48 / 40** scoring cells and **6 / 6 / 6 / 5** families scoring, in the registered roster order.\n\n*Arm favoured:* `local-gemma4-26b`\n*Why:* The `local-gemma4-26b` arm is the only one with 5 families scoring (due to the lost `glm-5.3` seat), while the other three arms have 6. This gives the local arm a smaller denominator (40 cells vs. 48), which could subtly favour it by making its scores appear more concentrated. The document discloses this, but the asymmetry could make the local arm's scores look more robust.\n*Rating:* **MATERIAL** (would change a reader's perception of score reliability).\n\n---\n\n### Top 3 Findings to Fix First\n1. **Finding 5 (MATERIAL):** The exam was authored by a Claude-family model, which could subtly favour the `cli-claude-fable-5-1` arm. This is a structural bias that cannot be \"fixed\" retroactively, but it should be disclosed more prominently in the establishing section, with a clear statement that no arm was given an advantage in the scoring process.\n2. **Finding 6 (MATERIAL):** The two frontier arms are scored by 7 families, while the recused arms are scored by 6. This asymmetry should be addressed by either:\n   - Disclosing that the frontier arms' scores are more \"complete\" and may appear more robust, or\n   - Computing all comparable figures on the 6-family panel (by dropping one family from the frontier arms) to ensure consistency.\n3. **Finding 11 (MATERIAL):** The CLI arm's prompt tokens were initially excluded from the context rule's field median, which could have affected the threshold for all arms. The amendment corrected this, but the initial exclusion should be disclosed as a potential source of bias in the establishing section."
 },
 "arms": {},
 "records": [
  {
   "label": "outside-read",
   "prompt_tokens": 41570,
   "completion_tokens": 3735,
   "cost_state": "plan-included",
   "shows_as": "$0.00*"
  }
 ],
 "withheld": {
  "rule": "a receipt publishes in the kit, so these fields are removed here rather than at kit time — and named, because a field quietly absent from a receipt is indistinguishable from a gate that never read it."
 },
 "note": "written by harness/cove/gates_cove.py. Every generation call rides the same send() the scored run uses; a gate that hand-rolled its own call would receipt a field the run never sent."
}
