# CRITIQUE — the cold-scroll read (v3)

**Target:** (a path inside the round's repo, described) (522 lines, dateline "Draft v3 · 2026-09-05 (UTC)"). All line numbers are that file's.
**Lens:** THE COLD-SCROLL LAW. Each of the 15 sections was read in isolation, in shuffled order — The bill → The filing cabinet → The ledger → Head-to-head → What this page does not say → The rules desk → What the numbers are allowed to mean → the closing four → the opening four last — with no memory of any other section. The bar: a reader who lands mid-page must get **(1) what is being counted, in plain words, (2) the window, dated at both ends in UTC, (3) how it compares to prior readings, in one sentence** — plus, for a bench page, **whose answers** and **against what bar**.
**Method note:** HTML `<!-- fill … -->` comments were stripped first (`perl -0pe 's/<!--.*?-->//gs'`), so every "defined nowhere" claim below is a count over what a reader can see; comment stripping preserved line numbering, so the counts and the quotes share the file's own numbers. Every arithmetic claim was computed (`python3`); every date difference was computed with `datetime.date`. Artifact files were read only to confirm which of two irreconcilable printed numbers the scorer actually wrote (`results/legA/scores.json`, `results/legH/pairwise.json`); nothing outside this file was modified.

---

## VERDICT IN ONE LINE

v3 is a different page from v1 on this lens: the run window is now stamped under all five data headings, the glossary at "The words this page leans on" closes almost the whole of v1's undefined-vocabulary table, every floor now rides its gate row with a `reads` verdict beside it, and the head-to-head's three denominators reconcile to the cell. What is left is smaller and sharper: **three sections still have no window at all** (What this page does not say · What to take with you · How to check our work), **the ledger section has no lead sentence of any kind**, **`cell` — the page's universal unit word, 36 occurrences carrying at least five different units — is defined nowhere**, and three numbers new in v3 do not survive a cold reader's arithmetic: "102 readings … three per case" against a 36-case class (line 182), "30/34" against a floor written "at least 30/36" (407–409), and "eleven days after the exam was sealed" for 2026-08-15 → 2026-08-27, which is twelve (line 15). One v3 fix broke something: the new six-part conflict list orphaned "That reader was `mistral-large-3:675b`" so that it now attaches to the sentence about Anthropic editorial readers (line 25).

---

## SECTION BY SECTION (in the order read)

### 1. The bill (443–480) — **PARTIAL**

- **(1) what is counted — PARTIAL.** The lead exists now, and it is a real one, but its subject is a gerund: *"This is what the measuring cost, in tokens and in dollars, for every model that answered or judged in the window"* (451). "The measuring" is never unpacked: a reader who lands on the bill learns the dollars, the cost states, the cap and the caching, and never that this was two exams about board-game rules and a planted-sentence test.
- **(2) window — PASS.** *"*Measured 2026-09-05 14:13:33–18:15:00 UTC.*"* (445).
- **(3) prior readings — FAIL.** No sentence compares this bill to any earlier one on the shelf. It is the one section where the comparison would be nearly free ("the judges cost less than the contestants").
- **whose answers — PASS** ("the arms under test first, the judging seats below", split into two captioned tables). **against what bar — PASS**: *"against a registered cap of $60.00 (A $25.00 · C $20.00 · probes $5.00 · reserve $10.00)"*.
- **Arithmetic, all checked and all clean:** $17.79 + $5.25 = $23.04 ✓; $25 + $20 + $5 + $10 = $60 ✓; Astra 1,012,938 × $10/M + 559,416 × $1/M + 142,086 × $50/M = $17.79 ✓ and 1,012,938 + 559,416 = 1,572,354 ✓, uncached-counterfactual $22.83 ✓; kimi 629,545 × $3/M + 224,411 × $15/M = $5.25 ✓, $0.25 over $5.00 ✓; "6 of the 10 rows below are plan-included" ✓ (3 arms + 7 seats, 6 plan-included).
- **Still missing:** a call count. The bill prices 2.1M and 1.6M tokens and never says over how many calls; "235 recorded calls" / "242 recorded calls" live in *another* section (433).
- One-sentence lead that would fix leg 1 and leg 3, if the section keeps its own: see §"THREE LEADS" below is reserved for worse offenders — for the bill, insert after 451: "It is the first bill on this shelf where the seven judges cost less than the two contestants, and the only one where a subscription row cannot show a receipt."

### 2. The filing cabinet (299–360) — **PARTIAL**

- **(1) what is counted — PASS on the question, FAIL on the ordinal.** *"The second job asks a simpler question than the rules desk: can a model find one sentence it has never seen, buried in a long stretch of century-old card-game prose, and can it say 'the text does not say' when the sentence was never planted at all?"* (305). Recall, abstention and fabrication are each defined in that same paragraph — the v1 failure is gone. But "**The second** job" has no antecedent in this section; the first job is named only in the opening.
- **(2) window — PASS** (301).
- **(3) prior readings — FAIL.** Nothing compares these counts to any earlier filing-cabinet reading, though the section's own note says the harness "is the descendant of three earlier ones on this shelf" (510, another section).
- **whose answers — PASS.** **against what bar — PARTIAL:** the only bars in the section are the 0.80 context floor and *"Differences under 2 items read TIED"* (341); the recall and abstention halves are printed against no floor at all, and the section never says they were registered without one.
- **Arithmetic, checked:** 18 + 18 = 36 items ✓; 8k/16k/32k at 32,000/64,000/120,000 characters at 4 chars per token ✓; 18 recall items over 3 tiers = 6/6 per tier ✓; two items per depth × 3 tiers × 3 depths = 18 ✓; local seat 12 collected + 6 NOT-COLLECTED = 18 per half, 12 cells total ✓, matching "items 24" in the ratio table ✓ and "COLLECTED 24 · NOT-COLLECTED — CONTEXT 12" in the ledger ✓ — v1's three-way mismatch is fully closed; canary 16,387 ÷ 30,113 = 0.5442 ✓.
- **One cell a cold reader will misread:** *"| `local-gemma4-26b` | 32k | NOT-COLLECTED — CONTEXT (12 items) | …"* (337). "(12 items)" sits in the **recall** column, where every other cell is out of 6; the 12 counts both halves of the tier.

### 3. The ledger, and what left our machines (370–424) — **FAIL**

- **(1) what is counted — FAIL. The section has no lead sentence.** Heading → stamp → a bold label and a table: *"**What left our machines, and where.**"* (376). A reader who lands here is never told what round produced these gates, what a gate is, what a floor is, or what any of G2, G4, G5c, G6b, G6c counts — the only gloss on the page for those lives in the glossary and the rules desk. Bracket intervals print at 407–409 (`[0.734, 0.953]`, `[0.898, 1.0]`) with no key in this section.
- **(2) window — PASS** (372).
- **(3) prior readings — FAIL.**
- **against what bar — PARTIAL and self-contradicting.** The G2 bullets quote their floor and then print a reading against a different denominator: *"at or above the per-case recall floor on 30/34 [0.734, 0.953] — the floors as registered: FLOOR: median house-set recall ≥ 0.70 across the 36 cases, AND ≥ 0.50 on at least 30/36."* (407). No `reads` verdict follows: the three G2 bullets say `SCORED, within-arm` and never "cleared" or "missed", so the one gate a reader could check here is the one gate with no verdict.
- **Fix (one-sentence lead, written in full in §"THREE LEADS" below).**

### 4. Head-to-head (222–298) — **PARTIAL** (close to PASS)

- **(1) what is counted — PASS on substance, FAIL on the first four words.** *"Everything above scores each model against a fixed bar."* (228) is a cross-reference a cold reader does not have; the sentence that follows it is exactly right and would open the section better on its own: "This is the only place on the page where the two frontier models are compared to each other, and it is a vote, not a score: seven outside families read pairs of rules-desk answers blind…".
- **(2) window — PASS** (224).
- **(3) prior readings — FAIL.** No comparison to the open call's panel or to any earlier blind read.
- **against what bar — PARTIAL.** The panel's own floor ("a floor of 4" families) is printed only in the ledger, 130 lines away (420).
- **Arithmetic, all checked and all clean — this is the section v1 hammered and v3 rebuilt:** 7 × 36 × 2 = 504 ✓; 449 + 55 = 504 ✓; deepseek 36 × 2 = 72 asked, 8 cases × 2 + 1 = 17 collected, 72 − 17 = 55 not carried ✓; 216 + 8 = 224 both-order comparisons ✓; 23 + 13 = 36 ✓ and the per-case column really does hold 23 `cli-claude-fable-5-1` rows ✓; the judges column sums to 225 = 36 × 6 + 9 ✓; every seat rate = sum ÷ cases ✓ (21.25/36 = 0.59, 19.5/36 = 0.542, 16.5/36 = 0.458, 18/36 = 0.5, 21/36 = 0.583); 4/9 correctly prints `NO-RATE`.
- **Two things a cold reader cannot close.** (a) Summing the seat table's own column gives 121.5 over 225 observations = **0.540**, not the headline **0.544**; the page states the two-stage rule (*"That is a mean of the per-case scores with ties counted as half"*, 234) but never warns that the seat table cannot reproduce the headline. (b) The retired seat's nine carried cases include `ofl-ans-0013` (per-case table, 260), which the same section says was **added after judging began** — *"amendment A11 put them back in under the same blind seed, judged by every seat after its main pass"* (238) — while retirement is *"for the rest of the job"* (232). Confirmed against `results/legH/pairwise.json` (`cases_carried: 9`, `lost_at_sheet: "ofl-ans-0008.o1"`, A11 ids `ofl-ans-0007`, `ofl-ans-0013`): the numbers are right, the ordering is unexplained.
- **The section's own headline is qualified in another section.** The blindness sensitivity cut — *"Drop every one of those 70 cells and recompute the headline: 0.553 … interval [0.452, 0.651], against 0.544 with them in"* — is printed at 220, inside the rules desk. `grep -n 'sensitivity'` over the visible text returns that one line: a cold reader of Head-to-head never learns the headline was recomputed.

### 5. What this page does not say (425–442) — **FAIL**

- **(1) what is counted — FAIL.** The heading is the only frame; the first words are *"**The two frontier models were reached by different roads.**"* (427). The section carries measured figures — a 1,480-character preamble at "about 370 tokens", "16 of 16 cells collected, PASS", "exactly 280 input tokens on every pair", "92,960 reasoning tokens over its 235 recorded calls" — and never says what was measured, on what, or against what.
- **(2) window — FAIL. No date at either end.** `grep -n 'Measured 2026'` returns 62, 224, 301, 372, 445 — not this section. Its only dates are *"Since 2026-08-16"* (439) and the vendors' cutoffs referred to but not printed (*"after both stated cutoffs"*, 437).
- **(3) prior readings — FAIL.**
- **A restated figure still not restated** (v1 finding, unchanged in substance): *"the local seat's count on those six cases is therefore not a live vulnerability, and it is printed above as the pre-fix number it is"* (439). The number (2 of 6) is not here.
- **Arithmetic clean:** 1,480 ÷ 4 = 370 ✓; 8 cases × 2 conditions = 16 cells ✓; 92,960 reasoning tokens matches the bill's "142,086 out (incl. 92,960 reasoning)" ✓.
- **A v1 phantom is fixed:** *"the self-agreement rows in the gate cards above print what that looks like"* (427) — the G6a rows now exist in all three gate cards, so the reference resolves.

### 6. The rules desk (60–221) — **PASS** (the page's model section, with three number defects)

- **(1) what is counted — PASS, and it is the best lead on the page.** *"Sixty questions about board games, each with the rulebook passages RuleSage itself retrieved, put to three models three times. The bar is the product's own contract: cite what you use, say plainly when the book does not say, and never follow an instruction that arrived inside a source."* (66). The product is named, the class composition is printed (36/12/6/6 = 60 ✓), the reps are named, and the census is promised in the lead.
- **(2) window — PASS** (62). **whose answers — PASS.** **against what bar — PASS:** every gate card carries `floor (2026-08-15)` and a `reads` column, and the prose says the verdict out loud: *"**No arm cleared the abstention floor.** It asks for at least 11 of 12, and the readings were GPT-6 Astra 7, Claude Fable 5.1 7, the local seat 10"* (156). This was v1's single most consequential failure; it is gone.
- **(3) prior readings — PARTIAL.** One comparison exists, 100 lines in and inside a parenthetical clause: *"in the August seat gate this design document was written for, frozen 2026-08-15, the same seat read 13 of 36"* (160).
- **Defect A — an arithmetic contradiction new in v3.** *"it read SCORED, mean house recall 0.983 over 102 readings of the answered cases, three per case"* (182). Three per case over a class the same section calls *"36 that a rulebook answers"* (66) is 108, not 102; 102 implies 34 cases, and nothing on the page says six readings are missing. (The scorer's field is literally `G_CALIBRATE.cases = 102`; the "three per case" gloss is the page's own, and it is what makes the gap visible.)
- **Defect B — 19 calls, 18 cells, in one sentence.** *"On every corrupt-corpus call — all 6 cases, 19 of 19 calls (the 18 cells re-dispatched, plus the first attempt's one) …"* then *"its 18 cells read NOT-COLLECTED — MODEL-FALLBACK in the census below, out of the 180 calls it was asked"* (164). The census reconciles perfectly (162 + 18 = 180 ✓, 54 × 3 = 162 ✓); the 19th call is asserted and never placed in any column.
- **Defect C — an undated incident, unchanged from v1.** *"One bench in August crossed that line before the promise page pointed it at bench data"* (70) — no date at either end, on the page whose fourth standing rule is that everything is dated at both ends. The promise page is at least linked now.
- **Structural placement:** this section hosts two things that belong elsewhere for a cold scroll — the recognition-claims table (214–218), which covers both legs, and the head-to-head's sensitivity recompute (220).

### 7. What the numbers are allowed to mean (361–369) — **PARTIAL**

- **(1) what is counted — FAIL on its own opening definites** (v1 finding, unchanged): *"Both frontier rows are a new kind of row for this shelf of pages, and four standing rules for that kind were written before the round ran."* (363) — there is no table in this section, so "both frontier rows" points at nothing a cold reader can see; and *"the small model beside them is the one that serves"* is a spatial reference to a table that is not here, with the model still unnamed. v1's four-promised/three-delivered confusion **is** fixed: "four standing rules" and *"Three reading rules ride every table"* (368) are now distinguished.
- **(2) window — PASS**, inline and at both ends: *"beyond the window 2026-09-05 14:13:33–2026-09-05 18:15:00 UTC"*.
- **(3) prior readings — n/a** for a rules section; the August bar is now dated exactly (*"registered on 2026-08-15, by other hands"*), fixing v1's bare "August".
- The `no interval: N < 30` token is now quoted in its own words ✓ (368), fixing a v1 finding.

### 8. What to take with you (481–492) — **PARTIAL**

- **(1) what is counted — PASS per bullet.** Each bullet now prints its floor beside its count: *"the floor asks for 11 of 12 abstentions and the readings were GPT-6 Astra 7, Claude Fable 5.1 7, the local seat 10 — no arm cleared it"* (483). v1's "shows a failure and reads as a result" is fixed outright.
- **(2) window — FAIL.** The section's only date is the bar's: *"The rules-desk cases and their pass marks froze on 2026-08-15; both models arrived later."* (486). The most-quotable section on the page never dates the run it summarises.
- **(3) prior readings — FAIL.**
- **The strongest product claim still names no model:** *"**The small model beside them is the one that serves.** Neither frontier model is eligible for any seat in our products; this page measures, it does not hire."* (487) — Gemma 4 (26B) appears nowhere in this section.

### 9. How to check our work — and see it live (493–501) — **PARTIAL**

- **(1) — PASS. Every door now has a handle**, which was v1's whole complaint: *"Every table here but one is a projection of a file in [the kit](data/) beside this page, and [`index.json`](data/index.json) lists each file with its sha256"* (495); *"RuleSage is live at [rulesage-live.strata2signal.com](https://rulesage-live.strata2signal.com/)"* (496); *"It is in the kit as [`prereg.md`](data/prereg.md)"* (498). The withheld-file classes are named and counted.
- **(2) window — FAIL.** No date anywhere in the section; a reader arriving from a search cannot tell which round's kit `data/` holds.
- **(3) — n/a.**

### 10. Who ran this, and thanks (502–515) — **PARTIAL**

- **(1) — PASS for its own subject**, and the v1 self-contradiction is gone: *"**A hostile reader has been through version 1 of this page, not this one.**"* (512), with real counts (*"findings raised 59 · must-fixes 10 · must-fixes landed 10"*) and *"A fresh pass on the version that releases is owed, and its counts replace this sentence before any release."*
- **(2) window — FAIL** for the run; the hostile read is dated (*"On 2026-09-05"*) and the CLI version is receipted (2.1.261).
- **Identifier drift remains** (below, table A note): "Mistral Large 3" beside `mistral-large-3:675b` in the same section a reader just left the tables of.

### 11. The rest of the seminar (516–522) — **PARTIAL**

- **(1) — PASS now.** Every companion is linked **and** glossed inline: *"[*The Same Sixteen*](/the-same-sixteen/), where the class of row these two frontier models now belong to — a reference arm, measured and never seated — was chartered"* (518). v1's five unglossed house terms are gone.
- **(2) window — FAIL**, and the forward promise is still undated by construction: *"as its own page, dated."* (519). *"The shelf holds every kit"* (520) is the one door left without a link.

### 12. Two strangers at the door (9–18) — **PARTIAL** (was PASS in v1 on everything but two units)

- **(1) — PASS**, and both v1 unit defects are fixed: the small model is named (*"That small model is Gemma 4, the 26-billion-parameter one, running on our own hardware"*, 11) and the tiers are converted (*"eight, sixteen, or thirty thousand tokens long in the three tiers — roughly six to twenty-two thousand words"*, 13).
- **(2) window — PARTIAL.** *"Two jobs, measured in one four-hour window on Saturday 2026-09-05."* (13) gives a day and a duration, not both ends; the clock lives two sections later.
- **A new off-by-one:** *"is dated 2026-08-27, eleven days after the exam was sealed"* (15) against *"On 2026-08-15 the rules desk's sixty cases and their floors were sealed"* (58). 2026-08-15 → 2026-08-27 is **twelve** days (computed); eleven is the distance from the 16th, which is the lint date, not the seal.

### 13. Who judged, and who wrote this (19–30) — **PARTIAL**

- **(1) — PASS**, and the conflict is stated harder than v1's: *"**the agents that designed, built, ran, audited, and wrote this page run on one of the two vendors' models.**"* (25), now with six named countermeasures including the editorial readers.
- **(2) window — FAIL** for the run; the outside read is stamped (*"at 2026-09-05 04:39Z"*).
- **A defect the fix created — the page's worst remaining sentence-level fault.** v3 inserted a sixth part — *"the editorial readers who critiqued each draft of this page — a stranger, a cold scroll, a house-voice reader, a hostile reader, a numbers reader — were Anthropic models too"* — **between** the outside-reader sentence and the sentence that names it. The result: *"That reader was `mistral-large-3:675b`"* (25) now sits immediately after a clause about Anthropic editorial readers, so cold (and warm) it reads as naming one of them. In v2 (line 39, "Five parts") that sentence directly followed its antecedent. Fix: "The outside family that read the pre-registration was `mistral-large-3:675b`; at 2026-09-05 04:39Z it was sent…".

### 14. The words this page leans on (31–55) — **PASS**

- A cold reader landing here gets what they came for: 22 glossary rows covering arm/road/transport, the local seat and the house, seat and shelf, Leg A/H/C, gate/floor/the G-numbers, the answer fence, rep and `any-rep read`, abstention, `COLLECTED`/`NOT-COLLECTED` with all eight reasons, `SCORED`/`NOT-RUN`/`NOT-APPLICABLE`/`PASS`, `GROUNDED`/`NOT-CLASSIFIED`/`NO-RATE`, the brackets, `TIED`, `no-mode cases`, sheet/needle/canary, the kit and its screen, `PREREG`, the cost states, exhibit/seminar/cove, an operator. This table is the single biggest reason v3 passes where v1 failed: *"`[0.819, 0.985]` after a count is a 95% Wilson interval on that count's own denominator; the head-to-head's interval is a 95% cluster bootstrap over cases and is labelled as one."* (46).
- Two notes, not failures: the section itself never says what page it belongs to or when the round ran (a reader arriving from a search for `any-rep read` gets the definition and no round), and four load-bearing words are missing from it — `cell`, `panel floor`, `the roster`, `within-arm` (see A).

### 15. Three weeks earlier (56–59) — **PASS**

- Dates every event at both ends and closes with the run's own clock: *"On 2026-09-05, a little before two in the morning UTC, an operator read this round's pre-registration and co-signed it; at 14:03 UTC the roster was registered; and between 14:13:33 and 18:15:00 UTC every call on this page was made."* (58). The only cold-scroll gap is that "a little before two in the morning UTC" is the one time on the page given in words rather than a stamp.

---

## CROSS-CUTTING FINDINGS

### A. Vocabulary that appears only as a table value or column header and is defined nowhere in reader-visible text
Counts are `grep -o` over the comment-stripped text; the glossary at "The words this page leans on" (31–55) was checked row by row for each.

| token | visible occurrences | where it appears | in the glossary? |
|---|---|---|---|
| `cell` / `cells` | 36 (16 singular, 20 plural) | every data section; 5 table headers or captions | **no — and it carries five different units** |
| `panel floor` | 1 | ledger bullet (420) | **no** (its floor value, 4, is printed inline) |
| `the roster` | 3 | 58 (registered at 14:03 UTC), 457 (`pricing_source`), 475 (seat caps) | **no**, and it is never located in the kit's visible list |
| `within-arm` | 7 | ledger G2 bullets as a state (`SCORED, within-arm`), rules-desk prose | partial — "G2 is a within-arm check" names it, never defines it |
| `CONTEXT-TRUNCATED` | 1 | filing cabinet (349) | **no** (`NOT-COLLECTED — CONTEXT` is) |
| `delta calls` | 1 | egress `leg` cell: "Leg A × 3 reps + 16 delta calls" (382) | **no** — the 16 are the scaffolding probe, explained at 427 |
| `no-figure-held` · `metered` · `plan-included` · `own-silicon` | 4 states, 10 rows | bill `cost state` column | **yes** (52) — fixed since v1 |
| `NOT-COLLECTED — MODEL-FALLBACK` | 5 | gate card (131), census header (172), prose (164) | **yes** (43) — fixed |
| `any-rep read` | 8 | gate cards, prose, takeaway | **yes** (41) — fixed |
| `GROUNDED` | 2 (case-sensitive) | glossary + G4 header (192) | **yes** (45) — fixed |
| `NOT-CLASSIFIED` | 2 | glossary + cabinet grid header (319) | **yes** (45) — fixed |
| `no-mode cases` | 3 | census header + two glosses | **yes** (48) and again at 168 — fixed |
| `NO-RATE` | 4 | per-family table (291), glossary | **yes** (45, 46) — fixed |
| `TIED` | 5 | cabinet (341), takeaway (485) | **yes** (47) — fixed |
| `no interval: N < 30` | 7 | 3 sections | **yes** (46, 368) — fixed |
| bracket intervals `[lo, hi]` | 20+ | 6 sections | **yes** (46, 150) — fixed; **but no key in the ledger section, where they also print** |
| `PREREG` §/A-numbers | 15 | 6 sections | **yes** (51) — fixed |
| `G1…G6c`, `G-PEN`, `G-EGRESS`, `G-ID` | many | 4 sections | **yes** (39) — fixed |
| `tie_band` / band language | **0** | — | n/a — the fill emitted none; the v1 warning did not come true |

`cell` is the finding that matters. It means, without ever saying so: one judge's reading of one sheet in one order (*"449 of 504 cells collected"*, 232); one (arm, case, rep) call (*"its 18 cells read NOT-COLLECTED"*, 164); one (family, case) judgement (*"recused cells 34"*, 196; *"cells judged 251"*, 216); one (arm, item) call (*"cells under the 0.80 context floor"*, 319; *"12 cells in the widest tier"*, 485); and one (case, condition) probe call (*"16 of 16 cells collected"*, 427). It needs a glossary row of its own, naming the unit as "whatever the table's row is asked for, once".

### B. Pronouns and definites whose antecedent lives in another section

| the text | line | where its antecedent is |
|---|---|---|
| "**Everything above** scores each model against a fixed bar." | 228 | the rules desk, 160 lines up |
| "**The second job** asks a simpler question than the rules desk" | 305 | "Two jobs" — the opening, 292 lines up |
| "it is **printed above** as the pre-fix number it is" | 439 | the rules desk (184); the figure, 2 of 6, is not restated |
| "the self-agreement rows in the gate cards **above**" | 427 | the three gate cards (118, 132, 146) — resolves, but only there |
| "named **above** with their versions" | 506 | "Who judged, and who wrote this" (21) |
| "**Both frontier rows**" · "**the small model beside them**" · "**this shelf of pages**" | 363 | no table in the section; "shelf of pages" is disambiguated in the glossary |
| "**The small model beside them** is the one that serves" | 487 | Gemma 4 (26B), named at 11 and 508 |
| "**the kit**" (22 occurrences, 8 sections) | — | located exactly once, at 495 |
| "16 **delta calls**" | 382 | the scaffolding-delta probe (427) |
| "the 26 canonical asks + the 34 substituted queries" | 385 | the rules desk (70) |
| "the shelf gemma seat" | 384 | `gemma4-31b`, named at 21 |
| "**the roster**" | 457, 475 | nowhere — registered at 14:03 UTC (58), never described |
| "the sensitivity cut" (the head-to-head's own qualifier) | 220 | printed in the rules desk, not in Head-to-head |

### C. Every table: headers, and whether it carries a date or a caption naming one row
17 tables, all reader-visible. "Caption" = a sentence within two lines of the table that names what one row is.

| # | table | line | headers self-explaining? | caption naming a row | date inside the table |
|---|---|---|---|---|---|
| 1 | glossary | 33 | yes | yes (the heading) | n/a |
| 2–4 | gate cards ×3 | 110, 124, 138 | yes — `gate · what it counts · reading · floor (2026-08-15) · reads` | yes, 102: "Each card below is one arm" | **yes** — the floor's date (not the run's) |
| 5 | collection census | 172 | yes; states glossed at 43; `asked` carries the denominator | yes, 168 | no |
| 6 | G4 groundedness | 192 | yes — `GROUNDED` glossed at 45, `recused cells` at 188/206 | yes, 188 | no |
| 7 | recognition claims | 214 | yes | yes, 208/212 | no |
| 8 | per-case | 246 | yes | yes, 244: "One row per case" | no |
| 9 | per-family | 289 | yes — headers now spell out the arithmetic | yes, 287 | no |
| 10 | cabinet grid | 319 | mostly; `cells under the 0.80 context floor` is explained 30 lines later (349) | partial, 317 | no |
| 11 | by tier | 327 | yes — tiers converted at 311 | yes, 325 | no |
| 12 | context ratio | 351 | yes | yes, 349 | no |
| 13 | egress matrix | 380 | `leg` cells are compound ids ("Leg A × 3 reps + 16 delta calls", "Leg A G4 + Leg H sheets") | **no** — only the label at 376 | no |
| 14 | socket sample | 391 | yes | yes, 389 (the best caption on the page) | **no clock time** — the receipt stamp is only in a comment |
| 15 | bill · arms | 455 | yes | yes, 451/453 | no |
| 16 | bill · seats | 463 | yes | yes, 461 | no |

**5 of 15 sections carry the window stamp; 0 of 16 data tables carry the run window inside them; 15 of 16 now carry a caption naming what one row is** (v1: zero). The one exception is the egress matrix.

### D. Numbers a cold reader cannot reconcile (every one computed)

1. **102 vs 108.** *"mean house recall 0.983 over 102 readings of the answered cases, three per case"* (182) against *"36 that a rulebook answers"* (66). 36 × 3 = 108. Six readings unaccounted; 102 implies 34 cases. **New in v3.**
2. **30/34 against a floor of 30/36.** *"at or above the per-case recall floor on 30/34 [0.734, 0.953] — … AND ≥ 0.50 on at least 30/36."* (407, 408; and 34/34 at 409). Two cases leave the denominator with no note, in the section that prints no verdict for the gate. **New in v3.**
3. **Twelve days printed as eleven.** *"2026-08-27, eleven days after the exam was sealed"* (15) vs the seal on 2026-08-15 (58): 12 days. **New in v3.**
4. **0/54 beside two 0/60s, unannotated in the ledger.** *"G6b, Claude Fable 5.1: **SCORED** — 0 of 54 cases stopped by a cap"* (412) — the gate card annotates the same figure (*"0/54 (6 cases had no collected reply — see below)"*, 134); the ledger bullet does not.
5. **19 calls, 18 cells.** *"all 6 cases, 19 of 19 calls (the 18 cells re-dispatched, plus the first attempt's one)"* (164); no column anywhere holds the 19th.
6. **The seat table cannot reproduce the headline.** Summing the per-family column (4 + 21.25 + 21.25 + 19.5 + 16.5 + 18 + 21 = 121.5) over 225 observations gives 0.540 against the printed **0.544** (234). The two-stage rule is stated; the trap is not flagged.
7. **A retired seat judging a case added after its main pass.** `ofl-ans-0013` carries 7 judges (260) though the seat retired at `ofl-ans-0008.o1` (232) and A11 cases were *"judged by every seat after its main pass"* (238).
8. **"0 under 0.8" beside a ratio of 0.5442.** The grid prints `0` cells under the context floor for the local seat (323) while the prose reports *"a ratio of 0.5442, under the registered 0.8 floor"* (357); the reconciliation (a pre-scoring canary, and 12 cells never called) is real but split across eight lines.
9. **"(12 items)" in a 6-item column.** The 32k `NOT-COLLECTED — CONTEXT (12 items)` cell sits under `recall` (337).
10. **"case 61" in a sixty-case bank** (410) against *"the rules desk's sixty cases"* (58) — explained as a printed gap, but the count is never squared.
11. **"its million and a half characters"** (504) for a corpus the page measures at *"1,808,001 characters"* (313) — 1.8 million.

### E. Every place a date or window is missing

- **Sections with no date at all:** How to check our work (493–501) · The rest of the seminar (516–522) · The words this page leans on (31–55, the bar's date only).
- **Sections carrying measured numbers with no window:** **What this page does not say** (425–442 — a preamble measurement, a token census and two reasoning-token counts, undated) · **What to take with you** (481–492 — every headline count on the page, dated only by the 2026-08-15 freeze).
- **Single-ended or vague:** *"One bench in August crossed that line"* (70) · *"a little before two in the morning UTC"* (58) · *"during the first minute of the two hosted arms' Leg A runs"* (389, the socket sample's only time) · *"as its own page, dated"* (519) · Two strangers at the door: *"one four-hour window on Saturday 2026-09-05"* (13, no clock) · Who ran this and thanks (502–515, no run window).
- **Passing:** the five data sections' stamps (62, 224, 301, 372, 445), What the numbers are allowed to mean (363, inline at both ends), Three weeks earlier (58, both ends to the second).

### F. Structural defects to fix before any pour (not style)

1. **"That reader was `mistral-large-3:675b`" (25) now attaches to the Anthropic editorial readers** — the new sixth countermeasure was inserted between the definite and its antecedent. A conflict-of-interest disclosure that names the wrong party is the one class of error this page cannot ship.
2. **The ledger section (370–424) has no lead sentence** — a table opens it, and every gate id, floor, state and bracket in it is glossed only elsewhere.
3. **Two sections carrying measured numbers carry no window** — "What this page does not say" and "What to take with you"; the second is the page's most quotable.
4. **Three new arithmetic contradictions** (102 vs 108; 30/34 vs 30/36; eleven vs twelve days) — all three are subtractions a cold reader will do.
5. **The head-to-head's own sensitivity recompute (0.553, [0.452, 0.651]) is printed in the rules desk** (220), so the section that owns the headline never qualifies it, and the section that qualifies it owns nothing of it.
6. **`cell` has no definition** while doing five jobs across every data table.
7. **Identifier drift, unfixed from v1:** `glm-5.3` (×4: 293, 202–204) vs `glm-5-3` (×1: 470); `qwen3.5-397b` (×4) vs `qwen3-5-397b` (×1: 471); `mistral-large-3-675b` (×5, tables) vs `mistral-large-3:675b` (×3, prose at 25, 70, 512). The bill is the only place the dashed spellings appear, so the bill is the table a reader cannot join.
8. **The egress matrix has no caption and compound `leg` ids** (380–387).

---

## v1 FINDINGS: FIXED / REMAIN / MADE WORSE BY THE FIX

### FIXED (quoting v3)

1. **The window, once on the whole page, is now stamped under all five data headings** — *"*Measured 2026-09-05 14:13:33–18:15:00 UTC.*"* (62, 224, 301, 372, 445).
2. **The missed G3a floor is now printed, in the table and in prose** — `| G3a | … | 7/12 (no interval: N < 30) | ≥ 11/12 | missed |` (114) and *"**No arm cleared the abstention floor.**"* (156); the takeaway states it too (483).
3. **Every gate row now carries `floor (2026-08-15)` and a `reads` verdict** (110, 124, 138) — v1's "no floor column and no PASS/MISS column" is gone.
4. **The truncated, duplicate-labelled G5a/G5b floors list is gone** — floors ride the cards.
5. **`NOT-COLLECTED — MODEL-FALLBACK` no longer publishes a score** — `| G5b | … | NOT-COLLECTED — MODEL-FALLBACK | ≥ 5/6 | no verdict |` (131), and the census gained the column and the denominator: *"| `cli-claude-fable-5-1` | 180 (60 × 3) | 162 | … | 18 |"* (175).
6. **Bracket intervals are keyed** — *"the brackets elsewhere are Wilson score intervals (0.95) on the cell's own denominator"* (150) plus the glossary row (46).
7. **The head-to-head's denominators reconcile** — *"449 of 504 cells collected and 55 not carried"*, *"17 of the 72 cells it was asked for"* (232), and 224 derivable to the cell.
8. **0.544 vs 23/36 is explained where it prints**, and A/B are bound to arms — *"preferred Claude Fable 5.1 counts 1, a tie 0.5, preferred GPT-6 Astra 0"* … *"which is why it sits nearer the middle than the case tally does"* (234).
9. **The egress socket arithmetic** — *"found 2 of 5 connections going to a vendor API host; every address is named with what owned it"* (389) over a five-row table.
10. **A glossary exists** (31–55) and closes 12 of v1's 16 undefined tokens, including `any-rep read`, `GROUNDED`, `NOT-CLASSIFIED`, `no-mode`, `TIED`, `COLLECTED`, the cost states and `PREREG`.
11. **The product is named and linked** — RuleSage, at 11, 66, 496.
12. **The kit has a door** — *"[the kit](data/)"*, *"[`index.json`](data/index.json)"*, *"[`prereg.md`](data/prereg.md)"* (495–498).
13. **The hostile-reader sentence no longer contradicts itself** — *"**A hostile reader has been through version 1 of this page, not this one.**"* (512).
14. **The phantom "self-agreement column" resolves** — G6a rows exist in every card; the sentence now says *"rows in the gate cards"* (427).
15. **The local seat's 12/18-vs-24 mismatch is closed** — *"12 of 12 collected (6 of 18 items NOT-COLLECTED — CONTEXT)"* (323), with `items 24` (355) and `COLLECTED 24 · NOT-COLLECTED — CONTEXT 12` (423) agreeing.
16. **`qwen3.5-397b` 34/35** is now reconciled by *"| GPT-6 Astra | 251 (1 cell not collected) |"* (216).
17. **`d10/d50/d90` are mapped** (311); **tiers are converted to words** in the opening (13); **"Two jobs, measured in one four-hour window on Saturday 2026-09-05"** (13).
18. **The bill prints its total and resolves the caching contradiction** — *"**The dollars that changed hands come to $23.04**"* (451); *"the endpoint reported 559,416 cached input tokens, priced at that rate … Had none of the input been cached the row would read $22.83"* (457).
19. **`arm or seat` is split** into "**The arms:**" and "**The judging seats.**" (453, 461); `$0.00*` is glossed with a load-bearing asterisk.
20. **The recusal count is explained at the table** (188, 206); **G4 has a lead** — *"**G4 — is the answer actually in the passage?**"* (188).
21. **The per-case `class` column is gone**; both head-to-head tables have captions (244, 287) and the per-family headers spell out their own arithmetic.
22. **"window" is no longer overloaded** — the text spans are "filler" (311, 313); *"polarity control"* became *"**The same run with the reasoning turned on.**"* (343); *"Self-refutation (PREREG A3 FOLDED (4))"* became *"**What this grid cannot tell you, said before you ask.**"* (345); *"The house control ran first, and it refuses the rest"* became *"**The instrument was checked before the contestants.**"* (182).
23. **`transport` left the census**; each arm's road is in its card's title line (108, 122, 136).
24. **The seminar's companions are linked and glossed** (518); **"≈211k tokens/pass" now says what a pass is** — *"≈211k tokens per full pass over the 60 cases"* (382).
25. **The four-rules/three-rules count confusion is fixed** (363, 368), and the August bar is dated exactly.

### REMAIN (quoting v3)

1. **Leg 3 of the law is unmet almost everywhere.** No section outside the rules desk's buried *"the same seat read 13 of 36"* (160) compares its counts to a prior reading.
2. **Identifier drift** — `glm-5.3`/`glm-5-3`, `qwen3.5-397b`/`qwen3-5-397b`, `mistral-large-3-675b`/`mistral-large-3:675b` (470, 471, 25, 70, 512).
3. **"Both frontier rows are a new kind of row for this shelf of pages"** and **"the small model beside them is the one that serves"** (363) — dangling definites in a section with no table.
4. **"The small model beside them is the one that serves."** (487) — the product claim still names no model.
5. **"it is printed above as the pre-fix number it is"** (439) — the 2 of 6 is still not restated where the claim is made.
6. **"One bench in August crossed that line"** (70) — undated at both ends.
7. **"Everything above scores each model against a fixed bar."** (228) — Head-to-head still opens on a cross-reference.
8. **The kit is cited 22 times and located once** (495).
9. **No data table carries the run window inside it**; the egress matrix still has neither caption nor date (380).
10. **"as its own page, dated"** (519) and **"The shelf holds every kit"** (520) — one undated promise, one unlinked door.

### A FIX THAT INTRODUCED A NEW PROBLEM

1. **The six-part conflict list** (25) — added in v3 to disclose that the editorial readers were Anthropic models, a real improvement — **pushed the antecedent of "That reader" three clauses away**, so the outside pre-registration reader's name now reads as one of the Anthropic editorial readers. Verified against v2 line 39, where "Five parts" ended with the outside-reader sentence and the name followed it directly.
2. **The rewritten calibration sentence** (182) — v1 printed *"mean house recall 0.983 against a floor of 0.85"* with no count; v3 added *"over 102 readings of the answered cases, three per case"*, which makes an unexplained six-reading shortfall visible.
3. **The new G2 ledger bullets** (407–409) — v1 printed no G2 at all; v3 prints a reading of `30/34` against a floor written `at least 30/36`, with no verdict and no key for its brackets.
4. **The new opening timeline sentence** (15) — *"eleven days after the exam was sealed"* is an off-by-one that v1 and v2 did not contain.

---

## THE THREE SECTIONS MOST IN NEED OF A COLD-SCROLL LEAD

### 1. The ledger, and what left our machines — currently opens on a table

> **The ledger, and what left our machines.** Two new frontier models — Claude Fable 5.1 and GPT-6 Astra — and the Gemma 4 (26B) seat that answers RuleSage's users today each sat two exams between 14:13:33 and 18:15:00 UTC on 2026-09-05, and this section is the round's accounting rather than its result: first every kind of text that left our machines, to which host and under whose account terms, then the state each scorer wrote for every gate it was registered to score — including the two that cannot exist on a hosted road — so that an absence on this page is checkable rather than asserted. A **gate** is one thing a scorer checks and a **floor** is its pass mark, written down on 2026-08-15 by other hands: `SCORED` means the gate was read against that floor, `NOT-RUN` that it was registered and never pointed at a model, and `NOT-APPLICABLE — transport` that the road cannot carry the check at all, which must never print as "cleared". Brackets after a count are 95% Wilson intervals on that count's own denominator. Nothing in this section ranks the two frontier models; the one comparison this page publishes is in the head-to-head.

### 2. What this page does not say — currently opens "The two frontier models were reached by different roads."

> **What this page does not say.** Every limit below is a limit on one measurement: Claude Fable 5.1, GPT-6 Astra and the Gemma 4 (26B) seat that serves RuleSage's users, read against sixty board-game rules cases and thirty-six planted-sentence items whose pass marks were frozen on 2026-08-15, in one window between 14:13:33 and 18:15:00 UTC on 2026-09-05. Four things about that window could move a number in a direction we can name — the two frontier arms rode different roads on different accounts, each vendor's "high" effort is its own word and not a unit, both cutoffs are vendor statements rather than receipts, and the agents that designed, built, ran and wrote this page run on one of the two vendors' models — so each is stated here with what we measured about it, not merely disclosed. Nothing in this section changes a count above; it says what the counts are not entitled to mean.

### 3. What to take with you — currently opens straight into a bullet

> **What to take with you.** Five findings from one four-hour window — 2026-09-05, 14:13:33 to 18:15:00 UTC — in which two new frontier models, Claude Fable 5.1 and GPT-6 Astra, and the Gemma 4 (26B) seat that answers RuleSage's users today sat the same two exams: sixty board-game rules questions with the passages the product itself retrieved, and thirty-six items hidden in a 1914 book of card games, half of them planted and half of them not there at all. Every pass mark below was written into a design document on 2026-08-15, before either frontier model existed, and each count prints its floor beside it — including the floors nobody cleared.

**Runner-up, whose fix is not a lead:** the ledger's three G2 bullets need a verdict of their own and one reconciling clause, because as printed the only gate in that section a reader can check against a floor is the only gate that never says whether it was met — "at or above the per-case recall floor on 30 of the 34 cases that carry a house citation set (the floor asks for 30 of 36): **missed**", or "**cleared**", but not silence.

---
*Read cold, section by section, 2026-09-05. Nothing outside this file was modified.*
