# CRITIQUE — THE STRANGER (version 2)

Reader: no prior context. I read only
`two-new-frontier-models-at-the-rules-desk-v2.md`, body from the first horizontal rule.
Line numbers below are that file's.

**What works, once:** the verdict is now on the page in words, in the same rows as the
scores, and again in the first takeaway bullet — that was the whole of my v1 complaint and
it is properly fixed. The glossary at lines 33–49 is the single biggest improvement: nine of
the ten words that locked me out of v1 are defined there in plain English. RuleSage is named
and linked in the first paragraph, the kit and the pre-registration have addresses, the
arithmetic that visibly didn't close in v1 now mostly closes, and lines 149 and 342 do the
thing I most wanted — they say what a number *means to a user* ("asserted an answer to 5 of
the 12 questions their own sources do not answer"; "the model that answers our users reads a
window this tier does not fit in"). This is a page I can now read. The remaining problems are
smaller and mostly local.

---

## 1. The sentence I would tell a friend, and the one I still cannot write

What I can now write, from the page, in one pass:

> "Two brand-new frontier models — OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1 —
> sat a board-game-rules exam whose pass marks were locked three weeks before either model
> existed, and neither cleared it: each asserted an answer to five of the twelve questions
> its own sources don't answer. Seven rival AI families, judging blind, couldn't tell the two
> apart. The small model already doing the job on the authors' own hardware abstained best of
> the three, and was the only one that obeyed an instruction hidden in a rulebook."

That is a real finding and it took me one read to build. v1 could not do this.

The sentence I still cannot write:

> "…and here is how Claude Fable 5.1 handled the corrupted rulebooks."

GPT-6 Astra has a number there (6 of 6). The local model has a number there (5 of 6). The
model that **wrote this page** has no number there, and in its place the page offers a
favourable guess — line 157: "the likeliest reading is that Claude Fable 5.1 declined the
corrupted passage", then "Declining a corrupted passage is, on this gate, arguably the right
behaviour", then "rather than a pass it might have earned or a miss it might have taken."
Three sentences of charity where the page's own method demands a measurement or a blank. And
the page never tells me why the number could not be obtained a second way — GPT-6 Astra was
reached through a plain vendor API; the page never says why the same model could not be. This
is the one place where a stranger's suspicion, rather than his confusion, is what stops him.

---

## 2. My v1 findings: fixed, remaining, and one fix that broke something

### FIXED (no further comment needed)

The missing verdict (now lines 147–155, 462, and a "reads" column in every gate card) ·
"citation contract" and "seven labs" in the dek · the local model unnamed until the end (now
line 11) · words vs tokens (line 13 gives both) · "blind" with no mechanism (line 21: "a
judge sees an answer, or two answers side by side, and never a name") · gate scores printed
before their floors (floors are now a column) · bracket notation unexplained (line 43) ·
"NOT-COLLECTED — MODEL-FALLBACK 6/6" ambiguity (line 124 now prints no number, and line 41
defines the state) · "no-mode cases" (line 45) · the broken G5a/G5b floor bullets that ended
mid-word on "by" · `stripped_count`, `is_abstention`, `answering.py:304-311` (all gone) ·
"house control … it refuses the rest" (line 173 runs forward) · "any-rep read" (line 39) ·
"leg" (line 37) · G4 with no lead sentence (line 179) · recusal explained before use (line
21) · 23–13 vs 0.544 reconciled (line 217) · 504 vs 449 vs 224 now closes · the order-flip
consequence drawn (line 217) · "70 judges recognised the author" — now checked against the
key with full counts and a sensitivity cut (line 203) · the 0/6 that meant "not attempted"
(line 310 prints the state) · the 207-vs-257 context contradiction (line 342 resolves it, and
well) · "polarity control" explained in place (line 328) · "Self-refutation (PREREG A3 FOLDED
(4))" as a heading (now line 330) · "Four of them" vs three (line 348 lists four, line 353
lists three) · the promised self-agreement column (G6a rows now exist) · 370 vs 280 tokens
(line 357 explains) · reasoning tokens for both arms (line 363) · the egress arithmetic (lines
382–390) · "70 environment names dropped" (line 392: 70 of 75, all five named) · "What was
NOT-RUN" over things that ran (now line 398, "Every gate, and what became of it") · no round
total (line 430: $23.04) · takeaway bullets in slugs (462–466 use human names) · no addresses
anywhere (RuleSage, the kit, `index.json`, prereg, three prior articles, the promise page are
all links now) · "an operator" undefined (line 49) · the hostile-reader sentence that
contradicted itself.

### REMAIN — quoted from v2

1. **Line 284 — the tier label still contradicts its own value.** "32k (30,000 estimated
   tokens, 120,000 characters)". A tier named 32k that holds 30,000 tokens. One clause fixes
   it; it has now survived two drafts.
2. **Line 284 — "forbidden-literal list" still unglossed.** "every drawn component was
   screened against the round's forbidden-literal list before a prompt existed."
3. **Line 193 — the denominator that changes mid-list.** "alibaba (qwen3.5-397b) 34/35".
   Every other family on that line is over 36. Nothing says why this one is 35.
4. **Line 143 — still written to an auditor.** "A scorer that did not score an arm leaves a
   missing card, never a renumbered one." I still do not know what renumbering would be.
5. **Line 65 — the two sentences about itself.** "those bytes are the publisher's, not the
   user's, and travel as sources by registration — the count prints here as the disclosure it
   is." Travel *where*? And the last clause is a sentence about its own printing.
6. **Line 357 — the scaffolding test still answers the wrong question.** The CLI preamble is
   measured in input tokens and prompt hashes only: "the endpoint's own tokenizer counted
   exactly 280 input tokens on every pair, with the user prompt's sha256 identical in both
   conditions on every pair". I asked in v1 whether the preamble changed the **answers**. It
   still is not said, and the 8 cases were run twice, so the comparison exists.
7. **Line 481 — "## The rest of the seminar"** is still the first use of "seminar" for the
   body of work, now at line 481 of 501.
8. **Line 483 — "trust contract" and "the cove" still unglossed**, in the same sentence as
   two more: "where the class of row these two frontier models now belong to — a reference
   arm, measured and never seated — was chartered; and [*The new kid*](/the-new-kid/), where
   the rules desk's trust contract first rejected three hosted models."
9. **Line 25 — "exhibit" still unglossed and not in the glossary**: "it is the conflict at
   the centre of this exhibit".
10. **The August incident is gone rather than told.** v1's line 238 disclosed that "One bench
    in August crossed that line" — a prior bench that used real users' questions. It appears
    nowhere in v2 (grep: zero hits). My complaint was that it raised an alarm and withheld the
    story. Deleting the alarm is the wrong repair: line 65's promise that "a typed question
    never leaves our machines, and a bench is not an exception" now reads as an unbroken
    record, and the page previously told me it wasn't one.

### One fix that introduced a NEW problem

**The glossary pins "the shelf" to one meaning, and the page then uses another meaning three
times.** Line 36: "**The shelf** is ollama.com's hosted model shelf, where all seven judging
seats ran". Then line 348: "Both frontier rows are a new kind of row for this shelf" — which
now reads as though the two frontier models are rows on ollama.com. Same at line 485, "The
shelf holds every kit", and line 495, "the descendant of three earlier ones on this shelf".
In v1 "shelf" was ambiguous; in v2 it is confidently wrong three times. Either gloss the
second sense in the same glossary row, or use a different word for this publication.

**A second, smaller one: the glossary now creates two different sets of "three jobs".** Line
13: "Two jobs, measured in one four-hour window on Saturday 2026-09-05." Line 17: "A third
job — the narrator's chair…". Line 37: "| the three jobs | … **Leg A** is the rules desk,
**Leg H** the head-to-head, **Leg C** the filing cabinet |". So the lead's three jobs are
rules desk / filing cabinet / narrator's chair, and the glossary's three jobs are rules desk /
head-to-head / filing cabinet. Two of the three overlap, which is what makes it hard to spot
and hard to recover from.

**A third: the glossary's list of G-numbers is not the page's list.** Line 38: "The G-numbers
(G1, G3a, G3b, G4, G5a, G5b, G6a, G6b) are the harness's own labels". The body also uses
**G2** (lines 402–404), **G5c** and **G6c** (lines 398, 411), and **G-PEN** (line 394). A
stranger who took that parenthesis as the complete set meets three surprises after line 390.

---

## 3. Every place I got lost in v2, in document order

**Line 3 (the dek) — the finding is not in the line most people read.**
> "*GPT-6 Astra and Claude Fable 5.1 sit an exam frozen in August — cite the page, or say the
> book does not answer — judged blind by models from seven families that built neither, beside
> the small local model that already does the job.*"

The body now states the verdict four ways. The dek still only promises a test. This is the one
line that travels — into a link preview, a chat message, a feed — and it still withholds the
result. **Replacement text:** see Fix 1.

**Line 13 — "your clause" is the only sentence that reaches outside board games, and it is
unsignposted.**
> "A model that can find the sentence can find your clause."

I was reading about card games; "clause" put me in a contract with no transition.
**Replacement:** "A model that can find one planted sentence in thirty thousand tokens can
find the one clause that matters in a long document of your own."

**Lines 13 / 17 / 37 — two different sets of "three jobs".** See §2 above.
**Replacement for line 37's first cell:** "| Leg A · Leg H · Leg C | the harness names the
three measured runs by letter, and the kit's files keep the letters: **Leg A** is the rules
desk, **Leg H** the head-to-head between the two frontier arms, **Leg C** the filing cabinet.
`C-think-true` is the filing cabinet run a second time with the local model's thinking channel
on |"

**Line 15 — 2026-08-27 is still a date doing unexplained work.**
> "GPT-6 Astra's model record on its maker's own API — the creation stamp its model listing
> carries — is dated 2026-08-27."

The gloss tells me what the stamp is; it does not tell me why *this* date and not the ship
date the previous paragraph gave me (the third of September). I worked out that it is the
earliest date the page can prove the model existed at all, which is what makes the 08-15 bar
safe — but I had to work it out. **Replacement:** "GPT-6 Astra's model record on its maker's
own API — the creation stamp its model listing carries, and the earliest date we can show the
model existed at all — is dated 2026-08-27, eleven days after the bar was sealed."

**Line 21 — "seat" is used fifteen lines before the glossary defines it**, along with "shelf"
(line 21), "arm" (line 25), "the kit" (line 25), "NOT-RUN" (line 17) and "an operator" (line
25). All six are defined at lines 35–49, i.e. *after* the three paragraphs that have to hook a
stranger. **Fix:** move the "The words this page leans on" table above the "Who judged, and
who wrote this" section, or add a single pointer at the end of line 13: "Words this page uses
in a particular way — arm, seat, shelf, gate, floor, the kit — are defined in one short table
below, before the numbers start."

**Line 21 — "the card it needed."**
> "a nineteen-gigabyte judge would have evicted a live product model from the card it needed"

"Card" for a graphics card, unglossed, in a sentence otherwise readable. **Replacement:**
"…would have evicted a live product model from the graphics card it needs to answer users."

**Line 23, and 32 more places — thirty-three blocks of machine scaffolding.**
> `<!-- fill \`panel-shape\` — prereg/panel.json: panel.seats, panel.seats[].seat / .family / .transport -->`

There are 33 of these and they are 25,440 of the file's 90,819 characters — 28% of what I
opened. Line 99 alone is a single 2,000-character block. I cannot tell whether they are part
of the page or not, and one of them (line 332) names files with two different timestamp
formats. If they vanish in rendering, say so somewhere a reader of the source can see; if they
do not, the page is unreadable. Either way this is the first thing I saw after the lead.

**Line 25 — the four findings that were not acted on are counted and never named.**
> "It raised 15 findings: 7 changed the instrument before the first call … 4 were misreadings
> answered in the text; 4 were disclosed and left standing."

The seven that changed things are described. The four "left standing" are the four an outside
model thought favoured one arm and the page decided to live with — the most interesting four
in the sentence — and they are a number. One clause each, or a pointer to the amendment that
holds them.

**Line 101 and 115 — "(house law 3)" is a door that does not open.**
> "this transport accepts none (house law 3)"

There is no house law 1 or 2 on the page, no list of house laws, and no address. This is
exactly the problem PREREG had in v1, reintroduced under a new name. Also note that the
glossary calls this thing a **road** (line 35, "on one named road") while the prose and the
states call it a **transport** (lines 101, 115, 398, 411) — two words for one thing again.
**Replacement:** "this transport accepts none — a standing rule of this workshop: where a road
will not take a sampling setting, we send none rather than send one it silently ignores."

**Line 106 — a met floor that reads "—" instead of "cleared".**
> "| G1 | forged citation markers, counted as their own failure | 0 markers | 0 | — |"

Floor 0, reading 0, verdict blank. Four lines later, line 109: "| G5a | hidden directives
followed (any-rep read) | 0/6 (any-rep read) | 0/6 | cleared |" — floor 0, reading 0, verdict
"cleared". I cannot tell what distinguishes them. If the forged-marker row is a component of
G1's own verdict rather than a gate, say so in the cell: **"— (counts into the G1 verdict
above)"**.

**Line 112 — "(the alternative floor)" with no verdict anywhere.**
> "| G6a | answered cases with the same citation set across the three reps | 27/36 [0.589,
> 0.862] | 36/36 (the alternative floor) | — |"

Alternative to what? Can an arm clear G6a by meeting it? Line 153 says "**No arm cleared the
self-agreement floor** (≥ 33/36 byte-identical…)" and never mentions the alternative again.
The local seat reads 33/36 on this row, which is the *other* floor's number — I spent a minute
deciding whether that meant something. **Replacement for the floor cell:** "36/36 — the second
of two registered ways to clear G6a; no arm met either".

**Line 127 — 0/54 among 0/60s.**
> "| G6b | cases stopped by an output cap | 0/54 | ≤ 2/60 | cleared |"

Every other arm's denominator is 60. Fifty-four is sixty minus the six corrupt-corpus cases
that never reached this model — which the page explains thirty lines later for a different
gate and never connects to this cell. **Replacement:** "0/54 (the 6 corrupt-corpus cases are
not on this road — see below)".

**Line 157 — "the other four case classes" is one class too many.**
> "Against that, 162 of 162 calls on the other four case classes and 36 of 36 on the filing
> cabinet came back from the model invoked."

Line 61 names four classes in total (36 answered · 12 unanswerable · 6 injection · 6 corrupt).
Take the corrupt class out and three remain — and 162 = 54 cases × 3 reps confirms three.
**Replacement:** "the other three case classes".

**Line 157 — 19 calls, 18 cells.**
> "all 6 cases, 19 of 19 calls across the first attempt and the re-dispatch" … "its 18 cells
> read NOT-COLLECTED — MODEL-FALLBACK"

Six cases × three reps is 18, and the census at line 168 shows 18. The nineteenth call is
presumably the re-dispatch, but the page sets 19 and 18 four sentences apart with nothing
joining them. **Replacement:** "19 of 19 calls — the 18 scored reps plus one re-dispatched
attempt".

**Line 157 — the page is soft on the model that wrote it, in the one place it cannot afford to
be.** Quoted in §1. Three consecutive sentences supply a favourable interpretation of a
missing measurement; a fourth ("a pass it might have earned or a miss it might have taken")
frames the blank as symmetric when the preceding sentences have already leaned one way. And
the obvious question — why not reach the same model through its maker's API, as the page did
for GPT-6 Astra, and get the number — is not asked or answered. **Replacement text:** see Fix
2.

**Line 173 — "over 102 cases" is bigger than the whole bank.**
> "it read SCORED, mean house recall 0.983 over 102 cases"

The page has 60 cases, 36 of which the book answers. 102 is not any of those numbers (it is 34
× 3, if this is counting reps). Whatever it counts, it is not "cases" as this page has used the
word everywhere else. **Replacement:** "mean house recall 0.983 over 102 readings (34 answered
cases × 3 reps)" — with the real unit substituted if that is not it.

**Lines 179 and 189 — the rule that drops each arm's weakest cases from its own denominator is
stated, and its direction is not.**
> "An arm whose reply on a case was an abstention has no answer to ground, so that case leaves
> its denominator"

GPT-6 Astra abstained on none, so it is judged on 36 of 36. Claude Fable 5.1 and the local seat
abstained on two each, so each is judged on 34 — their two worst cases removed. The rule is
defensible (there is no answer to ground), but it flatters the arms that abstained, one of which
is the author's, and the page never says so. **Add after line 189:** "That rule cuts in the
abstaining arm's favour: each arm's two weakest cases leave its own denominator, and the arm
that abstained on none is the only one read over all 36. Read against 36 with an abstention
counted as ungrounded, the readings would be GPT-6 Astra 36/36, Claude Fable 5.1 32/36, the
local seat 31/36."

**Lines 183–187 — the only scored table on the page with no floor and no verdict.**
> "| arm | judged cases | GROUNDED | recused cells | families carried | state |"

Every gate card carries a "floor (2026-08-15)" column and a "reads" column. G4 carries neither,
and its three readings (36/36, 32/34, 31/34) are the only place on the page where the two
frontier arms visibly differ on a quality measure. Was there a groundedness floor? Line 415
gives a *panel* floor ("7 families carried a verdict against a floor of 4"), which is about
judges, not answers. Either add the floor column, or say in one sentence that G4 was registered
without one and therefore draws no verdict.

**Line 203 — 251 where the same page's table says 252.**
> "GPT-6 Astra: 55 of 251 judged cells carried a claim"

Line 185 says Astra was judged on 36 of 36 cases by 7 families: 252 cells. One cell is missing
and nothing marks it. (Fable's 238 = 34 × 7 and the local seat's 204 = 34 × 6 both close
exactly, which is what makes the 251 stand out.)

**Line 203 — the recognition counts are printed and their meaning is not drawn.** Sixteen of
Astra's 55 claims named OpenAI; seven of Fable's 56 named Anthropic. That is the number that
either supports or damages the blindness claim, and the page moves straight to the head-to-head
sensitivity cut without saying which. One sentence: is 16 correct guesses out of 251 cells
consistent with chance, and does the page treat that as blind or not?

**Line 203 — a 120-word sentence.** "On the groundedness reading, where a sheet holds one arm's
answer, the claims can be checked against the key: GPT-6 Astra: 55 of 251 judged cells carried a
claim; of those, 16 named openai, 5 named another maker, and 34 named no maker at all — a human,
or the rulebook itself; Claude Fable 5.1: …". Three arms' worth of three-way splits in one
sentence with two levels of colon. This is a table.

**Line 217 — 9 cases × 2 orders is 18 cells, so 72 − 18 = 54, not 55.**
> "The DeepSeek V4 Pro seat carried 9 of the 36 cases … so its remaining 55 cells read
> NOT-COLLECTED — NOT-CARRIED."

The 55 is right against the total (449 + 55 = 504), so the reconciliation must be that one of
the nine cases carried only one of its two orders — which is the "1 pair missing one of its two
orders" mentioned two sentences earlier. Nothing joins them, and the reader who does the easy
multiplication gets a different answer than the page. **Replacement:** "…carried 9 of the 36
cases — 17 of the 72 cells it was asked for, one of those nine cases having only one of its two
orders, which is the missing half named above — so its remaining 55 cells read NOT-COLLECTED —
NOT-CARRIED."

**Line 217 — one paragraph, roughly 500 words, carrying eight separate facts.** It holds the
registered shape, what came back, a retired judge, the collapsing rule, the headline rate, the
case tally, the bootstrap interval, the order-flip rate and its consequence, and finally the
amendment. The last item is the one that most earns a stranger's trust — "The comparison set
changed once after judging had begun, against the author's own interest: the sheet builder had
left out 2 cases … because Claude Fable 5.1's reply on each was an abstention where the book
does answer — its two weakest cells — and amendment A11 put them back in" (65 words) — and it
is the last clause of the wall, where nobody will reach it. Break the paragraph after "…7
families carried a verdict."; give the amendment its own bolded paragraph.

**Line 217 / 219 / 463 — the headline interval never states its level.** Glossary line 43 gives
95% for the Wilson brackets and then says only "the head-to-head's interval is a bootstrap over
cases and says so". "[0.444, 0.638]" — a bootstrap at what level? **Replacement in 217:**
"puts that rate at [0.444, 0.638] — a 95% interval, and a wide one, because…"

**Line 217, 203, 264, 268 — every pooled figure on this page is written as a score *for the
author's own model*.** "the pooled preference rate is **0.544 for Claude Fable 5.1**"; "0.553
for Claude Fable 5.1 over 36 cases"; "23 for Claude Fable 5.1, 0 tied, 13 for GPT-6 Astra"; and
the seat table's own column heads, "sum of its scores for Claude Fable 5.1" and "share of its
cases favouring Claude Fable 5.1". Line 217 does offer "read the other way, 0.456 for GPT-6
Astra" once. The choice of numerator is arbitrary and the page chose itself, in every table and
every takeaway. A stranger notices. Either alternate the direction, or put both numbers in every
cell.

**Line 284 — numbers I could not read at all.**
> "Its spans are unique within a tier (unique fraction 8k 1.0 · 16k 1.0 · 32k 1.0) and overlap
> across tiers (8k×16k 0.477 · 8k×32k 0.819 · 16k×32k 0.941 of the smaller tier, measured),
> which is why this page draws no conclusion about depth across tiers."

I do not know what "unique fraction" counts, what "0.477 of the smaller tier" is a fraction of,
or why text shared between two tiers stops a conclusion about **depth**, which is a position
inside one tier. Three unreadable numbers and a non-sequitur, in the middle of an otherwise
clear paragraph. **Replacement:** "No two items inside a tier share any filler text (measured:
1.0 unique in all three tiers). Across tiers the filler does overlap — the 8k text is 0.477
contained in the 16k text, 0.819 in the 32k, and the 16k is 0.941 contained in the 32k — so the
three tiers are not independent samples of the book, and this page compares depths only within a
tier, never across them."

**Line 328 — "The polarity control" names nothing I can map onto what follows.** The sentence
after it is clear ("we ran its whole grid a second time with thinking on"), so the label is pure
friction. **Replacement heading:** "**The same run with the reasoning turned on.**"

**Line 342 — the number I most wanted to read, and could not.**
> "On that same 32k canary prompt it reported 16,387 prompt tokens against our chars÷4 estimate
> of 30,113 — a ratio of 0.5442, under the registered 0.8 floor — and the canary at the front was
> the one it did not find: on its registered window of 32,768 tokens, this model's tokenizer needs
> more room than the window has, and the front of the prompt is dropped."

16,387 is *half* of the 32,768-token window. If the prompt needed more room than the window has,
I would expect a count at or just under 32,768, with the overflow dropped. A count at half the
window says something else happened, and the sentence's explanation does not cover it. This is
the page's headline product finding — "the model that answers our users reads a window this tier
does not fit in" — and its one supporting number does not follow from its one supporting
sentence. I am not saying it is wrong; I am saying I cannot get from the number to the claim, and
this is the paragraph where a skeptical reader will stop.

**Line 348 — "no ordering is carried past the window" implies an ordering inside it.**
> "Second, no ordering is carried past the window 2026-09-05 14:13:33–2026-09-05 18:15:00 UTC."

Everywhere else the page says it draws no ordering at all (lines 219, 264, 326, 463).
**Replacement:** "Second, this page draws no ordering between the two frontier arms at all, and
nothing here should be read as one holding beyond the window 2026-09-05 14:13:33–18:15:00 UTC."

**Line 357 — a 143-word sentence, and the worst on the page.** It runs from "That harness adds
scaffolding of its own" to "the vendor's tokenizer", carrying the reason for the test, the
preamble's size in two units, a hash, the number of cases, the pass state, the measured delta,
the hash-identity check, the conclusion, and a parenthetical reconciliation of two earlier
numbers. Split it at "…as a proxy for what the CLI arm carries and cannot be run without." and
again after "PASS —".

**Line 430 — the plan whose price is withheld is the page's own comparison.**
> "Plan-included rows print `$0.00*`: no marginal charge, not free: a flat monthly plan whose
> price is the account's, not this round's"

The page's dollar finding is $23.04 for the metered work. Six of the nine models on the bill are
"$0.00*", and the Fable arm's $31.44 is itself an estimate against a subscription. So the one
sentence a cost-curious reader wants — what this would have cost without subscriptions — is
unavailable by construction, and the page does not say that plainly. Also: two colons in one
sentence made me reread it. **Replacement:** "Plan-included rows print `$0.00*`. The asterisk is
load-bearing: no marginal charge for this round, not free — a flat monthly plan whose price
belongs to the account, not to these calls. Because six of the nine rows below are plan-included
and a seventh is a subscription estimate, this bill is not a price for reproducing the round; it
is a receipt for the part of it that metered."

**Line 454 — a cap that was exceeded and did not fire, citing a rule with no address.**
> "`kimi-k3` reads $5.25 against its $5.00 seat cap, $0.25 over; the cap did not fire (no
> NOT-COLLECTED — CAP cell exists) — a cap allows at most one call's overshoot"

Every other rule on the page now carries a § I can look up. This one is asserted bare.

**Line 476 — the invitation and its two missing files are four clauses apart.**
> "The generator itself and the fixture builder are named in the kit's index with their shas but
> not shipped as they stand, because they name our own machines and the kit's screen refuses them
> — they ship on request, with those names described."

The bullet's promise is "Run the filing cabinet on your own machine"; by the end of it, two of the
files needed to do that are not in the kit. That is honest and I would rather have it than not —
but put it in the first clause, not the fourth. "The kit's screen" is also unglossed (it is the
publication filter mentioned at line 477).

**Line 497 — the one that made me distrust the page, and it is in the credits.**
> "A hostile reader went through version 1 of this page (READ, FOLDED — the pass on version 1 is
> folded into version 2; a fresh pass on the version that releases is owed): an Opus agent read
> version 1 in full … and named ten must-fixes in its own order; every one of the ten landed in
> version 2 — the page now prints the verdicts in words, prints no number for an uncollected cell,
> names the model that answered on the corrupt-corpus class, and points at a kit that ships every
> amendment and the receipts (findings raised 59 · must fixes 10 · must fixes landed 10)."

Four problems in one 117-word sentence. (a) **"an Opus agent"** — Opus is one of Anthropic's
models. So the adversarial reader of the page came from the same company as the arm that wrote
the page, and line 25 told me "No Anthropic or OpenAI model judges anything on this page." Both
statements are true of their own scope, and together they read as a boundary drawn exactly where
it was convenient. This belongs in the conflict paragraph at line 25, in the page's own voice,
not 470 lines later in a thanks section. (b) It reports its own grade — "every one of the ten
landed" — in the sentence that is supposed to be the check on it. (c) "findings raised 59 · must
fixes 10 · must fixes landed 10": forty-nine findings were raised and not landed, and they are a
subtraction I had to do. (d) "READ, FOLDED" is notation. **Replacement text:** see Fix 3.

---

## 4. Jargon still unglossed anywhere, named

The glossary retired most of v1's list. What is still used without ever being explained:

**House vocabulary:** **exhibit** (line 25) · **the cove** (483, 484) · **the seminar** (481) ·
**the open call** (483) · **reference arm** / "measured and never seated" / **chartered** (483) ·
**trust contract** (483) · **house law** 3 (101, 115) · **house answers** (61) · **answer ledger**
(61, inferable but never stated) · **re-derived** (61) · **the shelf** in its publishing sense
(348, 485, 495) · **the shelf gemma seat** (377) · **seat posture** (129) · **the kit's screen**
(476) · **forbidden-literal list** (284) · **the narrator's chair** (17, 484 — I still do not know
what job that is) · **transport** as a state value (398, 411) beside "road" in the glossary ·
**needle** (introduced only implicitly at 284, then a column head at 292 and 314) · **canary**
(342 — readable from context, never defined) · **NOT-CLASSIFIED** (292 — a state absent from the
state list at line 41) · **NO-RATE** (270 — same) · **polarity control** (328) · **sensitivity
cut** (203) · **the seal manifest** (474 — now with a half-gloss) · **collapsed**/**collapsed
observation** (217, 266) · **sheet** (217, 203 — the unit a judge fills in, never named as one).

**Numbers and notation:** the head-to-head interval's confidence level (never stated) ·
**cluster bootstrap** / **resamples** / **percentile** (217) · **unique fraction** and the
cross-tier overlap fractions (284) · **d10 / d50 / d90** (glossed in place at 284 — fine) ·
**chars÷4** (342, 357 — glossed by line 357's parenthesis, late) · **N ≥ 30** (still asserted as
a registered rule rather than justified, but it now has a § address, which is enough).

**Machine identifiers still printed at the reader:** truncated shas at 25 (×2), 284, 290, 342,
357, 497 — six of them, none with a stated use; `ofl-ans-0001` and the whole case-ID column (223,
227–262 — inferable); `ofl-ans-0008.o1` (217, 270 — the `.o1` order suffix is never explained);
`prompt_eval_count` / `eval_count` / `prompt_tokens` / `total_cost_usd` (338–340, 437, 438, 447);
`num_ctx` (129); v0.57.0 + commit 76169d2 (175); `claude-opus-5` (157); `G2`, `G5c`, `G6c`,
`G-PEN` (missing from the glossary's G-list); **33 `<!-- fill -->` blocks**.

---

## 5. The three fixes I would make first

### Fix 1 — Put the finding in the dek. Replaces line 3.

> *GPT-6 Astra and Claude Fable 5.1 sit an exam frozen in August — cite the page, or say the book
> does not answer — judged blind by models from seven families that built neither, beside the
> small local model that already does the job. Neither new model cleared it: each asserted an
> answer to 5 of the 12 questions its own sources do not answer, and the seven judges could not
> tell the two apart.*

### Fix 2 — Stop being kind to your own model on the one gate it has no result for. Replaces the span of line 157 running from "The stream's record class calls it a refusal fallback:" through "…a miss it might have taken."

> The stream's record class calls it a refusal fallback. That label is the tool's, not a
> measurement of ours: we do not know what Claude Fable 5.1 would have done with a corrupted
> passage, and this page is not entitled to guess in its own author's favour. It is worth saying
> which way the missing number could cut — refusing a corrupted passage would clear this gate,
> and answering one confidently would fail it — and then saying that we measured neither. So the
> single gate this page cannot score belongs to the model that wrote this page. Its 18 cells read
> NOT-COLLECTED — MODEL-FALLBACK in the census below, out of the 180 calls it was asked, and the
> verdict above reads *no verdict*. We did not get the number a second way, and the reason is
> worth printing: **[the reason the same model was not reached through its maker's plain API, as
> GPT-6 Astra was — a registered road, a different arm, or a cost, in one sentence].**

### Fix 3 — Name the vendor of your own hostile reader, and move it to the conflict paragraph. Replaces the first clause of line 497 (from "A hostile reader went through version 1" through "must fixes landed 10)."), with a pointer added at line 25.

At line 497:

> **The check on this page.** Version 1 was read in full, against the kit, the sealed
> pre-registration and the scorer output, by an adversarial agent running on an Anthropic
> model (the page calls it "an Opus agent") — so from the same company as one of the two arms; we had no outside model
> with the context to do this reading, and we are naming the limit rather than describing the
> reader as independent. It raised 59 findings and marked 10 as must-fix; all 10 are folded into
> version 2, and the other 49 are listed with their dispositions in the kit. The version that
> releases is owed a fresh pass, and has not had one.

And at the end of line 25, in the list of five procedures, add a sixth:

> And sixth: the adversarial read of this page itself was done by a model from Anthropic, because
> no outsider had the context — that one we could not close, and it is named again in the credits.

---

### Broken or unclosing text that must be fixed regardless (not counted among the three)

- **Line 157:** "the other four case classes" → three (162 = 54 cases × 3 reps).
- **Line 173:** "over 102 cases" — the bank has 60; name the real unit.
- **Line 203:** "55 of 251 judged cells" against the same page's 36 × 7 = 252.
- **Line 217:** "carried 9 of the 36 cases … its remaining 55 cells" — 9 × 2 leaves 54.
- **Line 193:** "alibaba (qwen3.5-397b) 34/35" in a list of 36s, second draft running.
- **Line 284:** "32k (30,000 estimated tokens…)", second draft running.
- **Line 342:** 16,387 prompt tokens reported into a 32,768-token window, explained as the window
  being too small.
- **Line 157 / 168:** 19 calls vs 18 cells.
- **Line 106 vs 109:** a met floor of 0 that reads "—" beside a met floor of 0 that reads
  "cleared".
- **Lines 23–499:** 33 `<!-- fill -->` blocks, 28% of the file's characters.
- **Line 36:** the glossary's "shelf" contradicts the page's own use of it at 348, 485, 495.
- **Line 38:** the G-number list omits G2, G5c, G6c, G-PEN, all of which the body uses.
- **Lines 13 / 17 / 37:** two different sets of "three jobs".
