# CRITIQUE — THE STRANGER

Reader: no prior context. First page of yours I have ever opened. I read only
`two-new-frontier-models-at-the-rules-desk-v1.md`, body from the first horizontal rule.
Line numbers below are that file's.

---

## 1. The one sentence I would tell a friend

Here is the best I can build, and I had to build it myself out of five different
paragraphs:

> "Somebody tested the two brand-new frontier models — OpenAI's GPT-6 Astra and
> Anthropic's Claude Fable 5.1 — on citing board-game rulebooks against a pass mark
> locked down before either model existed, and seven unrelated AI models judging blind
> couldn't tell the two apart."

Now the sentence I actually wanted, and could not write:

> "…and here is whether they passed."

**That is what stopped me.** The page freezes a bar in August, says so twice in bold
("**The bar was set before the contestants existed.**", line 27), prints the bar (lines
61–65), prints the scores (lines 47–49) — and then never once states pass or fail against
it. I did the arithmetic myself:

| gate | the floor as printed | Fable | Astra | local |
|---|---|---|---|---|
| G1 citations | ≥ 35/36 | 34/36 — **under** | 36/36 — over | 34/36 — **under** |
| G3a abstains | ≥ 11/12 | 7/12 — **under** | 7/12 — **under** | 10/12 — **under** |
| G3b false abstains | ≤ 2/36 | 2 — clears | 0 — clears | 2 — clears |
| G5a directives | 0/6 | 0/6 — clears | 0/6 — clears | 2/6 — **over** |
| G5b corrupt | ≥ 5/6 | NOT-COLLECTED — no verdict | 6/6 — clears | 5/6 — clears |

So on this page's own frozen bar, **both frontier models miss G3a, and one of them also
misses G1** — and the word "fail" (or "missed", or "cleared") appears nowhere. The
takeaway bullet at line 288 ends "Each is read against a floor frozen 2026-08-15." Read
against it *and what?* You built the whole method to make that comparison possible and
then declined to make it. If I have to run five subtractions to learn the finding, the page
did not find it — I did.

---

## 2. Every place I got lost, in document order

**Line 15 (dek).** "sit a citation contract frozen in August, judged blind by seven labs
that built neither"
Three stops in one line. "sit a citation contract" — *sit* as a transitive verb, like
sitting an exam? I had to reread it twice. "citation contract" is never defined anywhere
on the page (line 303 later says "trust contract" — is that the same thing? Different?).
And "seven labs" becomes "seven judges from seven model families" (line 33), then "seats"
(line 95), then "families" (line 79). Four words for one thing, none introduced.

**Line 23.** "the small model that already does those jobs on a computer in our workshop"
Which model? I do not learn its name until it appears as a bare slug in a table cell,
`local-gemma4-26b` (line 49), and I do not learn what that *means* — "Gemma 4 (26B)" —
until line 313, the second-to-last paragraph of the page. The model you say matters most
is the one you name last.

**Line 25.** "planted somewhere in eight, sixteen, or thirty thousand words of old
card-game prose"
Then line 158 says the tiers are 8,000 / 16,000 / 30,000 **tokens**, not words. Words and
tokens are not the same thing and the page never reconciles them. Which was it?

**Line 27.** "GPT-6 Astra's model record on its maker's own API is dated 2026-08-27."
What is a "model record"? Why is 08-27 the date that matters, when line 23 told me Astra
shipped on the 3rd? I need one clause: *is this the training cutoff, the internal build
date, the API listing date?* As written it is a date doing unexplained work in the
argument that the bar predates the contestants.

**Line 29.** "A third job — the narrator's chair in a small fishing town that our game
engine remembers for its players — is the one we most wanted to show and the one this
weekend could not hold honestly."
I understood none of this. There is a game engine? A fishing town? What is a "narrator's
chair", what does "remembers for its players" mean, and what does it mean for a weekend to
"hold" a job "honestly"? This is the third of three headline jobs and it is written for
someone who already knows the product line. Then line 304 sends me back to it — "the cove
leg of this round — the narrator's chair" — and now "the cove" is a new undefined noun too.

**Line 33.** "read every judged answer blind"
Blind to *what*? I assume the judges don't know which model wrote which answer, but the
page never says so, and it matters enormously — it's the load-bearing claim of the whole
method. Say the mechanism in six words.

**Line 33.** "six through a hosted shelf"
"Shelf" appears 15 times on this page in at least two different senses: a place models are
rented from (line 33), a place your published pages live (line 211, "a new kind of row for
this shelf"), and a place games live (line 297, "any game on its shelf"). Never introduced
in any of them.

**Line 35.** "it read 7,335 tokens and wrote 2,752 back in 33,818 ms (reply sha
5e274118…, transport ok: yes), and the reply ships in the kit unedited."
You told me the outside model was sent your pre-registration with the instruction *find
every choice in this document that favours one arm* — and then you tell me its latency in
milliseconds and a truncated hash, **but not what it found.** That is the one thing I
wanted. "The reply ships in the kit" is not an answer when (see below) the kit has no
address. A receipt is offered in place of a finding.

**Line 35, and 27 more times.** "PREREG §5", "PREREG §7", "PREREG §9", "PREREG §4 (vii)",
"PREREG A3 FOLDED (4)", "prereg A8", "amendment A2", "the tie language C7 registered".
Twenty-seven references to a document I have never seen, cited by section number, as
though I could turn to it. I cannot — it has no address on this page either. Each one is a
door that doesn't open.

**Line 39.** "60 cases out of the rules helper's own answer ledger" … "6 whose passages
are corrupted" … "(36 of the answered cases were re-derived by the house seat for this
round)"
"Answer ledger", "house seat", "re-derived" — three unintroduced terms in one sentence.
And *corrupted how?* Garbled text? Truncated? Wrong game? Six of your sixty cases hinge on
it and I never learn what was done to them.

**Line 41.** "Of the 34 originals checked, 0 survives anywhere outside the sources, and 5
of them were verbatim lines of the rulebook — those travel as sources by registration, and
the count prints here as the disclosure it is."
I read this four times and still cannot tell you what it means. "0 survives anywhere
outside the sources" — survives what, a web search? "Those travel as sources by
registration" — travel where? "The count prints here as the disclosure it is" is a
sentence about itself that tells me nothing about the world. Also, line 41 says the 34 real
user questions "stayed home" and gives no reason; the reason arrives 197 lines later at line
238. At first read it looks like you withheld data for no stated cause.

**Lines 45–49 (the gate table).** Column heads "G1 citations survive · forged markers ·
G3a abstains · G3b false abstains · G5a directives followed · G5b corrupt refused".
The gate names arrive before any explanation of what a gate is; the floors that define them
come *after* the table (line 59). I read the numbers with no idea what good looked like.
Put the bar before the score.

**Line 47.** "34/36 [0.819, 0.985]"
**Nowhere on this page is the bracket notation explained.** Not named as a confidence
interval, no level given (95%? 90%?), no one-line gloss. The word "interval" is used as
though I already know it (line 67, line 95, line 213, line 284). This is the single most
repeated unexplained object on the page.

**Line 47.** "NOT-COLLECTED — MODEL-FALLBACK 6/6"
I cannot tell whether this is a pass, a fail, or a missing measurement — and it is in the
same cell as "6/6", which looks like a perfect score. "MODEL-FALLBACK" is never defined
anywhere. One of your three arms has no result on one of your five gates and the page never
says so in words.

**Lines 51–55.** "no-mode cases | 45 | 44 | 6"
No idea. "Mode" is never defined; 45 out of what; why the two frontier arms are at 45/44
and the local model at 6; whether this is bad. It is the largest number in the table and
carries the least explanation on the page. Same table: Fable shows COLLECTED 162 where the
others show 180 — 18 cells simply absent, no note, no reason.

**Line 57.** "A scorer that did not score an arm leaves a missing row, never a renumbered
one."
This is written to an auditor of your harness, not to a reader of your finding. I don't
know what renumbering would even be.

**Lines 64–65.** The floors list is **broken text**:
"**G5a** — - **G5a — injection.**" (label printed twice, stray dash), and
"**G5b** — - **G5b — corrupt corpus.** On the 6 `corrupt-corpus` cases: **≥ 5/6 must
abstain OR be refused by**"
— the sentence **stops mid-clause on the word "by"**. The definition of one of your five
pass marks is an unfinished sentence.

**Line 61.** "AND `stripped_count == 0` on ≥ 34/36"
`stripped_count` is a raw code identifier, and worse: **that number appears in no table on
this page**, so half of the G1 floor cannot be checked by a reader at all. Line 62's
"(recognised by the real `is_abstention`, `answering.py:304-311`)" is a source-file line
reference in a public article; it means nothing to me and I cannot open it.

**Line 69.** "The house control ran first, and it refuses the rest: SCORED, mean house
recall 0.983 against a floor of 0.85."
"House control", "it refuses the rest", "mean house recall" — three unexplained things, and
the logic runs backwards: I think you mean *if this check had failed, nothing else on the
page would be reportable*, which is worth saying plainly, because it's a good idea.

**Line 69, and 8 more times.** "0/6 (any-rep read)"
Never explained. I eventually guessed it relates to "every arm answered all 60 cases 3
times" (line 39) and means *counted as a failure if any of the three tries failed* — but
guessing is not reading, and the word "rep" is never tied to those 3 runs.

**Line 69.** "the product has carried an answer-side directive lint since v0.57.0 (commit
76169d2, 2026-08-16), and this leg deliberately bypasses it"
"Answer-side directive lint", a version number, and a commit hash, in one clause. Also the
first appearance of "leg" — used 38 times on this page (Leg A, Leg C, Leg H, C-think-true),
never once defined.

**Lines 71–89 (G4).** "G4" is never introduced in prose at all — it appears as a table with
no lead sentence. Column "GROUNDED" — grounded in *what*, and judged by whom, and against
what floor? Column "recused cells" — recusal is not explained until line 89, and even there
only as "the google family's own seat"; I had to infer for myself that a Google model may
not judge a Google model. That is a *good* rule. Say it: one sentence, before the table.

**Line 82.** "alibaba (qwen3.5-397b) 34/35"
Every other family in that list has denominator 36. This one is 35 and nothing says why. A
denominator that changes mid-list with no note reads as an error.

**Line 95.** "23 favoured `cli-claude-fable-5-1`, 0 tied, 13 favoured `openai-gpt-6-astra`
— pooled preference rate 0.544 over 36 cases"
**These two numbers appear to contradict each other in the same sentence.** 23 of 36 is
0.639, not 0.544. I assume the rate is a mean of collapsed per-judge scores and the 23–13 is
a count of case winners, but the sentence sets them side by side as if they agree, and the
page never reconciles them. This is the headline number of the section.

**Line 95.** "changed a verdict on 75 of 224 (judge, case) comparisons"
Line 95 also tells me the design was "7 seats × 36 cases × 2 orders = 504 judged calls",
which is 252 comparisons, not 224. Line 154 then introduces a third denominator, "70 of 449
collected cells". 504 registered, 449 collected, 224 with both orders — three numbers, no
sentence connecting them, no statement of what went missing.

**Line 95.** "a property of the PANEL, not of the arms"
Swapping the two answers flipped the verdict **a third of the time (0.335)** and the page
names it, labels it a panel property, and moves on. If your judges flip a third of the time,
what does that do to the credibility of the headline they produced? That is the reader's
next question and it is never answered.

**Line 95.** "a cluster bootstrap over the 36 cases (2,000 resamples, percentile; t(35) —
36 clusters is barely above the N ≥ 30 line this round registered…)"
"Cluster bootstrap", "resamples", "percentile", "t(35)", "clusters" — five statistics terms
in one parenthesis with no translation. I cannot tell if this paragraph is being careful or
hedging.

**Line 146.** "deepseek | `deepseek-v4-pro` | 9 | 4 | NO-RATE…"
One of your seven judges carried **9 cases while the other six carried 36**, and the page
gives no reason. A quarter-strength judge in a seven-judge panel is a fact about the
instrument and it passes without comment.

**Line 147.** "21.25"
Fractional case counts, unexplained. "Favoured Fable (sum) = 21.25" reads like a typo unless
you already know about the collapsing rule from line 95.

**Line 154.** "70 of 449 collected cells carried such a claim."
Seventy judges said they recognised who wrote an answer — **and the page never says whether
they were right.** You raise the question that most threatens your blindness claim and
answer a different one (you re-run the headline without those rows). Were they correct?

**Line 158.** "32k (30,000 estimated tokens)"
The label contradicts its own value with no note. Is the tier 32,000 or 30,000?

**Line 158.** "screened against the round's forbidden-literal list"
Never introduced.

**Line 162.** "Fixture sha 76cef1d1…."
A truncated hash alone on a line, with no sentence around it. To me this is furniture.

**Line 164.** "cells under the 0.80 ctx floor"
"ctx floor" in a column head; the explanation lands two paragraphs later (line 207) as
"reported prompt tokens over our own estimate", which still doesn't tell me *why I should
care* — I eventually worked out it's a check that the model actually received the whole text.
Say that.

**Line 180.** "`local-gemma4-26b` | 32k | 0/6 | 0/6 | 0/6"
This reads as a total collapse — the small model can't handle long texts. But the real
reason is buried at **line 257**, seventy-seven lines later: "NOT-COLLECTED — CONTEXT 12".
It didn't fail; the prompt didn't fit. A zero that means "not attempted" must say so in the
cell.

**Line 207 vs line 257 — a contradiction I cannot resolve.** Line 207: "so the widest tier
sits inside the seat's own context, and the ratio is published rather than assumed." Line
257: "the filing cabinet, `local-gemma4-26b`: COLLECTED 24 · NOT-COLLECTED — CONTEXT 12."
One says the 32k tier fits; the other says twelve cells were lost to context. Both are on
the same page about the same model and neither mentions the other.

**Line 199.** "**The polarity control**, printed beside the grid and never instead of it"
I don't know what a polarity control is, what it would have shown if something were wrong,
or why the row is identical to the main one. The sentence tells me about the page's
typesetting policy instead of about the experiment.

**Line 205.** "**Self-refutation (PREREG A3 FOLDED (4)):**"
"Self-refutation" and "FOLDED (4)" are pure internal notation. As a heading, it made me
think the page was about to withdraw a claim.

**Line 207.** "The double canary planted at both ends of one 32k prompt"
"Canary" unintroduced. And the local model "did NOT find both" — a failed integrity check
stated and then dropped with no consequence drawn.

**Line 211.** "Both frontier rows are a new kind of row for this shelf, and the rules for it
were written before the round ran. **Four of them.**"
Four *what*? And the next paragraph (line 213) opens "**Three** reading rules ride every
table" and lists three. You promised four and delivered three, or you're counting two
different sets — either way I lost the thread of the section that is supposed to tell me
what the numbers mean.

**Line 217.** "the self-agreement column prints per arm with its sampler state in the same
cell."
**There is no self-agreement column anywhere on this page.** I went looking for it.

**Line 217.** "measuring an input-token delta of exactly 280 tokens on every pair"
Line 217 also says the preamble is "370 estimated tokens". 370 vs 280, unremarked. And the
question I actually have about scaffolding — did it change the *answers*, not the token
count — is never addressed.

**Line 221.** "We publish the reasoning-token counts each produced so a reader can see what
the word bought"
You publish one of them, buried inside a cost-basis cell (line 268, "incl. 92,960
reasoning"). Fable's reasoning-token count appears nowhere. The promised comparison is not
on the page.

**Line 236.** "2 of 5 sampled connections matched a vendor API host — api.anthropic.com — 1
address · api.openai.com — 2 addresses · ollama.com — 1 address · statsig.anthropic.com — 0
addresses; … The 3 addresses that matched no vendor API are named…"
The arithmetic does not close in front of me: "2 of 5", then a list summing to 4 addresses,
then "3 addresses that matched no vendor API" (2 + 3 = 5, but the list shows 4 matched).
Connections vs addresses vs hosts are three different units used as one.

**Line 236.** "70 environment names dropped and 5 allowed through"
No frame. Is 70 thorough or alarming? Was 5 the minimum? Compared to what?

**Line 238.** "One bench in August crossed that line before this page pointed the promise at
bench data; the promise page carries a dated addendum about it, in an operator's words."
Something went wrong with real users' questions once, and I'm told in a subordinate clause,
with no detail, no date, no link, and no named person ("an operator"). This raises an alarm
and withholds the story — the worst combination.

**Lines 244–257.** Heading: "What was NOT-RUN, with its reason:" — and then six of the
twelve bullets say **SCORED** and three say **COLLECTED**. Things that ran are listed under
a heading that says they didn't. The first bullet is also circular: "case 61 … is in the
bank but no rows were collected for it" — that restates NOT-RUN; it is not a reason.

**Line 267.** "cost state | no-figure-held … | $31.44"
The cell says no figure is held and the next cell holds a figure. The basis text explains
it's a list-rate estimate, but at a glance the row contradicts itself. Also: no total for
the round anywhere, and six seats priced "$0.00" as "plan-included" — a plan someone pays
for, which the page doesn't disclose.

**Line 273.** "$5.25 … priced at the UNCACHED input rate (cached input bills at a tenth…)"
So this is an upper bound, and the USD column doesn't say so.

**Line 288 onward (the takeaway bullets).** Every bullet is written in slug names —
`cli-claude-fable-5-1`, `openai-gpt-6-astra`, `local-gemma4-26b` — not in the human names
you introduced in paragraph one. The section a hurried reader reads first is the one that
reads most like a config file. And none of the five bullets contains a verdict.

**Line 296–299 (How to check our work).** "The kit" is invoked 13 times across this page and
**never once given an address.** No URL, no filename, no "it's at …". Same for "Ask the
rules desk yourself. The rules helper is live" — live *where*? The product has no name and
no link on this page. Same for the pre-registration I'm told to "read before the results",
and for `index.json`. There is exactly one address on this entire page: an email at line
305. The page's central trust claim — everything is re-derivable — is unreachable.

**Line 303.** "Behind this page: *A kid, an elder, and a tired parent walk into the cove*
(the open call, where the blind-panel-with-recusal method was built…), *The same sixteen*
(where the reference-arm class was chartered), and *The new kid*…"
Three titles, no links, no dates, and four more undefined terms in the parentheses ("the
open call", "the cove", "reference-arm class", "chartered", "trust contract"). Also the
section is called "The rest of the seminar" — first use of "seminar" for the body of work,
on line 301 of 320.

**Line 317.** "A hostile reader went through this page before it went out (PENDING — this
draft pours to the drafts shelf before the hostile reader's pass…): no hostile pass has run
yet on this version"
**The sentence asserts a thing happened and then says it did not happen.** As printed, this
is a false statement followed by its own retraction, inside one set of parentheses, in the
credits. "Pours to the drafts shelf" is jargon too.

---

## 3. Jargon used before (or without) explanation — named

House vocabulary, never introduced: **arm** (first used line 35, defined never) ·
**leg** / Leg A / Leg C / Leg H / C-think-true · **seat** (and "house seat", "shelf gemma
seat", "seat posture", "eligible for any seat") · **shelf** (three different senses) ·
**transport** (and its values agent-harness-cli / openai-api / local-ollama) · **the kit** ·
**exhibit** (front matter "exhibit: forty") · **the cove** · **the estate** · **an
operator** · **the house** / house control / house recall · **floor** (as a noun) ·
**gate** · **rep** / "any-rep read" · **pour** / "pours to the drafts shelf" · **the
seminar** · **the open call** · **reference-arm class** · **chartered** · **citation
contract** (dek) and **trust contract** (line 303) · **answer ledger** · **one-tap
question** · **needle** · **canary** / double canary · **fixture** · **tier** ·
**window** (line 207, meaning a slice of filler) · **polarity control** ·
**self-refutation** · **the seal manifest** · **forbidden-literal list** ·
**answer-side directive lint** · **posture** · **NOT-COLLECTED — MODEL-FALLBACK** ·
**no-mode cases** · **collapsed observation** · **recusal** / recused cells (explained 4
lines after first use) · **GROUNDED** · **blind** (mechanism never stated).

Statistics and notation, never glossed: the **bracket pairs** `[0.819, 0.985]` (never named
as confidence intervals, no level stated) · **interval** · **cluster bootstrap** ·
**resamples** · **percentile** · **t(35)** · **clusters** · **pooled preference rate** ·
**sensitivity cut** · **N ≥ 30** (asserted as a rule, never justified) · **order flip
rate** · **ctx floor** / context ratio · **estimated tokens** and **chars÷4** ·
**num_ctx**.

Raw machine identifiers printed at the reader: `cite_survival` · `stripped_count` ·
`is_abstention`, `answering.py:304-311` · `prompt_eval_count` / `eval_count` ·
`total_cost_usd` · G1 / G3a / G3b / G4 / G5a / G5b / G5c / G6b / G6c / G_CALIBRATE /
G-EGRESS / G-PEN · `ofl-ans-0001` · `ofl-ctx-24k` · v0.57.0, commit 76169d2 ·
truncated shas (`6d57ffbb…`, `5e274118…`, `aaa6d35a…`, `76cef1d1…`, `631c2f98…`) ·
27 × PREREG §/A section references · amendment A2 · C7 · A3 FOLDED (4) · A8.

Model slugs used in prose in place of the human names the page introduced:
`cli-claude-fable-5-1`, `openai-gpt-6-astra`, `local-gemma4-26b` — every takeaway bullet.

---

## 4. Where the page talks past me, or makes me feel dumb

1. **It withholds the verdict it spent the whole page earning.** Floors frozen in August,
   bolded twice as "the whole method of this workshop", then five numbers printed and no
   comparison performed. Two of three arms miss G1; all three miss G3a. The page never says
   it. Either you didn't notice, or you decided not to say — and a stranger will suspect the
   second, which costs you exactly the trust the freezing was meant to buy.
2. **It answers audit questions and skips reader questions.** I learn the outside reviewer's
   latency in milliseconds but not its finding. I learn that 70 judges claimed to recognise
   an author but not whether they were right. I learn the order-flip rate is 0.335 but not
   what that does to the headline. Three times the page produces a receipt where a sentence
   was owed.
3. **The receipts talk to a machine.** Truncated hashes, `answering.py:304-311`, "transport
   ok: yes", "A scorer that did not score an arm leaves a missing row, never a renumbered
   one", "the count prints here as the disclosure it is". I am being shown that a process
   was followed, in the vocabulary of the process. It reads like a build log that learned to
   write prose, and it makes me feel like an intruder in someone else's audit.
4. **Every route out of the page is closed.** 13 mentions of "the kit", 27 citations to a
   pre-registration, an invitation to try the rules helper, an instruction to run the
   filing cabinet myself, three prior articles named — and **not one URL on the page.** The
   page says "Take the rows and re-derive our figures" and gives me no rows. That is the
   moment I stopped believing the transparency was for me.
5. **Numbers that visibly don't add up go unremarked**, so I can't tell whether I'm confused
   or the page is wrong: 23–13 alongside a rate of 0.544; 504 vs 449 vs 224; "2 of 5" against
   four listed addresses and three unmatched ones; 370 estimated vs 280 measured tokens; "32k
   (30,000 tokens)"; a 34/35 denominator in a list of 36s; 162 vs 180 collected. Each one is
   probably explicable in one clause. None gets it.
6. **The page contradicts itself in three places and never notices**: the 32k tier "sits
   inside the seat's own context" (207) vs "NOT-COLLECTED — CONTEXT 12" (257); "What was
   NOT-RUN" listing nine things that ran (244–257); "A hostile reader went through this page"
   vs "no hostile pass has run yet" (317). Plus a promised "self-agreement column" that does
   not exist, and "Four of them" followed by three.
7. **Broken text shipped in the load-bearing part.** The G5b floor sentence ends on the word
   "by" (line 65) and the G5a/G5b bullets print their own labels twice. Those are the
   definitions of the pass marks. If the bar itself is a truncated sentence, the "frozen in
   August" argument has nothing to stand on for a reader.
8. **The insider frame is assumed from the first line.** "Exhibit forty", "the cove", "the
   narrator's chair in a small fishing town that our game engine remembers for its players",
   "the rest of the seminar" — the page is written as episode 40 of a serial. I came in on
   episode 1 of what I read. Nothing on the page tells me who "we" are, what the products are
   called, or where they live.
9. **The star of the page is the last thing named.** "The small model that already answers
   our users" carries two of your five closing bullets and is identified as "Gemma 4 (26B)"
   in the third-from-last paragraph.

---

## 5. Three fixes I would make first

**Fix 1 — State the verdict against the frozen bar. Insert immediately after line 67**
("Every count above prints its denominator…"), as its own bold paragraph:

> **Against those floors: `openai-gpt-6-astra` cleared four of the five gates and missed
> G3a (7/12 abstentions against a floor of 11/12). `cli-claude-fable-5-1` cleared two,
> missed G1 by one case (34/36 against 35/36) and missed G3a (7/12), and has no result on
> G5b because its rows were not collected. `local-gemma4-26b` cleared two, missed G1
> (34/36) and G3a (10/12), and followed 2 of the 6 hidden directives. Every one of those
> marks was written in August, by other hands, before either frontier model existed — this
> is the comparison that freezing the bar was for, and nothing on this page moved it.**

Then change the takeaway bullet at line 291 from "Nothing on this page set a threshold for
either of them." to: "Nothing on this page set a threshold for either of them — and neither
of them cleared all five."

**Fix 2 — Give the page an address, and a name for the thing under test. Replace line 296's
first sentence** ("**The kit.** Every table on this page is a projection of a file in the
kit…") with:

> **The kit — every file behind every table on this page — is at
> `<https://…/kits/two-new-frontier-models-at-the-rules-desk/>`, and `index.json` there
> lists each file with its sha256, what it answers, and what was scrubbed from it: the gate
> counts, the head-to-head's per-case rows and its key, the filing cabinet's cells, the
> counting rules, the seats, the bill, the seal manifest, the pre-registration as a public
> copy, and the three scorers as code.**

And in the same section replace "**Ask the rules desk yourself.** The rules helper is live;"
with: "**Ask the rules desk yourself.** The rules helper is `<product name>`, live at
`<https://…>`;". (Same treatment for the pre-registration at line 299 and the three prior
articles at line 303 — every one of them is currently a reference with no address.)

**Fix 3 — Reconcile the head-to-head's two headline numbers in the sentence that prints
them. In line 95, replace** "— pooled preference rate 0.544 over 36 cases, and a cluster
bootstrap over the 36 cases" **with**:

> "— that is a count of which arm won each case outright. The rate the interval is built
> from is a different unit and a smaller number: each judge's two readings of a case are
> averaged first (preferred = 1, tie = 0.5), then the 36 case averages are pooled, which
> gives 0.544 — near-even, because most cases were won narrowly rather than unanimously.
> A cluster bootstrap over the 36 cases"

And in the same paragraph, after the order-flip sentence, insert: "**A third of the
comparisons flipping on order is the reason this page reports a range that includes 0.5 and
declines to rank the two arms: a panel that unstable cannot separate answers this close.**"

---

### Broken text that must be fixed regardless (not counted among the three)

- **Line 65** ends mid-sentence: "**≥ 5/6 must abstain OR be refused by**" — finish it.
- **Lines 64–65** print "G5a" and "G5b" twice each with a stray "- " between.
- **Line 317** asserts a hostile pass happened and then says it hasn't. Delete the claim or
  the parenthesis; as printed it is a false sentence in the credits.
- **Line 217** promises a "self-agreement column" that does not exist on the page.
- **Lines 244–257**: a heading reading "What was NOT-RUN" over nine entries marked SCORED
  or COLLECTED.
- **Line 211** says "Four of them"; line 213 lists three.
