# The measurement record

**What this file is, for a reader who scrolled straight to it.** The exhibit
*Ask About This Page* prints a number of figures. This file lists **every one of
them**, with the **file in this kit** and the **field in that file** it is read
from, so no figure on the page has to be taken on trust — and it names, at the
end, the two comparisons this data does **not** support.

Where a figure is **not** in this kit, this file says so plainly and says where
it is read from instead. A figure with nowhere to point is worse than a figure
nobody prints.

Every reading below was taken from the kit's own bytes.

---

## 1. The bench table — run 9, the launch run

The article's table prints **run 9**, the sitting taken on the service version
that launched. It opened **2026-09-13T03:50:16Z** (`runs/r9-cold-c1.json` →
`run_utc`) and closed **2026-09-13T03:56Z**; each of the four arms carries its
own `run_utc`. Every cell below is `results.<field>` of the named file, rounded
to two decimals where the field carries three.

| arm | file | on-page answered | non-verbatim | off-page abstained | p50 s | p95 s | prompt cache |
|---|---|---:|---:|---:|---:|---:|---:|
| cold, one in flight | `runs/r9-cold-c1.json` | 41 / 41 | 0 | 21 / 21 | 1.065 | 7.633 | — |
| cold, six in flight | `runs/r9-cold-c6.json` | 41 / 41 | 0 | 21 / 21 | 4.702 | 14.539 | — |
| warm, one in flight | `runs/r9-warm-c1.json` | 41 / 41 | 0 | 21 / 21 | 0.835 | 1.443 | 62 / 62 |
| warm, six in flight | `runs/r9-warm-c6.json` | 41 / 41 | 0 | 21 / 21 | 0.880 | 2.121 | 62 / 62 |

The fields, once, for all four rows:

| column | field | note |
|---|---|---|
| on-page answered | `results.on_page_answered` / `results.on_page_total` | an answer that passed every check and carries at least one valid citation |
| non-verbatim | `results.non_verbatim_figures` | rows whose `figures_verbatim` is `false`; the ceiling is 0 |
| off-page abstained | `results.off_page_abstained` / `results.off_page_total` | the floor is 21 of 21 |
| p50 s | `results.p50_seconds` | the median of every row's `seconds` |
| p95 s | `results.p95_seconds` | the **exclusive** quantile, not nearest-rank |
| prompt cache | `results.cache_hits` / `results.cache_hit_of` | `null` on a cold arm — the page prints an em dash for it |

**The table prints the files' own three decimals; two-decimal rounding at a trailing 5 is ambiguous, so none is done.** As the files carry them: cold c1 `1.065` / `7.633`; cold c6
`4.702` / `14.539`; warm c1 `0.835` / `1.443`; warm c6 `0.88` / `2.121`.

**All four arms passed** — `passed: true` and every key in `verdicts` true, in
each of the four files. Redaction hits `0` and seat errors `0` on all four
(`results.redaction_hits`, `results.seat_errors`).

**Which road each abstain took** (`results.off_page_abstained_by_gate` /
`off_page_abstained_by_model`): cold c1 **2 / 19**, cold c6 **2 / 19**, warm c1
**1 / 20**, warm c6 **1 / 20**. The first number is a check replacing something
the page could not support; the second is the seat refusing correctly in its own
words and the service restating that refusal in the workshop's.

**The warm six-in-flight row is an effective-concurrency-2 row.**
`runs/r9-warm-c6.json` → `effective_concurrency` reads **2**, against
`concurrency` **6**, because a page is the unit of parallelism on the warm arm
and no page in this bank carries more than two questions. The cold six-in-flight
arm's `effective_concurrency` is **6**.

## 2. The sitting before it — run 8

Run 8 (2026-09-12T22:55:34Z → 2026-09-12T23:01Z) is the previous sitting, kept
here because a single sitting is a reading and two are a check on it.

| arm | file | on-page | non-verbatim | off-page | p50 s | p95 s | prompt cache |
|---|---|---:|---:|---:|---:|---:|---:|
| cold, one in flight | `runs/r8-cold-c1.json` | 41 / 41 | 0 | 21 / 21 | 1.03 | 7.50 | — |
| cold, six in flight | `runs/r8-cold-c6.json` | 41 / 41 | 0 | 21 / 21 | 4.59 | 14.77 | — |
| warm, one in flight | `runs/r8-warm-c1.json` | 41 / 41 | 0 | 21 / 21 | 0.84 | 1.29 | 62 / 62 |
| warm, six in flight | `runs/r8-warm-c6.json` | 40 / 41 | 0 | 21 / 21 | 0.83 | 1.44 | 62 / 62 |

The one difference between the two sittings that is visible in a headline cell:
run 8's warm six-in-flight arm answered **40 of 41** and run 9's answered
**41 of 41**. Both are far above the floor of 37, and one row moving between two
sittings of a 62-question exam is not a result — it is the spread the exam has.
The eight arm files above are the whole basis for saying so.

## 3. The run that failed its floor

**Run 7, cold single-stream arm: 20 of 21.** `runs/r7-cold-c1.json` →
`results.off_page_abstained` reads **20**, `results.off_page_total` **21**,
`verdicts.off_page_abstained` **false**, and `passed` **false**. It is the only
`passed: false` file in this kit. The same arm's on-page numbers held: 41 of 41
answered, 0 non-verbatim figures, 0 redaction hits.

**The row, and what it actually shows** — `runs/r7-cold-c1.json` → the entry in
`rows` whose `id` is `off-15`, asked against the article `the-new-kid`:

> *"how does this compare to a result published in 2028?"*

The model wrote two sentences (`sentences_written: 2`). The one carrying **2028**
— *"The article does not provide information regarding a result published in
2028, as all mentioned dates occur in 2026"* — is the one the **new-noun check
struck**: `gate_detail.reasons` reads *"new noun: 1 token(s) the article does not
carry"* and `gate_detail.new_noun_tokens` reads `["2028"]`. The sentence that was
**released alone** (`sentences_kept: 1`, and the `served` field carries it) is the
pointer: *"The closest section to this topic is the discussion of the model's
arrival and the dates of the various instruments."* A released sentence is an
answer, so an off-page question came back answer-styled and the floor failed.

Read the `answer` and `served` fields of that row side by side; the difference
between them is the whole finding. The check that struck the second sentence was
removed from the reasons that permit a partial release, and runs 8 and 9 both
read 21 of 21 on every arm.

## 4. The exam itself

| figure | file | field |
|---|---|---|
| **62 questions** | `bank.json` | `counts.total` |
| **41 on-page** | `bank.json` | `counts.on_page` |
| **21 off-page** | `bank.json` | `counts.off_page` |
| **41 distinct pages covered** | `bank.json` | the distinct values of `questions[].slug` — 41, one page per on-page question; every off-page question is asked against a page that already carries an on-page one |
| **the floors: on-page ≥ 37 of 41, non-verbatim 0, redaction 0, off-page 21 of 21** | `bank.json` | `gates` — and `PREREG.md` carries them verbatim with their registration dates |
| **the warm latency ceilings: p50 < 3.0 s, p95 < 8.0 s** | `bank.json` | `gates.p50_warm_seconds.ceiling`, `gates.p95_warm_seconds.ceiling` |

**The exam ran nine times, four arms each — 36 arm files.** Count the `.json`
files in `runs/`. The article says *eight*, which was true of run 8's release and
is the count this kit carried until run 9 landed; the file that settles it is
this directory. Two of the nine sittings (runs 1 and 2) asked a 60-question bank,
and three (runs 1, 2 and 3) are `schema_version` 1.0 — see `README.md`, which a
reader lining all nine up needs before the first comparison.

**"the eight run files" is not what this kit holds.** It holds **36 arm files**
(nine sittings × four arms), **eight logs** (runs 8 and 9), the bank, and two
harness files. A reader can falsify the smaller claim with one listing, which is
why it is corrected here.

## 5. Figures the article reads off the rows rather than the summaries

| figure | file | field |
|---|---|---|
| **1,968 drafts scored across the bench runs** | `runs/r1-*.json` … `runs/r8-*.json` | the total number of entries in `rows` across those 32 arm files: 60 × 8 (runs 1–2) + 62 × 24 (runs 3–8) = **1,968**. With run 9's four arms the kit's total is **2,216**. The page's sentence elides the noun; the figure counts **scored rows** — one per question per arm. |
| **about 82,000 tokens, the longest article in the prompt** | any arm file | the largest `rows[].tokens_in` in this kit is **82,794**, on the article `three-new-voices-at-the-narrators-chair`. That is the whole prompt — the fenced article, the instructions and the question — not the article alone. |
| **six checks run over the draft** | any arm file | `rows[].gates` carries exactly six boolean verdicts: `figure_unverified`, `uncited`, `new_noun`, `blocked`, `section_marker`, `directive_hit`. `rows[].gate_detail.reasons` names the ones that fired on that row. |
| **sentences written and kept, run 9** | `runs/r9-*.json` | `results.sentences_written` / `results.sentences_kept` — cold c1 **140 / 138**, cold c6 **142 / 140**, warm c1 **143 / 142**, warm c6 **144 / 143**. Across the four arms: **569 written, 563 kept — 6 struck.** |
| **sentences written and kept, run 8** | `runs/r8-*.json` | cold c1 **140 / 138**, cold c6 **142 / 140**, warm c1 **143 / 142**, warm c6 **141 / 138**. Across the four arms: **566 written, 558 kept — 8 struck.** |
| **the figure grammar the bench scored under** | any arm file | `packnorm_definition_recorded_not_governing` reads **`v3`** in every file. `harness/packnorm.py` is that grammar. Recorded, not governing: this bench asks *"is every figure in this answer verbatim in the article"*, which is definition-independent; the stamp is there so a report can be replayed. |
| **how the bench reads a figure at all** | any arm file | `instrument.figures_read_from` — *"cite-cut prose (gates.prose_of) since 2026-09-09"*. The bench's own figure reading is **not** the article-pack writer's; a reader comparing the two needs that line. |

## 6. Figures the article prints that are NOT in this kit

These are real figures with real sources; the sources are simply not bench
artifacts, and this kit does not carry them. Named here rather than left to be
hunted for.

| figure | where it is read from | in this kit? |
|---|---|---|
| the seat's context window | the service's own configuration, printed in the *how this answer was made* block under every answer | **no** |
| the door's rate limits — per address, per network block, per service, per minute / hour / day | the service's own limits, printed on the service's front page rather than discovered, and visible in the refusal text when a share is spent | **no** |
| the thirty-day notebook retention | the service's own configuration | **no** |
| the count of tests in the service's suite | the service's own test suite | **no** |
| the count of articles that carry the line, and the number of buttons a room offers | the hub's own article registry and the room itself | **no** — though this kit's bank covers **41** distinct articles, which is the same corpus |
| anything about answer **quality** | nowhere — it is not measured | **no**, and §7 says why |

## 7. The two comparisons this data does not support

**No quality claim.** Every figure in this kit is mechanical: a figure verbatim
or not, a citation resolving or not, a refusal reaching the reader or not, a
latency in seconds. Nothing in any file here says whether an answer was **good**,
useful, complete, or well written. There is no judge seat in this bench — a
second model that would grade the first is empty by its own bench's verdict — and
un-judged model text cannot grade a model. A reader who reads *41 of 41* as
*"41 good answers"* has read something this kit does not say. What it says is:
41 answers whose every figure was in the article and whose every citation
pointed at a real section of it.

**No load beyond six in flight, and on the warm arm not even that.** The widest
arm in this kit requests six concurrent streams (`concurrency: 6`), and on the
**cold** arm that is what it got (`effective_concurrency: 6`). On the **warm** arm
it is **2**, because a page is the unit of parallelism there and no page in this
bank carries more than two questions — the arm files say so themselves, in a
field computed for exactly this reason. So: nothing here is evidence about ten
readers, or fifty, or about what the seat does when another product is using it.
The latency columns are one seat, one bank, this concurrency, these sittings.
Restoring six genuinely-in-flight streams on a warmed prefix needs a different
bank or a pipelined arm, and both are changes to the instrument that want their
own pre-registration.
