# The kit behind "Reading the answer instead of writing it"

*What this directory is, what was held back from it and why, and what was rewritten on
the way out of the private record. Written 2026-09-22 (UTC) for the draft of that page,
refreshed the same day when the gap card's rungs and arm 12 landed, and again on 2026-09-23
(UTC) at the release cut, when one rewrite rule was added (below) and one sentence of
`TABLES-ADDENDA.md` was corrected at its generator, which now counts the items that changed
answer instead of estimating them. The bench ran 2026-09-21T16:51:34Z to 2026-09-22T10:43Z — the
pre-registration's own stamp at one end, and arm 12's last scored row at the other
(`rows-gaps-card/kev-4b.kperm.jsonl`, 10:43:21Z). This note said 09:49Z until the release
cut, which is no stamp in the record: the gap card's last rows are 10:06:13Z and arm 12's
10:43:21Z.*

---

## What is here

296 files, 12,070,484 bytes. They are the bench's own record, copied out of the tree that
ran it:

- **`README.md`** — the pre-registration, written and committed before any row was
  measured, with the results appended under it as they landed. **`GAPS.md`** — what the
  first eleven arms did not measure and could not settle. **`GAPS-CARD.md`** — the gap
  card: each of those gaps with the rung that would close it, the pre-registration for
  the rungs that then ran, and what each one found.
- **The tables** — `TABLES.md`, `TABLES-ARM2.md`, `TABLES-ADDENDA.md`, `TABLES-GAPS.md`
  and `TABLES-GAPS-CARD.md`, every one derived from the row files by a generator that
  ships beside it.
- **The harness, as it ran** — `run.py`, `run_addenda.py`, `run_arm8.py`, `run_arm11.py`,
  `run_gaps_card.py` (the gap card) and `run_kev.py` (arm 12), the five `tables*.py` generators, `phb_allreduce.py` (the
  two-card link measurement), `rerank_baseline.py`, `build_ownjev_tasks.py`,
  `arm3_window.sh`, and the two scripts that stood arm 12 up on the bench box,
  `kev-serve.sh` and `kev-setup.sh`.
- **`kit/`** — the frozen task states for the first eleven arms, with a sha256 per state
  in `kit/manifest.json` and the builder that froze them. Every row records the sha256 of
  the prompt it was asked.
- **`rows/`, `rows-arm3/`, `rows-addenda/`, `rows-gaps/`, `rows-gaps-card/`** — every
  published arm's scored rows, its own report, the residency census each proved before it
  was allowed to run, and the 1 Hz power samples the joules are integrated from.
- **`receipts/`** — what the run itself printed: the substrate each arm found, the serving
  logs, and each arm's own console output.

`index.json` is the machine-readable front door and `provenance.json` is the record of
every rewrite, rule by rule, file by file, with both digests of each file a rule touched.

## What is held back, and why

**The texts of three task sets are not here, and neither is any arm's per-item run of
them.** The doorman's planted hostile lines, the rulebook excerpts behind the field exam
and the judge set's items are not this workshop's to publish. Their **scores are** —
`TABLES-ADDENDA.md` carries arms 4, 5 and 6, and `TABLES-GAPS-CARD.md` carries arm 12's
runs of the same three sets — and their per-item row files are not. That covers the
original arms' `.d`, `.f` and `.j` rows, arm 12's `.kd`, `.kf` and `.kj` rows, and the two
arm-11 held-out splits taken over the same sets. The page says the same thing in its own
last section.

**What is NOT held back from the gap card**: the shuffled-option re-run
(`*.creorder.*`) and the two screenshot tasks (`*.shot_*`) ship in full, because both are
asked over this workshop's own published pages and the long table's own public wall.

**The outage runbook is not here.** The operator runbook for the 45-minute window on the
large card records what was stopped and restarted on live machines and carries a table of
this workshop's own product traffic by local hour. It measures nothing on the page. The
window it plans is stated at both ends in `README.md`, which is here.

**The screenshots themselves are not here.** The two screenshot tasks were asked over 36
PNG captures of this workshop's own published pages. The pages are public and the rows
record the sha256 of the capture each decision was made over; the image files are the
print surface rather than the measurement, and a rule-cut kit publishes text it can
re-derive, not bytes no rule can read.

**The weights are not here.** The checkpoints are downloads from their publishers, named
with their licences in the page's own licence section. `kev-setup.sh` and the setup
receipt name the upstream revision arm 12's server was pinned to, so a reader can build
the same thing.

## What was rewritten, and by what rule

Every rewrite below is a named mechanical rule applied by one program
(`tools/kit_sanitise.py` in the hub repo) in a fixed order, to every file, so a holder of
the private file can re-derive the public byte. **170 of the 296 files are byte-identical
to the private record** — no rule fired on them at all — and `provenance.json` says which,
and what fired on the other 126.

| what it was | what it is here | times |
|---|---|---|
| A path into a private home directory | `/workshop` | 643 |
| The machine that ran every arm but one | `benchbox` | 164 |
| The machine that holds this workshop's live seats | `inferencebox` | 55 |
| The machine that serves this workshop's public pages | `the public host` | 2 |
| The machine the kit builder runs on | `this laptop` | 1 |
| The large board's pet name, as a variable | `LARGECARD` | 12 |
| The large board's pet name | `largecard` | 986 |
| The inference box's other board, by its pet name | `the 3090 board` | 7 |
| The private network's name | `a private network` | 3 |
| A directory listing's owner and group columns | `operator operator` | 24 |
| The operator's own name, in a ruling or a note | `an operator` | 24 |
| The runtime's default port, outside a documented URL | `its default port` | 4 |
| A private network address in a startup banner | `<a private address>` | 9 |
| A planned window restated on a local clock | the same window, named | 1 |
| An instant carrying a local offset | the same instant, in UTC | 2 |
| A wall clock with no zone at all, in the four receipts of the one arm run on the inference box | the same instant, in UTC | 267 |

The two instants are an upstream project's own commit time, quoted by the bench when it
pinned the revision it built. They are not this workshop's clock, and they are converted
rather than cut for the same reason any instant here is: deleting an offset would silently
re-label a local clock as UTC and move the moment by hours.

**The 267 zone-less clocks are the inference box's own.** Arm 3 is the one arm that ran on
that box, whose clock is local, and its four receipts (`receipts/arm3-bf16.log`,
`receipts/arm3-fp8.log` and the two `receipts/vllm-serve-arm3-*.log` beside them) printed
that clock with no zone marked at all — the serving runtime's `INFO MM-DD HH:MM:SS` prefix
and its tuning library's `YYYY-MM-DD HH:MM:SS,mmm` — so neither the gate nor the two rules
above could see them. They are converted rather than cut, like every other instant here,
and the rule proves the conversion on every build: the harness's UTC stamp and the server's
own stamp of the same event, the shutdown of the same process, land in the same second, and
every converted instant must fall inside the window its own run's UTC stamps bracket or the
cut refuses. Every other receipt's zone-less stamps are the bench box's clock, which is UTC,
and they are left exactly as written.

**26 row files were RENAMED** as well as rewritten: the large board's pet name was in their
filenames, and a kit's filenames are printed on the listing page beside this note. The rule
that rewrites every citation of those names is the same one that renames them, so nothing
in the kit now cites a file that is not here.

**No figure moved.** The tool's own fence re-reads every numeric token on both sides of the
cut and refuses on any difference: **730,088 numeric tokens, identical in order and in
value**, 269 instants converted and each proved to be the same moment, and 0 role words glued
to a digit run. The only spans cut out of that comparison are the name spans above whose two
sides disagree about digits, and the fence prints every one of them with its count on every
run.

## The one place a reader should not expect a sha to match

The operator rule fires inside **two frozen task states** and the rows and logs that carry
them: task (b)'s option list quotes a ruling that names the operator. So an item of task
(b) rebuilt from this public copy will not reproduce the `prompt_sha256` the row beside it
records. **Task (a) and task (c) are byte-identical here** — `kit/task_a.json`,
`kit/task_c.json` and `kit/articles.json` left this cut exactly as they went to the model —
and task (c), the 108 six-way items, is the one the page's reproduction claim is about.
`provenance.json` lists all 19 files that rule fired in.

## Licence

CC BY 4.0, as attributed in `index.json`. The models, corpora and frameworks this bench
stood on carry their own licences and the page's licence section names each one.
