# The kit behind "Small open deciders, measured"

This folder holds the rows, receipts, pre-registrations and recount behind the page's figures that
[the Reading the answer kit](https://research.strata2signal.com/reading-the-answer/data/) does not
already hold. Arms 1, 2, 3, 10 and 12 and gap 4 of the Jev bench are there, with the six-way's 108
items (`kit/task_c.json`). Everything else the page counts is here, except the sets held back below.
Their counts and figures are in `recount-v2.json` and on the page.

Cut 2026-09-29 (UTC) from the benches' record at one commit, and re-cut the same day four times:
after a privacy audit and an accuracy audit of the first cut, to fold the page's own checks into the
recount, after a final privacy pass, and to print three pre-registration seals again once a ruling
kept them. Re-cut on 2026-09-30 (UTC) for the page's release. Nothing in this kit changes a figure
on the page.

## What is here

The folders keep the record's own paths, less the leading `bench/`, so every path the page, the
evidence file or the recount prints can be found as written.

| Folder or file | What it holds |
|---|---|
| `jev-2026-09-21/` | Arm 13, OpenJev 27B's 4-bit GGUF through ollama on card B: its pre-registration, the six-way and six-order rows, with each task's report and telemetry, and the counted run's record |
| `cost-of-intelligence-2026-09-23/` | Mistral Small 3.2 24B's six-way rows from the cost series, on card A and on the workstation card, and `KIT.lock.six-way.json` |
| `deem-2026-09-27/` | deem-0.8-v1's six-way rows (D0), the three T1 controls, the word counter whole, and two pre-registrations |
| `lev-2026-09-28/` | Lev-4B on a laptop's processor (L-1) and on card B (L-2) |
| `apus-2026-09-28/` | APUS-OpenJev-v1 9B (A-1) and 4B (A-1b), high and low |
| `imajev-2026-09-28/` | imajev-4B on a laptop's processor (I-1), and the pre-registration of both its runs. Its run on card B (I-2) read only held-back sets and is held back whole (below) |
| `opendecider-2026-09-28/` | OpenDecider nano (O-1), OpenDecider small (O-2) and its base model (O-2c) |
| `doorman-planted-2026-09-13/` | The 2026-09-13 door run's pre-registration |
| `EVIDENCE-v2.md` | The evidence file the page was written from, cell by cell, with each row file and the commit it was read at |
| `READS-v2.md`, `work-v2/reads/manifest.tsv` | The reads behind the page's dates, revisions, card quotes and ollama digests, and each read's URL, time, bytes and sha256 |
| `recount_v2.py`, `recount-v2.json` | The recount as it ran over the full record, and its output, with the page's own checks (`checked`) among them |
| `render_evidence_v2.py` | The script that writes the evidence file's tables from the recount's output: `python3 render_evidence_v2.py recount-v2.json > tables.md` prints every table line of `EVIDENCE-v2.md`. It needs the full record's output; over `recount_on_the_kit.py`'s it stops at the first held-back set |
| `work-v4/stat/` | The three scripts that first computed the page's checks (`checks.py`, `cluster.py`, `pairs.py`), as they ran over the private record; `recount_v2.py` now re-derives every figure the page prints from them |
| `recount_on_the_kit.py` | Runs the recount over this kit (below) |
| `index.json`, `provenance.json` | Every file's bytes and sha256, and every rule that touched it |

The citations in `EVIDENCE-v2.md` and `recount-v2.json` end in `@` and a commit of the private
record. The file each one names is here or in the Reading the answer kit, at the path given, or it is
held back (below). Three kinds of cited file are none of those: the benches' working notes
(`RECON-*.md` and the `results*.md` notes), which are not shipped; and the cost series'
`KIT.lock.json`, whose six-way part ships as `KIT.lock.six-way.json` (below). The evidence file
also names working folders of the lanes that wrote it (`work-v2/`, `work-v4/`, `work-v5/`); of those, only
`work-v2/reads/manifest.tsv` and `work-v4/stat/` ship.

A file name ends in the question it holds: `.kc`, `.c` and `.s0` the six-way; `.ka` long pages;
`.kperm`, `.creorder` and `.k1-5` the six-way under other option orders; `.s0.P.rep2` a repeat;
`.s0.S` the six-way's registered second rendering; `.fold1` to `.fold4` the word counter's folds.

## Where the Reading the answer kit's files go

Two things here read files that kit publishes:

- **The recount** reads arms 1, 2, 3, 10 and 12 and gap 4 from `jev-2026-09-21/`. You do not need to
  move anything: `recount_on_the_kit.py` takes that kit's folder as its argument.
- **The word counter** imports the six-way's builder and reads its 108 items from
  `jev-2026-09-21/kit/`. Copy that kit's `kit/` folder to `jev-2026-09-21/kit/` here.

To fetch that kit whole: `wget -r -np -nH --cut-dirs=2 -P rta https://research.strata2signal.com/reading-the-answer/data/`,
then pass `rta` to the recount. Its `index.json` lists every file with its bytes and sha256.

## Run the recount

```
python3 recount_on_the_kit.py <the Reading the answer kit's data/ folder>
```

It runs `recount_v2.py`, unchanged, statement by statement, over the files the two kits hold, and
compares every figure it re-derives with `recount-v2.json`, the page's own checks (`checked`)
among them. Standard library only, no model and no
network. It changes three things and says so: a commit it cannot read becomes "kit"; the cost series'
kit lock becomes its six-way part; the cost series' own series file, which is held back, becomes an
empty series. It builds a tree of symbolic links to both kits' files in your system's temporary folder
(on Windows, turn on Developer Mode or run it under WSL) and leaves it there with its output,
`recount-on-the-kit.json`, whose path it prints first; delete that folder when you are done.

On 2026-09-29, over this kit and the Reading the answer kit as published, it printed 1,842 figures
equal, none different, 47 held back, 15 not run and 6 partial:

- **Held back and not run:** the held-out lines, the doorman, the field exam and the judge question,
  whose rows are not here, and the sets built on them: the six-way minus the held-out lines, the door
  by kind, and everything read from imajev-4B's run on card B, which read only held-back sets: its
  times, which are timed on the held-out lines, and its row of the memory table (grew at load and
  highest), whose receipts are held back with those sets; and three of the page's own checks,
  which read those sets: the bound on APUS-OpenJev-v1 9B's 0 of 12 harmless door lines, the
  held-out intervals' widths, and the 9B against the live door. Their counts and figures are in
  `recount-v2.json`.
- **Partial:** three memory fields count every file of a run, including the held-back sets'
  telemetry and rows: how many telemetry files or samples a highest reading was read over, and each
  run's longest input overall. From the public files they count fewer. Every other memory figure the
  page prints (grew at load, highest, headroom and the longest input by the highest reading)
  re-derives exactly.

## Run the word counter

The word counter imports the six-way's builder from the Reading the answer kit and checks its
training lines against the sha256 its manifest recorded. That digest is withheld here (see "What the
rules replaced"). So, from this folder, copy that kit's `kit/` folder in, then set the digest to the
published pool's own sha256:

```
cp -r <the Reading the answer kit's data/ folder>/kit jev-2026-09-21/kit
python3 -c "import hashlib,json;p='deem-2026-09-27/t1-tune/kit/';m=json.load(open(p+'t1-pool.manifest.json'));m['fingerprints']['t1_pool_json_sha256']=hashlib.sha256(open(p+'t1-pool.json','rb').read()).hexdigest();open(p+'t1-pool.manifest.json','w').write(json.dumps(m,indent=1))"
python3 -B deem-2026-09-27/t1-tune/baseline_nb.py --set s0 --print-only
```

It prints 68 of 108 right, 73 with the named-guest exclusion, from 227 training lines, the pool
fingerprint `e4593840551294fa`, and a `rows_sha256` equal to the sha256 of
`deem-2026-09-27/t1-tune/rows/t1-baseline-nb.s0.P.rep1.jsonl`: the rows re-derive byte for byte.
`--set fold1`, `fold2` and `fold4` do the same. `fold3`'s rows re-derive in every field but
`set_sha256`, which is withheld here. Without `--print-only` the script rewrites its rows file.
`--set h48` needs the held-out lines, which are not here. The second command above edits one file of
this kit, so `index.json` then lists that file's old bytes. The fingerprint the run prints is read
from the manifest; to recompute it from the 227 lines themselves (it prints `e4593840551294fa`):

```
python3 -B -c "import sys,json;sys.path.insert(0,'deem-2026-09-27/t1-tune');import t1common as C;print(C.labelled_fingerprint(json.load(open('deem-2026-09-27/t1-tune/kit/t1-pool.json',encoding='utf-8'))))"
```

## The word counter's 227 training lines

Every one is a guest's line from 32 test dinners of the long table, run by the bench's harness on
2026-09-12 between 01:43:16 and 01:55:08 UTC, a week before the long table's door opened to everyone.
Checked for this kit on 2026-09-29 against the dinners' own records:

- each of the 227 is the reply of one logged model call (HTTP 200, not a stub) by its labelled
  guest's chair model, word for word (223 exactly, and 4 once runs of whitespace are collapsed):
  Gemma 4 26B wrote 128 and Mistral Small 3.2 24B 99;
- every host line at those dinners, 80 in all, is one of six preset lines written by the workshop
  before the dinners ran, and no training line quotes one;
- the pool shares no line with the six-way or the held-out lines.

None is a visitor's words, so none was held. The pool names one of its four test tables after the
workshop's operator; the rules rename it `operator` (below). The word counter reads only a line's
text and label, so the rename changes no figure, and the pool's (text, label) fingerprint is
unchanged. Ten held-out lines share a run of eight or more words with a training line: phrases two
chat models reached for playing the same six guests. No training line repeats a held-out line.

## What is held back, and why

Of the files the page's sources table and the recount reach, 169 are held back. Every file of a
held-back set is held with it, whatever its kind: rows, report, telemetry, summary, pass receipt and
log; and so is every receipt of a run that read only held-back sets. The benches' harnesses, dry runs
and working notes are not part of this kit either: every count the page prints is read from the rows,
receipts and telemetry here or in the Reading the answer kit, except the held-back sets' (their
counts and figures are in `recount-v2.json`) and what the page lists as quoted, not recounted.

| Held back | Files | Why |
|---|---:|---|
| The 48 held-out lines (`.h48`, `.kh`) | 38 | Publishing them would spend the control, and a line's hash identifies it on the long table's public wall |
| The doorman (`.d`, `.kd`, `.door`) | 39 | The planted lines are hostile by design |
| The field exam (`.kf`) | 14 | Its excerpts are not ours to publish |
| The judge question (`.kj`) | 14 | Its items are not ours to publish |
| imajev-4B's run on card B (I-2): its console log and its telemetry | 2 | That run read only the held-out, field-exam and judge sets, and its log prints each held-out line's prompt token count, which with the public wall could identify the line |
| The voice-control menu (sets `N`, `NM`, `R`) | 54 | The page does not report it, and no figure reads it |
| The sets themselves and the cost series' own files | 8 | The held-out lines and their manifest; the planted lines, the field exam and the judge set; the 2026-09-13 door run's verdicts; the cost series' pre-registration and series file |

**What holding the held-out lines back does not hide.** Deem's pre-registration ships as it was
registered, and its recipe for the draw (the source, the eligibility rules, the pool per guest, the
seed and the order of the draw) lets a reader rebuild one guest's eight held-out lines from the public
wall, because that guest's eligible lines number exactly eight. One more is found through the word
counter's near-duplicate check: its manifest and T1's pre-registration print the pool's largest
overlap with the six-way and held-out lines, 0.0282, and one eligible wall line alone reaches it. The
other 39 cannot be rebuilt from anything published: the draw read a frozen copy of the wall that is
not published, and the live wall reorders its entries on every read. The recipe, and the prompt-token
totals the receipts give over the held-out lines, still make a guess at them better than chance.

The Reading the answer kit held back the doorman, field exam and judge sets on the same grounds.
The cost series' pre-registration and series file print prices, account details and the series'
readings of hosted models, none of which this page reports, and the rest of its kit lock carries
per-line hashes of the planted lines and the judge set. The six-way part of that lock ships as
`cost-of-intelligence-2026-09-23/KIT.lock.six-way.json`: the lock's `sets.c` with only
`options_in_kit_order`, `labels`, `author_family` and `items_list`, written with sorted keys and a
one-space indent, then passed through the same rules as every other file. No rule fired on it.

The 51 read responses behind `READS-v2.md` (model cards, file trees, registry manifests) are other
people's documents; `work-v2/reads/manifest.tsv` gives each one's URL, time, bytes and sha256, so
each can be fetched and checked. The model cards at their pinned revisions and the registry manifests
re-fetch byte for byte; the Hugging Face API's repository, commit and file-tree listings carry live
fields (downloads, security-scan status), so compare their fields, not their sha256. The model
weights are their publishers' downloads.

## What the rules replaced

149 files here were copied from the record and rewritten only by the named rules in
`provenance.json`, in a fixed order, by the hub's `tools/kit_sanitise.py`; 63 are byte-identical to
the record. `provenance.json` gives every rule's firing count on every file. The strings replaced:
machine names, which become the roles the Reading the answer kit gives the same machines
(`benchbox`, `largecard`, `this laptop`); the operator's name in prose, in paths, and as the word
counter's table name; paths into a home directory, absolute or written from it with a tilde,
which take the neutral root `/workshop`; one agent session's scratch and transcript folders,
which become `/agent-scratch` and `/agent-transcripts`; the private network's name and daemon; a
configuration directory; a sandbox's home directory, renamed `sandbox-home`, and the home root in a
sandbox's own view of it; the runtime's default port, in one list of co-resident models, and the
name of its parallel-requests variable, whose value is unchanged; the hyphen in two UTC time ranges;
the held-out set's fingerprints; and two kinds of sha256, below. Every file name that carried a
machine name is renamed to match.
The tool's fence re-reads every number on both sides and refuses on any difference: no figure moved.

**Withheld digests.** A digest of a file published with words replaced lets anyone test guesses for
what was replaced, offline, and so do its last few characters. So every sha256 of a file that this
kit or the Reading the answer kit publishes rewritten, whole or shortened (to its start, or to its
start and end around an ellipsis, which is replaced whole), reads
`<withheld: rewritten file, see README.md>`. The pool's own digest in its manifest is one. Three
pre-registration seals are kept on purpose: the sha256 of imajev's registration as pushed for I-1 and
for I-2 (`imajev-2026-09-28/PREREG-imajev.md`) and of Lev's archive copy (`lev-2026-09-28/PREREG-lev.md`).
Each is over a text that also holds a value no file here or on the shelf prints (the held-out set's
own sha256, or an agent session's id), so none can confirm a guess, and those values stay unpublished. The
held-out set's sha256, its text fingerprint, the fingerprints of its prompts in each rendering, and
the sha256 of any run's rows or receipts over it read `<withheld: held-out set, see README.md>`: with
the public wall and a guess at which 48 lines the set holds, any of them would confirm the guess.
`provenance.json` gives a file's original digest only where no rule touched it.

## Licence

CC BY 4.0 covers the records in this folder: the rows, reports and receipts, the word counter, the
recount, the evidence and the pre-registrations. Some rows carry the outputs and scores of models
that are not ours, under their own licences, among them OpenJev 27B's 4-bit build, whose card says
Apache-2.0 while the OpenJev README puts the weights under CC BY-NC 4.0. They are reproduced as measurements so a reader can check the page;
this kit grants no rights in any model. The page's licence section says what each repository
carries. The 227 training lines were written by Gemma 4 26B and Mistral Small 3.2 24B at our test
dinners.

Corrections are welcome at [hello@strata2signal.com](mailto:hello@strata2signal.com).
