# PREREG-T1: teaching deem-0.8-v1 the six-way speaker question

*Registered 2026-09-27T15:30:56Z on this laptop by the T1 lane, BEFORE any training step and before any scored read
(no baseline row, no fold read, no tuned read exists at this commit). The gate below is section 5 of
`PLAN-deem-T1-tune-v2.md` (sha256 `11efe4b2750347f22bdd70eae77bb76ea9a9b0c7d8c6b8226f75c3e0d38d01ae`), copied VERBATIM by line range (lines 386 to 532).
Only dated amendments are appended below it; no registered text changes (5.8).*

## Pins at registration

| what | pin |
|---|---|
| the plan | `bench-archive/plans-2026-09-26/deem/PLAN-deem-T1-tune-v2.md` sha256 `11efe4b2750347f22bdd70eae77bb76ea9a9b0c7d8c6b8226f75c3e0d38d01ae` |
| base | `LibertAIDAI/deem-0.8-v1` @ `8cbabbb2c4a7ef13c6b43f0ef3ae4157983c6d21`, `model.safetensors` sha256 `80438125177c0681e855158d46720cbbd4ff07ea620cd55b0f456aafbfcf778d` |
| runtime clone | `Libertai/deem` @ `6755b30bf6bbd9a81f8db6ba42cc0fd62c9f4719` |
| D0's runner (v0.3, as pinned) | `run_deem.py` sha256 `78fe9562a7de88ed9a30fa03bb3b2c769b99d496301ae6d8972a2e6edc09b2a6` |
| the runner patch v0.4 (5.0) | `run_deem.py` sha256 `754b9a18b510f2f48a4942c2e20bdb0f197c07414f695b374d3fdb14a4da3d8f` at this commit; **pinned by amendment once its equivalence receipt is on file** (T-V5) |
| `tables_deem.py` | `751630d95c178ca2b67724ee46e8da7e6b921fe0fac46adc285169c8dcd6d8ff` |
| `hostos.py` (imported) | `a6a74bba57250be1936825466c6614876d757cf81373fcf6fa0b48587cf0e3ad` |
| S0 | `bench/jev-2026-09-21/kit/task_c.json` sha256 `61a0b1d6e7cf66fd91992c0c3a4572f6fbae90e6471048fd7dc1eb41ff2d7f39` |
| H48 | `bench/deem-2026-09-27/kit/h48.json` sha256 `<withheld: held-out set, see README.md>`, fingerprint `<withheld: held-out set, see README.md>` |
| the voice kit | `bench/deem-2026-09-27/d1-voice/kit/voice-items.json` sha256 `cea9e12d7983a3c52a204c1ed3e2d6c709e7504066f7ff37c1f86153bcf95d3c` |
| `NAME_TOKENS` | imported from `bench/jev-2026-09-21/kit/build_kit.py` (sha256 `<withheld: rewritten file, see README.md>`) |
| the 32 PoC dinner records + the judge file | sha256 of each file in `kit/t1-pool.manifest.json` → `sources_sha256` (master `34d44069`) |

## The training pool, T-poc (3.1), as built by `build_t1_pool.py`

- **288 guest lines → 227 eligible.** Drop order as applied and printed:
  1 short (< 80 chars) (1) → 2 self-naming (51) → 3 exact duplicate (9).
- Per guest {"darwin": 41, "einstein": 37, "hypatia": 37, "ibn_sina": 37, "sagan": 40, "socrates": 35}; per chair family {"gemma": 128, "mistral": 99};
  per table {"ancients": 64, "cross": 47, "operator": 63, "span": 53}; PoC-judge `break` rows kept and flagged: 10;
  rows naming another guest: 214.
- **The refusing dedup** (text sha256): pool ∩ S0 = 0,
  ∩ H48 = 0, ∩ voice kit = 0,
  ∩ every wall read on master (4 files) = 0; a planted S0, H48 or voice text makes
  the builder exit non-zero naming the set (receipt: `receipts/tests.log`).
- **Near-duplicates** (normalized 8-word shingles against S0 ∪ H48): max Jaccard 0.0282,
  pairs ≥ 0.1: 0, removed at ≥ 0.5: 0.
- **Fingerprints:** text set (S0's convention) `df69dd1eba91341c`; (text, label) `e4593840551294fa`;
  `kit/t1-pool.json` sha256 `<withheld: rewritten file, see README.md>`; **`kit/folds.json` sha256 `<withheld: rewritten file, see README.md>`**.
  `build_t1_pool.py --check` reproduces all 7 kit files byte for byte.
- **The folds:** 4, leave one TABLE out (fold 1 ancients 64 · fold 2 cross 47 · fold 3 operator 63 · fold 4 span 53).
  Each held-out table's guests still appear in training through their other table.

## The code at registration (sha256)

| file | sha256 |
|---|---|
| `t1common.py` | `<withheld: rewritten file, see README.md>` |
| `build_t1_pool.py` | `bd3f14ec4064dfe2d66b425cc654e05472817a1b6dab92892fda0757e3819280` |
| `render.py` | `ad2bdeae50cf285cf1c3998d2fa9149e4a92c16c10e914f95769deb43379388f` |
| `readout.py` | `021cf0fcda4b3166868553e57277d78d1a5b67c13a24708d919fe1e287ac749e` |
| `train_letter_slot.py` | `ca4637206352c98d2b084bbc79f3d8710310d46b5fa5ae7a2ed5b90199feb0ce` |
| `baseline_nb.py` | `7d7fbe1a6f1c9a2419323cbc8a33289fbb99a9b43a22fc948ef56158b9942033` |
| `serve_t1.sh` | `3f264f954634319d6fa17fa8a99f631ba9dc6cc9018bce6d8e04931c063d430f` |

The trainer is `train_letter_slot.py` above. Its loss is cross-entropy over the six valid-letter logits at
the answer slot through `readout.slot_logits`, the server's `_forward` line for line (4.2).

## The baseline (G-BASE), registered here with its disclosure

The spec and the disclosure are 5.2's, verbatim below. The implementation is `baseline_nb.py` (stdlib only;
an unseen token carries no evidence, as in any vocabulary-restricted multinomial NB). **Its registered
S0 count is whatever it prints from the committed pool, appended as a dated amendment with the rows
file's sha256; nothing about it is chosen on S0 or H48.**

## Stated at registration: what the hour-one arm does and does not read

- SMOKE-CPU (6.2) trains on folds 2 to 4's tables (163 lines), 1 epoch, the first two cyclic shifts of
  each line's seeded base order (326 renderings), lr 1e-5, swap off, seed 20260927, and is read on the
  held-out `ancients` table only (in-process on the saved bf16 weights; and through the server with the
  runner v0.4 `--kit`). **The smoke never reads S0 or H48**; its parity/round-trip probe uses ten
  held-out-table rows, not S0 rows. It is labelled SMOKE in every table and changes no registered bar.
- In-process reads of a held-out table use the cyclic shifts of kit order (shift 0 IS kit order), so
  every guest sits at every letter once; "under rotation" counts a row right when its modal choice over
  the six shifts is the label (a tie in the count goes to the higher mean probability).

---

## 5. The gate, registered before training (PREREG-T1 carries this section verbatim)

### 5.0 The instrument: D0's server, D0's runner patched to v0.4, D0's box and cores

v1 promised "the SAME server, runner and validity checks D0 used". **As pinned, D0's `run_deem.py`
(sha256 `78fe9562…b2a6`) refuses every tuned checkpoint** (verified: it hard-codes `MODEL_ID`,
`CHECKPOINT`, `MODEL_SHA256` and `SERVER_CPUS = "0-7"`; its preflight exits on a `/health` model-id
mismatch, a weights-sha mismatch, a `DEEM_CHECKPOINT` mismatch, and on any non-empty
`CUDA_VISIBLE_DEVICES`). So:

- **A registered runner patch, v0.4:** `--manifest <T1-MANIFEST.json>` supplies the model id, the
  checkpoint path and the weights sha; V-5 reads the manifest's output sha (this folds v1's T-V2 into it);
  `--kit <path>` accepts a kit file in S0's shape for the smoke's held-out table. Everything else is D0's
  code. **An equivalence receipt comes before any tuned read:** the patched runner, pointed at the base,
  reproduces D0's S0 P rep 1 rows bit for bit (108 probability vectors). The new sha is pinned in PREREG-T1.
- **Every gated read runs on this laptop's CPU, P-cores 0-7, as D0's did,** whatever device trained the arm
  (the CPU bf16 path and the CUDA autocast path differ numerically; D0's exact ties, R-1, would land
  differently). 108 rows take about a minute. No GPU read is a gated read.
- Server environment as `serve_deem.sh` has it: `DEEM_CHECKPOINT=<the tuned dir>`,
  `DEEM_MODEL_ID=deem-0.8-t1-<arm>-<seed>`, `DEEM_BATCH_SIZE=1`, no calibration, raw T = 1, one question
  per request. **One server per arm, D1's memory guard inherited** (restart between passes above 8 GiB
  VmRSS, restarts recorded; D0's audit S-3 saw 16.1 GB after ~420 requests, D1 restarted at 8.84 GB; a
  headline arm is ~1,070 requests and the laptop's swap reads 10 of 17 GB used).

### 5.1 Validity: an arm is VOID, never "low", unless every check holds

D0's V-1 to V-7 (`PREREG-v0.md` §4, with R-1's tie rule: first in presented order on an exact tie,
D-20260927-018), through the same code, plus:

- **T-V1 disjointness:** the served checkpoint's manifest names a pool whose texts ∩ (S0 ∪ H48 ∪ the voice
  kit ∪ the wall) = ∅, re-checked from the kit files at read time, receipt printed.
- **T-V2 provenance:** V-5 as patched: the manifest's output-weights sha256 equals the sha256 of the file
  the server loaded (the preflight recomputes it).
- **T-V3 parity on file:** `receipts/parity-<arm>-<seed>.json` exists and passes for that checkpoint.
- **T-V4 the null arm's read** (5.4) is on file before any tuned figure is called a result.
- **T-V5 the equivalence receipt** (5.0) is on file before any tuned read.

### 5.2 The bars on S0 (n = 108), each read at min(kit-order count, mean count over k = 1..5)

Every tuned arm is read S0 P × 1 (kit order, k = 0) and S0 P × 5 more (k = 1..5 with gap 4's schedule,
`random.Random(20260921 + k)`), 648 rows. **A bar is read at the smaller of the kit-order count and the
mean count over the five rotated orders** (Tare's own conditioning; TM-10: at k = 0 socrates is always F,
so "F fixed" and "socrates learned" are one event, and a kit-order pass alone is not a pass).

| bar | rule | why this number |
|---|---|---|
| **G-BASE** (new) | (a) at kit order, the tuned arm beats the registered baseline on the same 108 items by an **exact one-sided McNemar test on the discordant items, p < 0.05**; and (b) the tuned arm's mean rotated count exceeds the baseline's S0 count | A tuned model is supervised on in-domain labels; its comparator is a supervised model on the same labels. **The baseline, fixed now:** multinomial naive Bayes, unigram tokens `[a-z']+`, Laplace α = 1, empirical prior, no masking, plain argmax, trained on the T-poc pool fingerprint; `baseline_nb.py` (stdlib only) writes rows in the runner's shape (`uid`, `text_sha256`, `chosen`, `probs`) so `tables_t1.py` prints it like any arm; the named-guest-exclusion variant is printed beside it, not gated. **Disclosure:** these reads are exploratory and have been seen: at critique time (2026-09-27T15:02Z, the critic's rebuild of T-poc) the baseline read **68/108** plain and **73/108** with exclusion on S0, 27/48 and 29/48 on H48; the registered count is whatever `baseline_nb.py` prints from the committed pool, expected 68. No refinement of the baseline is chosen on S0; a variant, if any, is chosen on the table folds |
| **G-FLOOR** | count **≥ 34 / 108** (31.5 %) | the smallest count with P(X ≥ k) < 0.05 against the task's informed floor of **24.0 %** (`bench/jev-2026-09-21/TABLES.md`: 103 of 108 lines name another guest, never the speaker). **Verified under the exact Poisson-binomial null** (per item p = 1/(6 − guests named)): P(X ≥ 34) = 0.046, P(X ≥ 33) = 0.071. Against chance the same rule gives 26; the floor is the honest one |
| **G-KEV** | count **≥ 72 / 108** (66.7 %) | Kev-9B zero-shot, the Apache comparator (`rows-gaps-card/kev-9b.kc.jsonl`). Equalling is passing |
| **G-SEAT-TALK** | count **≥ 88 / 108** (81.5 %) | gemma4 K0 on benchbox's 3090 (COI row `r0-bba46e51`), D0's registered number |

**Across seeds (TM-4):** the headline is seed 20260927; a bar passes iff **the headline passes AND at least
two of the three seeds pass**; min, median and max are printed; seeds are never pooled into one binomial.

**What each verdict means, written now:**

- **Below G-FLOOR (the null hypothesis):** "227 lines from four fixed trios do not teach the six-way to
  deem-0.8-v1 under this recipe." No threshold table; the spread over seeds is printed. The next rung is
  **data** (new dinners with new trios, in the wall's frame, which also mint S1; the data rung's plan), never
  a hyperparameter hunt on S0. T2's tuning is not earned; D2 (the 9B zero-shot) stays its own question.
- **G-FLOOR, not G-BASE:** "it learns the task; a bag of words on the same lines does better." T2 is not
  earned on this evidence; the data rung is (a bigger pool is the cheapest test of whether the LLM's
  advantage appears). No seat talk.
- **G-BASE (with G-FLOOR):** "tuning added something a bag of words cannot get from the same lines." **T2
  (the 9B LoRA) and the data rung are earned.**
- **G-KEV with G-BASE:** "a 0.8B Apache checkpoint tuned on lines we own matches Kev-9B zero-shot and beats
  a bag of words on the same lines." A candidate COI row (readout mode, its own rendering id and
  fingerprint) and a candidate tuned response set for the six-voices dataset, **both confirmed on S1 (5.9)
  before anything is public.** Latency is printed in 5.7's table only, never in a verdict sentence (TM-16).
- **G-SEAT-TALK with G-BASE:** a seat-exam DISCUSSION opens (not adoption).

### 5.3 Consistency and the letters (from the 648 rows; reported, never gated)

- **Tare's consistency line:** items whose answer moved, of 108; order-pairs disagreeing, of 1,620; the
  rule "moved ≤ (1 − acc) + 0.05" printed as Deem's `eval/tare` prints it, **reported, not gated** (at acc
  0.33 it would let ~78 of 108 items move). Kev's reorder figures (33/108 items, 15.43 % of pairs) beside it.
- **The F test with an equivalence margin (TM-10):** accuracy when the label sits at F minus accuracy when
  it sits at A to E, 95 % item bootstrap (10,000 resamples, seed 20260927). "The wall is gone" iff the
  interval lies within ±10 points; "F is still suppressed" iff the upper bound is below 0; otherwise "not
  shown". Reported.
- **Accuracy by label letter, A to F**, with a Cochran-Armitage trend test (D1's audit M-3 found a
  gradient A > B > C > D > E, then 0 at F; a test of F against A to E hides it). Reported.
- **Letter shares** over the 540 rotated rows, beside 1/6. Reported.

### 5.4 The self-refutation arm: shuffled labels (one per seed, run before the final reads)

Train the chosen recipe on T-poc's rows with the labels permuted across rows by one fixed
`random.Random(seed).shuffle` (class counts preserved), one null per seed (20260927, 20260928, 20260929).
Read S0 P × 1 through the server. **Each null must score ≤ 25 / 108** (P(X ≥ 26 | 1/6) = 0.031, verified).
Above it, the instrument leaks and **every tuned figure is VOID** until the leak is named and fixed. The
null's named-guest share is printed too.

### 5.5 The contamination control, read as a difference in differences (TM-8)

H48 P × 1 is read for the headline (and for every seed that passes G-FLOOR). S0 and H48 differ in
difficulty even for a model that cannot be contaminated (the baseline reads 67.6 % on S0 against 60.4 % on
H48 with exclusion; H48 is the 8 per guest S0 left, 28 of 48 from the 7 newer dinners, mistral 28 /
gemma 20). So the reading is **(tuned S0 − tuned H48) − (baseline S0 − baseline H48)** with a paired item
bootstrap, printed beside the family-stratified gap. The verdict wording is "contamination not excluded"
or "no signal", never a pre-written cause. (The base's own exposure to S0's texts is D0's question, already
answered: 5/108, no memorisation; the critic notes S0 has been public since 2026-09-23 and the Deem repos
date from 09-24, dates not re-verified here; the control above does not depend on them.)

### 5.6 The forgetting probe (reported, never gated; CPU, no fan; runs when the headline passes G-FLOOR)

Deem's own mixture cost 7.3 JevBench-hard points against its frozen base; a specialist on 227 rows may
forget more. Through the same server, base rows on master `34d44069` vs the tuned headline:

- **D1's route set R at rotation 0** (154 rows; `d1-voice/rows/d1-cpu-0.8.R.rep2.jsonl`; read with
  `run_d1.py --set R --rep 2` from master against the tuned checkpoint): agreement with the grammar's
  hit/miss and the act-risk count.
- **D1's N set under every rotation** (68 × 19 = 1,292 rows, ~7 minutes): **the transfer read for letters
  F to S** (TM-11). T1 never renders more than six options (letter ids 38 to 50 are never a target), so
  nothing in T1 can be presumed to fix D1's F to S wall; this read says whether anything moved there.
- **The 8 factual questions** (`diag_letters.py`): the base answered 4 of 6 kit-order questions at
  p ≥ 0.818 and missed the Sagan and Socrates ones; both questions whose answer sat at F failed. The tuned
  answers are printed beside them.
- JevBench public `original.jsonl` (72 items, MIT): T1.1, only if the checkpoint is ever shown outside.

### 5.7 Latency, energy, the OS

As D0 and D1: median and p95 seconds per decision at concurrency 1 over measured, uncontended rows (the
v0.2 contention gate); a 10 s idle read before each pass; RAPL package joules on CPU; net J per decision
printed only for a clean pass; three discarded warm-ups; residency by growth; `fla` installed or not; the
OS, kernel and CPU model on every row. The training runs' board watts are printed with the
`scribe-embed.timer` state (6.3). A tuned 0.8B has the base's weights shape, so its latency is D0's
family; the number is printed, never claimed as a seat's, and never in a verdict sentence.

### 5.8 Registration order (the receipts are `git log` and the rows' UTC)

1. PREREG-T1 (this section verbatim; the pool fingerprints; `folds.json` sha; the trainer's sha; the
   baseline's spec and disclosure) and `LICENCE-CHAIN.md` committed and pushed. 2. Tests 4.6, the parity
   receipt against the served base, the baseline's rows. 3. **SMOKE-CPU** (6.2): a held-out TABLE, never
   S0. 4. The runner patch v0.4 and its equivalence receipt. 5. The 8 selection runs; their table and the
   chosen (swap, epochs) pushed as a dated amendment. 6. The three nulls, then the three finals, trained.
   7. Reads, CPU: the nulls' S0 P × 1; the headline's S0 P × 1, S0 P × 6, H48 P × 1; the other seeds' S0
   P × 1 (and their × 6 and H48 iff the headline passes G-FLOOR); the forgetting probe iff G-FLOOR.
   8. `results.md`; results notes below the line; the audit. Only dated amendments are appended; no
   registered text changes.

### 5.9 S1: the next held-out set is minted by the data rung, before S0 wears out (TM-12)

S0 has been read in 4 D0 passes and will be read by ~7 T1 arms; the verdict tree routes the next rung on
it; it is public (exhibit fifty-eight). **The first new dinners (the data rung, T1-05 in 8.3) mint a sealed
S1**: 18 lines per guest under S0's builder rules, new trios and chips, hashed and frozen before any tuned
arm reads it, never published while claims are open. **Public claims (a COI row, a tuned response set) are
confirmed on S1; S0 stays T1's test of record.** Registered as intent here; no T1 work.

---

## Dated amendments

### A-1 (2026-09-27T15:35Z): the trainer's disk floor is checked before every save

The registered `train_letter_slot.py` (sha256 `ca463720…eb0ce`) checked the 6 GB disk floor only at start
and saved a checkpoint after every epoch. The first launch of test 4.6.4 (30 epochs of 4 rows) wrote
2 GB of per-epoch checkpoints in its first two minutes and would have filled the laptop's 16 GB; the
lane stopped it at 15:33Z and deleted them (disk back to 16 GB free). The fix: a disk check before every
save (exit 4 if a checkpoint would leave under 6 GB) and `--save-epochs all|last` (default `all`, the
plan's 4.5). No change to the loss, the data, the schedule or the optimizer. **The trainer's sha256 is
now `409fa5b61e650353e64c91e9acd9581405a2617508cf900b94871c83f1ba7a52`**; every T1 checkpoint is trained by it.

### A-2 (2026-09-27T15:31Z): the baseline's registered counts

`baseline_nb.py` as registered, from the committed pool (fingerprint `e4593840551294fa`), first run
2026-09-27T15:31:17Z. **S0: 68 / 108 plain (G-BASE's registered count), 73 / 108 with the named-guest
exclusion; H48: 27 / 48, 29 / 48.** Exactly the disclosed critique-time figures (68 / 73, 27 / 29).
Rows: `rows/t1-baseline-nb.s0.P.rep1.jsonl` sha256 `fd92574b50c288a1d5dfdb7d3f39a0e172f3367a42d89fb7eadf4e9d1a6e01a9`,
`rows/t1-baseline-nb.h48.P.rep1.jsonl` sha256 `<withheld: held-out set, see README.md>`.
The table folds (trained on the other three tables): fold 1 ancients 0 / 64, fold 2 cross 4 / 47, fold 3
operator 20 / 63, fold 4 span 10 / 53 (the empirical prior under a held-out table; 3.3's warning, measured).

### A-3 (2026-09-27T15:39Z): the runner v0.4 is pinned

`run_deem.py` v0.4 sha256 **`754b9a18b510f2f48a4942c2e20bdb0f197c07414f695b374d3fdb14a4da3d8f`**. Its equivalence
receipt (`receipts/equivalence-base.json`, T-V5): through `--manifest` pointed at the base, it reproduces
D0's S0 P rep 1 on **108 of 108 items, the same choice and bit-identical probability vectors (max
difference 0.0)**. The parity receipt for the base (`receipts/parity-base.json`): 10 of 10 S0 rows
bit-identical between the server and the in-process bf16 readout.

### A-4 (2026-09-27T15:43Z): test 4.6.4 (overfit-4) FAILED as registered

30 steps on 4 rows (one each of darwin, einstein, hypatia, ibn_sina), lr 1e-4, a fresh seeded base order
every step, read on the saved bf16 weights. **The mean loss went 4.945 → 1.875, not below 0.05**
(per step: 4.945, 4.505, 1.734, 13.884, 3.353, 1.933, 3.292, 2.174, 3.206, 1.615, 4.396, 6.086, 2.428, 1.594, 1.415, 2.636, 1.884, 2.213, 2.122, 1.816, 2.289, 1.95, 1.916, 2.026, 1.336, 2.651, 1.56, 2.385, 2.531, 1.875); pre-clip gradient norms 24 to 784, clipped to 1.0. The loss jumped
to 13.884 on the step after the schedule reached its 1e-4 peak. What stands beside it: the read path is
proven bit-identical to the server (A-3), and test 4.6.6 shows one step moves the weights (two seeds, two
different output shas, `receipts/test-seeds-weights.json`), so the loss reaches the weights. What the failure
does not separate: a learning rate ten times Deem's own full-fine-tune rate (1e-5) destabilising the run,
versus a defect in how the loss meets the served readout. **SMOKE-CPU was launched at 15:40:14Z, before this
test finished (15:42:40Z)**; its loss log at the registered 1e-5 is read before anything is said about it.
A re-specified overfit test (the same 4 rows at lr 1e-5) is run as a DIAGNOSTIC, labelled so; it does not
replace 4.6.4. Whether the GPU window may proceed on it is the lead's ruling.

### A-5 (2026-09-27T16:03Z): the trainer records what the plan's 4.5 asks of a GPU run

After SMOKE-CPU finished (trained by `409fa5b6…7a52`, A-1), three changes, none to the loss, data,
schedule or optimizer: the trainer's own sha256 and its modules' are read at START (they were read at
the end, so an edit during a run would have been recorded as the code that ran); each step logs the
CUDA VRAM peak (4.5 "VRAM peak or VmRSS"); the manifest records the board (name, UUID, enforced power
limit, driver) for a CUDA run; and `float(loss.detach())` replaces `float(loss)` (a warning, not a
number: the logged value is identical). **The trainer's sha256 is now
`6be547428ff4dd1f61c73df1d199d351f1b87915ff00a2feddcb714ea09cbe93`**; the selection runs, nulls and finals
are trained by it.

### A-6 (2026-09-27T16:25Z): the A-4 diagnostic, and what it separates

The DIAGNOSTIC named in A-4 (test 4.6.4's four rows, 30 steps, one seeded order per step, read on the saved
bf16 weights, at **lr 1e-5** instead of the registered 1e-4; `receipts/diag-overfit4-lr1e-5.json`,
`logs/diag-overfit4-lr1e-5-20260927.steps.jsonl`): the mean loss went **4.945 → 0.0**, first under 0.05 at
step 7; **all four rows right under all six cyclic shifts** ({'p000': 6, 'p001': 6, 'p002': 6, 'p065': 6}).
So the loss reaches the served readout; the registered test failed on its learning rate, not on the wiring.
4.6.4 stays FAILED as registered (nothing is re-scored). **Proposed for the lead's ruling, not applied:**
re-specify 4.6.4 at lr 1e-5 (the recipe's own rate) before the GPU window.

### A-7 (2026-09-27T16:40Z): test 4.6.4 re-specified at lr 1e-5 (ruling D-20260927-019 (1))

The lead's ruling on A-4 and A-6: the test as registered ran a full fine-tune at lr 1e-4, ten times the
recipe's own rate, and FAILED (loss 4.95 → 1.88 with a 13.9 spike; each row right under 1 of 6 shifts;
`receipts/test-overfit4.json` and `logs/test-overfit4-20260927.steps.jsonl` stay on the record, unchanged).
That is a spec error in the test, not a trainer fault. **4.6.4 now reads: "30 steps on 4 rows at lr 1e-5
drive the mean loss below 0.05 and make all four right at kit order and under rotation"**; everything else
as before (the first row of each of the first four guests in pool order, batch 4, one seeded base order
per step, read in-process on the saved bf16 weights under the 6 cyclic shifts of kit order). It is re-run
under this text on the CPU **before any GPU arm**, by the trainer as it now stands (A-5, `6be54742…be93`);
receipt `receipts/test-overfit4-lr1e-5.json`. A-6's diagnostic stays a diagnostic.

### A-8 (2026-09-27T16:40Z): the recipe fixed a priori; no selection runs (ruling D-20260927-019 (2))

**T1-14 = (b).** The recipe for the nulls and the finals is fixed now, before any of them trains:
**swap off, 2 epochs, lr 1e-5**, six cyclic shifts per line per epoch (1,362 renderings, 114 steps per
epoch, 228 steps), effective batch 12 (micro-batch 4 × 3 on the GPU), everything else as 4.3. Seeds: the
headline **20260927**, then 20260928 and 20260929 (5.2's across-seeds rule unchanged); one shuffled-label
null per seed (5.4). **The 8 table-fold selection runs of 4.4 are not run, and 5.8 step 5 is struck:**
fold 1 reads 0 / 64 for both SMOKE-CPU and the registered baseline (the leave-one-table-out trap 3.3
names: fold 1's held-out guests keep one table each in its training), so the folds cannot discriminate
and a selection on them would be a selection on noise. The swap variant stays registered for T1.1; it
is not chosen or read in T1. The smoke's own reading stays as measured (0 / 64 on `ancients`; the
named-guest answers 45 / 60 → 0 / 60). No bar in 5.2 changes.

### A-9 (2026-09-27T16:36Z): the window and the finals' reads after it, registered before either runs

**The window** (`sweep.py`, as of this commit): preflight (A-8 on file; the A-7 overfit-4 receipt passing;
the 6 GB disk floor after a checkpoint; `scribe-embed.timer` paused and stamped, restarted at the end; a
1-step CUDA probe on 2 rows, saved and deleted, since the CUDA path cannot be dry-run on the CPU), then
null-20260927, null-20260928, null-20260929, final-20260927, final-20260928, final-20260929 on the fixed
recipe of A-8. Each null is read on the CPU beside the next GPU run: T-V1, T-V3 parity on 10 S0 rows,
S0 P x 1 through `serve_t1.sh` and `run_deem.py` v0.4 `--manifest` (P-cores 0-7, the runner on core 23:
D0's instrument), then its weights are deleted (6.1); a failed null read keeps its weights and never
stops the GPU runs. Training runs from CPU cores 16-22.

**The finals' reads** (`reads_t1.py`, CPU only, fan off), in this order (5.8 step 7):
1. It refuses to start unless all three nulls' S0 P x 1 rows are on file (T-V4).
2. **The headline 20260927:** T-V1 re-checked at read time; T-V3 parity on 10 S0 rows (bit-identical);
   **S0 P at kit order (k = 0, the gated read)**; **S0 P at gap 4's orders k = 1..5**, one order per pass;
   **H48 P x 1**. The 648 S0 rows live in one file, `rows/t1-final-<seed>.s0.P.rep1.jsonl`, told apart by
   `order_k` (this replaces the plan's separate `.s0.P.k6.jsonl` name; the rows are the same).
3. **20260928 and 20260929:** T-V1, T-V3, S0 P at kit order.
4. **G-FLOOR on the headline**, computed from its rows as 5.2 defines it: min(kit-order count, mean over
   k = 1..5) >= 34. **Iff it passes**, 20260928 and 20260929 get S0 k = 1..5 and H48 P x 1.
5. The two non-headline checkpoints are deleted (6.1); only final-20260927 is kept.
6. After every pass the server's VmRSS is read and above 8 GiB the server is restarted (D1's guard);
   every restart is recorded in `receipts/reads-t1.json`.

The forgetting probe (5.6, iff G-FLOOR) stays registered as the plan has it; it is scripted only if the
headline passes G-FLOOR. The verdicts are computed by `tables_t1.py` from these rows and 5.2's rules; no
figure is read before the nulls' rows are on file.

*Correction of record: A-7 and A-8 carry the stamp 16:40Z; they were committed and pushed at
2026-09-27T16:33:12Z (`02b93eeb`). The stamp was written ahead of the commit; the text is unchanged.*

### A-10 (2026-09-27T17:19Z): the H48 reading's cut (plan section 5.5), ruled before the window and before any read

Plan section 5.5 names the two wordings of the contamination control ("contamination not excluded" or
"no signal") and not the cut between them. **The lead ruled the cut on 2026-09-27, with no T1 read on
file** (no null or final has trained; `rows/` holds no `t1-null-*` or `t1-final-*` file; the checkpoint
directory is empty). The rule, exactly as `gate_t1.py` `did()` implements it (committed at `f717fce6`):

- The quantity is 5.5's difference in differences at kit order:
  **(tuned S0 − tuned H48) − (baseline S0 − baseline H48)**, the baseline being the registered naive
  Bayes (A-2).
- Its interval is a **paired item bootstrap**, **10,000 resamples, seed 20260927**. S0 items and H48
  items are resampled independently, and the tuned and baseline readings share every draw. The
  interval is the 2.5th and 97.5th percentiles.
- **"contamination not excluded"** when the interval's **lower bound is above 0**; **"no signal"**
  otherwise.

The reading stays what 5.5 made it: reported beside the family-stratified gap, never gated, and never
worded as a cause. Nothing else changes.

### A-11 (2026-09-28T14:15:01Z): the GPU window moves from the RTX 5090 Laptop to benchbox's RTX 3090 at 300 W, registered before the window's first call (ruling D-20260928-011)

*Cold scroll: this amendment moves ONE thing, the card and box that TRAIN the three nulls and three finals of A-8
and A-9, and changes nothing about how they are trained or read. The window never ran on the laptop: no
`t1-null-*` or `t1-final-*` row exists, the checkpoint directory is empty (A-10's own reading), and the 5090 was
untouched (`receipts/gpu-untouched.txt`, 2026-09-27T16:29:03Z). So nothing has to be re-read or re-scored.*

**1. What moved, and why.** The ruling D-20260928-011 (an operator, ~08:43Z, verbatim: "for any and all benches", "so
recs to all", "and keep rolling") moves the window to benchbox's RTX 3090: "Deem T1's GPU window moved from the
laptop 5090 to benchbox's 3090 (no turbo-fan ask; the recipe fixed a priori stands, the card change registered as
a numbered amendment to PLAN-deem-T1-tune-v2 before its first call)". D-011 expected the card on 2026-09-29; it was
seated a day early, and an operator said on 2026-09-28: "benchbox, the 3090 has 2 separate 8 pin pci power cables hooked
up, and the fans are cranked, let's roll with all the benches you need to run on there now". The window runs on
2026-09-28, after the lev L-2 lane left the card (~13:22Z). The turbo-fan ask T1-02 (plan 8.1) is moot: D-011 says
"no turbo-fan ask", and the fan window of 2026-09-28 was skipped (D-20260928-002, rec 5).

**The board, read on benchbox at 2026-09-28T13:27:44Z** (`date -u` on this laptop and on benchbox read the same second;
`nvidia-smi` by query, `lscpu -e`, `/proc/driver/nvidia/version`, `/etc/os-release`). *Correction of the
planner's draft of this amendment: it named the board taped "46890836" (`GPU-46890836-…`). That board is not in
benchbox. The board seated is the one below.*

| | registered in A-9 (never run) | from this amendment |
|---|---|---|
| box | this laptop (Intel Core Ultra 9 290HX Plus) | **benchbox**: AMD Ryzen 7 3700X, 8 cores / 16 CPUs (CPU n and CPU n+8 share core n), 30,990 MiB RAM, Ubuntu 26.04 LTS, kernel `7.0.0-31-generic` |
| board | NVIDIA GeForce RTX 5090 Laptop GPU, 24 GB, `GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff`, enforced 150 W, driver 595.91.07 | **EVGA GeForce RTX 3090 XC3 Ultra ("the 3090 board")**, 24,576 MiB, `GPU-fde82ed0-be70-eb1f-2d4d-5538c5b82166` (fingerprint `sha256(uuid)[:12]` = **`33a0b4acb8e8`**, the form every committed log carries), PCI `00000000:2B:00.0`, the x16 seat (`pcie.link.width.max` 16, gen 4), subsystem `0x39753842` (EVGA), VBIOS `94.02.42.C0.10`, **alone in the box** (`nvidia-smi -L`: one board). **`power.limit` 300.00 W** (enforced 300.00 W; default 350.00 W), set at boot by benchbox's `/etc/workshop/gpu-power-caps.conf` row `GPU-fde82ed0-be70-eb1f-2d4d-5538c5b82166 300` (an operator's paste, 2026-09-28; conf md5 `d3affa79364227dbec2df311697726ed`); **persistence Enabled** (the house 3090 posture, R-12). Driver **595.84** (open kernel module) |
| air | the laptop's turbo fan, by hand (T1-02) | none asked: benchbox's open frame, an operator's USB fans "cranked" |
| the CUDA venv | this laptop `/workshop/bench-deem-2026-09-27/.venv-cuda` (python 3.14.4; torch `2.13.0+cu130`, CUDA 13.0; transformers `5.17.0`) | **the same venv, copied byte for byte to the same absolute path on benchbox** (benchbox has `/usr/bin/python3.14`, 3.14.4). Its tree hash (t1transfer.tree_hash: sha256 over the sorted `<path>\t<sha256>` lines, symlinks by target) on this laptop is `e32f2389d77801a9e38fc33c01ce3276e0d96a6c1b52a41c31d92f44c46bb931` (21,991 files, 4,913,422,077 B); benchbox's must equal it or no arm trains. The same for `weights/` (`f3617d4e63eef3650cc3dace7efd4c6d76ba4700fd97cab9261e93e025f5bc8e`) and the deem clone (`1351230d4b2dea5b45a78a040949f0ccf97ed6a782d584d94982b88d2a0cfe6c`) |
| the trainer's CPU side | `taskset -c 16-22`, 7 threads | **`taskset -c 1-7`, 7 threads: CPUs 1 to 7 on 7 distinct cores (1 to 7)**, checked from `/sys/…/topology/core_id` at the window. Core 0 (CPUs 0 and 8) carries the window's own process, the stop guard and this laptop's remote reads |
| where checkpoints are written | this laptop `/workshop/bench-deem-2026-09-27/t1/ckpt/` | **benchbox's disk, the same absolute path.** The trainer's 6 GB floor (A-1) applies there |
| where every gated read runs | this laptop's CPU, P-cores 0–7, the runner on core 23 | **unchanged** (section 3) |

**2. What changes in the window's mechanics.** In `sweep.py` (registered at A-9's commit; sha256 at this commit
`99286672a2133e217b887527ff6968e05843e7c5d1490357099be076c0a9f335`), `reads_t1.py` (`3b551074bd6c18c20f393baedd31ccad5415cbff5831ad3e7830a1558690e402`), and four new files: `t1transfer.py` (`60cf37f39073f836e74b9b271dfa69dc3d41861e0d6ae37dc519a135661f72b4`), `guard_t1.py`
(`0085e1f32d562c6401e6f89c76ac97e2a996665b21b9ffc121fb3fdd829018dc`), `cuda_parity_file.py` (`47f17d9351e39563bc89279451ad39bca3cc0a5f20a78cefd9c32b32dec59032`) and `dryrun_split.sh` (`666a108dee5c53c2b3b6c3bf5a7b94615a12409fa33074fbd820fc00e926f05b`). The trainer does not
change (section 4). The single-box path of A-9 stays in `sweep.py` as it was.

1. **Split.** `sweep.py --side benchbox` trains the six runs in A-9's order (null-20260927, null-20260928,
   null-20260929, final-20260927, final-20260928, final-20260929) and logs the same `trained` events; it never
   reads, never serves, never calls a timer. `sweep.py --side this laptop` learns that a run finished by reading
   benchbox's `logs/sweep.jsonl` over ssh, and does every null read; `reads_t1.py --fetch` does the finals'.
2. **Transfer, verified** (`t1transfer.fetch`). this laptop rsyncs the finished run's directory (`model.safetensors`,
   the tokenizer files, `T1-MANIFEST.json`) over the private network into the lane's RAM scratch. Before a server starts,
   benchbox's `sha256sum` of the file (read at fetch) and this laptop's sha256 of the copy must BOTH equal the
   manifest's `output_model_sha256`; otherwise the copy is deleted, nothing is read, and the reads stop. **T-V2 is
   unchanged** and runs on that copy. One receipt per run, `receipts/transfer-<run>.json` (benchbox's sha after
   save and at fetch, this laptop's sha before serve, bytes, seconds, one entry per fetch).
3. **One path on both boxes.** The manifest's `ckpt`, the server's `DEEM_CHECKPOINT`, runner v0.4's preflight and
   its residency proof (the server's `/proc/<pid>/smaps` must map the whole file under that path) must all name
   one path. benchbox writes `/workshop/bench-deem-2026-09-27/t1/ckpt/<run>`. On this laptop the read side runs in a
   private mount namespace (`bwrap --dev-bind / / --bind <RAM scratch>/ckpt <that path>`;
   `t1transfer.ensure_ram_namespace` re-execs the command under it), so the same path holds the RAM copy while
   this laptop's disk directory stays a real, empty mountpoint. **Why RAM:** this laptop's disk read 7,612,125,184 B
   free at 2026-09-28T08:49:59Z and 13,274,386,432 B at 13:28Z; the sweep's own guard needs free minus 1.6 GB at
   least 6 GiB, which the first reading fails and the second clears by one checkpoint, and the laptop's disk moves
   with its local backup repo. Plan 6.1's "disk, never `/dev/shm`" governs where a checkpoint is **kept**, and the
   kept copy is on benchbox's disk; this laptop's copy is a transient read copy.
4. **Read-side guards.** Before each pull, the RAM scratch must have the checkpoint plus 4 GiB free and
   `MemAvailable` must be the checkpoint plus 6 GiB (`MEM_FLOOR_BYTES`); otherwise the reader waits and logs.
   **A read is never skipped.**
5. **Deletions.** Each null goes after its S0 P × 1 is on file: this laptop's copy (read_null, unchanged) and
   benchbox's (`deleted on benchbox`). A failed null read keeps benchbox's copy and drops this laptop's (A-9: "a failed
   null read keeps its weights"). Each final's this laptop copy is dropped when its pass ends, so one final's copy at
   most sits on this laptop at a time. The two non-headline finals go from benchbox after their last read.
6. **Pinned by UUID.** The trainer runs with `CUDA_VISIBLE_DEVICES=GPU-fde82ed0-be70-eb1f-2d4d-5538c5b82166`
   (`--cuda-uuid`). The window refuses to start unless the caps gate reads 300 ± 1 W and persistence `Enabled`
   on the host (never inside a sandbox, where persistence reads falsely Disabled), exactly one board in the box,
   no compute process on it, and the CUDA venv's torch sees exactly one `NVIDIA GeForce RTX 3090` with that UUID.
   **One bench on the card at a time** (D-011; the lead's card coordination, 2026-09-28: PAIR-3090 holds the card
   first): the window starts only after `benchbox:/workshop/CARD-LOCK` is absent and the card reads 1 MiB with no compute
   process (read on the host), then takes the lock (`deem-t1 <utc>`) before anything reads or uses the card, and
   releases it when the card is left empty at its posture. The benchbox side refuses to start, and refuses each
   run, unless the lock names `deem-t1`.
7. **`scribe-embed.timer`.** benchbox has no such timer; the window logs `scribe-embed.timer: not on this box` and
   never calls it. The trainer's manifest still records whatever `systemctl --user is-active` reads.
8. **The package-temperature guard** reads Intel's `x86_pkg_temp` zone; a Ryzen has none, so on benchbox it reads
   `None` and never pauses. **The stop guard replaces it for the card and the box:** `guard_t1.py`, its own user
   unit beside the window's, reads each second the UPS (NUT `pr1500`), the free disk and the card's 2 Hz sample,
   and stops the window's unit when the UPS load reads ≥ 80 %, the card's core ≥ 83 °C, the free disk < 15 GiB,
   the UPS is unreadable for 10 s, or its sampler exits (lev L-2's guard, the same limits). A trip is recorded
   first. **No CUDA run starts without the guard's heartbeat fresh (≤ 10 s).** A stopped window is resumable: a
   rerun skips the runs on file, and a run cut part-way restarts from the base (the trainer has no mid-run
   resume), recorded.
9. **Energy, descriptive only** (5.7): the board at 1 Hz in `logs/gpu-1hz.csv` (6.3, unchanged) and at 2 Hz in the
   guard's telemetry; benchbox's UPS `ups.realpower` at 1 Hz in the guard's reads (a step function with a
   one-minute floor); RAPL is root-only on benchbox, so "not recorded".
10. **The descriptive CUDA parity.** No Deem server runs on benchbox. `cuda_parity_file.py` reads the base in bf16
    on the 3090 on S0's first ten rows and hands the served base's probabilities on file (`receipts/parity-base.json`,
    this laptop CPU, 10 of 10 bit-identical, A-3) to `probe_parity.cuda_parity`, unchanged: choices match and
    |dp| ≤ 1e-3. **Never a gate and never blocking.** Receipt: `receipts/parity-cuda-base-3090.json`.
11. **The window's closing event** is "GPU FINISHED". There is no fan to switch off.
12. **Files come home to be committed, with the card's UUID fingerprinted.** Everything benchbox's side writes
    (`logs/sweep.jsonl`, `logs/manifests/*.json`, `logs/*.steps.jsonl`, `logs/*.epochs.jsonl`, `logs/gpu-1hz.csv`,
    the guard's files, `receipts/benchbox-window.json`, `receipts/parity-cuda-base-3090.json`) is pulled back into
    this laptop's worktree by `t1transfer.pull_home`, which replaces the UUID with `GPU-fp:33a0b4acb8e8` and any
    LAN or private-network IPv4 with `<lan-ip>`, then refuses if either survives. Nothing is committed on benchbox.
    **The manifest the reads use is the fingerprinted one.** `t1transfer.fetch` writes it to
    `logs/manifests/<run>.json` with that one substitution (the receipt records benchbox's manifest sha and the
    reads' sha); runner v0.4 reads it, so every row names the committed manifest's sha, as `gate_t1.validity`
    requires. `output_model_sha256`, `ckpt` and every other field are benchbox's bytes.
13. **What the manifest says about the box.** The trainer writes `"box": "this laptop"` as a constant (A-5's
    sha stands, so it is not edited). On benchbox that field misnames the box; the box of record for these six
    runs is benchbox, read from the same manifest's `device.cpu_model` (AMD Ryzen 7 3700X), `device.gpu`, `os`,
    and from `receipts/benchbox-window.json`. benchbox's copy of the estate tree is a shallow clone at this
    amendment's commit, so the manifest's `git_commit` names it.
14. **Every process carries the phone-home switches**: `ORT_DISABLE_TELEMETRY=1`, `HF_HUB_DISABLE_TELEMETRY=1`,
    `HF_HUB_OFFLINE=1`, `TRANSFORMERS_OFFLINE=1`, `VLLM_NO_USAGE_STATS=1`, `GRADIO_ANALYTICS_ENABLED=False` (never
    `DO_NOT_TRACK`), set in each user unit's environment; the benchbox side refuses to run without them. Every
    step longer than a smoke runs in its own `systemd-run --user` unit with a `MemoryMax`, on both boxes.

**3. The reads: the same instrument, the same set, the same order.** Every gated read still runs on this laptop's
CPU: D0's server through `serve_t1.sh` (taskset 0–7, bf16, `DEEM_BATCH_SIZE=1`, no calibration), with runner v0.4
(sha256 `754b9a18b510f2f48a4942c2e20bdb0f197c07414f695b374d3fdb14a4da3d8f`, A-3) on core 23. Plan 5.0 says why the
reads cannot follow the card: "Every gated read runs on this laptop's CPU, P-cores 0-7, as D0's did, whatever device
trained the arm". The equivalence receipt of A-3 stays valid, because the instrument has not moved. A-9's reads run
in A-9's order, steps 1–6, through `reads_t1.py --fetch`; each checkpoint is fetched when its step needs it (the
seeds 20260928 and 20260929 twice if G-FLOOR passes), and the headline stays on benchbox's disk (plan 6.1's "only
the headline's checkpoint is kept" moves there). If the forgetting probe runs (5.6, iff G-FLOOR), it fetches the
headline again.

**4. What this amendment cannot change, and does not.**
- **The recipe (A-8):** swap off; 2 epochs; lr 1e-5; six cyclic shifts per line per epoch (1,362 renderings, 114
  steps per epoch, 228 steps); effective batch 12 as **micro-batch 4 × 3 on the GPU**; AdamW (0.9, 0.999), eps
  1e-8, weight decay 0.0; grad clip 1.0; linear warmup over the first 10 % of steps, then linear decay; fp32 master
  weights with bf16 autocast on CUDA; gradient checkpointing on; the last checkpoint is scored, no early stopping.
  **A failure of this recipe on the 3090 stops the window**, an out-of-memory at micro-batch 4 included. It is never
  answered by a smaller micro-batch, a shorter sequence or another precision without its own amendment,
  registered before the retry.
- **The seeds:** 20260927 (the headline), 20260928, 20260929, and one shuffled-label null per seed, in A-9's order.
- **The pool and the folds:** T-poc, 227 lines; (text, label) `e4593840551294fa`; text set `df69dd1eba91341c`;
  `kit/t1-pool.json` `<withheld: rewritten file, see README.md>`; `kit/folds.json`
  `<withheld: rewritten file, see README.md>`.
- **The trainer:** `train_letter_slot.py` sha256 `6be547428ff4dd1f61c73df1d199d351f1b87915ff00a2feddcb714ea09cbe93`
  (A-5). Unchanged also: `readout.py` `021cf0fc…749e`, `render.py` `ad2bdeae…388f`, `serve_t1.sh` `e8f9623e…b749`,
  `run_deem.py` `754b9a18…3d8f`, `gate_t1.py` `0c65dedc…2942`, `served.py` `56dc02c1…f488`, `probe_parity.py`
  `c989d981…a061`, `t1common.py` `<withheld: rewritten file, see README.md>`, `tables_t1.py` `8949c17b…9af`.
- **The gates:** 5.1's validity (D0's V-1 to V-7 under R-1, T-V1 to T-V5); 5.2's bars read at min(kit-order count,
  mean count over k = 1..5): **G-BASE** (exact one-sided McNemar p < 0.05 against the registered naive Bayes, 68/108
  by A-2, and the rotated mean above it), **G-FLOOR ≥ 34/108**, **G-KEV ≥ 72/108**, **G-SEAT-TALK ≥ 88/108**; the
  across-seeds rule; **the null bar ≤ 25/108** (5.4), or every tuned figure is VOID; the contamination DiD with
  A-10's cut; 5.3, 5.6 and 5.7 reported and never gated; A-9's "no figure is read before the nulls' rows are on file".
- **The reads' instrument, set and order** (section 3).

**5. What the card change means for the readings, said before any run.** The trained weights are the 3090's: the
same recipe and seed on another card would take other CUDA kernels and would not reproduce them bit for bit. No
such weights exist, because the window never ran, so no T1 figure is ever set beside a figure from another card.
Every gated number is a this laptop CPU read of the bytes the 3090 produced. The window's minutes, watts and VRAM
peak are the 3090's at 300 W, measured and printed (4.5, 5.7); the plan's 5090 minutes were estimates for another
card and are never printed as a comparison.

**6. The receipts this amendment adds** (in `receipts/`): `benchbox-window.json` (the board by UUID: name, PCI bus,
link width, VBIOS, `power.limit`, persistence, driver; `uname -r` and `hostos.read_os()`; the three tree hashes on
both boxes; the python, torch and CUDA versions torch reports; the trainer's cpuset and its cores; the idle
`memory.used` baseline, three reads over 2 s, median; the UPS line; the switches; the caps gate's verdict),
`transfer-<run>.json` per checkpoint pulled, `parity-cuda-base-3090.json` (descriptive), `split-dryrun.json`
(the dry run below), and `pull-home.json` (what came home and what was fingerprinted).

**7. The dry run, before this commit.** `dryrun_split.sh` ran the whole split on this laptop's CPU with `--tiny
--device cpu` (1 step per run; two held-out-table items standing in for S0 and two for H48; "benchbox" a second
local directory bound under bwrap over the RAM scratch so the one checkpoint path resolves to each side's copy),
then `reads_t1.py --tiny --fetch`, then the pull home: 2026-09-28T13:55:25Z to 14:12:57Z, exit 0 (`receipts/split-dryrun.json`). The train side's preflight passed (A-8 and A-11 registered, the A-7 receipt, the disk floor, the three tree hashes equal, the cpuset 16-22 on 7 distinct cores, the CPU probe) and it trained the six runs in A-9's order; the read side fetched each null, read it (T-V1, parity 2 of 2 bit-identical, the stand-in S0 pass) and deleted both copies; `reads_t1.py --fetch` read the three finals in A-9's order, fetched seeds 20260928 and 20260929 again on the forced G-FLOOR branch, and deleted them from "benchbox" at the end; all 8 fetches verified (benchbox's sha, this laptop's sha and the manifest's agreed); 8 files came home with nothing identifying left. **A first attempt at 13:43:45Z was stopped by the lane:** with this laptop's checkpoint root as a symlink into RAM, runner v0.4's residency proof mapped 0 of 1,504,827,608 B (`/proc/<pid>/smaps` names the resolved path, not the manifest's), so the first null read failed and kept its copy; the bwrap bind of item 2.3 replaced the symlink, and a test pins it. The dry run executed `sweep.py` at `189d767b344cacb30038826eb760b436f46a799abc2f03c3d93dc4826c503db2`; the committed `sweep.py` adds only the card-lock check of item 2.6, which acts on CUDA runs alone. The tests: `tests/test_fast.py` 7 of 7, `tests/test_tables.py` 5 of 5, `tests/test_transfer.py` 7 of 7 (a planted byte flip in the copied weights refuses the read and leaves no copy; a faked full `/tmp` waits and never skips; a file mapped under the bind is named by the registered path in `/proc/<pid>/maps`). The CUDA path cannot be
dry-run without the card; its first minute on the day is the preflight's 1-step probe, **the window's first
call**, which comes after this commit is pushed.

### A-12 (2026-09-28T14:58:35Z): the read side's ssh skips the system config inside its namespace (a transfer fix; nothing read yet)

*Cold scroll: this amendment changes one line of how this laptop reaches benchbox, and nothing about what is trained,
read or gated. At its commit the window is training on benchbox (A-11 as pushed at 2026-09-28T14:15:11Z), and no
T1 checkpoint has been read.*

- **What happened.** The window started at 2026-09-28T14:56:35Z (the card lock `deem-t1 2026-09-28T14:56:35Z`
  taken after PAIR-3090 released the card; the caps gate passed at 14:56:49Z; **the 1-step CUDA probe, the
  window's first call, finished ok at 14:57:01Z**, 12.0 s). The descriptive CUDA parity against
  `receipts/parity-base.json` read **6 of 10 rows agree (same choice, |dp| <= 1e-3) -> FAIL (descriptive)** at
  14:57:09Z; as registered (6.3, A-11 2.10) it is never a gate and never blocking, and the window went on. The
  read side (`sweep.py --side this laptop`, started 14:56:42Z) stopped at once, before any fetch: inside its bwrap
  namespace, root-owned files map to `nobody`, and ssh refused `/etc/ssh/ssh_config.d/20-systemd-ssh-proxy.conf`
  ("Bad owner or permissions", rc 255). The dry run (A-11 7) used a local stand-in for benchbox and never ran ssh.
- **The fix.** `t1transfer.ssh_base()`: every ssh, and rsync's transport, runs as `ssh -F /workshop/.ssh/config -o
  BatchMode=yes`, which skips the system-wide file and keeps the user's own config and known_hosts. Proven live
  inside the namespace before this commit: benchbox's `logs/sweep.jsonl` read (8 events), the card count read (1),
  and one receipt rsynced. `tests/test_transfer.py` gains a test that ssh carries `-F` (8 of 8 pass).
  `t1transfer.py` sha256 is now **`f12f734a337f0b4ea337100e32ac92959318ff2d1f48e862d2c6bba763b7f0c1`** (A-11 registered `60cf37f3…72b4`). Nothing else changes: not the
  fetch's verification, the guards, the order, the instrument, the recipe, or any gate.

### A-13 (2026-09-28T15:23:59Z): the window stopped on the null bar (5.4); the finals are not read; one receipt's name

*Cold scroll: this amendment records what ran and what did not after the registered null bar refused on
2026-09-28, and moves one descriptive receipt to a name the pinned tables code does not misread. It changes no
rule, no bar and no figure. The ruling on the leak (5.4: "the instrument leaks and every tuned figure is VOID until
the leak is named and fixed") is the lead's.*

- **The nulls, read on this laptop's CPU through D0's instrument (runner v0.4, A-3), each at kit order:**
  null-20260927 **18 / 108**, null-20260928 **29 / 108**, null-20260929 **24 / 108**. The bar is ≤ 25 / 108 each.
  `gate_t1.null_check` returns `ok: false` (seed 20260928 above the bar). Every read is valid (V-1 to V-7, T-V1,
  T-V2, T-V3 10 of 10 bit-identical each, T-V5; V-6 by `check_tokens_t1.py` after the reads). The rows are
  `rows/t1-null-<seed>.s0.P.rep1.jsonl`; the tables are `results.md` sections 4, 5 and 11, as `tables_t1.py` prints them.
- **The stop.** The lane's brief says to stop and report when a registered gate refuses. The null-20260928 read
  landed at 15:16:24Z, while final-20260927 was training. The lane read the gate with `gate_t1.null_check` and
  **stopped the window unit by hand at 2026-09-28T15:21:43Z** (`logs/sweep.jsonl`: `WINDOW STOPPED`, signal 15).
  - **Trained:** null-20260927, null-20260928, null-20260929 and final-20260927 (228 steps each).
  - **Cut:** final-20260928, at step 127 of 228 (the trainer saves only at the end, so no checkpoint exists).
  - **Not trained:** final-20260929.
  - **Not read:** the finals (A-9's reads, `reads_t1.py --fetch`), and no final figure exists.
  - **Kept:** final-20260927's checkpoint (sha256 `d1cb11f1a261c9615f81de647dbadbac91b60ebb88366bad36f3165c5c26d1d2`),
    on benchbox's disk at the registered path.
  - The window is resumable by its log. A rerun skips the four trained runs and restarts final-20260928 from the base.
- **The card after the stop:** empty (0 compute processes, 1 MiB, the idle baseline B0 read at the window's start was
  also 1 MiB), 300.00 W, persistence Enabled, read on the host. The lane's lock was released at 15:22:05Z. The stop
  guard never tripped (0 of 1,392 reads; core max 72 °C; UPS load max 40 %).
- **One receipt's name.** `tables_t1.py` (sha pinned, unchanged) reads every `receipts/parity-*.json` as a CPU parity
  receipt and fails on the descriptive CUDA one (`KeyError: 'bit_identical'`). A-9's own `cuda_parity` would have
  written `parity-cuda-base.json` and failed the same way. So the descriptive receipt that A-11 2.10 and 6 name
  `receipts/parity-cuda-base-3090.json` is stored as **`receipts/cuda-parity-3090.json`** (the planner's draft
  name). The bytes are unchanged, and `receipts/pull-home.json` records the move. It reads 10 of 10 same choice and
  **6 of 10 within |dp| ≤ 1e-3** (max |dp| 0.016), so FAIL (descriptive, never a gate, as registered).

---

*Results go below this line (5.8 step 8), and in `results.md` and `rows/`, only after the first row.*

## Results notes (written after the rows; post hoc, dated)

### N-1 (2026-09-28T15:39Z): the independent check of PR #186, wording and disclosure only

*Cold scroll: T1's three shuffled-label nulls were read on this laptop's CPU from 2026-09-28T15:02:52Z to
15:19:47Z (S0 P × 1 at kit order): 18, 29 and 24 of 108, against the registered bar of ≤ 25 each (5.4). The
verdict is **VOID** (5.4, T-V4). The lead ruled on 2026-09-28 (D-20260928-027): T1 closes VOID as registered, the
two unfinished finals are not resumed, and no tuned checkpoint is read. So A-13's "resumable by its log" describes
the code, not a plan. This note changes no rule, bar, figure, row, receipt or pinned file.*

- **The only verdict is VOID.** The pinned `tables_t1.py` (`8949c17b…89af`) draws `results.md`. Its fixed wording
  was written before the nulls were read, so read three places this way:
  - **The title and header.** "hour one" and "S0 and H48 are read here only by the baseline … and … by the
    untuned base" are registration-day words. The three nulls also read S0 (sections 4, 5 and 10).
  - **Section 6, "across seeds".** Every bar's cell prints "fail". No final row is on file (each seed reads "not
    on file"), and the pinned code prints a bar with no rows as "fail". **These cells are not verdicts.** The
    verdict is section 6's verdict line: VOID.
  - **Section 11.** The training losses, final-20260927's included, are losses on each run's own training
    renderings. None is a read, and all fall under section 5's "every tuned figure below is VOID".
- **Descriptive, not registered** (no bar reads it, and it names no leak): over its last 10 % of steps,
  null-20260928's training loss was 0.000, so it fit its shuffled labels. null-20260927 ended at 1.792 (ln 6, the
  uniform guess) and null-20260929 at 0.255 (`results.md` section 11).
- **Steps after the pull, disclosed.** `t1transfer.pull_home` (A-11 2.12) scrubs card UUIDs and IPv4s only. After
  it ran, the lane did three things by hand, outside the pinned code:
  - It replaced the PCI bus id in `receipts/benchbox-window.json` with `<pci>`, using lev L-2's pattern. So the
    "PCI bus" that A-11 6 lists reads `<pci>` in that receipt; the value is in A-11's own board table.
  - It renamed the descriptive CUDA parity receipt (A-13).
  - It added four blocks to `receipts/pull-home.json`: `manifests`, `manifest_bytes_equal_to_the_reads`,
    `pci_pass` and `renamed_after_pull`. `pull_home`'s own receipt is the `pulled` block.
- **Errata of record** (the registered text is not edited):
  - **A-11 4 mis-abbreviates `gate_t1.py`'s pin** as `0c65dedc…2942`. The file's sha256 is
    `0c65dedc0bdb7e98d6b2d0f18a35563a7307e47947a29704cc8078dc16942590`, unchanged since `f717fce6` (A-10). Only
    the tail of the abbreviation is wrong.
  - **`logs/sweep-this laptop.jsonl` has no refusal event.** It records the read side's two starts (14:56:42Z and
    14:58:47Z) but not the first start's refusal. The refusal is in A-12.
