# Deem arm D0: pre-registration v0

*Written 2026-09-27 between 11:40Z and 11:55Z on this laptop, and committed and pushed BEFORE the first
model call (`git log` is the receipt: the first row's `utc` is later than this file's commit). The
install and the weights download ran before this commit, as did the renderer's `--dry` pass and the
H48 build. None of them is a model call. Only dated amendments are appended; no registered text changes.
an operator gave the bench go at about 11:35Z ("let's do some more research / benches on this, and see what
we find, compared to our current seats, etc, and what these models can do / how we can tune them
ourselves"). The recipe is `/workshop/bench-archive/plans-2026-09-26/RECON-deem-decision-models.md`, sections
0 and 7. The lead's three binding additions (the stub trap, the H48 contamination control, and the
per-author split) are folded in below; this file was not yet committed when they arrived.*

---

## 1. The question: a measurement arm, not an adoption verdict

**What does deem-0.8-v1 score on our six-way S0 (108 items) on this laptop's CPU, and at what latency?**

deem-0.8-v1 is LibertAI Labs' Apache-2.0 0.8B "decision model" (Qwen3.5-0.8B, full fine-tune), and
this arm runs it exactly as shipped:

- Deem's own Python reference server, at the direct single-pass readout;
- no calibration, so every answer is a raw softmax at T = 1.0;
- bf16 on this laptop's 8 P-cores, one question per request.

It also asks whether the same model does as well on 48 wall lines that were never in any kit (H48,
section 7), which is the contamination control.

It is **not** an adoption verdict, a seat exam, or a claim about deem-9b-v1. Section 9 says what it
may and may not say.

## 2. The pins

| what | pin |
|---|---|
| model | `LibertAIDAI/deem-0.8-v1` at revision `8cbabbb2c4a7ef13c6b43f0ef3ae4157983c6d21` (the v1.1 weights, repo head 2026-09-25T11:42:30Z), `v1.0/*` excluded |
| weights | `model.safetensors` 1,504,827,608 B, sha256 `80438125177c0681e855158d46720cbbd4ff07ea620cd55b0f456aafbfcf778d` (verified at 11:39:57Z, `receipts/install.log`); tokenizer sha256 `06b9509352d2af50381ab2247e083b80d32d5c0aba91c272ca9ff729b6a0e523` |
| runtime | `github.com/Libertai/deem` at `6755b30bf6bbd9a81f8db6ba42cc0fd62c9f4719`, `serve/deem_server.py` (TorchBackend); torch `2.14.0+cpu` from the PyTorch CPU index, transformers `5.17.0`, Python 3.14 in a uv venv; `torch.cuda.is_available()` = False (`receipts/install.log`) |
| where | this laptop (Intel Core Ultra 9 290HX Plus; P-cores 0-7 at 5.4-5.5 GHz max; no AVX-512); weights, venv and clone in `/workshop/bench-deem-2026-09-27/` on disk, outside git |
| S0 | `bench/jev-2026-09-21/kit/task_c.json`, 108 items, sha256 `61a0b1d6e7cf66fd91992c0c3a4572f6fbae90e6471048fd7dc1eb41ff2d7f39` (the COI series set `c`) |
| H48 | `bench/deem-2026-09-27/kit/h48.json`, 48 items, sha256 `<withheld: held-out set, see README.md>`, text fingerprint `<withheld: held-out set, see README.md>` (section 7) |
| instrument | `run_deem.py` `d29b2105…760f`, `serve_deem.sh` `5222f951…e145`, `check_tokens.py` `e6f092e5…7fbf`, `tables_deem.py` `2cea785d…d6ea`, `build_h48.py` `a7c8dd7f…3ea8`, `install_deem.sh` `32b19b4c…4912` (sha256, this commit) |
| the OS | read on every row by the estate's `hostos.read_os()`, imported from its canonical copy `bench/render-grid-2026-09-24/hostos.py` (sha256 `a6a74bba…e3ad`), never copied. Read at 11:4xZ: `{"pretty_name": "Ubuntu 26.04.1 LTS", "id": "ubuntu", "version_id": "26.04", "kernel": "7.0.0-34-generic", "arch": "x86_64", "source": "/etc/os-release"}` |
| seed | 20260921: the kit's seed and the H48 draw's. No sampling happens at inference (a readout at the answer slot), so no seed reaches the model |

**Why the server's `/health` alone is not enough (the stub trap, recon section 0, trap 8).** An empty
or unset `DEEM_CHECKPOINT` makes the server serve a stub that returns uniform logits, silently. Every
answer would then be option A, at confidence 0. The runner therefore refuses before the first row
unless ALL of these hold:

- `/health` reads `{"status": "ok", "model": "deem-0.8-v1", "backend": "torch"}`;
- the server process's own environment names `DEEM_CHECKPOINT=` the verified download
  directory, `CUDA_VISIBLE_DEVICES=` empty, and CPU list `0-7`;
- the weights' sha256 is recomputed and matches the pin;
- the clone is at the pinned commit;
- `import deem` fails in the venv from a neutral directory, so no PyPI `deem` is installed;
- run with the server's own fallback logic from `serve/`, `deem` resolves to
  `/workshop/bench-deem-2026-09-27/deem/src/deem/__init__.py`.

During the pass, any row whose six probabilities are all equal stops the pass and voids it (V-4).

## 3. The two renderings, frozen

Both renderings go through Deem's own `build_prompt` (`src/deem/format.py`). The state is always
serialized as canonical JSON. The options are `"<key>: <name>"` in the kit's own order, which is
alphabetical by key (darwin, einstein, hypatia, ibn_sina, sagan, socrates = A to F) for every item.

- **P, the primary (the headline).** This is arm 12's split, word for word. The state is
  `"A dinner table with six guests: Charles Darwin, Albert Einstein, Hypatia of Alexandria, Ibn Sina, Carl Sagan, Socrates of Athens."`
  The instructions are the Jev bench's `run.py` `INSTR_C % text`, verbatim. This is the rendering
  that makes the row comparable with Kev's.
- **S, the secondary (reported, never gated, never swapped in after replies are seen).** This is the
  Deem-native split. The state is the line itself, and the instructions are
  `"Six people are at a dinner table. Which of them said the line in the state?"`

| set / rendering | set fingerprint (sha256 over the sorted distinct prompt sha256s, first 16 hex) | prompts |
|---|---|---:|
| S0 P | `9044636cdef22cdf` (equals the recon's) | 108 |
| S0 S | `51d85c79f334c089` (equals the recon's) | 108 |
| H48 P | `<withheld: held-out set, see README.md>` | 48 |
| H48 S | `<withheld: held-out set, see README.md>` (built, not run in v0) | 48 |

**Item 0 of S0 (`c000`, label darwin, chair gemma), rendering P.** Prompt sha256
`028cd141528bb2d431e82341b68dc6ab23cf42282b8ef5e3323f54d9b2affe80` (equals the recon's):

```text
<state>
"A dinner table with six guests: Charles Darwin, Albert Einstein, Hypatia of Alexandria, Ibn Sina, Carl Sagan, Socrates of Athens."
</state>

Question 1: Six people are at a dinner table. Which of them said the line below?
Line: It has been a most stimulating evening, and I find myself profoundly humbled by our discourse. To Hypatia, I offer my deepest admiration for your vision of a cosmic harmony, and to Socrates, my gratitude for your relentless pursuit of truth through inquiry. As for myself, I shall return to my studies with a renewed sense of the vast, interconnected web of life that binds us all, however much we may struggle to comprehend its full complexity.
Options:
(A) darwin: Charles Darwin
(B) einstein: Albert Einstein
(C) hypatia: Hypatia of Alexandria
(D) ibn_sina: Ibn Sina
(E) sagan: Carl Sagan
(F) socrates: Socrates of Athens
Answer 1: (
```

**Item 0, rendering S.** Prompt sha256 `8ffb1f475b0f997352d943295b928c76278288ee92fe726f5c576a6e4aa2a279`:

```text
<state>
"It has been a most stimulating evening, and I find myself profoundly humbled by our discourse. To Hypatia, I offer my deepest admiration for your vision of a cosmic harmony, and to Socrates, my gratitude for your relentless pursuit of truth through inquiry. As for myself, I shall return to my studies with a renewed sense of the vast, interconnected web of life that binds us all, however much we may struggle to comprehend its full complexity."
</state>

Question 1: Six people are at a dinner table. Which of them said the line in the state?
Options:
(A) darwin: Charles Darwin
(B) einstein: Albert Einstein
(C) hypatia: Hypatia of Alexandria
(D) ibn_sina: Ibn Sina
(E) sagan: Carl Sagan
(F) socrates: Socrates of Athens
Answer 1: (
```

The request body is `{"state": <state>, "questions": {"q": {"type": "choice", "instructions": <instructions>, "options": [six strings]}}}`.
It never carries a `dataset`, so no calibration temperature can apply.

## 4. Scoring and validity

**Scoring.** The server's `choice` (the argmax Choice), mapped back to its guest key, is compared
with the registered label: correct or incorrect, strict, with **no lenient pass**. Deem's answer is
typed output (a probability per option, read at the answer slot), so **there are no format failures
and no parse step**. Nothing can be "unparseable", and none is reported. The Brier score is multi-class
(the Jev bench's `run.py` definition) on the raw T = 1 probabilities, so it is comparable in kind
with Kev's `-raw` rows only. It is not a calibrated Brier.

**Validity: a pass is VOID, never "low", unless all seven hold** (`tables_deem.py` `validity`):

- **V-1** Every item answered once: 108 distinct `(uid, text_sha256)` rows for S0, 48 for H48.
  The kit's `id` is not unique (108 lines under 50 ids), so rows are never keyed on it.
- **V-2** The pass's prompt set fingerprint equals the registered one (section 3).
- **V-3** The server's choice equals the argmax of its own probabilities on every row.
- **V-4** No row has six equal probabilities (the stub).
- **V-5** Every row carries the pinned revision and weights sha256.
- **V-6** The server's `usage.prompt_tokens` equals a local count of the re-rendered prompt with the
  checkpoint's own tokenizer on every row (`check_tokens.py`, run after the passes, loads the
  tokenizer only). This is the evidence that the server tokenized the prompt this kit rendered.
- **V-7** No failed request. A failed request is never a row and never a wrong answer: it goes to
  `<stem>.errors.jsonl` and the pass stops.

**The one label the arm may apply (the recon's contender bar, section 7).** On S0 P rep 1:

- at or above **72/108**, deem-0.8-v1 is "a CPU contender for the six-way", at least Kev-9B's
  zero-shot score (Kev-9B is the Apache-2.0 comparator);
- below 72/108 it is "not a contender as shipped", with **no threshold table** at all.

Latency and energy are reported either way. Neither case is an adoption verdict.

## 5. Latency, pinning, energy

- **Server:** `taskset -c 0-7` (the P-cores), `OMP_NUM_THREADS=8`, `DEEM_DEVICE=cpu`,
  `DEEM_BATCH_SIZE=1`, `DEEM_N_ORDERS` unset (1), no `DEEM_CALIBRATION`, bound to
  `127.0.0.1:8300` (`serve_deem.sh start`).
- **Runner:** `taskset -c 23` (one E-core), so the client never competes with the server's cores.
- **GPU hidden and checked:** `CUDA_VISIBLE_DEVICES=` (empty) for the server and the runner. The
  torch wheel is CPU-only anyway. `nvidia-smi --query-compute-apps` must list no process of ours
  before, during and after (`receipts/gpu-check.txt`). It listed none at 11:50Z.
- **Residency by growth (the CPU form):** `serve_deem.sh` samples the server's `VmRSS` from exec
  until `/health` answers. The growth must be at least the 1,504,827,608 B of weights, or the start
  refuses (`receipts/residency.json`).
- **Latency:** the client's wall time per decision on loopback, with one request in flight. Report
  the median and p95 (nearest rank) per pass, over every row.
- **Warm-up:** three requests before each pass, not kit items (`{"state": "warm-up", …
  "options": ["first", "second"]}`), never rows. Their times go in the pass receipt.
- **Energy:** RAPL is readable without root through `/run/rapl/energy`, the 1 Hz root sampler's
  world-readable file. The sysfs counters themselves are root-only, and this arm never reads them.
  So CPU energy **is** measured:
  - the package-0 and psys joules over each pass, wrap-safe at 262,143,328,850 µJ;
  - a 10 s idle read immediately before each pass, with the server loaded and idle;
  - net J per decision = (pass J − idle W × pass s) ÷ decisions.
  The package counter covers everything the laptop runs, and the pass receipts carry a load and
  temperature snapshot before and after. At 11:4xZ a peer's `unittest` process held one core at
  100 % and the package read 99 °C. Both are named here as conditions, never subtracted by hand.
- **this laptop throttles** (the ear bench measured a 101-104 °C package on 2026-09-25). Its latency is
  a first reading on this laptop, not the timing baseline for a CPU seat (that is the recon's D1, on
  the mini node).

## 6. The run, in order

Box: this laptop for every step. These are lane commands, not an operator paste.

1. `bash serve_deem.sh start`. It refuses unless RSS grows by the weights.
2. `CUDA_VISIBLE_DEVICES= taskset -c 23 python3 -B run_deem.py --set s0 --render P --rep 1 --server-pid <pid>`
   is **the headline**.
3. The same with `--rep 2`. **Determinism:** S0 P twice, same server, same bytes. The report gives
   the identical-choice agreement (x of 108), bit-identical probability rows, and the largest
   probability difference. There is no seed to fix: a readout is a single forward pass with no
   sampling.
4. `--set s0 --render S --rep 1`: the secondary rendering.
5. `--set h48 --render P --rep 1`: the contamination control (section 7).
6. `bash serve_deem.sh stop`.
7. `/workshop/bench-deem-2026-09-27/.venv/bin/python -B check_tokens.py` (V-6), then
   `python3 -B tables_deem.py` writes `results.md`.

Not in v0, named so its absence is not a silent skip: gap 4's reorder (S0 P × 6 orders, 540 more
rows), H48 under S, a mini-node latency of record (D1), and deem-9b-v1 (D2, which needs an operator). Each is
its own dated amendment, written before its first call.

## 7. The contamination caveat, and the held-back slice H48

**The caveat, as the recon states it (section 1).** Nothing in Deem's named training sources touches
our long table. The anchors are 2015-2023 public NLP sets, and everything else is generated. The
timeline:

- the six-way kit went public with exhibit fifty-eight at 2026-09-23T13:17:23Z;
- Deem's HF repos were created 2026-09-24T18:12Z;
- the wall lines have been public at `longtable.strata2signal.com/api/wall` since before the kit
  was built (COI `SERIES-SETS.md`: the six-way set public since 2026-09-19 on the wall).

So the window is about 29 hours for the kit and longer for the wall lines, and the stated recipe gives
no path by which either entered. **Risk: low, not provable.** Deem publishes weights only, with no
data manifest (the HF cards say `datasets: []`). The paper's "teacher distillation as augmentation"
is the only open door, and a teacher would not reproduce our speaker labels. Under the COI's rule R-1,
a model released after a set's public date carries the contamination-possible flag, so every S0 cell
of this arm carries it.

**The control: H48 (the lead's addition, the recon's section 7 construction).**

- **The source.** The wall as read 2026-09-27, frozen in the kit as `kit/wall-5f4a4e1c.json`
  (sha256 `5f4a4e1c3f096e7dcc284781d2d761166e19f583b549a59398796961692734d8`).
- **Eligibility.** S0's own rules, imported from `bench/jev-2026-09-21/kit/build_kit.py`: no host
  lines, no withheld lines, at least 80 characters, and a line naming its own speaker is dropped.
  Every line whose text is in S0 is then removed.
- **The pool.** darwin 15, einstein 12, hypatia 30, ibn_sina 8, sagan 27, socrates 18: **110 lines**.
  The recon printed "99" beside these same per-guest counts; the counts sum to 110. This is recorded,
  not corrected by hand.
- **The draw.** 8 per guest, guests in key order. Each guest's pool is sorted by
  `(seat, course, line_index)`, and one `random.Random(20260921)` draws `sample(pool, 8)` per guest
  (build_c's own pattern).
- **The result.** 48 items, text fingerprint **`<withheld: held-out set, see README.md>`** (equals the recon's), chairs
  mistral 28 / gemma 20 (equals the recon's). `build_h48.py` refuses to write any other fingerprint.
- **Never in a kit.** A search of every file in this repo at `c67a0fc6` for each H48 line found them
  only in the COI's three raw wall snapshots (`KIT.lock.receipts/wall-*.json`), never in a kit, a
  row file or a prompt.
- **No comparator has read H48.**

**What it controls and what it cannot.** H48 was public on the wall, so it controls for **kit**
leakage, not wall leakage. The wall has not grown since the kit, so no post-release slice exists yet.

**The registered reading.** On rendering P, rep 1: S0 minus H48, with Newcombe's hybrid-score 95 %
interval.

- If the interval's lower bound is above zero, the file prints **"CONTAMINATION SIGNAL"** as a
  finding. It is not interpreted away.
- Otherwise it prints "no signal", with the reminder that the control cannot prove absence.
- H48 above S0 is not a contamination signal.

The two sets also differ in guest balance (8 against 18 per guest) and chair mix (mistral 28 / gemma
20 against S0's 57 / 51). Section 2 of the results therefore prints both sets split by chair.

**Per-author split (the lead's addition; the COI's out-of-family gate).** Every score, on every pass,
is also reported split by the chair model that wrote the line, as counts (n < 60).

- S0's family per item is read from `bench/cost-of-intelligence-2026-09-23/KIT.lock.json`
  (`sets.c.items_list[].author_family`: gemma 51, mistral 57, matched to the wall's `chairs`).
- H48's family is the wall's `chairs` for each line's entry, carried in `kit/h48.json`.

## 8. The comparators, each read from its file (never re-run, never from memory)

`tables_deem.py` re-reads every row below at build time, with its gates.

- **Jev bench rows:** it recounts the rows file, which must equal that file's own `report.json`
  summary, with 0 errors, and prints any floored rows.
- **COI rows:** the `series.csv` row must be `headline` true, `row_status` complete, `rate_status`
  rate, on a verified set fingerprint with 0 instrument artefacts and no hold.

The values below were read from these files at 11:44Z.

| comparator | on S0 | file | gates as read |
|---|---:|---|---|
| OpenJev-FP8 readout, arm 2 (CC BY-NC 4.0, bench-only) | **101/108**, 93.5 %; Brier 0.1219; 0.35 / 0.42 s | `bench/jev-2026-09-21/rows/openjev-fp8-readout.c.jsonl` (+ `.report.json`) | recount = report, errors 0, floored 0 |
| Kev-9B as served, arm 12 (Apache-2.0) | **72/108**, 66.7 %; Brier 0.4803; 0.11 / 0.14 s | `bench/jev-2026-09-21/rows-gaps-card/kev-9b.kc.jsonl` | recount = report, errors 0, floored 0 |
| Kev-9B raw logits (the Brier comparable in kind to Deem's T = 1) | 72/108 | `…/rows-gaps-card/kev-9b-raw.kc.jsonl` | the same |
| Kev-4B as served / raw | **60/108**, 55.6 % | `…/rows-gaps-card/kev-4b.kc.jsonl`, `kev-4b-raw.kc.jsonl` | the same |
| gemma 4 26B FP8 readout, arm 10 (the docent seat's checkpoint, read out on arm 2's runtime) | 87/108, 80.6 % | `bench/jev-2026-09-21/rows-addenda/gemma4-fp8-readout.c.jsonl` | recount = report, errors 0, **qualified: 103 of 108 rows floored** (a letter outside the served top-20; the Jev bench's own caveat) |
| **the seat's bytes** `gemma4:26b-a4b-it-q4_K_M@5571076f` on largecard beside the live seats (COI B-1) | **86/108**, 79.6 % strict | `bench/cost-of-intelligence-2026-09-23/series/series.csv` row `r0-6e7d38ff` | headline, complete, rate, fingerprint `1feac3433cb4217b` ok, 0 artefacts, 0 format failures, contamination clean |
| **the seat's bytes** `mistral-small3.2:24b@5a408ab5` on largecard (COI B-1; the classifier seat's model) | **74/108**, 68.5 % strict | the same file, row `r0-4643b046` | the same |
| gemma4 `@5571076f` on benchbox's 3090, the COI's B-seat reference row | 88/108, 81.5 % strict | the same file, row `r0-bba46e51` | the same |
| gemma-4-26B-A4B-it-FP8-dynamic generate on vLLM (the docent seat's checkpoint, Jev pre-0 rows) | 87/108, 80.6 % strict | the same file, row `rpre0-f6b766b2` | headline, complete, rate, fingerprint ok, 0 artefacts, **4 format failures**, contamination unverified |

The "seat's bytes" wording is PLAN-v2 line 289 ("the same-bytes pairs"). Recon section 5 names
mistral-small3.2 as the classifier seat's current model.

**Protocols differ, and the table says so every time.** The Jev rows are letter readouts under the
OpenJev helper's layout, and Kev runs under its own server. The COI rows are generative answers under
a strict parser, on another rendering of the same 108 items (fingerprint `1feac3433cb4217b`). Deem's P
is arm 12's split, so Kev is the nearest in protocol. The latencies were measured on GPUs, this arm's
on a laptop CPU, so they are never a ratio.

## 9. What the arm may and may not say

**It may say:**

- deem-0.8-v1 (`8cbabbb2`), as shipped (single pass, raw T = 1, bf16, Python reference server),
  scored k/108 on S0 under rendering P on this laptop's CPU, and k/48 on H48;
- its latency on this laptop, with the load and temperature conditions stated;
- whether it reached the contender bar (section 4);
- the per-chair split, the S0-minus-H48 figure with its interval, the P-vs-S agreement, and the
  determinism count.

**It may not say:**

- anything about deem-9b-v1, the paper's gated or chain-of-thought modes, or the Rust int8 runtime;
- "adopt", "seat-ready", or any threshold for a seat;
- a calibration claim (T = 1 is uncalibrated);
- a latency figure for any box but this laptop, or a latency ratio against a GPU comparator;
- that Deem is better or worse than OpenJev, Kev or gemma "in general": this is one task, 108 items,
  and a difference under about ±9 points is inside binomial noise at this n;
- that the model is uncontaminated. The control can show a signal, never absence.

## 10. Deviations from the recon's recipe, each named

- **Kit and runner location.** The kit and runner live in the repo at `bench/deem-2026-09-27/`
  (the lane brief). The runner reads S0 from `bench/jev-2026-09-21/kit/` rather than a copy. It
  reads the deem source and weights from `/workshop/bench-deem-2026-09-27/`.
- **`--rep` in the resume key.** The recon's runner resumes on `(arm, render, order_k, uid,
  text_sha256)`, so a second P pass would have been skipped. `--rep` joins the key. `--set` joins it
  too, for H48.
- **More columns on every row.** Each row also carries the model revision, the weights sha, the deem
  commit, torch and transformers, the server CPU list, and the OS from `hostos.read_os()`. The box is
  named by its estate name (this laptop), never by `platform.node()`, because the hostos law keeps the
  unix hostname out of public objects.
- **The added checks** (section 2 and V-1 to V-7): the preflight, the stub void, failed requests
  kept off the rows, and the token check.
- **Three warm-ups** per pass (recon section 7) and a pass receipt.
- **The runs in v0 are P, P, S and H48-P, per the lane brief.** The recon's gap-4 leg
  (`--orders 6`) is not in v0; see section 6.
- **The H48 pool counts sum to 110, not the recon's printed 99** (section 7).
- **The install ran as one script** (`install_deem.sh`). `git clone` then `checkout` of the pinned
  commit; the venv, torch and transformers exactly as section 0.1; the download exactly as section 0.2.
  Disk: `/` read 18 GB free before and 15 GB after (`/bin/df`, `receipts/install.log`).

---

## v0.1: residency is proved after the warm-ups (2026-09-27T11:53Z, before any model call)

*A dated amendment written after v0 was pushed (`64bcdc6e`, 11:52:05Z) and before any request
reached the model. No warm-up and no row had been sent. The server that failed the check below
never answered a `/v1/systemone` request and was stopped. The code half of this amendment landed
in `b797fed6` (11:53:30Z); the text landed one commit later, because the first attempt to write it
failed on a format character. Both commits precede the first call.*

**What happened.** `serve_deem.sh start` at 11:52:10Z brought the server up with `/health` =
`{"status": "ok", "model": "deem-0.8-v1", "backend": "torch"}` in 8.6 s. Its RSS then grew
**890,056,704 B**, which is less than the 1,504,827,608 B of weights, so the v0 check refused
(`receipts/residency-start1-refused.json`).

**Why.** `/proc/<pid>/smaps` (`receipts/residency-start1-smaps.json`) shows the server maps the
**whole** weights file (1,504,829,440 B mapped, page-rounded), with only **306,241,536 B (20.4 %)**
resident.

- transformers 5.17 loads bf16 safetensors zero-copy, as a file mapping.
- A page is resident only once a forward pass touches it, and `/health` answers before any forward.
- The 508,559,360 B tied embedding (`embed_tokens`, also the LM head, `tie_word_embeddings: true`)
  is the largest untouched block.

v0's form of the law (growth by `/health` ≥ weights) therefore cannot pass on this loader however
healthy the load is. The fault is in the check's form, not in the model.

**The re-registered check (replaces v0 section 5's residency bullet).**

- After the three warm-ups and before the first row of every pass, the server's mapping of
  `model.safetensors` must be **the whole file** (mapped ≥ 1,504,827,608 B).
- The same mapping must be **resident** (Rss ≥ 99 % of the file).
- `run_deem.py` `weights_residency()` reads `/proc/<pid>/smaps`, refuses the pass otherwise, and
  records the result in each pass receipt as `residency_after_warmup`.

A warm-up forward touches every layer's weights, and the tied LM head projects over the full
248,320-row vocabulary, so every page of the file is read. The RSS growth by `/health` is still
recorded in `receipts/residency.json`, as a fact, not a gate.

**Instrument after the amendment:** `run_deem.py` `13515b0c9134769cb6a905467a94f17f80a8f126f335cf000b50458414d0e9e4`, `serve_deem.sh` `7673831ad78b66ea51be227695a4118e92ee8a2b3019d888b91eff3e56f19ffd` (sha256).
Nothing else changes: not the renderings, the sets, the scoring, the comparators, or the order of
the run.

---

## v0.2: the contention gate (2026-09-27T12:05Z, after S0 P rep 1 and 4 rows of rep 2, before any further call)

*A dated amendment. The rows already on file are not edited or dropped: S0 P rep 1 (108 rows,
11:54:09Z to 11:54:57Z) and the first 4 rows of S0 P rep 2 (its segment started 12:00:17Z; the rows are stamped 12:01:38Z to
12:03:17Z).*

**What happened.** From about 12:00Z, peer lanes' CPU-bound processes held some of the server's
P-cores:

- a `unittest` run on CPUs 2 and 3;
- a mutation run's short-lived children on CPUs 0, 5, 6 and 7.

These processes are not this lane's to move. A decision that took 0.441 s (median) in rep 1 took
**28.0 to 37.7 s** in rep 2. Deem's CPU path is thousands of small parallel regions (transformers
logs that the gated-delta-rule and causal-conv1d kernels fall back to their reference PyTorch
implementations), and each region waits for its slowest thread, so a core shared with a peer stalls
every region. At that rate the three remaining passes would take over two hours, their latencies
would measure the peers rather than Deem, and our spinning threads slowed the peers in turn. The
lane stopped the chain at about 12:03:30Z. The runner resumes by key, so no recorded row is re-asked; c004 was in flight at
the kill, and the server answered it once unrecorded (a BrokenPipe in `server.log`) and again on
resume.

**What contention changes and what it does not.** The answers do not depend on timing. The same 8
threads run the same static partition of the same arithmetic, whoever else is on the core. So the
determinism test (rep 1 against rep 2), S and H48 all stay valid under contention. Latency and energy
do not.

**The gate and the instrument (`run_deem.py` `gate()`, `foreign_cpu_s()`).**

- **Gate before every row, and before the idle read.** The runner waits until a 1.0 s window shows
  less than **0.15 CPU-seconds** of foreign work on CPUs 0-7. It polls every 2 s. Foreign work is
  every process except the server and the runner, read from `/proc/<pid>/stat` utime + stime. A
  process counts when its last CPU was 0-7 at either end of the window; a process's main thread
  stands for it, which is an approximation. The 0.15 s line comes from a reading at 12:04Z: the
  laptop's own background ran 0.00-0.10 CPU-s per second on those cores.
- **The gate cap.** A pass that has waited more than **1,800 s in total** stops. It is resumable,
  and it never runs dirty.
- **Per row:** `gate_wait_s`, and `foreign_cpu_s` during the request. The row is `contended` when
  foreign CPU during it is at least **25 %** of its wall time.

**How it is reported (replaces v0 section 5's latency and energy bullets for every pass from here).**

- **Latency** is the median and p95 over the rows that were measured and uncontended, printed with
  their count, the contended count, and the unmeasured count.
- **Rep 1** predates the instrument. Its latency is printed over all 108 rows, labelled "not
  instrumented", with the receipt's load snapshots as the only evidence. Its rows run 0.27 to
  0.65 s, with no outlier of the rep-2 kind.
- **Rep 2's first 4 rows** are unmeasured and excluded from its latency. Their answers count.
- **Energy** is printed only for a clean pass: one segment, zero gate wait, zero contended rows.
  Otherwise it prints "not a clean reading" with the reason. The package counter would hold the
  peers' work and the waiting.
- **Pass receipts become a list of segments,** one appended per start or resume, written at the
  start and again at the end. Rep 2's first segment was killed before its end-of-pass receipt was
  written, so it is reconstructed from the chain log and marked so.

**Unchanged:** the renderings, the sets, the scoring, the validity rules, the comparators, the
determinism test and the run order. The server keeps running (pid unchanged, same environment, same
cores, weights resident).

**Instrument after v0.2:** `run_deem.py` `656c9543cc039b440766ca89b74e21c7909c573f4e0f567dc80ee29be7ac6cec`, `tables_deem.py` `17b0fc5ad43b3b55bdcc66494964aca0705a9a1faf087545c5d9ceece534e6e3` (sha256).

---

## v0.3: one post hoc instrument diagnostic, registered before its calls (2026-09-27T12:06Z)

*Written after S0 P rep 1's rows were seen. It is labelled post hoc, and it is never rows and never a
score.*

S0 P rep 1 scored 5/108, below chance. It never chose socrates, the guest at letter F in every S0
item, and gave F about 0.001 on every row. Two explanations are possible:

- **the model:** a last-letter aversion, or what it takes Socrates to sound like;
- **the instrument:** letter F read from the wrong vocabulary row.

**The diagnostic (`diag_letters.py`, sha256 `20756b2307924812f845f5becfd654639f3e41ca1c579bbad04ad14c73366752`).** Eight factual questions with obvious
answers, one request each, under P's state, run after the four registered passes:

- one question per guest in kit order, so the answers sit at letters A to F;
- then the Socrates (hemlock) and Darwin (Origin of Species) questions with the order reversed, which
  puts Socrates at A and Darwin at F.

**The reading.**

- **The F readout works** if the hemlock question answers socrates at F, or the reversed Darwin
  question answers darwin at F, with p above 0.5.
- **A readout fault is suspected** if F stays near zero on questions whose answer is F, while the
  same content is answered at another letter. That would void nothing already measured (the
  server's own readout is what "as shipped" means), but it would go in the report as an instrument
  finding for the recon's Rust or 9B legs.

The result goes to `receipts/diag-letters.json` and a labelled note in `results.md`.

---

*Results go below this line, and in `results.md` and `rows/`, only after the first row.*

## Results notes (written after the rows; post hoc, dated)

**R-1 (2026-09-27T12:15Z): V-3's code was stricter than its registered words; the words stand.**

- **The registered wording** (section 4) is "the server's choice equals the argmax of its own
  probabilities on every row".
- **What the runner's `argmax_agrees` field recorded** is a stricter test: a UNIQUE argmax equal to
  the choice. Eleven rows over the four passes have an **exact tie** at the top: 3 in S0 P rep 1,
  3 in S0 P rep 2, 3 in S0 S, and 2 in H48.
- **Why ties happen.** The CPU path reads bf16 logits, whose 8-bit mantissa lets two or three letters
  land on one value.
- **What the server does on a tie.** On every one of the eleven rows, it chose the first tied option
  in presented order. That is Deem's own documented rule (`src/deem/format.py` `read_answers`, lines
  418-419: "ties broken by first option in original order").
- **The tables now test the registered words.** `tables_deem.py` `choice_is_argmax()` passes the
  choice when it is the argmax, first in presented order on a tie. The rows' strict field is kept
  and printed beside it.
- **Tie-break sensitivity** (results.md section 1):
  - S0 P (both reps): no tie includes the label, so the headline is 5/108 whatever the tie-break;
  - S0 S: 6 as served, 6 to 7 over tie-breaks;
  - H48 P: 4 as served, 4 to 5 over tie-breaks.

**R-1, ratified (2026-09-27T12:36Z, the lead's ledger row D-20260927-018 on master at `69781572`).**

- **The ruling.** R-1 is ratified, labelled as after the fact. Under the registered `tables_deem.py`
  (`2cea785d`, then `17b0fc5a`), V-3 is a unique-argmax test, and all four passes print VOID.
- **The label.** Every D0 figure is therefore printed as **"VOID by the instrument as registered;
  valid only under R-1 (D-20260927-018)"**, never as a clean pass. This applies to `results.md`
  sections 1, 4 and 11, the page's status line, and the PR body.
- **The instrument pins.** R-1's `tables_deem.py` at `5ed1e246` was `d43bc1b4…`. After the audit
  fold (R-3) it is `751630d95c178ca2b67724ee46e8da7e6b921fe0fac46adc285169c8dcd6d8ff`.

**R-2 (2026-09-27, the audit's M-4): the wall grew after the kit.** Section 7's "the wall has not
grown since the kit, so no post-release slice exists yet" was false. It was carried from the recon's
line 298.

- **The growth.** The COI's wall reads of 2026-09-22 (`KIT.lock.receipts/wall-*.json`, fetched 23:07
  to 23:59Z) hold 18 entries. The 2026-09-27 read that H48 was drawn from (`kit/wall-5f4a4e1c.json`)
  holds 25.
- **The seven newer entries**, identified by `order_seed`, are seats 0, 1, 4, 8, 16, 17 and 21.
  28 of H48's 48 lines come from them, including all 8 ibn_sina lines.
- **The split.** H48 P rep 1 scores 1/28 on the newer lines and 3/20 on the older. That is
  meaningless at these n, but D2 needs it (`results.md` section 3).
- **What cannot be settled here.** The wall carries no dates, so whether those dinners postdate
  Deem's weights (revision 2026-09-25T11:42:30Z) needs the long table's own records. A dated
  post-release subset may exist inside H48; date the 7 entries from the long table's store before D2.
- **No number changes.** The recon's author should be told that line 298 is the source.

**R-3 (2026-09-27, the audit fold; `AUDIT-DEEM-D0-PR171.md`, LAND AFTER FIXES).** No model call and
no re-run. What changed, by finding:

- **M-2 (letter F).** `results.md` section 10 now prints v0.3's registered outcome verbatim: by its
  own rule, **a readout fault is suspected** (hemlock 0.002 at F against 0.940 at A; reversed Darwin
  0.107 at F).
  - Post hoc, beyond v0.3: `check_letter_ids.py` (`6974725e2146e36e2b606677a655ab504fce3acc6b7d6783c011638d1f5d4ff8`, tokenizer only) writes
    `receipts/letter-ids.json`. The server's letter ids are 32 to 37: single, distinct, and matching
    the answer slot's own context. So the wrong-row fault is ruled out, and F suppression is the
    model's as shipped, with position and guest confounded at k = 0. It is noted for the Rust and 9B
    legs, as v0.3 requires.
  - **Withdrawn:** the claim, in commit `5ed1e246`'s bullet 4 and note line and in the PR body, that
    "letters E and F are suppressed" and that the model "reads the first four letters". E was chosen
    24 of 108 times on S0 P with mean p 0.212, and it has the highest mean p on H48 (0.279). Only F is
    supported.
- **M-3 (energy).** Section 6 labels S0 P rep 1 as not instrumented for contention. Its idle read was
  23.6 W, against 11.6-12.7 W on the three gated passes. The page prints the audit's sensitivity at a
  12.2 W idle (56.5 J per decision) and the gated passes' own 44-53 J beside it. No pass is a clean
  energy reading under v0.2.
- **M-5 (chance, latency).**
  - The exact one-sided binomial p values are printed. S0 P's 5/108 is below chance (p = 1.3e-4).
  - **H48's 4/48 is not** below chance (p = 0.080; Wilson 3.3-19.6 % contains 16.7 %).
  - The latency headline is the gated 0.421 s (S0 P rep 2); rep 1's 0.441 s is labelled not
    instrumented.
- **S-1 to S-7.**
  - S-1: the p values, against chance and against the Jev bench's 24.0 % informed floor, which the
    page cites as the prior finding of the named-guest effect.
  - S-2: the gemma4 readout's "103 of 108 floored", spelled out in plain words, with its Brier marked
    "not comparable (floored)".
  - S-3: the server's memory growth by receipt, 890,840 kB to 16,117,520 kB VmRSS over 388 answered
    requests. The audit's "0.89 to 16.1 GB over about 420" reads /proc's kB as 1,000 B, and its count
    is higher than the receipts give.
  - S-4: the OS line.
  - S-5: section 9's comparator join is by text. Kev joins on `text_sha256`; OpenJev and gemma join on
    the Jev prompt sha rebuilt from S0's text with the Jev `run.py` `readout_prompt`.
  - S-6: the contention counts are lower bounds, and the blind spot is named.
  - S-7: the uniform named-guest baseline (30.9 % on S0, 27.9 % on H48) is shown beside Deem's
    85.4 % and Kev-4B's 30.1 %.
- **Errata, in place and listed here.** L-10: the header's "Nothing above the results line changes
  after the first row" now reads "Only dated amendments are appended; no registered text changes".
  - L-4: section 1's "Section 8 says what it may and may not say" now says section 9.
  - L-4: `run_deem.py`'s docstring "section 9 names each" now says section 10. This is a docstring
    only; its sha goes from `656c9543…` (the instrument that wrote the rows) to `78fe9562a7de88ed9a30fa03bb3b2c769b99d496301ae6d8972a2e6edc09b2a6`.
  - L-1: the stamps are now at or before their commits. v0.2 went from "12:06Z" to "12:05Z"
    (committed 12:05:17Z), v0.3 from "12:07Z" to "12:06Z" (12:06:31Z), and R-1 from "12:16Z" to
    "12:15Z" (12:15:16Z).
  - L-2: v0.2's rep 1 range, "0.33 to 0.65 s", is now 0.27 to 0.65 s (c012 is 0.273 s).
  - L-3: v0.2's "(12:00:17Z onward)" now says the segment started at 12:00:17Z and the 4 rows are
    stamped 12:01:38Z to 12:03:17Z. "Nothing is re-asked" now says that c004 was in flight at the
    kill, and the server answered it once unrecorded (a BrokenPipe in `server.log`) and again on
    resume.
  - L-5: `receipts/gpu-check.txt`'s "during/after the passes" line is relabelled "after the last
    call". The proof that the GPU was unused is the CPU wheel.
  - L-6: section 5 says that the receipts' `runtime.threads` = 1 is the runner's probe, not the server
    (17 threads).
  - L-8: the pin above.
  - L-9: section 5's percentile footnote is reworded.
  - L-7: two earlier commits lack a "note" line; this is moot on a squash merge.
- **Info.**
  - **Never publish `kit/h48.json` or the wall copy while H48 is a live control** (I-2). This is
    stated on `results.md` and in `kit/README.md`.
  - I-1: a second Deem server started at 12:29:19Z (`server-d1.pid`) is not this arm's. This lane did
    not touch it.
  - I-3: rep 2's four contended rows are bit-identical to rep 1 (section 7).

