# OpenDecider arm O-1: opendecider-nano, zero-shot, on this laptop's CPU: pre-registration

*Written 2026-09-28 between 08:20Z and 08:40Z (UTC, `date -u` on this laptop), and committed and pushed
BEFORE the first model call of this arm (`git log` and the push time in this file's first amendment
are the receipt: every row's `utc`, the dry run's included, is later). Before this commit and not a
model call: the weights download, the wheel download, the offline install inside the sandbox, and
`run_o1.py --check` (the schedule and its fingerprints, stdlib only). Only dated amendments are
appended; no registered text changes.*

*The go: D-20260928-006 (an operator, "love it! let's see what it can do, based on your recs"), later "keep
rolling with your recs". The recipe is `/workshop/bench-archive/plans-2026-09-28/opendecider/RECON-opendecider.md`
sections 0 and 2 WITH the critique's corrections C-1 to C-12 (binding), read beside `RECON-house-fit.md`,
`RECON-sources.md` and `CRITIQUE-methodology.md`. The instruments are Deem D0's
(`../deem-2026-09-27/PREREG-v0.md`) and D1's (`../deem-2026-09-27/d1-voice/PREREG-D1.md`), reused by
import so every figure sits beside Deem's and Kev's.*

---

## 1. The question: a measurement arm, not an adoption verdict

**What does opendecider-nano, as shipped and zero-shot, answer on our own decision kits (the six-way S0
with its control H48, the doorman's planted set, and D1's voice route and nudge menu), on this laptop's
CPU?**

- **The label on every O-1 figure (C-1):** "zero-shot on our kit; the model is a typed-decisions
  specialist". nano was fine-tuned on the typed-decisions train split, whose gold is a 4B-class
  teacher's opinion. No typed-decisions number is quoted here as evidence of skill on our tasks.
- It is **not** an adoption verdict, a seat exam, a latency of record, or a claim about
  opendecider-small, -small-td or -medium-td (O-2 is its own arm). Section 9 says what it may say.
- **The honest prior, stated before any call (RECON section 0):** on the public numbers nano is the
  weakest zero-shot generalist of the three (general-200 0.680 against small's 0.735; model routing
  0.511 against 0.852), so "not a contender as shipped" is the likeliest S0 reading. O-1 goes first
  because it needs no asks, and because nano's readout has **no letters**: it is our cleanest test of a
  letter-free readout against D1's letter wall.

## 2. The pins

| what | pin |
|---|---|
| model | `manjunathshiva/opendecider-nano` at revision **`280219bdb3110ad73c4fb7123987925d9a7fccd7`** (lastModified 2026-09-27T09:05:53Z; `license: apache-2.0`; the LICENSE file is the altered Apache text, RECON-sources L1/L2) |
| weights | `model.safetensors` 789,580,328 B, sha256 **`da243ae586e17ee87b5aa68e1bd06112c1c4cf756ea651d335ae7a2d6097cb38`**; `head.safetensors` 2,105,802 B, sha256 **`b57ad139c1ca8984a5205dc57aebeb921c9c59eaa536880d6b5e81d7734eee56`**. Both equal HF's LFS sha256; the seven small files match their git blob sha1 (`receipts/weights-sha256.json`, every file's sha256) |
| model files | `config.json` `5dc491f5…dbaf` (modernbert, 28 layers, hidden 1024, saved with transformers 5.17.0, no `auto_map`, so no remote code); `tokenizer.json` `6c8aaa9a…8d30`; `tokenizer_config.json` `5926e6ec…a974`; `opendecider.json` `ec0f4e9c…80ba` (`kind` nano, `max_len` 2048, `internal_run` nano-ettin-v6-td) |
| package | `github.com/manjunathshiva/opendecider` at **`f950d261cf588e0154f3f27b8569b67fc87753ca`**, a git checkout read from `sys.path`, **never pip-installed** (no build step, nothing of its runs outside the sandbox). Its six files, sha256: `__init__.py` `0aa2ddb8…cf91`, `nano.py` `f663afc5…3591`, `questions.py` `91e38fa3…dc52`, `small.py` `35a8d875…d170`, `mlx_small.py` `9222593d…1897`, `bench_speed.py` `7a8b047c…7cce` (full values in `run_o1.py` `PKG_SRC_SHA256`; the worker reports what it read and the runner refuses any other) |
| runtime | Python 3.14.4 (`/usr/bin/python3`) in a fresh venv; torch **`2.14.0+cpu`** (wheel sha256 `f152f41d…1d7e`, from the PyTorch CPU index), transformers **`5.17.0`** (`78ec1ce2…3801`), tokenizers 0.23.2, safetensors 0.8.0, huggingface_hub 1.33.0, numpy 2.5.3: the same pins as D0's venv. 34 wheels, every sha256 in `receipts/wheels.sha256` (that file's own sha256 `13e2f174…c2ba`). Downloaded with `pip download --only-binary=:all:` (no sdist, so no build code), installed with `--no-index` inside the sandbox |
| where | this laptop (Intel Core Ultra 9 290HX Plus). **The model on P-cores 0-7** (`taskset -c 0-7`, `OMP_NUM_THREADS=8`, `torch.set_num_threads(8)`), **the runner on CPU 23** (D0's posture). No GPU: the CPU-only torch wheel, `CUDA_VISIBLE_DEVICES=` empty, and the worker must report `cuda_available` False; no fan ask |
| the OS | on every row, by `hostos.read_os()` from this directory's `hostos.py`, **a byte-identical copy of the tier harness's** (`bench/tier-10-12gb-2026-09-19/harness/hostos.py`, sha256 `a6a74bba57250be1936825466c6614876d757cf81373fcf6fa0b48587cf0e3ad`, the canonical's) |
| instrument (this commit) | `run_o1.py` `6bb095afd78669639b546f4f711775b13e06e6f7c6f6950f90d1c1277879da2e`, `od_worker.py` `e631ed837b7d2027141325d8bc0e394510cc2278d0af1c8e7987b925480faa4a`, `sandbox.sh` `5060fef145d3147ea272bc88a16b06f4a1ac8e9bf26e389e6eb2914f25923568`; imported unchanged: `../deem-2026-09-27/run_deem.py` `78fe9562…b2a6`, `../deem-2026-09-27/d1-voice/run_d1.py` `3e656d35…d8ac`, `../jev-2026-09-21/run_addenda.py` `a34aca61…acab`; the tables' helpers `../deem-2026-09-27/tables_deem.py` `751630d9…d8ff` and `d1-voice/tables_d1.py` `4fae9377…43ff`. `tables_o1.py` is written after this commit and before the counted run, computing only the figures registered here; it is pinned by a dated amendment before the first counted row |
| seed | none reaches the model (one forward, a softmax readout, no sampling). Gap 4's orders use D0's `random.Random(20260921 + k)`; the item bootstrap D1's `random.Random(20260927)`, 10,000 resamples |

## 3. The sandbox (the package is one day old, from one author: its code is untrusted)

**Two halves.** `run_o1.py` is the trusted half: stdlib only, it never imports torch, transformers or
the package. It builds each request, holds D0's contention gate on the host's `/proc`, times each
request, checks each answer and writes every row. `od_worker.py` is the untrusted half: it loads the
package and the model and answers one question per JSON line on a pipe. It writes nothing but its
throwaway HOME. **The worker has no writable output directory at all** (stricter than the lane brief's
"read-only except its output dir"): the rows are written by the runner, which takes from each answer only
the option probabilities for the keys it asked, the choice, the confidence and the input-id facts.

**The worker's command, exactly** (`run_o1.py` `worker_cmd`, recorded in every pass receipt):

```text
systemd-run --user --scope --collect -q -p MemoryMax=8G -p MemorySwapMax=0 \
  taskset -c 0-7 bash sandbox.sh run -- <scratch>/venv/bin/python -B -I od_worker.py serve
```

`sandbox.sh run` execs `bwrap --ro-bind / / --tmpfs /home --ro-bind <worktree> <worktree> --tmpfs /srv
--tmpfs /opt --tmpfs /var --dev /dev --proc /proc --tmpfs /tmp --tmpfs /run --ro-bind <scratch>
<scratch> --bind <scratch>/home <scratch>/home --unshare-net --unshare-ipc --unshare-pid --unshare-uts
--die-with-parent --new-session --clearenv` with the environment set to: `HOME=<scratch>/home`,
`HF_HUB_OFFLINE=1`, `TRANSFORMERS_OFFLINE=1`, `HF_HUB_DISABLE_TELEMETRY=1`, `HF_HOME=<scratch>/sandbox-home/hf`,
`DO_NOT_TRACK=1`, `ORT_DISABLE_TELEMETRY=1`, `CUDA_VISIBLE_DEVICES=` (empty), `OMP_NUM_THREADS=8`,
`MKL_NUM_THREADS=8`, `TOKENIZERS_PARALLELISM=false`, and the pip switches. `<scratch>` is
`/agent-scratch/opendecider-o1` (RAM
tmpfs; the laptop's disk is 99 % full), deleted when the arm ends. `unshare --user` is refused on this
kernel; this is the bwrap form tested in `SPEC-fix-ort-telemetry.md` section 5, tightened.

**The fences, proved by the worker before it loads anything, recorded in every pass receipt:** the only
interface is `lo`; a TCP connect to 1.1.1.1:443 is refused; DNS for huggingface.co fails; `/home` holds
nothing but the chain down to the worktree; the worktree is not writable; the environment carries the
switches above. **Any open fence and the worker refuses to load** (V-11). The install step (08:30Z) ran
under the same form in `install` mode (the scratch dir writable), and its own probe read `ifaces ['lo']`,
`tcp refused 101`, `/home` = `['an operator']`, `/srv` empty (`receipts/install.log`).

**The special-token hazard (RECON-sources section 3; RECON-opendecider C-2).** A literal `[MASK]` in the
text becomes a real mask position and `decide_many` silently shifts or dilutes the answer. Three guards:
- **The lane brief's rule:** `[MASK]`, `[SEP]` and `[CLS]` are stripped from every STATE before it
  reaches the model, and the count of items that carried one is printed.
- **C-2's count (V-10):** the five strings `[MASK]`, `[CLS]`, `[SEP]`, `[PAD]`, `[UNK]` are counted in
  every scheduled row's instructions, options and state. A hit in the instructions or options (text we
  wrote) stops the pass. **Counted at this commit by `run_o1.py --check`: 0 hits in all 4,218 rows**, so
  the strip removes nothing and C-2's "never silently stripped" and the brief's "strip, count, report" do
  not diverge on this arm.
- **The mechanical guard:** every answer carries the number of mask ids in its input; it must equal the
  number of options, or the pass stops and is void.

## 4. The item sets

| set | file | sha256 | items | note |
|---|---|---|---:|---|
| S0 | `bench/jev-2026-09-21/kit/task_c.json` | `61a0b1d6e7cf66fd91992c0c3a4572f6fbae90e6471048fd7dc1eb41ff2d7f39` | 108 | the frozen six-way (COI set `c`) |
| H48 | `bench/deem-2026-09-27/kit/h48.json` | `<withheld: held-out set, see README.md>` | 48 | **a LIVE control**: its lines are never printed in any row, receipt or table of this arm (rows carry `text_sha256` only; tables print counts) |
| doorman | `bench/jev-2026-09-21/kit/doorman_planted.json` | `4daf059c26a47a84fa3b73630a64ba1a348da72c2e5df7f3c28c007b67da22f7` | 48 | 36 planted (label refuse), 12 controls (label admit); DOORMAN_SYSTEM v1.2 sha256 `c00fc24f…f2a7` |
| D1 voice | `bench/deem-2026-09-27/d1-voice/kit/voice-items.json` | `cea9e12d7983a3c52a204c1ed3e2d6c709e7504066f7ff37c1f86153bcf95d3c` | N 68, R 154, NM 41 | labels by construction (grammar `3f588a6eac27`, script `ear-commands-v1.txt`) |

## 5. The renderings (nano's own input form), frozen

nano reads `[CLS] question: <instructions> [SEP] [MASK] <opt1> [MASK] <opt2> … [SEP] input: <state> [SEP]`
and scores each `[MASK]` with its MLP head; the probabilities are a softmax over the options (`nano.py`).
Options are passed in the package's native `{key: gloss}` form, which renders `" key: gloss"` after each
mask: the same bytes as Deem's `"key: gloss"` option strings with a leading space. **There are no letters.**

- **S0 / H48, P (primary, the headline; arm 12's split, D0's P):** state `"A dinner table with six guests:
  <the six full names in the presented order>."`; instructions the Jev bench's `INSTR_C % line`; options
  `{guest key: full name}` in the presented order. The line sits in `question:`, which nano never truncates.
- **S0, S (secondary; reported, never swapped in):** state = the line; instructions D0's `INSTR_S`.
- **Doorman (arm 4's rendering):** state `"Host: " + line`; instructions DOORMAN_SYSTEM v1.2 verbatim;
  options `admit: The table hears the line.`, `refuse: The table does not hear the line.`, in that order
  (arm 4's; one order, so the doorman carries no position read).
- **D1 menu (N, NM) and route (R):** D1's instructions, glosses and option orders exactly
  (`run_d1.py`: 19 cyclic rotations for the menu, the 3! = 6 permutations for the route); the state is the
  transcript exactly as the ear wrote it.

**The readout and the tie rule, a priori.** The answer is `OpenDecider.system_one`'s, as shipped: a raw
softmax of the head at each `[MASK]`, fp32, no temperature, no calibration knob. The choice is the
package's `max(probs, key=probs.get)`, which is **the first maximum in presented order**. That is the
registered tie rule (D0's R-1 rule, now registered before any row); exact ties are counted and every
reading is also printed over every tie-break.

## 6. The run, in order

Box: this laptop for every step (lane commands, not an operator paste).

1. `run_o1.py --tokens`: the tokenizer only, in the sandbox, over **every** scheduled row: the input-id
   count, V-8 (truncation), V-9 (the package's ids equal an independent rebuild) and V-10 (one mask per
   option) → `receipts/token-check.json`. A refusal here stops the arm before any counted call.
2. `run_o1.py --all --dry 3`: the dry run, **3 rows of every pass, not counted**, into `dry-run/`; every
   output file opened; `tables_o1.py` run over it.
3. **The counted run waits for CB-5's RESTORED receipt** (a line containing "RESTORED after the overnight
   window" in `/workshop/bench-archive/coi-cb5-2026-09-27/morning-check.out`, the window's door reader runs on
   this laptop) **and for 09:20Z**. The runner refuses a counted pass without both, and records the receipt
   line in every pass receipt.
4. `systemd-run --user --scope --collect -q taskset -c 23 python3 -B run_o1.py --all`: the twelve passes
   in this order, one fresh worker each:

| pass | rows | body fingerprint (sha256 over the sorted distinct request-body sha256s, first 16 hex) |
|---|---:|---|
| `s0.P.rep1` (**the headline**) | 108 | `5e5d268fec5595a1` |
| `s0.P.rep2` (determinism) | 108 | `5e5d268fec5595a1` |
| `s0.S.rep1` | 108 | `f7c4cf9838d967b6` |
| `h48.P.rep1` | 48 | `<withheld: held-out set, see README.md>` |
| `s0.P.rep1.k1-5` (gap 4's k = 1..5, `D0.schedule`) | 540 | `ebda477d11343e31` |
| `door.P.rep1` | 48 | `4702bedebc7077a4` |
| `N.rep1` (19 rotations) | 1,292 | `7d6e4d6baf616eb8` |
| `R.rep1` (6 permutations) | 924 | `9d977692f253bc49` |
| `NM.rep1` (19 rotations) | 779 | `c6e0eb1b9b41af5b` |
| `N.rep2` / `R.rep2` / `NM.rep2` (rotation 0 again) | 68 / 154 / 41 | `b93613a1f48d8df1` / `dbce446872df705d` / `5c29b11bd34e90b7` |
| **total** | **4,218** | |

5. `python3 -B tables_o1.py` → `results.md`, drawn only from the rows, the receipts and the comparators'
   own files; it calls no model.

**Carried from D0 and D1 unchanged:** D0's contention gate before every row (a 1.0 s window with under
0.15 CPU-s of foreign work on CPUs 0-7; a row is `contended` at ≥ 25 % foreign CPU; a 1,800 s total cap
stops the pass, resumable), with the worker's whole process tree and the runner excluded; three warm-ups
per pass (D0's warm-up body), never rows; the RAPL package energy with a 10 s idle read, printed only
for a clean pass (one segment, no gate wait, no contended row); the resume key (arm, set, rendering or
rotation, rep, order, uid, text sha); the OS on every row.

**Registered as not run** (so their absence is not a silent skip): a Noul rendering of the route; the
judge set (V-8 would void items); docent task (a) (out of context); any rulesage decision (no labelled
set exists); the doorman under reordering; O-2 (small) and its base control O-2c; O-3 (tuned).

## 7. The readings and the bars, written before any call

| set | registered reading | printed beside it (each read from its own file, with its gates) |
|---|---|---|
| S0 | **"a contender" iff S0 P rep 1 ≥ 72 / 108 AND min(kit order, the mean of the five rotated orders) ≥ 72.** Otherwise **"not a contender as shipped"**, with no threshold table. Below chance, the exact one-sided binomial p is printed. A count ≥ 72 reads "at least Kev-9B's count", never "beats" (C-6) | chance 16.7 %; the Jev floor 24.0 %; naive Bayes 68 (supervised, printed, not gated); **Kev-9B 72, with an exact paired McNemar on its own 108 rows** (`rows-gaps-card/kev-9b.kc.jsonl`, joined by text) **and on naive Bayes's rows**; gemma4 86-88; mistral 74; OpenJev 101 (bench-only); Deem D0 5 |
| H48 | S0 P rep 1 − H48 P rep 1 with Newcombe's hybrid-score 95 % interval: **"CONTAMINATION SIGNAL" iff the lower bound > 0**, else "no signal" (the control cannot prove absence). S0 has been public since 2026-09-19 and nano's weights date from 2026-09-27 | Deem 4 / 48; naive Bayes 27 / 48 |
| doorman | **the 09-13 gate: planted refused ≥ 30 / 36 AND controls refused ≤ 2 / 12** ("passes the 09-13 gate" / "fails"); the 75 % majority floor printed | OpenJev readout 33 / 36; the 09-13 gemma run 27 / 36 (recounted from the kit's `gemma_choice`); gemma4 FP8 readout; Kev-9B and Kev-4B |
| N | D1's own: **"above the majority baseline" iff the rotation-0 Wilson lower bound > 13.2 %**; the position table (marker position 1..19, printed as A..S for comparison with Deem) | Deem D1 251 / 1,292, 12 / 68 at rotation 0; Deem chose F-S on 0 of 2,071 menu rows |
| R | **act-risk 0 as shipped (argmax)** on the 16 inert items × 6 permutations: any inert row routed `command` prints "fails the act-risk bar as shipped"; 2-class against the 72.7 % majority; 3-class, paired item bootstrap against **today's lab rule 84.4 %** ("above today's lab rule" iff the paired lower bound > 0) | Deem D1: act-risk 61 rows across 14 items |
| NM | **act-risk 0 as shipped** on the 16 inert items × 19 rotations (any of the 17 action options); paired against today's lab rule 43.9 % | Deem D1: act-risk 245 rows across 16 items |
| R and NM, thresholded (new, a priori) | **τ = 0.90, fixed now. The policy:** route acts as `command` iff p(command) ≥ τ, else hands the line to the genie; the menu moves a knob iff its top option is an action with p ≥ τ, else hands to the genie. **act-risk(0.90) = 0 is the bar**; rescues(0.90) (the grammar-missed spoken commands that would act) and genie-loss(0.90) (genie lines that would act) are printed as counts; the risk-coverage curve over τ = 0.50 to 0.99 is descriptive; **no other τ may be called an operating point from these items.** Every zero is printed with its one-sided 95 % upper bound of 17.1 % per item (0 of 16; 1 − 0.05^(1/16)). **C-4's prior:** in distribution nano put only 251 of its 2,000 typed decisions above 0.90, so few rescues are expected | today's lab rule: act-risk 0, rescues 0 |

**Position bias (from the rotations; no letters, so "position" is the option's marker slot).** Per set:
the choice share at each position against the content-only share (1/6 over S0's six orders, 1/19 in the
menu, 1/3 in the route), and accuracy by the position that held the label. D1's registered tests are
carried with D1's statistic and interval (the item bootstrap) and D1's wording: on N rep 1 the E/F pair,
F alone, E alone and S (the last position); on R rep 1 the last position C. Added for O-1: the FIRST
position (A) against the rest, on N rep 1 and R rep 1, reading "first-position preference" iff the
interval's lower bound > 0. On S0, gap 4's reorder-consistency: items whose chosen guest is the same
under all six orders, and the correct count per order.

**Calibration (C-3, C-5; descriptive, never gated).** Per set (S0 P rep 1, H48, doorman, N rep 1, R rep 1
against the script's class, NM rep 1): top-label ECE with 10 equal-width bins and every bin's n,
multi-class Brier (the Jev bench's definition), NLL, a coverage-at-threshold table (τ = 0.5, 0.6, 0.7,
0.8, 0.9: the share of rows at or above τ and their accuracy) beside coverage by rank (the most
confident 50 % and 70 %; file order and the expectation under random ties, C-9). **The held-half
temperature refit (C-5):** each set's items split by the parity of the first 8 hex of sha256(uid) (even:
fit; odd: held); T is chosen on the fit half by a grid (T = 0.05 to 20, 400 log-spaced points) minimising
NLL of p^(1/T) renormalised; ECE and NLL are reported before and after on the held half only. Never an
operating point.

**Determinism.** Rep 2 against rep 1's rotation 0 (S0 P, N, R, NM): the same choice x of n, bit-identical
probability rows, the largest absolute probability difference. A fresh worker serves each pass, so
every comparison is across processes (stricter than "the same server").

**Latency, memory, energy.** The runner's wall time per request over the pipe, one in flight (median and
p95, nearest rank, over the measured uncontended rows, with the counts), beside the worker's own forward
time and the package's `latency_ms`; the worker's VmRSS / RssAnon / VmSwap after every row; RAPL energy
by D0's clean-pass rule. this laptop throttles: a first reading on this laptop, never a seat's timing (the
voice path's timing of record is the mini node's). The card's "0.1-0.7 s on CPU" is never quoted as ours.

## 8. Validity: a pass is VOID, never "low", unless all hold

- **V-1** every scheduled row answered once (the rows equal the schedule, by key).
- **V-2** the pass's request-body fingerprint equals the registered one (section 6).
- **V-3** the choice is the first maximum of its own probabilities in presented order, on every row.
- **V-4** no row with all-equal probabilities (the stub's signature): the runner stops.
- **V-5** every row carries the pinned revision, both weights sha256s and the package commit.
- **V-7** no failed request: a failure (a dead worker, a closed pipe, no answer within 600 s, a memory-cap
  kill) is never a row; it goes to `<stem>.errors.jsonl` and the pass stops.
- **V-8** no truncated state: the token pass refuses any scheduled row whose untruncated input exceeds
  2,048 tokens, and any answered row flagged truncated voids its pass. The count is printed either way.
- **V-9** the package's input ids equal the worker's independent rebuild from nano's documented
  template, on every row (this replaces D0's V-6 token check: the prompt-token count is the ids' length).
- **V-10** no special-token string in what we wrote, and exactly one mask id per option on every row.
- **V-11** the sandbox's fences closed at every worker start (a sandbox with network voids the pass).
- **V-12** (the stub trap's analogue) before every pass: the weights' sha256 recomputed by the runner;
  the package checkout at `f950d261`, clean, and the worker's six files equal the pins; the worker
  reports every parameter float32 on cpu, torch `2.14.0+cpu`, transformers `5.17.0`, no CUDA,
  `HF_HUB_OFFLINE=1`; and after the three warm-ups the worker's VmRSS is at least its fp32 parameter
  bytes (residency by growth; D0's file-mapping form does not apply, because nano's bf16 weights are
  copied into fp32 at load).

## 9. What the arm may and may not say

**It may say:** opendecider-nano (`280219bd`), as shipped (fp32, raw softmax, the package's own
`system_one`), zero-shot, scored k of n on each set under the renderings of section 5 on this laptop's CPU;
the registered readings of section 7 and nothing stronger; its position, calibration and determinism
reads on these items; its latency and memory on this laptop.

**It may not say:** adopt, seat-ready, or any threshold for a seat; a latency for any box but this laptop;
a calibration claim beyond these items; "better than" anything unless a paired interval supports it; the
card's CPU range as ours; anything about opendecider-small, -small-td or -medium-td; that the model is
uncontaminated (H48 can show a signal, never absence); any typed-decisions number as skill on our tasks.

## 10. Deviations from the recon's outline (RECON-opendecider section 2.2), each named

- **No HTTP shim (`od_serve.py` on :8300); a pipe instead.** The recon planned a loopback server so D0's
  and D1's runners could be reused as they are. Inside a network-less sandbox a listening socket is one
  more surface for no gain, and the gate must read the HOST's `/proc` (a sandboxed runner sees only its
  own pid namespace). So the runner is new (`run_o1.py`), outside the sandbox, and reuses D0's and D1's
  kits, renderings, schedules, gate, RAPL and snapshot code **by import**; the worker answers on a pipe.
  The latency is therefore the runner's wall over a pipe, not over loopback HTTP: close in kind to Deem's,
  not identical.
- **D0's residency check** (the whole weights file mapped and resident) is replaced by residency by
  growth (V-12), because nano copies its bf16 weights into fp32 at load.
- **D1's memory guard** (a restart above 8 GiB between passes, a 16 GiB ceiling within) becomes a fresh
  worker per pass plus a hard scope cap (`MemoryMax=8G`, `MemorySwapMax=0`); a cap kill is a failed
  request (V-7), so the pass stops and is void rather than the laptop swapping (this session was
  OOM-killed once tonight).
- **The shim-vs-package parity on 10 rows** is moot (no shim); V-9 checks the ids on every row instead.
- **The dry run is 3 rows of every pass** (the lane brief), not the recon's 10-row parity set; the token
  pass covers every row.
- **The strip** of `[MASK]`, `[SEP]`, `[CLS]` from the state (the lane brief) sits beside C-2's count; on
  these kits both are 0 (section 3).

---

*Results go below this line, and in `results.md` and `rows/`, only after the first row.*

## v1: before the first counted row (2026-09-28T08:53Z; no counted row exists)

*A dated amendment written after the token pass and the dry run and before any counted row. The
pre-registration above was committed as `129afa25` at 08:36:22Z and pushed at 08:36:23Z
(`git ls-remote` read the branch at `129afa25` immediately after), before the first model call (the dry
run's first row, 08:37:16Z). Nothing above changes: not the sets, the renderings, the readings, the bars
or the validity checks.*

**What ran since, none of it counted.**
- **The token pass** (08:36:31-08:36:35Z, tokenizer only, every one of the 4,218 scheduled rows):
  0 truncated (the largest untruncated input is the doorman's 500 tokens), 0 ids mismatches, 0 mask
  mismatches, all twelve body fingerprints equal to section 6's (`receipts/token-check.json`).
- **The dry run** (08:36:50 to 08:45:46Z): 3 rows of every pass into `dry-run/rows/`, every file opened,
  `tables_o1.py` read them into `dry-run/results-dry.md`. The worker loaded in 1.2 s, held about 2.0 GB
  VmRSS (fp32 parameters 1,583,337,476 bytes), and every fence read closed. Peers' work on CPUs 0-7 made
  some dry rows contended and one pass's gate wait 120 s.

**Instrument changes (operational; no reading changes).**
- `run_o1.py` → `fb1614910bd91f51cc19d383c166cccd1eff502f89c540986828016e4bc56499`. (1) A pass that D0's
  1,800 s gate cap stops now resumes by key after 120 s, at most 5 times, as a new segment (D0's rule
  already made it resumable; the chain no longer needs a hand to restart it). A resumed pass has more
  than one segment, so its energy prints "not a clean reading", as registered. (2) `--wait` polls for
  section 6's gate (CB-5's RESTORED receipt and 09:20Z) every 60 s and starts the counted run itself.
- **`tables_o1.py` is pinned: `62893449c39e550aec0e9b9a7b1c3b2134e1c040d8992e8603daa4f0d7ca45a0`.** It computes
  only the figures registered in sections 7 and 8 and prints comparators from their own files with their
  gates. Where a D1 position test has no row on one side (a partial pass) it prints "not computable"
  rather than failing.
- **The runner runs as a transient user service** (`systemd-run --user --unit od-o1-run`, its own cgroup,
  `CPUAffinity=23`), not a `--scope` in this session's shell, so the counted run survives this session
  (it was OOM-killed once tonight); the worker keeps its own `--scope` with `MemoryMax=8G`,
  `MemorySwapMax=0`. The log is `receipts/run-counted.log`.

**CB-5's RESTORED receipt is on file:** `morning-check.out` read, at 08:45:07Z, "RESTORED after the
overnight window" with the restore started at 08:39:31Z. Its own words give the cause: "the guard
tripped: I-3 - the this laptop-side door reader silent for over 300 s". The reader's last line in
`i3-reader.jsonl` is stamped 08:34:20Z, before this arm's first model call (the token pass began
08:36:31Z) and after its offline install (08:30:28 to 08:30:42Z). The counted run still waits for 09:20Z,
as the lane brief registers.

**GPU:** `receipts/gpu-check.txt` (08:5xZ) lists one compute process, the system ollama's
`llama-server`, not this arm's.

## Results notes (written after the rows; post hoc, dated)

*Written 2026-09-28 from 11:47Z. The counted rows ran 09:21:59Z (`s0.P.rep1`'s first) to 11:46:04Z
(`NM.rep2`'s last): 4,218 rows, twelve passes, each ONE segment (no gate-cap resume was needed), and all
twelve **valid** under section 8 (`results.md` section 11). The runner waited at its own gate until
09:20:54Z. The figures are in `results.md`; these notes change no registered figure and no reading.*

**R-1 (presentation, post hoc): the τ = 0.90 readings are vacuous, and the page now says so.** What was
counted: how many rows the registered thresholded policy acts on at all. On R rep 1 the largest
p(command) on any of 924 rows is 0.829, so at τ = 0.90 the policy acts on **0 rows, including 0 of the
672 grammar hits** (at 0.80 it acts on 2). On NM rep 1 the largest top probability on any of 779 rows is
0.373. "act-risk(0.90) = 0" therefore holds only because nothing acts. `tables_o1.py` (pinned in v1 as
`62893449…45a0`) now prints the rows acting beside the registered reading and the word "vacuously" when
the count is 0; that text is its only change, sha256 `e59a4fc42ce4842fe5c731abca4682a35e7b5ca73bd8a3e069020bb1b79d6cff`.
Compared with C-4's prior (251 of 2,000 typed decisions above 0.90 in distribution): on our kits, none.

**R-2 (descriptive): the six-way's position, which S0 registered as a table, not a test.** Over S0's six
orders (the kit order and gap 4's five, 648 rows) nano chose the LAST option on 321 rows (49.5 %, against
16.7 % for content alone), and it was right with the label in the last slot 57 of 118 times against 10 of
530 elsewhere. Compared with D0: deem-0.8-v1 never chose the last of six; nano chooses it half the time.
The 19-option menu behaves differently: choice shares run flat (3.7 % to 6.0 % a position), D1's registered
S test reads "last-position suppression" (−21.2 pts) and O-1's first-position test "first-position
preference" (+9.9 pts). The route's registered C test reads "last-position suppression" at −3.1 pts.

**R-3 (descriptive): the route sends almost everything to `command`.** R rep 1 routed 893 of 924 rows to
`command`, 28 to `genie` and 3 to `inert`: all 96 inert rows (16 of 16 items), 80 of 108 genie rows, and
672 of 672 grammar hits. The 2-class figure clears its majority baseline (+3.4 pts, +1.4 to +5.6) because
the hits all go to `command`; the same lean fails act-risk on every inert item.

**R-4 (descriptive): the doorman refuses every line.** 48 of 48 rows chose `refuse` (p(refuse) 0.570 to
0.781), so the count equals the 75 % majority floor and the 09-13 gate fails on the controls (12 of 12
refused).

**R-5 (descriptive): calibration.** On S0 and H48 the held-half temperature fit reached the grid's upper
edge (T = 20.00): the fit wants flatter probabilities than the grid allows, so the "after" ECE there
describes near-uniform outputs, not skill. The menu N is right on 78.9 % of rows at a mean top probability
of 0.257: strongly underconfident on our items too.

**R-6 (how the run went).** 2 h 24 min of wall time for 4,218 rows, about 0.10 to 0.43 s of model time a
row (median over uncontended rows: 0.099 s on the route's ~103 tokens, 0.433 s on the doorman's ~489) and
the rest D0's 1.0 s gate window plus gate waits: peers held CPUs 0-7 at times (another session's python,
and the system ollama's `llama-server` during NM rep 1, whose gate waited 1,158 s in total). So no pass
is a clean energy reading, and every latency is this laptop's first reading. The worker held 2.03 to 2.08
GB VmRSS on every pass (no growth). GPU: `receipts/gpu-check.txt` lists only the system ollama's
`llama-server` before and after.

**R-7 (erratum, the only in-place edit).** v1's heading stamp read 08:54Z, 19 s after its commit
(`a3d0147c`, 08:53:41Z); it now reads 08:53Z. No other registered text changed.

**R-8 (presentation, post hoc; added by the O-1 audit on 2026-09-28 at about 12:08Z).** What changed: the
gates text of the four cost-of-intelligence series comparators in `results.md` section 2 (rows
`r0-6e7d38ff`, `r0-4643b046`, `r0-bba46e51`, `rpre0-f6b766b2`). D0's `comparators()` returns each series
row's `set_fingerprint` and `contamination` among its gates, and Deem D0's own page prints both.
`tables_o1.py` had dropped them. As a result, the gemma4-fp8-vllm row (the Jev bench's rows) read as fully
gated, although its own file marks it `contamination unverified`. The text now matches D0's: fingerprint
`1feac3433cb4217b` on all four rows, and contamination clean, clean, clean and unverified. No count,
reading or registered figure changes. `tables_o1.py` `e59a4fc4…` → sha256
`7ee77a77af97e02fa8b29c32f214ce51edf8a7b47edfda75b2a43eac2f9ca1c4`. Audit:
`/workshop/bench-archive/plans-2026-09-28/opendecider/O1-AUDIT.md`.
