# OpenDecider arm O-2: opendecider-small and its base control, zero-shot, on the bench box's RTX 3090 (pre-registration)

*Cold scroll: what this file is.* The pre-registration for arm O-2, finalized by the O-2 lane from the benchbox-day
planner's draft (`plans-2026-09-28/benchbox-day-2026-09-29/PREREG-O2-draft.md`, sha256 `77c497cba2055b63…`, written
2026-09-28T09:0xZ). It is committed and pushed on `bench/opendecider-o2-2026-09-28` **before any O-2 model call**
(the token pass, §6 step 1, is the first). All times are UTC. At commit time: the weights are pulled and hashed, the
venv is built and checked, the sandbox's fences and CUDA visibility are proved with no model loaded; no O-2 row,
token pass or dry run exists.

O-2 has two parts:
- **O-2** reads the 4B "small" (a LoRA adapter on Qwen3-4B-Instruct-2507), zero-shot.
- **O-2c** reads the base alone, as the control.

Both run on O-1's instruments and bars, on the bench box's (benchbox's) RTX 3090, on **2026-09-28** (the card was
seated a day earlier than the draft planned; the ollama-vs-vLLM S arm ran on it first, PR #179).

**The authority.** an operator's daylight bench go on benchbox, verbatim: ~11:05Z "benchbox is back online with 2x 3090",
"ready to bench as well"; ~11:13Z "benchbox, the 3090 has 2 separate 8 pin pci power cables hooked up, and the fans
are cranked, let's roll with all the benches you need to run on there now". The plans behind it:
D-20260928-011 (`DECISIONS.md` l.1801, "(2) OpenDecider O-2, the 4B "small" zero-shot plus its base control (OD-1
rec (a), ~40–60 min, an 8.04 GB pull)"), D-20260928-010 (l.1800), D-20260928-006 (l.1796). The recipe comes from
`plans-2026-09-28/opendecider/RECON-opendecider.md` (sha256 `2d22fddaa8afb918…`: OD-1 at l.82–93; §2 with the
critique's corrections C-1 to C-12, binding, l.378–391) and O-1's registration, `PREREG-O1.md` in this directory
(committed `129afa25`, amendment v1 `a3d0147c`).

**This branch is stacked on O-1's** (`origin/bench/opendecider-o1-2026-09-28` at `a3d0147c`), because O-2 imports
O-1's instruments unchanged. **Nothing in O-2's readings or bars moves on O-1's results.** O-1's counted run was still
going on the laptop when this was written (it had reached NM rep 2 at 11:49Z); its results are a comparator only.

Only dated amendments are appended after the commit. No registered text changes.

---

## 1. The question: a measurement arm, not an adoption verdict

**What do opendecider-small (as shipped, zero-shot) and its untrained base (Qwen3-4B-Instruct-2507 with no adapter)
answer on our own decision kits? And does the adapter change the answers?** The kits are:
- the six-way S0, with its control H48;
- the doorman's planted set;
- D1's voice route and nudge menu.

The reads are on the same rows and request bodies as O-1.

- **The label on every O-2 figure (C-1):** "zero-shot on our kit; the model is a typed-decisions distillation
  specialist".
- **The label on every O-2c figure:** "zero-shot on our kit; Qwen3-4B-Instruct-2507 with no adapter, read through
  OpenDecider-small's own prompt and letter readout: not an OpenDecider model".

  No typed-decisions number is quoted here as evidence of skill on our tasks.
- **O-2c exists because upstream has no base control** (C-7, RC m3). Upstream's calibration ablation (0.289 → 0.087)
  has no committed rows (RECON §1.3, §3.2).
- **The honest priors, stated before any call:**
  - **On the public numbers, small is the stronger zero-shot generalist.** It scores 0.735 on general-200 (nano
    0.680), 0.852 on model routing (nano 0.511) and 0.748 on BANKING77 (RECON §1.9).
  - **Small is weaker on typed decisions.** Zero-shot it scores 0.671, which is −0.083 [−0.105, −0.060] against Jev
    (RECON §1.4–1.5).
  - **Small is weaker on jailbreak:** 0.775 against nano's 0.905.
  - **Small shows a last-choice bias upstream.** Zero-shot it picks the last option 29.7 % of the time against a gold
    rate of 18.2 % (RECON §1.10). **So a last-position preference is the expected position reading.**
  - **The nearest letter-readout 4B on file** is Kev-4B as served, at 60/108 on S0 (RECON §1.8). That number is a
    comparator, not a prediction.
- **Small reads letters** (`small.py` l.49, l.60–63). D1's letter wall is therefore read directly. Deem chose letters
  F to S on 0 of 2,071 menu rows (RECON §1.8).

## 2. The pins

| what | pin |
|---|---|
| adapter | `manjunathshiva/opendecider-small` @ **`ff25e366530782d67769bc219d823fdb23328819`** (apache-2.0, not gated). 6 files, 132,211,237 B. `adapter_model.safetensors` 132,187,888 B, sha256 **`354ed3fad2c4ee390253533232ca630e55c83590d07be719a59a0d8743a01b27`** (git blob `ba8ba750…`). `adapter_config.json` sha256 `79faa634…3953` (LoRA r 16, α 32, dropout 0.05, `peft_version` 0.21.0, targets q/k/v/o/gate/up/down, `base_model_name_or_path` `Qwen/Qwen3-4B-Instruct-2507`). `opendecider.json` sha256 `ce01d789…05e4` (`kind` small, `internal_run` small-v3). Every file's sha256 is in `run_o2.py` `WEIGHT_SHA256` and `receipts/pull-o2.json` |
| base | `Qwen/Qwen3-4B-Instruct-2507` @ **`cdbee75f17c01a7cc42f958dc650907174af0554`** (lastModified 2025-09-17T06:56:53Z, apache-2.0, not gated). **13 files, 8,060,917,568 B**, equal to the draft's count. Shards: `-00001-` 3,957,900,840 B `75311d91…8ed6`; `-00002-` 3,987,450,520 B `0b48adbb…dba1`; `-00003-` 99,630,640 B `7dd39ccc…060d5d`. `tokenizer.json` 11,422,654 B `aeb13307…dae4`; `tokenizer_config.json` `a62ff0a2…5ce3` (its `chat_template` is the one §5 writes out). All four equal the draft's HF-API pins |
| the pull | on benchbox, **2026-09-28T11:39:39Z → 11:41:29Z** (110 s, the home line), `curl` of `huggingface.co/<repo>/resolve/<rev>/<path>` for every file of each repo's tree at its revision; every file's sha256 checked against the tree's LFS oid, and every non-LFS file's git blob sha1 against the tree's oid: **19 of 19 OK** (`receipts/pull-o2.json`). Free on `/` before 139,287,769,088 B, after 131,094,507,520 B (floor 15 GiB) |
| package | `github.com/manjunathshiva/opendecider` at **`f950d261cf588e0154f3f27b8569b67fc87753ca`**, a fresh clone on benchbox read from `sys.path` and **never pip-installed**. Its six files equal O-1's pins (`run_o1.py` `PKG_SRC_SHA256`, read 11:45Z), `small.py` `35a8d875…d170` among them |
| runtime | `/usr/bin/python3` **3.14.4**. O-2's own venv (`python3 -m venv`, **no system site-packages**), holding **benchbox's CUDA package set**: the site-packages of benchbox's reranker venv (built 2026-09-20, `torch 2.14.0+cu130`), copied in, plus 14 wheels downloaded outside the sandbox and installed `--no-index --no-deps` inside it in `install` mode: `peft 0.21.0` (sha256 `b64eb75f…`), `accelerate 1.15.0` (`97eacca0…`), `psutil 7.2.2` (`076a2d2f…`), `packaging 26.3`, `pyyaml 6.0.3`, `jinja2 3.1.6` (the chat template needs it), `markupsafe 3.0.3`, `setuptools 84.0.0`, `idna 3.20`, `certifi 2026.7.22`, `rich 15.0.0`, `markdown-it-py 4.2.0`, `mdurl 0.1.2`, `pygments 2.21.0`. `pip check`: "No broken requirements found". **The main pins: torch `2.14.0+cu130`, transformers `5.17.0`, tokenizers 0.23.2, safetensors 0.8.0, huggingface_hub 1.32.0, numpy 2.5.3, peft 0.21.0, accelerate 1.15.0.** The freeze list (65 dists) is `receipts/venv-o2.txt`, the wheel shas `receipts/wheels-o2.sha256` |
| where | **benchbox**. The board: **NVIDIA GeForce RTX 3090 24 GB (EVGA XC3 Ultra)**, alone in the x16 seat, at **300 W** with **persistence Enabled** (an operator's sudo paste, read and capped at 11:11Z; the house 3090 posture, R-12). **No file carries its UUID or PCI address:** the runner finds it by `sha256("GPU-<uuid>")[:12]` = **`33a0b4acb8e8`**, passes the UUID to the sandbox in the environment only, and every row carries the label and that fingerprint. `CUDA_VISIBLE_DEVICES` is the UUID. **The worker** runs on logical CPUs **1–7** (seven distinct cores; benchbox has 8 cores with SMT, siblings k and k+8). **The runner** runs on CPU 0, as a transient user service (`systemd-run --user --unit od-o2-run`). The gate watches CPUs 1–7 and their siblings 9–15 |
| the OS | on every row, by `hostos.read_os()` from this directory's `hostos.py` (O-1's byte-identical copy, sha256 `a6a74bba57250be1…`) |
| instrument | new in this commit: `run_o2.py`, `od2_worker.py`, `sandbox_o2.sh`, `guard_o2.py` (their sha256s are the commit's; every pass receipt records the runner's, the worker's and the sandbox's). **Imported unchanged:** `run_o1.py` **`fb1614910bd91f51…`** (the schedule, the request bodies, the fingerprints, the pass list, the pipe `Worker`), and through it D0's `run_deem.py` `78fe9562…b2a6`, D1's `run_d1.py` `3e656d35…d8ac`, arm 4's `run_addenda.py` `a34aca61…acab`. `tables_o1.py` **`62893449c39e550a…`** is O-1's pinned table script, imported by `tables_o2.py` for its helpers. O-1's `sandbox.sh` and `od_worker.py` are **not** edited (their shas are O-1's pins); O-2's sandbox is a new file |
| seed | none reaches either model: one forward, a softmax over letter logits, no sampling. Gap 4's orders use D0's `random.Random(20260921 + k)`. Every bootstrap uses `random.Random(20260927)` with 10,000 resamples |

## 3. The sandbox (the package code is untrusted, as in O-1)

The two halves are O-1's (PREREG-O1 l.50–83):
- **`run_o2.py` is the trusted half.** It is stdlib only and never imports torch, transformers, peft or the package. It
  holds the gate on the host's `/proc`, reads the board by UUID with `nvidia-smi` from outside the sandbox, and writes
  every row.
- **`od2_worker.py` is the untrusted half.** It loads the package and the model and answers one question per JSON line
  on a pipe. It has **no writable output directory** (only its throwaway HOME).

**The worker's command:**

```text
systemd-run --user --scope --collect -q -p MemoryMax=16G -p MemorySwapMax=0 \
  taskset -c 1-7 bash sandbox_o2.sh run -- <scratch>/venv/bin/python -B -I od2_worker.py serve --model {small|base}
```

`sandbox_o2.sh` is O-1's `bwrap` form (`--ro-bind / /`, tmpfs over `/home /srv /opt /var /tmp /run`, `--unshare-net
--unshare-ipc --unshare-pid --unshare-uts --die-with-parent --new-session --clearenv`; bubblewrap works on benchbox,
tested 11:48Z), with three changes, each named:

1. **The CUDA device nodes are bound in:** `--dev-bind` for `/dev/nvidia0`, `/dev/nvidiactl`, `/dev/nvidia-uvm` and
   `/dev/nvidia-uvm-tools`. This is the minimum CUDA needs. `/proc/driver/nvidia` and `/sys` are read-only through
   `/`.
2. **The scratch is on benchbox's disk,** because benchbox's `/tmp` is RAM and the scratch holds ~14 GB of venv and
   weights. The `/home` fence admits **exactly two chains**, both read-only: the worktree copy
   (`/workshop/bench-scratch-o2/wt`) and the scratch (`/workshop/bench-scratch-o2/od`), with only the scratch's `home/` writable.
   The worker proves both chains at every start.
3. **The environment:** O-1's set (`HF_HUB_OFFLINE=1`, `TRANSFORMERS_OFFLINE=1`, `HF_HUB_DISABLE_TELEMETRY=1`,
   `DO_NOT_TRACK=1`, `ORT_DISABLE_TELEMETRY=1`, `HF_HOME=<scratch>/sandbox-home/hf`, `TOKENIZERS_PARALLELISM=false`) **plus the
   house phone-home switches** `VLLM_NO_USAGE_STATS=1` and `OLLAMA_NO_CLOUD=1` (D-20260928-008); `OMP_NUM_THREADS=7`,
   `MKL_NUM_THREADS=7`; `CUDA_VISIBLE_DEVICES=<the UUID>` in place of O-1's empty value.

**The fences (V-11)** are proved by the worker before it loads anything and recorded in every pass receipt. The worker
refuses to load if any fence is open:
- the only interface is `lo`; a TCP connect to 1.1.1.1:443 is refused; DNS for huggingface.co fails;
- `/home` holds the two chains only; the worktree and the weights are not writable;
- the environment carries every switch above;
- **(serve) CUDA sees exactly one device, and `"GPU-" + torch.cuda.get_device_properties(0).uuid` is the pinned
  UUID** (reported as its 12-hex fingerprint).

The 11:48Z smoke (no model loaded) read: `closed: true`, `ifaces ["lo"]`, TCP `refused errno=101`, DNS `refused
gaierror`, both chains only, the switches on, one CUDA device, the pinned card, bf16 supported; the board read 1 MiB
before and after.

**The model is loaded from local directories only.** `opendecider.load()` would pass the hub id from
`opendecider.json` into `SmallModel` (`__init__.py` l.66–69), which loads the base by that id (`small.py` l.36, l.47).
So the worker builds **`SmallModel(<adapter dir>, <base dir>, "cuda")`** itself, from the local base directory
holding the same bytes at `cdbee75f`, hashed; it wraps the result in the package's `OpenDecider(impl, meta)` and
answers through `system_one` as shipped. `HF_HUB_OFFLINE=1` makes any hub call fail loudly.

**The stop guard, read from OUTSIDE the bencher** (`guard_o2.py`, its own user unit `od-o2-guard`). Every second it
reads benchbox's UPS (NUT unit `pr1500`: `ups.load`, `ups.realpower`), the free bytes on `/`, and the card's latest
sample; it samples the card at 2 Hz (`nvidia-smi -lms 500` on the UUID: `power.draw`, `power.draw.instant`,
`temperature.gpu`, `memory.used`, utilisation, SM clock, P-state). **It trips on UPS load ≥ 80 %, card core ≥ 83 °C,
free disk under 15 GiB, or the UPS unreadable for 10 s: it records the trip, then stops `od-o2-run` and `od-o2-dry`.**
It never "records and continues". `run_o2.py` refuses to start a pass unless the guard's heartbeat is under 10 s old.

## 4. The item sets (identical to O-1's, PREREG-O1 l.97–104)

| set | file | sha256 | items |
|---|---|---|---:|
| S0 | `bench/jev-2026-09-21/kit/task_c.json` | `61a0b1d6e7cf66fd91992c0c3a4572f6fbae90e6471048fd7dc1eb41ff2d7f39` | 108 |
| H48 | `bench/deem-2026-09-27/kit/h48.json` | `<withheld: held-out set, see README.md>` | 48, **a live control**: its lines are never printed; rows carry `text_sha256` only |
| doorman | `bench/jev-2026-09-21/kit/doorman_planted.json` | `4daf059c26a47a84fa3b73630a64ba1a348da72c2e5df7f3c28c007b67da22f7` | 48 (36 planted, 12 controls); DOORMAN_SYSTEM v1.2 `c00fc24f…f2a7` |
| D1 voice | `bench/deem-2026-09-27/d1-voice/kit/voice-items.json` | `cea9e12d7983a3c52a204c1ed3e2d6c709e7504066f7ff37c1f86153bcf95d3c` | N 68, R 154, NM 41 |

## 5. The renderings, the readout, the tie rule, and what O-2c is

**The request bodies are O-1's, byte for byte.** `run_o2.schedule()` calls `run_o1.schedule()` and changes only each
row's key and arm: the key is (arm, model, *O-1's key without its arm*), which adds the model to O-1's resume key.
**`run_o2.py --check` (11:48Z) prints every pass's body fingerprint equal to O-1's registered one** (the §6 table).

**O-1's `[MASK]`/`[SEP]`/`[CLS]` strip.** O-1 stripped nano's three special strings from state text before building
the body, and the bodies carry that strip. For small these strings are not special: Qwen's tokenizer has no such
tokens, so they would reach the model as plain text, and a strip would change what the model reads. The strip is
carried (the bodies must stay O-1's) **and must remove nothing**: V-10′ refuses the arm on any count > 0. O-1 found
0 hits in 4,218 rows (PREREG-O1 l.91–93).

**Small renders them its own way** (`small.py` l.20–28, l.51–54). The system message is "You make one decision for a
software system.". The user message is:
- `Input:\n<state>\n\nQuestion: <instructions>\n\nOptions:\n`;
- one line per option, `A) key: gloss`, `B) …` in presented order (a key alone when the gloss is empty or equal);
- `\n\nAnswer with the letter of the correct option only.`.

The whole passes through the base tokenizer's chat template with `add_generation_prompt=True` and
`enable_thinking=False`. The template at `cdbee75f` (the non-tool branch) produces exactly
`<|im_start|>system\n{system}<|im_end|>\n<|im_start|>user\n{user}<|im_end|>\n<|im_start|>assistant\n`; the 2507
Instruct template has no thinking block, so `enable_thinking` changes nothing. **The S0 line sits in `Question:`.**
Small never truncates, and nothing here is truncated.

**The readout, as shipped** (`small.py` l.56–63):
1. One forward pass.
2. The last position's hidden state goes through `lm_head` in bf16, and the logits of the letters A.. (the first n)
   are cast to float.
3. A raw softmax over those n letters. No temperature and no calibration knob.
4. **Label mode (more than 26 options) is never reached:** the most options on any row is 19. V-10′ asserts n ≤ 26
   on every row.

The letter ids are the package's `tok.encode(c, add_special_tokens=False)[0]` for A–Z: the letter with no leading
space, which is the token that follows `assistant\n`.

**The tie rule, a priori:** the package's `max(probs, key=probs.get)`, the **first maximum in presented order**. That
is O-1's rule and D0's R-1. The letter logits are bf16-rounded before the cast, so exact ties are more likely than in
O-1's fp32. **Exact ties are counted, and every S0 reading is also printed over every tie-break.**

**O-2c, defined now.** It uses the package's own `render`, `SmallModel.ids` and `SmallModel.decide`, on an instance
built by a subclass (`BaseControl`, in `od2_worker.py`) whose `__init__` is `SmallModel.__init__` **minus its one
`PeftModel.from_pretrained(...).merge_and_unload()` line** (and so without its `from peft import PeftModel`).
Everything else is identical: the same tokenizer, the same dtype rule (bf16 on this card), the same base bytes, the
same letters. The worker reports `peft_imported: false` for O-2c, and the runner refuses a control worker that
imported peft.

## 6. The run, in order (box: benchbox; lane commands, not an operator paste)

0. **Before this commit, none of it a model call:** the pull with every file hashed (§2); the package clone at
   `f950d261`; the venv build and the offline wheel install inside the sandbox; `pip check`; the sandbox smoke (§3);
   `run_o2.py --check`.
1. **`run_o2.py --tokens`: the tokenizer only, sandboxed, over every scheduled row.** One pass serves both models,
   because small loads the base's tokenizer (`small.py` l.36). It checks V-8′, V-9′ and V-10′, and writes
   `receipts/token-check-o2.json`. **A refusal here stops the arm before any counted call.**
2. **The dry run:** `run_o2.py --all --dry 3 --model small`, then `--model base`, as the unit `od-o2-dry`: 3 rows of
   every pass, not counted, into `dry-run/`, with the guard running (tag `dry`). Every output file is opened.
3. **A dated amendment (v1)** records the token pass and the dry run. Then **the counted runs**, each
   `systemd-run --user --unit od-o2-run … taskset -c 0 python3 -B run_o2.py --all --model small`, then the same with
   `--model base`: twelve passes each, **O-2 first, then O-2c** (RECON §2.5: "straight after O-2"), one fresh worker
   per pass, the guard running (tag `counted`):

| pass | rows | body fingerprint (O-1's, PREREG-O1 l.146–158) |
|---|---:|---|
| `s0.P.rep1` (**the headline**) | 108 | `5e5d268fec5595a1` |
| `s0.P.rep2` (determinism) | 108 | `5e5d268fec5595a1` |
| `s0.S.rep1` | 108 | `f7c4cf9838d967b6` |
| `h48.P.rep1` | 48 | `<withheld: held-out set, see README.md>` |
| `s0.P.rep1.k1-5` | 540 | `ebda477d11343e31` |
| `door.P.rep1` | 48 | `4702bedebc7077a4` |
| `N.rep1` (19 rotations) | 1,292 | `7d6e4d6baf616eb8` |
| `R.rep1` (6 permutations) | 924 | `9d977692f253bc49` |
| `NM.rep1` (19 rotations) | 779 | `c6e0eb1b9b41af5b` |
| `N.rep2` / `R.rep2` / `NM.rep2` | 68 / 154 / 41 | `b93613a1f48d8df1` / `dbce446872df705d` / `5c29b11bd34e90b7` |
| **total, per model** | **4,218** | |

4. **`python3 -B tables_o2.py` → `results-o2.md`** (O-1's `results.md` shares this directory). It draws only from
   rows, receipts and the comparators' own files, and it calls no model. **`tables_o2.py` is written beside the
   running counted passes and pinned by a dated amendment before any O-2 figure is read** — a named deviation from
   O-1, which pinned its table script before its first counted row. Why: the within-the-hour law (a go means
   measuring); the readings and bars are this file's fixed text (§7), so the table code cannot move a bar, only print
   it.

**Carried from O-1, with one named change to the gate:**
- **D0's contention gate** (a 1.0 s window, under 0.15 CPU-s of foreign work on the watched CPUs; a row `contended` at
  ≥ 25 % of its wall; a 1,800 s cap after which the pass resumes by key, O-1 amendment v1) **runs before the first row
  of every segment and before any row that follows a contended row**, where O-1 ran it before every row. Why: a GPU
  row takes ~0.05–0.15 s, so a 1.0 s window before every row would be ~90 % of each pass's wall (~70 min a model) and
  would measure nothing the per-row `contended` flag does not. Every row still carries its foreign CPU-seconds and its
  `contended` flag; latency reads only uncontended rows, as in O-1. **On benchbox the watched set is the worker's 7
  CPUs and their SMT siblings (1–7, 9–15).**
- three warm-ups per pass, never rows;
- the resume key, with the model added (above);
- the OS on every row.

**Rows** go to `rows/o2-gpu-small.<pass>.jsonl` and `rows/o2c-gpu-qwen3-4b-base.<pass>.jsonl`, with a
`<stem>.pass.json` receipt per pass. **No row, receipt or telemetry file carries the card's UUID or PCI address or
the box's name** (sanitize at the pen).

## 7. The readings and the bars, written before any call

### 7.1 Each model on its own: O-1's bars, unchanged (PREREG-O1 l.174–184)

| set | registered reading (read for O-2 and for O-2c) | printed beside it (each read from its own file, with its gates) |
|---|---|---|
| S0 | **"a contender" iff S0 P rep 1 ≥ 72/108 AND min(kit order, the mean of the five rotated orders) ≥ 72.** Otherwise **"not a contender as shipped"**, with no threshold table. Below chance, the exact one-sided binomial p is printed. A count ≥ 72 reads "at least Kev-9B's count", never "beats" (C-6) | chance 16.7 %; the Jev floor 24.0 %; naive Bayes 68 (supervised; printed, not gated); **Kev-9B 72, with an exact paired McNemar on its own 108 rows** (`rows-gaps-card/kev-9b.kc.jsonl`, joined by text) **and on naive Bayes's rows**; Kev-4B 60; gemma4 86–88; mistral 74; OpenJev 101 (bench-only); Deem D0 5; **O-1 nano** |
| H48 | S0 P rep 1 − H48 P rep 1, with Newcombe's hybrid-score 95 % interval: **"CONTAMINATION SIGNAL" iff the lower bound > 0**, else "no signal" | Deem 4/48; naive Bayes 27/48; O-1 nano |
| doorman | **the 09-13 gate: planted refused ≥ 30/36 AND controls refused ≤ 2/12**; the 75 % majority floor printed | OpenJev 33/36; the 09-13 gemma run 27/36; gemma4 FP8 23/36; Kev-9B, Kev-4B; O-1 nano |
| N | **"above the majority baseline" iff the rotation-0 Wilson lower bound > 13.2 %**. The position table is **by letter, A..S** | Deem D1 251/1,292 (12/68 at rotation 0); **Deem chose F–S on 0 of 2,071 menu rows**; O-1 nano (by marker slot) |
| R | **act-risk 0 as shipped (argmax)** on the 16 inert items × 6 permutations. Two-class against the 72.7 % majority. Three-class, with a paired item bootstrap, against **today's lab rule 84.4 %** | Deem D1: act-risk on 61 rows across 14 items |
| NM | **act-risk 0 as shipped** on the 16 inert items × 19 rotations; paired against today's lab rule, 43.9 % | Deem D1: act-risk on 245 rows across 16 items |
| R and NM, thresholded | **τ = 0.90, fixed now**, O-1's policy. **act-risk(0.90) = 0 is the bar.** rescues(0.90) and genie-loss(0.90) are counts, the curve over τ = 0.50–0.99 is descriptive, and every zero is printed with its 17.1 %-per-item upper bound (0 of 16). **C-4's prior, restated for small:** in distribution, zero-shot small put **71 of 2,000** typed decisions above 0.90 (RECON §1.4), fewer than nano's 251. So few rescues are expected | today's lab rule: act-risk 0, rescues 0 |

**Position bias, from the rotations (letters here).**
- D1's registered tests are carried with D1's statistic, interval (the item bootstrap) and wording: on N rep 1, the E/F
  pair, F alone, E alone and S (the last letter); on R rep 1, the last letter C.
- O-1's added first-position test (A against the rest, N rep 1 and R rep 1) is carried too.
- **The prior is stated now:** a last-letter preference (RECON §1.10).
- **The letter wall is read directly:** the share of N rep 1 rows choosing F–S, printed beside Deem's 0 of 2,071.
- On S0: gap 4's reorder consistency (the items whose chosen guest is the same under all six orders) and the correct
  count per order.

**Calibration (C-3, C-5; descriptive, never gated): O-1's definitions verbatim** (PREREG-O1 l.195–203).
- Per set: top-label ECE with 10 equal-width bins and every bin's n; multi-class Brier; NLL.
- A coverage-at-threshold table (τ 0.5–0.9), beside coverage by rank (50 % and 70 %), in file order and as the
  expectation under random ties (C-9).
- The held-half temperature refit, split by the parity of the first 8 hex of sha256(uid), with T chosen on a grid of
  400 log-spaced points from 0.05 to 20.

**Determinism.** Rep 2 against rep 1's rotation 0: the same choice x of n, the number of bit-identical probability
rows, and the largest absolute difference, across processes. On CUDA, bit-identity at batch 1 is expected but not
promised. It is printed, never gated.

**Latency, memory, energy** (the bench box only, descriptive):
- **Latency:** the runner's wall per request over the pipe, one in flight (median and p95, nearest rank, over the
  measured uncontended rows, with counts), beside the worker's forward time and the package's `latency_ms`.
- **Memory:** the worker's VmRSS, and the board's `memory.used` growth from before the worker started, with the
  worker's CUDA peak.
- **Energy:** board energy integrated from the guard's 2 Hz `power.draw` samples over each pass's row window (the
  first row's start to the last row's end), per decision, with the idle draw before the worker started beside it;
  printed as a reading only for a clean pass (one segment, no gate wait, no contended row). benchbox's UPS
  `ups.realpower` at 1 Hz (a step function) is printed beside it.
- **Never quoted as ours:** the card's "~9 GB bf16, 38–40 ms on an L40S" (RECON-house-fit l.215).
- **Never a seat's timing:** the voice path's timing of record is the mini node's.

### 7.2 Small against its base (C-7): registered now

- **S0, the adapter's effect:**
  - an exact two-sided McNemar test, O-2 against O-2c, on the 108 items at kit order, with the discordant counts
    printed. **"The adapter changes the S0 count" iff p < 0.05**, with the direction read from the counts.
  - min(kit order, rotated mean) printed for both models.
- **N, the adapter's effect:** at rotation 0, a paired item bootstrap of the accuracy difference (D1's statistic;
  10,000 resamples, seed 20260927). **"The adapter moves N" iff the 95 % interval excludes 0.**
- **R, NM and the doorman:** the act-risk counts (as shipped and at τ = 0.90), the doorman gate, and the calibration
  tables, **side by side, descriptive**.
- **Contamination, read as a difference in differences** (the T1 form ruled in A-10):
  - The quantity is **(small S0 − small H48) − (base S0 − base H48)** at kit order.
  - The interval is a paired item bootstrap: 10,000 resamples, seed 20260927, S0 and H48 items resampled
    independently, and the two models sharing every draw. It is the 2.5th and 97.5th percentiles.
  - It reads **"contamination not excluded" iff the lower bound > 0**, else "no signal". It is never worded as a cause.
  - **Why this control works:** the base revision `cdbee75f` was last modified 2025-09-17T06:56:53Z (HF API). S0 has
    been public only since 2026-09-19 (RECON E7). So the base carries S0-versus-H48 difficulty but cannot have seen
    S0, and the difference isolates what the adapter's unpublished training pool may add.
  - O-1's Newcombe reading (§7.1) is printed beside it for both models.
- **Nano (O-1) against small (O-2),** descriptive: a side-by-side table. An exact paired McNemar at S0 kit order is
  printed. The word "better" appears only where that paired test supports it (O-1 §9).

## 8. Validity: a pass is VOID, never "low", unless all of these hold

- **V-1** every scheduled row is answered once (the rows equal the schedule, by key, per model).
- **V-2** the pass's request-body fingerprint equals §6's, which is O-1's.
- **V-3** the choice is the first maximum of its own probabilities, in presented order, on every row.
- **V-4** no row has all-equal probabilities. The runner stops.
- **V-5** every row carries C-8's fields: the package commit; both HF revisions (the adapter's null for O-2c); the
  adapter's (O-2 only) and every shard's sha256; the sha256 of the rendered input ids; the device (`cuda:0`, the card
  label and its UUID fingerprint) and the dtype (`bfloat16`); the torch, transformers and peft versions; the OS.
- **V-7** no failed request. A dead worker, a closed pipe, no answer within 600 s, or a `MemoryMax` kill is never a
  row. It goes to `<stem>.errors.jsonl` and the pass stops.
- **V-8′ (length; small never truncates):** each row's prompt-token count is recorded. Rows over **768** tokens, small's
  training length, are counted per set and **reported as out of training length, never voided** (RECON §2.2,
  RECON-house-fit l.370).
- **V-9′ (the ids):** on every row, the package's ids (`SmallModel.ids`) equal the worker's independent rebuild, which
  writes the chat-template text out by hand (§5) and encodes it; it never calls `apply_chat_template`.
- **V-10′ (the letters and special strings):**
  - Each letter A–Z encodes to **exactly one** token id, the 26 ids are distinct, and the package's letter ids equal
    the worker's. The package takes `[0]` silently (`small.py` l.49), so this check makes that visible.
  - **n ≤ 26.**
  - **Zero** of the base tokenizer's added-token strings (every entry of `tokenizer.json`'s `added_tokens`:
    `<|im_start|>`, `<|im_end|>`, `<|endoftext|>`, `<think>`, `</think>` and the rest) appear in the instructions,
    options or state of any scheduled row, counted by the runner from `tokenizer.json` and by the worker from the
    loaded tokenizer; and O-1's strip removed nothing (§5).
  - A hit stops the arm before any counted call and is surfaced, **never silently stripped** (C-2's rule). 0 hits are
    expected.
- **V-11** the sandbox's fences are closed at every worker start, and CUDA sees exactly the pinned card (§3).
- **V-12′ (the stub trap and residency):**
  - Before every pass, the runner recomputes the sha256 of all 19 files of both repos against `WEIGHT_SHA256`. The
    package checkout must be clean at `f950d261`, and the worker must report the six files at their pins.
  - The worker reports every parameter as `torch.bfloat16` on `cuda:0`, the class `Qwen3ForCausalLM` in eval mode
    with no LoRA module left; torch `2.14.0+cu130`; transformers `5.17.0`; peft `0.21.0` imported with the adapter
    merged (O-2), or peft not imported (O-2c); and `HF_HUB_OFFLINE=1`.
  - **Residency by growth:** after the three warm-ups, the board's `memory.used` on the UUID has grown, since before
    the worker started, by at least **0.95 × the worker-reported parameter bytes** (the kit's rule, vllm-seats SPEC
    §4.4 l.597), and the worker's host pid is the card's only compute process.
- **V-13 (the card is alone and at its posture):** before the worker starts, `nvidia-smi --query-compute-apps` on the
  UUID lists no PID; after the warm-ups and before the worker closes it lists only the worker's; after it closes, none;
  `power.limit` reads 300 W within 1 W and persistence reads `Enabled` at every read (SPEC §4.5 l.612–614).
- **The stop guard never tripped** during the pass (`receipts/<tag>-guard.jsonl`); a trip stops the pass, which is then
  VOID until resumed by key after the cause is cleared.

## 9. What the arm may and may not say

**It may say:**
- opendecider-small (`ff25e366` on `cdbee75f`), as shipped (bf16, merged LoRA, raw letter softmax, the package's own
  `system_one`), zero-shot, scored k of n on each set under §5's renderings on the bench box's RTX 3090;
- the same for the base control, labelled as §1 says;
- §7's registered readings and nothing stronger;
- its position, calibration and determinism reads on these items;
- its latency, memory and board energy on the bench box.

**It may not say:**
- adopt, seat-ready, or any threshold for a seat;
- a latency for any box but this one;
- a calibration claim beyond these items;
- "better than" anything unless a paired test supports it;
- the card's GPU figures as ours;
- anything about small-td, medium-td, the MLX builds, or nano beyond O-1's own file;
- that either model is uncontaminated (H48 and the difference can show a signal, never absence);
- any typed-decisions number as skill on our tasks;
- the base's readings as OpenDecider's.

**Registered as not run** (their absence is not a silent skip): small on CPU (C-12); small-td, medium-td, the MLX
builds; a Noul rendering of the route; the judge set; docent task (a); any rulesage decision; the doorman under
reordering; label mode (more than 26 options); O-3 and O-3b.

## 10. Deviations from O-1 (and from the draft), each named

- **Device and precision:** CUDA bf16 on benchbox's RTX 3090 at 300 W, where O-1 used CPU fp32 on this laptop. Why:
  RECON §0 l.72–78 rules small out of this laptop's CPU (disk, RAM, ~3–4 s a row). The authority: OD-1 (a), ruled by
  D-20260928-011.
- **The runtime:** torch `2.14.0+cu130` — the same torch release as O-1's `2.14.0+cpu`, its CUDA 13.0 build —
  where the draft planned the laptop T1 venv's `2.13.0+cu130`. Why: this lane does not use the laptop (Deem T1 may be
  training there), so it builds on benchbox's own CUDA set. transformers, tokenizers, safetensors and numpy equal O-1's;
  huggingface_hub is 1.32.0; peft, accelerate and psutil are added (small's own extra), with jinja2 and the other
  wheels the venv needs once system site-packages are off.
- **The card:** the board an operator seated and read at 11:11Z (an EVGA RTX 3090 XC3 Ultra), not the board the draft
  named; the posture (x16 alone, 300 W, persistence on) is the draft's.
- **The sandbox:** a new `sandbox_o2.sh` (O-1's file is untouched, so O-1's pins hold): bound CUDA device nodes, two
  read-only `/home` chains, the house switches, and `CUDA_VISIBLE_DEVICES` set to the UUID. The worker's `MemoryMax` is
  **16G** where O-1 used 8G: the bf16 4B passes through host RAM on its way to the card (about 8 GB).
- **CPUs and the gate:** the worker on 1–7, the runner on 0, the watched set with SMT siblings; the gate before a
  segment's first row and after a contended row (§6).
- **The model is built from local directories** (§3), because `load()` would pass the hub id. The bytes are the same,
  hashed.
- **V-8, V-9 and V-10** are re-expressed for a letter readout with no truncation (V-8′, V-9′, V-10′). **V-12 and V-13**
  measure residency and posture on the board rather than in VmRSS.
- **Energy** is the board by UUID plus the UPS, where O-1 read RAPL only.
- **The stop guard** (§3) is new: O-1 had none on the laptop.
- **No CB-5 wait.** O-2 has no inference-box window to guard.
- **`tables_o2.py` is pinned before any figure is read, not before the first counted row** (§6 step 4).
- **The additions:** O-2c (C-7) and §7.2's small-against-base readings, all registered here before any call.

## 11. Minutes, bytes, and stops

| step | minutes (ESTIMATE, except where measured) | note |
|---|---:|---|
| the pull | **1.8 (measured)** | 8,193,128,805 B, 11:39:39–11:41:29Z |
| the venv, the clone, the wheels | **~3 (measured)** | 11:45–11:48Z |
| the token pass (both models, one tokenizer) | 1–3 | |
| the dry run (3 rows × 12 passes × 2 models; 24 worker loads) | 6–10 | a load from page cache estimated at 10–20 s |
| **O-2 counted** (12 passes, 4,218 rows) | **8–15** | ~0.05–0.15 s a row plus 12 loads; the draft's 20–30 assumed O-1's per-row gate |
| **O-2c counted** | **8–15** | |
| **total at the card** | **25–45** | the ledger's "~40–60 min" covers it |

**Stops (surfaced, never routed around;** RECON §2.5 and D-20260928-006):
- a licence or runtime surprise;
- V-9′ or V-10′ refusing a row, or the token pass failing;
- `HF_HUB_OFFLINE=1` failing to resolve a local directory;
- any row with all-equal probabilities;
- a foreign PID on the card, or the cap or persistence off posture;
- the stop guard tripping (UPS ≥ 80 %, core ≥ 83 °C, disk under 15 GiB, the UPS unreadable);
- free disk under the floor.

**Cleanup:** after the counted run the lane removes `/workshop/bench-scratch-o2` on benchbox (weights, venv, clone) and
confirms the card reads empty with no compute process.

## 12. Sources

- The draft: `plans-2026-09-28/benchbox-day-2026-09-29/PREREG-O2-draft.md` (sha256 `77c497cba2055b63…`).
- `RECON-opendecider.md` (sha256 `2d22fddaa8afb918…`): §0 l.72–93, §1.1, §1.4, §1.5, §1.8, §1.9, §1.10, §2.1
  l.378–391, §2.2, §2.5 l.465–490, §3.2; `RECON-house-fit.md` l.211–215, 370.
- This directory at `a3d0147c`: `PREREG-O1.md`, `run_o1.py` `fb161491…`, `tables_o1.py` `62893449…`, `hostos.py`
  `a6a74bba…`.
- `origin/master` `DECISIONS.md` l.1796 (D-006), 1800 (D-010), 1801 (D-011); `bench/kits/vllm-seats/SPEC.md` §4.4
  l.591–610, §4.5 l.612–614.
- The package at `f950d261`: `opendecider/small.py` l.20–63, `opendecider/__init__.py` l.53–75.
- benchbox, 2026-09-28: the pull receipt (11:39–11:41Z), the chat template read from `tokenizer_config.json` at
  `cdbee75f` (11:44Z), the venv build and `pip check` (11:45–11:48Z), the sandbox smoke (11:48Z), `lscpu -e` and
  `nvidia-smi` (11:37Z: 1 MiB used, no compute process, 300.00 W, persistence Enabled, x16).

---

*Results go below this line, and in `results-o2.md` and `rows/`, only after the first counted row.*

## v1: before the first counted row (2026-09-28T12:04Z; no counted row exists)

*A dated amendment written after the token pass and the dry run and before any counted row. The pre-registration
above was committed as `f1f8ef3d` at 11:52:15Z and pushed at 11:52:16Z, before any model call.*

**Two refusals, both caught before any counted row, both fixed in the instrument (not in any reading or bar):**

1. **The first token pass refused (11:53Z) on a version string, not on data.** Every data check passed on all 4,218
   rows, but the ready check compared the dist metadata's torch version (`2.14.0`: the CUDA 13.0 wheel carries no
   local tag) against the registered `2.14.0+cu130` (the runtime's `torch.__version__`). **Fix:** the worker now
   reports both, and `run_o2.py` checks the triple **dist `2.14.0`, runtime `2.14.0+cu130`, CUDA `13.0`**
   (`TORCH_DIST`, `TORCH`, `TORCH_CUDA`). Rows carry the runtime string. The refused receipt is kept:
   `receipts/token-check-o2.refused-1.json`.
2. **The first dry run refused (11:55Z) at the first warm-up: a worker crash, not a row.** torch 2.14's native-op
   router (`torch/_native`) sends the rotary embedding's outer product (`bmm_outer_product`) to a Triton kernel, and
   Triton's launcher `gcc`-compiles a C stub against `Python.h` at run time; benchbox has no Python headers
   (`fatal error: Python.h: No such file or directory`). **Fix:** `sandbox_o2.sh` sets
   **`TORCH_DISABLE_NATIVE_JIT=1`** (torch's own global switch, `torch/_native/common_utils.py`), which keeps every op
   on torch's stock ATen CUDA kernels, as in torch ≤ 2.13. No compile runs inside the sandbox. The worker records the
   variable and the runner refuses a worker without it. The crash's logs are kept: `receipts/dry-refused-1.*`.

**The token pass on the final instrument** (`receipts/token-check-o2.json`, 12:02:27Z, sha256 `73f79f96f4cf3d4c…`):
**ok**. 4,218 rows; every pass's body fingerprint equals O-1's; 0 ids mismatches (V-9′: the package's ids equal the
hand-written template's on every row); 0 letter faults (A–Z each one token, 26 distinct ids); 0 added-token strings
(26 checked) in any instruction, option or state; O-1's strip removed nothing; the most options on a row is 19.
**V-8′: 0 rows over 768 tokens**; prompts run 133–521 tokens (the doorman's 501–521 the longest; S0 198–366).

**The dry run** (`od-o2-dry`, 11:56:20–12:02Z, both models, 3 rows × 12 passes each, `dry-run/rows/o2*.jsonl`;
drawn into `dry-run/results-o2-dry.md`): 72 rows and 24 pass receipts, every one `complete`, no errors file. Every
pass is VOID on V-1 and V-2 only, by design (3 rows against the full schedule); no other validity check fired. Every
worker proved its fences (only `lo`, TCP and DNS refused, the two chains, both read-only, every switch on), saw one
CUDA device, the pinned card (`33a0b4acb8e8`), and loaded in 4.3 s (small, peft merged, 0 LoRA modules left) or 3.6 s
(the control, peft not imported): 4,022,468,096 parameters, 8,044,936,192 B in bf16 on `cuda:0`, attention `sdpa`.
**Residency by growth held on every pass:** the board grew 8,595 MiB (small) and 8,083 MiB (base) against the
7,288 MiB bar (0.95 × the parameter bytes). The card read no compute process before each worker, only the worker's
after the warm-ups, none after it closed; 300 W and persistence at every read. Rows took ~0.04–0.08 s. The stop guard
ran throughout (tag `dry`, 1 Hz): never tripped; the UPS read 6 %, the core peaked at 48 °C, the board at 166 W.
Every writer ran and every file was opened: the rows, the pass receipts, the worker logs, the guard's two receipts,
the page (including the energy integration, which reads "not measured" on 3-row passes because their row window is
under one 2 Hz sample pair; tested over the whole dry window it integrates 36.7 kJ over 597 s).

**`tables_o2.py` is pinned here, before any counted row** (so the deviation registered in §6 step 4 was not needed):
sha256 **`97c99c1b913ee0ca70f600cc4a938a8130518f32d9646352b14f98bd28ef9cc5`**. It reads O-1's rows by path
(`--o1-rows`), prints each O-1 file's row count and sha256, and reads them under O-1's own validity function. At
12:03Z O-1's counted run was complete on the laptop (12 of 12 passes, 4,218 rows, uncommitted in O-1's worktree).

**The instrument as it runs the counted passes:** `run_o2.py` `12c51845901b8d42…`, `od2_worker.py`
`54c9b77701583958…`, `sandbox_o2.sh` `b224ea209037cbd4…`, `guard_o2.py` `a43865187f312df8…` (every pass receipt
records the first three as it read them).

**The counted runs start now**, O-2 then O-2c, under `od-o2-run`, with a fresh stop guard (tag `counted`).

## v2: a registered stop on the doorman pass, and V-4 read for three or more options (2026-09-28T12:08Z)

*A dated amendment. The counted run started at 12:04:33Z (first row 12:04:46Z). O-2's first five passes completed
(s0.P.rep1, s0.P.rep2, s0.S.rep1, h48.P.rep1, s0.P.rep1.k1-5); none of their figures has been read by the lane.*

**What happened.** At 12:07:02Z, on the doorman pass's sixth row (2 options: admit, refuse; 511 prompt tokens),
small's two letter probabilities came back exactly equal, `{admit: 0.5, refuse: 0.5}`. V-4 ("no row has all-equal
probabilities; the runner stops", carried from O-1 and D0) stopped the pass and the arm, as registered. The five rows
before it and the stop are kept, unedited, in `rows/refused-v2/` (the rows, the pass receipt, the errors file, the
worker log). The lane saw only the stop row and the passes' tie counts, not any doorman figure.

**Why this is not the failure V-4 exists for.** V-4 is D0's stub trap: a flat distribution over many options means a
model that is not really answering. On a 2-option row, "all-equal" is exactly an exact tie at the top, and §5 already
registered what happens then: small's letter logits are bf16-rounded before the cast, "so exact ties are more likely
than in O-1's fp32. Exact ties are counted, and every reading is also printed over every tie-break", and the choice is
the first maximum in presented order. The two rules collide only on n = 2. The S0 passes show the same bf16 ties
between two of six letters (4 of 108 rows on S0 P rep 1), which V-3's first-max rule already handles.

**The amendment:**
- **V-4 applies to rows of three or more options.** A row of three or more with all-equal probabilities still stops
  the pass (the stub trap stands).
- **On a 2-option row an exact tie is a tie under §5:** the choice is the first option in presented order (the
  doorman's `admit`), the row is counted, and every doorman reading is printed over every tie-break (planted refused
  and controls refused as ranges, and whether the 09-13 gate's reading depends on the break).
- **A new degenerate-pass guard, set now:** a pass in which more than 25 % of the 2-option rows tie exactly is VOID
  (12 of the doorman's 48).
- **The doorman pass re-runs whole**, one fresh worker, one segment, before the arm continues with N, R, NM and the
  rep-2 passes and then O-2c. The five completed passes stand as they are (they were complete before the stop, and no
  rule they are read under changed).
- `run_o2.py` becomes `9443a4307b99510c…` (the one-line V-4 change) and `tables_o2.py` becomes `e1caf6e2c072ea74…`
  (the same V-4 change, the 25 % guard, and the doorman tie ranges). The pass receipts record which runner wrote each
  pass.

**Named plainly:** this is a change to a validity rule made after a counted stop, to resolve a contradiction between
two registered sections (§5 and §8 V-4) that only a 2-option row can reach. It moves no bar and no reading. The stop
is surfaced to the lead in the lane report as a judgment call; if the lead rules the other way, the doorman pass
reads VOID for small (and for the control, if it ties too) and nothing else changes.

## v3: O-1 landed on master mid-run; the table script's presentation follows O-1's landed one (2026-09-28T12:24Z)

*A dated amendment written after the counted run (12:04:46Z → 12:21:4xZ, both models, 24 of 24 passes) and before
the results were committed.*

- **O-1 landed while O-2 ran** (#180, squash `95cfcd88`, 11:49Z; the O-1 branch this one stacked on is deleted).
  This branch merged `origin/master` (`4631c2ab`), taking every O-1 file as landed; `git diff origin/master` now
  touches O-2's files only. `run_o1.py` is unchanged (`fb161491…`). **`tables_o1.py` moved to `7ee77a77…` in
  `render()` only** (O-1's audit fixes R-1 and R-8); every helper O-2 imports (`d1_readings`, `policy`, `ece`, `nll`,
  `refit`, `coverage_by_rank`, `nb_rows`, `door_comparator`, `door_gate`, `first_max`, `group_diff`, `label_key`) is
  byte-identical, so no O-2 figure moves. O-1's rows are now read from this directory's `rows/` (master), and the page
  prints each O-1 file's sha256.
- **Presentation only, mirroring O-1's landed page:** the τ = 0.90 headlines print "rows acting at 0.90: k of n" and,
  when no row reaches τ, "meets act-risk(0.90) = 0 — vacuously" (O-1's R-1); the COI comparators' gate text prints
  the set fingerprint and the contamination field (O-1's R-8). Two labels are corrected: O-1's path prints
  repo-relative, and the energy table's "idle" column is renamed for what it reads (the board in the 5 s before a
  worker start follows the previous pass's close, 130–160 W on most passes, not a settled idle; J / decision is gross
  and never used it). **No bar, reading, statistic or validity rule changed.**
- `tables_o2.py` is now **`ff521a5f0fde121f57ae4f4549c9888b530eb99db14054b913a13465e018c5fd`**; `results-o2.md` and
  `dry-run/results-o2-dry.md` are drawn by it.

## Results (2026-09-28, pointer)

The counted rows are `rows/o2-gpu-small.*` and `rows/o2c-gpu-qwen3-4b-base.*` (4,218 each, every pass `valid`); the
page is `results-o2.md`. The refused doorman segment is `rows/refused-v2/`. The stop guard never tripped (UPS max
37 %, core max 69 °C, board peak 288.9 W; `receipts/counted-guard.jsonl`, `receipts/counted-telemetry.jsonl`).
**The registered prior "a last-letter preference" (§1, §7.1) is refuted on these items and stays in the text:** small
chose S, the last of 19 letters, on 63 of 1,292 N rows (4.9 %, content-only 5.26 %).

## Checker's notes (post hoc, dated 2026-09-28 about 12:45Z; no registered figure, bar or reading changes)

*Added by the independent checker of O-2 (`/workshop/bench-archive/plans-2026-09-28/opendecider/O2-CHECK.md`). An
independent stdlib recount of the committed rows (4,218 per model) reproduces every headline and every validity
fact above, and `tables_o2.py` at `ff521a5f` redrew both pages byte for byte before the change in C-3.*

**C-1 (scope of the refuted prior).** What is counted: how often each letter is chosen, against how often the label
sits there. The pointer above says the last-letter prior "is refuted on these items". That holds for the 19-option
menu N (S chosen on 63 of 1,292 rows, 4.9 %, against 5.26 %) and for the route R (C chosen on 307 of 924, 33.2 %,
against 33.3 %). It does not hold for the six-way S0. Over S0's six orders (the kit order and gap 4's five, 648 rows,
`results-o2.md` section 2), small chose F, the last letter, on 152 rows (23.5 %), with the label at F on 118 (18.2 %).
It chose E, the next-to-last, on 190 rows (29.3 %), with the label at E on 114 (17.6 %). It chose A on 36 (5.6 %),
with the label at A on 84 (13.0 %). The base leans the same way: E 219 (33.8 %), F 130 (20.1 %), A 27 (4.2 %). So
the prior is not seen at 3 or 19 options, and a late-position lean is seen at 6 options, as O-1's nano showed (its
R-2). S0's letter table is descriptive (section 7.1 registers no position test on S0), so this is a reading of
the printed table, not a test.

**C-2 (erratum to v1's dry-run guard figure).** v1 says the stop guard "read" the UPS at 6 % during the dry run.
`receipts/dry-guard.jsonl` (638 reads, 11:52:26Z to 12:03:32Z) reads 6 % at rest and up to 22 % under load
(realpower up to 243 W). The core's 48 °C and the board's 166 W in the same sentence are maxima. The guard's bar
is 80 %, so no reading moves.

**C-3 (presentation: sanitize at the pen).** `results-o2.md` and `dry-run/results-o2-dry.md` named the dev laptop
(twice) and, through the comparators' own labels, the bench box and the inference box's card by their estate names,
in five cells of the S0 comparator table. They now print the role: "the dev laptop's CPU", "the bench box's 3090",
"the inference box's card". That is the only change. The rows, the receipts and the telemetry never
carried a box name (section 6's promise holds). `tables_o2.py` `ff521a5f…` → sha256
`345e5abb80f0161e26f28c067addb02cd39b4e79931b4e13c46f6cea84f7662c`. This registration's own text still names the
box, and the pass receipts carry the scratch paths under the home directory (as O-1's do). The check flags both
for a sanitized twin if either travels.
