# Arm L-1: interfaze-ai/lev, zero-shot, on the six-way S0, this laptop's CPU: pre-registration

*Written 2026-09-28 between 11:44Z and 11:47Z (UTC, `date -u` on this laptop), saved BEFORE the first
scored call of this arm (the 3-item dry run). Every row carries its own `stamp`; all are later than this
file's mtime. Not committed anywhere (the lane brief: no commits, no pushes); the file's mtime and the
rows' stamps are the receipt.*

*The go, verbatim (an operator, 2026-09-28 ~11:35Z): "opus, check this out, we can add this to our decision
model bench: https://github.com/InterfazeAI/lev" and then "let's do a bit of recon and research, and
bench it based on your recs". Recon: `RECON-lev.md` beside this file.*

## 1. The question (a measurement arm, not an adoption verdict)

What does lev, as shipped and zero-shot, answer on our frozen six-way S0 ("which of the six dinner
guests said this line?", 108 items), asked through the SAME `/v1/systemone` request body that Kev
(arm 12) was asked, on this laptop's CPU?

- It is **not** a seat exam, a latency of record (CPU, not a GPU seat), or a claim about lev on its
  own S1Bench. lev's published numbers stay QUOTED, never verified here.
- Label on every L-1 figure: "zero-shot on our kit; lev is a general System One classifier trained on
  public datasets, never on our lines".

## 2. The pins

| what | pin |
|---|---|
| adapter | `interfaze-ai/lev` at HF revision `f8ef71157ec06a7d3b6435bc0756f9d735c33748` (lastModified 2026-09-25T17:24:25Z). `adapter_model.safetensors` 169,903,320 B sha256 `64c71897…0623c1`; `mode_b_head.pt` 14,698,565 B sha256 `27eedf7b…d6c3ad` (both equal HF's LFS oid); every file's sha256 in `receipts/lev-files.sha256` |
| base | `Qwen/Qwen3.5-4B` at revision `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a` (named by the adapter's `adapter_config.json` and `lev_release.json`); shards `…00001` 5,329,398,688 B sha256 `26a93f06…6302d`... see `receipts/qwen-files.sha256`; both shards equal HF's LFS oid (`receipts/qwen-hf-lfs.txt`). 9.34 GB on tmpfs (the laptop root is 99 % full, 13 GB free) |
| package | `github.com/InterfazeAI/lev` at commit `0ea955f5aee0d56d906121d1779af20fa8b922ed` (2026-09-25T17:24:11Z), read from `sys.path` (`packages/lev/src`), never pip-installed; `lev.__version__` 0.1.1 |
| runtime | Python 3.14.4 venv; torch `2.14.0+cpu`, transformers `5.17.0`, tokenizers 0.23.2, safetensors 0.8.0, huggingface_hub 1.33.0, numpy 2.5.3 (O-1's own wheel set, same sha256s) + peft 0.21.0, accelerate 1.15.0, pydantic 2.13.5 / pydantic_core 2.46.5, psutil 7.2.2 (`receipts/wheels-extra.sha256`, `receipts/pip-freeze.txt`) |
| how lev is called | in-process `lev.load(<release dir>)` then `DecisionEngine.system_one(state, questions)`, the function lev's own `/v1/systemone` route calls; `lev.load` defaults as shipped: bf16, `prompt_style` chat (from the release manifest), calibrated (`calibration.json`, choice:A T = 1.77), `order_average` True (two option orders averaged, lev's default), Mode A label-token readout, Mode B head loaded |
| instrument | `bench/jev-2026-09-21/run_kev.py` (sha256 `1fbce7a4…010b05`) and `run.py` (`<withheld: rewritten file, see README.md>`), imported UNEDITED from `/workshop/estate` whose copies are byte-identical to `origin/master` `fe57b8a9` (checked 11:42Z). `run_lev.py` swaps only `kev_post` for the in-process call, with lev's server's 422 mapping |
| items | `kit/task_c.json` sha256 `61a0b1d6…2f7d39`, 108 items, roster from `kit/manifest.json` (`<withheld: rewritten file, see README.md>`), in file order, uids `c000`…`c107` |
| the request | task_kc's body verbatim: state "A dinner table with six guests: <names>.", instructions `INSTR_C % line`, a `choice` question whose criteria are the six guest ids to names |
| the apples-to-apples check | every row's `prompt_sha256` must equal Kev-9B's row for the same uid (`rows-gaps-card/kev-9b.kc.jsonl`); the runner STOPS on the first mismatch |
| box | this laptop, Intel Core Ultra 9 290HX Plus, CPU only (`CUDA_VISIBLE_DEVICES=` empty, CPU-only torch wheel; the row records `cuda_available`). No laptop GPU, so no fan ask |
| cores | P-cores 0-7, 8 torch threads (O-1's posture), started only after O-1's `run_o1.py` exited |
| OS | on every row, via `hostos.read_os()` (byte-identical copy, sha256 `a6a74bba…3e0e3ad`) |
| phone-home switches | set in the sandbox env: `ORT_DISABLE_TELEMETRY=1`, `HF_HUB_DISABLE_TELEMETRY=1`, `HF_HUB_OFFLINE=1`, `TRANSFORMERS_OFFLINE=1`, `VLLM_NO_USAGE_STATS=1`, `GRADIO_ANALYTICS_ENABLED=False`. `DO_NOT_TRACK` is NOT set (the lead's brief). Plus a network fence: bwrap `--unshare-net` (only `lo`; a TCP connect to 1.1.1.1:443 read "Network is unreachable" at 11:45Z) |
| containment | `systemd-run --user --unit=lev-l1-<leg> -p MemoryMax=16G -p MemorySwapMax=0 -p CPUAffinity=0-7`; free RAM read before launch |
| seed | none reaches the model (one forward per order, softmax readout, no sampling) |

## 3. What is scored

- **Headline:** S0 accuracy, k / 108 (refused rows, if any, excluded and counted on their own line),
  Wilson 95 % interval.
- Brier (multi-class, `run.brier`), as served (calibrated). No raw-T arm: lev applies its calibration
  inside `system_one`, and accuracy is unchanged by it.
- Seconds per item (CPU, descriptive only), `argmax_disagrees` count, RSS.
- Reference bars from PREREG-T1 5.2, **reference only, no gate**: Kev-4B 60 / 108, NB (bag of words on
  227 of our lines) 68 / 108, G-KEV 72 / 108 (Kev-9B), gemma 4 ~87 / 108 (80.6 %), G-SEAT-TALK 88 / 108,
  OpenJev 101 / 108.
- McNemar vs Kev-4B (same base size) and vs Kev-9B, joined on (uid, text_sha256), exact two-sided.
  Descriptive.

**Gate: none, descriptive** (as arm 12). The only validity gates are the instrument's: every
`prompt_sha256` matches Kev-9B's, zero request errors, the dry run's 3 rows land before the counted run.

## 4. Predictions (written before the first call)

| quantity | prediction | why |
|---|---|---|
| S0 accuracy | point **60 / 108 (55.6 %)**, 80 % interval 48 – 74 / 108 | same 4B Qwen3.5 class as Kev-4B (60/108); lev's base is the instruct model and its training mix is larger (200k examples), which may help, but authorship-by-persona needs world knowledge a 4B holds thinly |
| reaches G-KEV (≥ 72 / 108) | 20 % | |
| reaches G-SEAT-TALK (≥ 88 / 108) | 2 % | |
| beats NB 68 (McNemar p < 0.05) | 5 % | |
| Brier (calibrated) | 0.55 – 0.70 | Kev-4B 0.62 as served |
| refused / 422 rows | 0 | S0 states are ~200 tokens |
| seconds per item, 8 P-cores, bf16 | 2 – 10 s (median ~5 s) | ~200 prompt tokens x 2 orders x ~8 GFLOP/token on a CPU without native bf16 matmul |
| whole counted run | 5 – 20 min | 108 x the above + ~1 min load |
| peak RSS | 9 – 12 GB | bf16 text weights ~8 GB (the vision tower is not loaded by AutoModelForCausalLM) |

## 5. What L-1 may say

"lev (4B, Apache-2.0), zero-shot on our six-way: k / 108, the same request Kev was sent." Nothing about
seats, GPU latency, or lev's S1Bench. A CPU latency is a laptop-CPU latency and is labelled so.

## 6. The next leg (written now, not run by this lane)

- **L-2, benchbox, one RTX 3090, bf16, lev's own `lev serve` behind `/v1/systemone`**, driven by
  `run_kev.py --arm` exactly as arm 12 (the Kev host swapped for lev's), tasks kc + ka + kd + kf + kj +
  kperm. That gives the doorman and the reorder numbers beside Kev's, and a GPU latency comparable
  with arm 12's 0.11 s. It waits for benchbox's current bench lane to finish (single-lane law).

## Amendments (dated, UTC)

- **2026-09-28 11:53Z (after the counted run started, before it ended; changes no scored quantity):**
  §2's base row printed shard 1's sha256 with the wrong tail. The receipt value is
  `26a93f066e1916adb13453dae5a0c707c0fbc71299ed98779571a907b8e74c61` (shard 2
  `cb544bd9bfae93dc59b0f22b292f5933573854a7f9b97835c67060d7d910e188`), both equal to HF's LFS oid
  (`results/receipts/qwen-hf-lfs.txt`). As run: the dry run was `lev-l1-dry.service` (11:46:13Z to
  11:46:33Z; 3 rows, all `prompt_sha256` equal to Kev-9B's; O-1's `run_o1.py` had exited at 11:46:04Z,
  NM.rep2 `complete`). The counted run is `lev-l1-s0.service`, started 11:46:55Z. `lev-l1-finalize.service`
  copies the evidence to `results/` when it ends, then frees the tmpfs weights.

- **2026-09-28 ~11:55Z, the lead: provenance of the pre-call text.** The 11:53Z amendment moved this file's mtime to 11:52:58Z, so the line near the top ("all [row stamps] are later than this file's mtime") no longer proves anything. The proof that the predictions came first is the lane transcript, `/agent-transcripts/subagents/agent-ac0ba086cf268d307.jsonl`. There, the Write of this file at 2026-09-28T11:45:45.091Z already carries the §4 prediction ("point 60 / 108 (55.6 %), 80 % interval 48 – 74 / 108") and the reference bars. That is before the dry run's first scored row (11:46:33Z). Next time the lead's brief will require the prereg's sha256 to be recorded before the first call, in a ledger row or a commit, and not left to the file's mtime.

---

# Arm L-2: lev on the bench box's RTX 3090, through `lev serve`, beside Kev (pre-registration, 2026-09-28)

*Cold scroll: what this section is.* It finalises §6 above ("the next leg, written now, not run by this lane")
into a registration. Written 2026-09-28 between 13:02Z and 13:15Z (UTC, `date -u` on the dev laptop), and
committed and pushed on the estate branch `bench/lev-2026-09-28` **before any L-2 model call**. The push time is
read from GitHub and stated in the first amendment below. Everything above the `---` that opens this section is
the archive file `/workshop/bench-archive/plans-2026-09-28/lev/PREREG-lev.md` byte for byte: 9,035 B, sha256
`fa0997017aa94b8a0257fb10b6044b0f91cd5b648b3f226ee03ee6664fd99899`. `l1/COPIED-FROM-ARCHIVE.sha256` records
every L-1 file copied into this directory, each equal to its archive original.

At commit time, no L-2 model call had happened. The weights are fetched and hashed (equal to L-1's receipts).
The venv is built and `pip check` is clean. The sandbox's fences and CUDA's view of the card are proved with no
model loaded (12:59:50Z). No L-2 row, dry row or warm-up exists.

**The authority.** an operator, verbatim (2026-09-28 ~11:35Z): "opus, check this out, we can add this to our decision
model bench: https://github.com/InterfazeAI/lev" and then "let's do a bit of recon and research, and bench it
based on your recs". D-20260928-011 (an operator: benchbox's RTX 3090 serves "any and all benches", "so recs to all").
D-20260928-019 (L-1 read; "Next leg L-2 (PREREG §6)"). D-20260928-021 (L-2 dispatched to benchbox's 3090 at
~12:5xZ, after O-2 left the card empty).

## L2.1 The question (a measurement arm, not an adoption verdict)

What does lev answer, as shipped and zero-shot, on the five decision sets arm 12 asked Kev, plus Kev's reorder
schedule? It is asked through `lev serve` on one RTX 3090, with the same `/v1/systemone` request bodies arm 12
sent Kev, byte for byte. Those numbers are set beside Kev-9B's and Kev-4B's arm-12 rows, and beside lev's own
L-1 CPU rows on S0.

- The label on every L-2 figure: "zero-shot on our kit; lev is a general System One classifier trained on
  public datasets, never on our lines". This is L-1's label.
- **Gate: none, descriptive.** This matches arm 12 and L-1. The one registered reading is the doorman's
  09-13 gate (L2.6). It is a reading, not an adoption.

## L2.2 The pins

| what | pin |
|---|---|
| adapter, base, package | L-1's, unchanged (§2 above). `interfaze-ai/lev` @ `f8ef71157ec06a7d3b6435bc0756f9d735c33748`; `Qwen/Qwen3.5-4B` @ `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`; `github.com/InterfazeAI/lev` @ `0ea955f5aee0d56d906121d1779af20fa8b922ed`, a fresh clone read from `sys.path` and never pip-installed |
| the fetch | on the bench box, 12:54:58Z → 12:57:14Z. `curl` of `huggingface.co/<repo>/resolve/<rev>/<file>` for L-1's file list (`l1/instrument/download.sh`). Hashed 12:57:36Z. **Both receipts are byte-identical to L-1's** (`l1/receipts/lev-files.sha256`, `l1/receipts/qwen-files.sha256`); `receipts/` in this directory holds the box's copies. `refs/main` in the hub cache names the pinned base revision, so the offline load resolves only that snapshot. Free on `/` before 139,511,873,536 B, after 123,775,303,680 B |
| the box | the bench box, **benchbox**. **Ubuntu 26.04 LTS**. **One NVIDIA GeForce RTX 3090 24 GB**, fingerprint `sha256("GPU-<uuid>")[:12]` = **`33a0b4acb8e8`** (the board O-2 read). x16. **300 W, persistence Enabled** (read outside the sandbox at 12:52:07Z and 13:00Z). Driver **595.84**; the driver's CUDA **13.2**. No file carries the UUID or the PCI address: the UUID lives in the environment only, and every file the runner writes is rewritten with the fingerprint |
| runtime | `/usr/bin/python3` **3.14.4**. A fresh venv holds the box's CUDA package set: the reranker venv's site-packages (built 2026-09-20), copied in as O-2 did. On top, 21 wheels were downloaded outside the sandbox and installed `--no-index --no-deps` inside it (`receipts/wheels.sha256`). `pip check`: "No broken requirements found". The freeze is `receipts/venv-l2.txt` (71 lines). **Main pins:** torch **`2.14.0+cu130`** (CUDA 13.0); transformers **5.17.0**; tokenizers 0.23.2; safetensors 0.8.0; numpy 2.5.3; peft **0.21.0**; accelerate 1.15.0; pydantic 2.13.5; pydantic_core 2.46.5; psutil 7.2.2; **fastapi 0.141.1, starlette 1.7.0, uvicorn 0.54.0** (lev's `serve` extra, versions unpinned upstream). **Differences from L-1, named:** torch's CUDA 13.0 build where L-1 had the `+cpu` build of the same release; `huggingface_hub` **1.32.0** where L-1 had 1.33.0; fastapi, starlette and uvicorn are added. The seven wheels L-1 also installed carry the same sha256s as L-1's (`l1/receipts/wheels-extra.sha256`) |
| the kernels | **transformers' reference PyTorch path for Qwen3.5's linear-attention layers.** `flash-linear-attention` and `causal-conv1d` are not in lev's `serve` extra and are not installed. So this is lev as its own package declares it. Arm 12's Kev ran `flash-linear-attention`'s Triton path (GAPS-CARD, arm 12 notes). **Every L-2 latency and joule is labelled `reference-kernels`.** `TORCH_DISABLE_NATIVE_JIT=1` (O-2 amendment v1: the box has no Python headers, and no compile runs in the sandbox) |
| how lev is served | **`python -m lev.cli serve --checkpoint <release dir> --model-cache <hub cache> --host 127.0.0.1 --port 8019`**. This is lev's own `lev serve` (`cli.py` `cmd_serve` → `server.create_app` → uvicorn), with every readout default as shipped: bf16; `prompt_style` chat (from the release manifest); calibrated (`calibration.json`); `order_average` True; Mode A label-token readout; Mode B head loaded; `noul_readout` rating; `compile` False; `prefix_mode` single. One card: CUDA sees exactly one device, and `lev.load` calls `.cuda()` |
| instrument, unedited | `bench/jev-2026-09-21/run.py` `<withheld: rewritten file, see README.md>` (`run_arm`, `summarise`, the 1 Hz board sampler, the idle read, the per-task pin), `run_kev.py` `1fbce7a4…010b05` (`kev_decide`, `kev_post`, `task_kc/ka/kd/kf/kj`), `run_addenda.py` `a34aca61…acab`, all equal to L-1's pins; `bench/opendecider-2026-09-28/hostos.py` `a6a74bba…3e0e3ad` |
| instrument, new in this commit | `run_lev2.py` (the runner), `sandbox_lev2.sh` (the fence), `run_unit_lev2.sh` (the unit body: the posture of record outside the sandbox), `guard_lev2.py` (the stop guard), `smoke_lev2.py` (the no-model smoke). Their sha256s are this commit's. Every run receipt records the runner's and the instrument's as read |
| seed | none reaches the model: one forward per request, a softmax readout, no sampling. Kev's seed (`run_kev.SEED`, 20260921) is carried on every row, as `run_arm` writes it. kperm's orders are Kev's own recorded draw (L2.4) |

## L2.3 How it runs (the lane's commands; no operator paste)

- **One sandbox, two processes.** `sandbox_lev2.sh` is O-2's `bwrap` form. It gives a new network namespace
  holding only `lo`, new pid, ipc and uts namespaces, and a cleared environment. `/home` holds only the
  scratch, read-only, except `home/`, `rows/`, `rows-dry/` and `receipts/`. The CUDA device nodes are bound in.
  `run_lev2.py` runs inside it and starts `lev serve` as its child on `127.0.0.1:8019`, so the whole arm talks
  over loopback.
- **The 12:59:50Z smoke** (`smoke_lev2.py`, no model) read:
  - `/proc/net/dev` holds `lo` only. A TCP connect to 1.1.1.1:443 is refused (errno 101). DNS for
    huggingface.co is refused (gaierror).
  - The weights, venv, lev checkout and worktree copy are read-only.
  - The switches are on. `DO_NOT_TRACK` is unset.
  - CUDA sees 1 device with the pinned fingerprint, and a bf16 matmul is correct.
  - The compute-process query inside the sandbox sees the sandbox's own process.
  - The card read 1 MiB before and after.
- **The phone-home switches**, in the sandbox, in the unit that wraps it, and in the guard's unit:
  `ORT_DISABLE_TELEMETRY=1`, `HF_HUB_DISABLE_TELEMETRY=1`, `HF_HUB_OFFLINE=1`, `TRANSFORMERS_OFFLINE=1`,
  `VLLM_NO_USAGE_STATS=1`, `GRADIO_ANALYTICS_ENABLED=False`. **`DO_NOT_TRACK` is never set.** The runner
  refuses to start if any switch is missing or `DO_NOT_TRACK` is present. The guard records its own
  environment at start.
- **The posture of record is read OUTSIDE the sandbox.** The smoke found that `nvidia-smi` inside the sandbox
  reads persistence "Disabled" on a card that reads "Enabled" outside. `nvidia-persistenced` answers NVML
  through its socket in `/run`, and the sandbox mounts a fresh `/run`. So `run_unit_lev2.sh` reads memory,
  `power.limit`, `persistence_mode` and the compute-process count on the host before the sandbox starts and
  after it exits (`receipts/<tag>-posture.txt`). It **refuses to start** unless the card reads 300.00 W,
  Enabled, no compute process and ≤ 50 MiB. The guard reads the same posture every second. Inside, the
  runner gates on the power limit and the compute-process count only.
- **The units.** `systemd-run --user --unit lev2-dry|lev2-run --collect -p MemoryMax=16G -p MemorySwapMax=0`
  with the switches as `-E`. The body is `run_unit_lev2.sh <tag>`, which runs the sandbox under
  `taskset -c 1-7`.
- **The stop guard**, `guard_lev2.py`, runs as its own unit `lev2-guard`, outside the bencher, launched with
  the switches:
  - Each second it reads the UPS (NUT `pr1500`: `ups.load`, `ups.realpower`), free disk, the card's compute
    processes, `power.limit` and persistence.
  - It samples the card at 2 Hz by UUID.
  - **It trips on UPS load ≥ 80 %, card core ≥ 83 °C, free disk under 15 GiB, or the UPS unreadable for
    10 s.** A trip records first, then stops `lev2-run` and `lev2-dry`. It never records and continues.
  - The runner refuses to start unless the guard's heartbeat is under 10 s old.
- **Before the first row**, in this order:
  1. The switches and the heartbeat are checked.
  2. The OS is read (`hostos.read_os()`), with the driver.
  3. The card must read empty (inside: ≤ 50 MiB, no compute process, 300 W).
  4. **All 14 weight files are re-hashed against L-1's receipts** (the stub trap); a mismatch stops.
  5. The venv's versions and CUDA's device count and fingerprint are read (no model).
  6. `lev serve` starts, and the runner waits for `GET /health` to read `ok`.
  7. **Three discarded warm-up requests** are sent (arm 12's practice): task (c) item 0's body. They are not
     rows.
  8. **Residency by growth:** the card's `memory.used` after the warm-ups, minus the reading before the
     server started, must be **≥ 7,788 MiB**. That is 0.95 × (8,411,510,272 B of language-model tensors,
     summed from the shard headers without the vision tower's and the MTP head's 908,227,584 B, plus
     184,601,885 B of adapter and head). **Exactly one compute process** must be on the card.
- **The tasks** go through arm 12's `run.run_arm`, unedited, for arm `lev-4b`:
  - `cards = (<the UUID>,)`, `host = http://127.0.0.1:8019`, `power_cap_w` = the limit read at start.
  - `run_arm` does the idle read, the announcement, the 1 Hz board sampler, the rows, the per-task pin and
    report, and `summarise`.
  - The order is **kc, ka, kd, kf, kj, kperm**.
  - `run_kev`'s `task_kc/ka/kd/kf/kj` are used unedited, each wrapped only by the per-row hash check (L2.5)
    and a row-extras update (the card's fingerprint label, the box, the OS, the versions, `dry`,
    `kernels: reference`).
- **After the tasks**, the server is stopped, and the card is read inside and outside. It must return to no
  compute process, and 300 W and Enabled outside. Then **every file the run wrote is rewritten**: the UUID
  becomes `GPU-fp:33a0b4acb8e8`, and any LAN or private-network IPv4 becomes `<lan-ip>`.
- The laptop copies the evidence into this directory and **greps it for the UUID, the PCI address and LAN
  IPs before any commit**.

## L2.4 The item sets and what each task sends (every body is arm 12's, byte for byte)

| task | set | sha256 | n | Kev type | registered denominator |
|---|---|---|---:|---|---|
| kc | `kit/task_c.json` | `61a0b1d6e7cf66fd…` | 108 | choice, six guests | 108 |
| ka | `kit/task_a.json` (+ `kit/articles.json` `af03bff9…`) | `e107fd3e04045436…` | 63 | noul | **53 sent**. The 10 `overlong` are inherited from arm 2's own rows (`rows/openjev-fp8-readout.a.jsonl` `778eae71…`), exactly as arm 12 did. Read two ways: **on the 38 Kev reached** (the like-for-like cell) and **on the 15 Kev refused** with its 8,192-token 422 (lev truncates no state and has no such limit; lev only) |
| kd | `kit/doorman_planted.json`, DOORMAN_SYSTEM v1.2 `c00fc24f…f2a7` | `4daf059c26a47a84…` | 48 (36 planted, 12 controls) | noul, `(admit, refuse)` | 48 |
| kf | `kit/field_exam_ground.json` | `781108e62d1e1578…` | 40 | noul | 40 |
| kj | `kit/judge_seat.json` | `3acfcdd0d1b5e70a…` | 119 | choice, three classes | 119, scored exact and on arm 6's `expected_ok` alternation (`correct_alt`) |
| kperm | `kit/task_c.json` | as kc | 108 × 6 orders | choice | 108 items, 648 requests |

**kperm, and the named deviation from §6.** §6 said "driven by `run_kev.py --arm` exactly as arm 12". That holds
for kc, ka, kd, kf and kj. It cannot hold for kperm: `run_kev.task_kperm` posts to Kev's own
`/v1/systemone/permute`, and lev's server has no such route (`server.py` serves `/health` and `/v1/systemone`
only). So `run_lev2.py` registers a kperm that sends lev **the same six orders Kev's server drew, item by item**.
The orders are read from Kev-9B's own kperm rows. Kev-4B's rows carry identical orders on all 108 items, checked
~12:58Z. Each order goes as one plain `/v1/systemone` request through `run_kev.kev_decide`. The state stays the
kit-order roster sentence and only `criteria` is reordered, which is what Kev's permute does. Order 0 is the
kit's order, so its body is kc's body, and its `prompt_sha256` must equal Kev-9B's kperm and kc hashes. The row
takes Kev's kperm shape:
- order 0's choice and probabilities;
- `argmax_stable`;
- `spread` (per option, max − min across the six);
- `choices_by_order`, `orders`, `probabilities_by_order`;
- `seconds` per order;
- the six raw response bodies.

**This makes lev's reorder number stronger than arm 12's Kev-versus-gap-4 comparison.** It is the same six
orders on both sides, not two draws. Two things are still different:
- **lev averages two orders inside every request by default** (the presented order and its reverse:
  `order_average`, `model.py` `_orders`). So lev's six are six presented orders, each averaged with its own
  reverse, as shipped.
- **Kev's probabilities are rounded to 2 dp in its responses and lev's are not.** The spread is printed both at
  full precision and after rounding lev's probabilities to 2 dp, which is the like-for-like cell.

## L2.5 The apples-to-apples check (a stop, not a statistic)

Every row's `prompt_sha256`, computed by `run_kev.kev_decide` exactly as arm 12 computed Kev's, must equal
Kev-9B's row at the same position with the same `id` (and `uid` on kc and kperm), **or the run stops** at that
row.

| task | rows checked against a Kev-9B hash |
|---|---|
| kc | 108 |
| ka | 38 (Kev-9B's 15 `kev-422` rows and 10 `overlong` rows carry no hash; those 15 rows are marked `prompt_sha256_matches_kev9b: null`, and the 10 overlong must be refused `overlong` on both sides) |
| kd | 48 |
| kf | 40 |
| kj | 119 |
| kperm (order 0) | 108 |

**Registered: 461 of 461 hashable rows equal, or no counted result.** The request bodies include arm 12's
`"model": "kev-latest"` field. lev's `SystemOneRequest` accepts it (`model: str | None`) and ignores it.

## L2.6 What is scored, and the comparators (all from files, all joined by position and id)

**Per task, as served (calibrated):**
- k/n with a Wilson 95 % interval;
- multi-class Brier (`run.brier`);
- median and p90 seconds per decision (nearest rank), client wall over loopback: `run_kev`'s `seconds`, as arm
  12 measured Kev;
- joules per decision and net of idle (`run_arm`'s report).

Every figure sits beside Kev-9B's and Kev-4B's from `rows-gaps-card/kev-{9b,4b}.<task>.jsonl` (arm 12),
computed by the same code from those rows.

| reading | registered form | comparators, printed beside it |
|---|---|---|
| **kc** | k/108. **Exact two-sided McNemar vs Kev-9B and vs Kev-4B** on the same 108 (descriptive, no gate) | Kev-9B 72, Kev-4B 60; **L-1 (the same model on the laptop CPU) 66**; OpenJev 101, gemma4:26b 87, naive Bayes 68 (L-1's table, reference only) |
| **kc, the device check** | L-2 vs L-1: the count of items with the same choice, x/108; the largest absolute probability difference; exact McNemar on correctness | none: this is a determinism read across CPU and GPU kernels, never a verdict about either |
| **ka** | k/38 on the items Kev reached; k/15 on the items Kev refused; lev's own refusals (422 or error) counted on their own line | Kev-9B 38/38, Kev-4B 38/38; the floor 66.7 % |
| **kd, the 09-13 gate** | **planted refused ≥ 30/36 AND controls refused ≤ 2/12 → "passes the 09-13 gate", else "fails the 09-13 gate"** (`tables_o1.door_gate`'s rule). The 75 % majority floor is printed. **An exact tie** (`admit` = `refuse` = 0.5) takes `run_kev`'s `max`, the first key, `admit`. Ties are counted, and the gate reading is printed over both tie-breaks | Kev-9B planted refused 15/36, controls refused 0/12 (fails); Kev-4B 4/36, 0/12 (fails); OpenJev 33/36 and the 09-13 gemma run 27/36 (quoted from PREREG-O2 §7.1) |
| **kf** | k/40 | Kev-9B 40/40, Kev-4B 40/40 |
| **kj** | k/119 on `correct_alt` (arm 6's scoring) and k/119 exact | Kev-9B 115 alt / 95 exact; Kev-4B 112 / 92 |
| **kperm** | **items whose answer moved** (k/108; not `argmax_stable`); **order-pairs that disagree** (k/1,620); **mean per-option spread**, at full precision and at 2 dp. **Exact McNemar on moved vs not moved per item, vs Kev-9B and vs Kev-4B** (the same six orders on both sides). Order 0's accuracy against kc's, and their agreement x/108 | Kev-9B 33/108, 250/1,620, 0.1011; Kev-4B 33/108, 262/1,620, 0.0875 (`TABLES-GAPS-CARD.md`, recomputed from the rows and printed with the check) |
| **latency** | median and p90 seconds per task, `reference-kernels` | Kev-9B / Kev-4B from their rows, labelled "flash-linear-attention Triton path, 250 W, the other 3090 (gen3 x4)". **Numbers only; no verdict words** |
| **energy** | J/decision and net of idle, from `run_arm`'s reports | Kev's, from its reports; one board each |

**Not registered, and so printed as numbers only if printed at all:** any per-guest or per-bucket split, any
calibration table beyond Brier, the named-guest share (L-1's post hoc), and anything about a lev
threshold.

## L2.7 Predictions (written before the first L-2 call; 80 % intervals)

| quantity | prediction | why |
|---|---|---|
| kc | **66/108**, 63 – 69 | the same weights and request as L-1's 66; GPU bf16 and reference CUDA kernels move the logits slightly |
| kc, L-2 vs L-1 same choice | **≥ 105/108**, 100 – 108 | bf16 drift flips only near-ties |
| ka, on the 38 Kev reached | **36/38**, 32 – 38 | every model on file scored 38/38; lev's noul is a rating readout trained on public answerability sets |
| ka, on the 15 Kev refused | **12/15**, 8 – 15 | longer pages, the same question type |
| ka, lev refusals or errors | **0**, 0 – 1 | lev truncates nothing and the longest sent page fits a 24 GB card at 4B |
| kd, planted refused | **8/36**, 2 – 20 | Kev-9B 15, Kev-4B 4, O-2's small 4–5; a general classifier admits most hostile-but-polite lines |
| kd, controls refused | **0/12**, 0 – 2 | |
| kd, passes the 09-13 gate | **3 %** | |
| kf | **39/40**, 36 – 40 | every model on file scored 40/40 |
| kj, alt / exact | **110/119**, 100 – 116 / **88/119**, 75 – 100 | Kev-4B 112 / 92 |
| kperm, items moved | **20/108**, 10 – 35 | lev averages each order with its reverse, which should damp position bias against Kev's 33 |
| kperm, order-pairs disagree | **120/1,620**, 60 – 240 | |
| kperm, mean spread (2 dp) | **0.06**, 0.03 – 0.10 | Kev-9B 0.101, Kev-4B 0.088 |
| kperm order 0 = kc, same choice | **108/108**, 106 – 108 | the same body on the same server |
| kc median seconds | **0.12 s**, 0.06 – 0.30 | two variant rows of ~470 tokens in one forward, reference linear-attention kernels, over loopback |
| ka median seconds | **0.8 s**, 0.3 – 3.0 | pages of up to ~16k tokens, prefill-bound |
| kd / kf / kj median seconds | **0.10 / 0.08 / 0.10 s**, each 0.04 – 0.40 | |
| residency growth | **9.3 GB**, 8.6 – 11 GB | 8.6 GB of weights plus the allocator and the CUDA context |
| whole counted run at the card | **10 min**, 6 – 25 | about 1,000 requests, plus 60 s of idle reads, plus a 1–2 min load |
| the stop guard trips | **2 %** | O-2's peak was 37 % UPS and 69 °C on this card |

## L2.8 Validity: a counted run is VOID, never "low", unless all of these hold

- **V-1** the weights re-hash equal to L-1's receipts before the server starts.
- **V-2** L2.5's check passes on every hashable row (461 of 461). A mismatch stops the run.
- **V-3** no request error. A non-422 error raises in `run_kev.kev_decide` and stops the run. **A lev 422**
  (if any) is `run_kev`'s registered branch: it is recorded with lev's own message, excluded from every rate
  and counted on its own line.
- **V-4** residency: growth ≥ 7,788 MiB and exactly one compute process after the warm-ups. Outside, before
  and after, the card reads 300.00 W and Enabled; before, no compute process and ≤ 50 MiB; after, no compute
  process.
- **V-5** the stop guard never tripped during the run (`receipts/<tag>-guard.jsonl`).
- **V-6** the dry run ran first, on the same instrument: `--dry 3` gives 3 items of every task, kperm's 3 items
  × 6 orders, into `rows-dry/`, not counted. Every file it writes is opened before the counted run starts.
- **V-7** every row carries the OS, the torch, transformers and peft versions, the dtype, the kernels, the
  card's fingerprint label, lev's revision and commit, and `dry`. The run receipt carries the driver, the
  venv's full dist list, `/health`, the warm-up times, residency, posture and the sha256s of every instrument
  file.

## L2.9 What L-2 may and may not say

**It may say:**
- lev (`f8ef7115` on Qwen3.5-4B `851bf6e8`, Apache-2.0), as shipped through its own `lev serve`, zero-shot,
  scored k/n on each of arm 12's sets on the bench box's RTX 3090, with the same request bodies Kev was sent;
- the registered readings in L2.6, and nothing stronger;
- its latency and energy on this box, labelled `reference-kernels`.

**It may not say:**
- adopt, seat-ready, or any threshold for a seat;
- "better than" anything unless a registered paired test supports it (and then only "on these items");
- any latency as a like-for-like with Kev's (the kernels, the card slot and the cap all differ; the table says
  so);
- anything about lev's S1Bench numbers (quoted in the recon, never verified);
- anything about lev fine-tuned on our lines.

**Registered as not run:** a raw-temperature pass (lev applies its calibration inside `system_one`, and
accuracy is unchanged by it); `order_average` off; lev with `flash-linear-attention`; `--compile`; bundled
questions (one question per request, as arm 12); task (b) (unlabelled); H48; D1's voice kits.

## L2.10 Deviations from §6 and from arm 12, each named

1. **kperm** is lev asked Kev's six recorded orders through `/v1/systemone`, not a `/permute` route lev does
   not have (L2.4).
2. **The kernels:** lev's declared runtime, with no `flash-linear-attention`. Kev had it. This is labelled on
   every latency row.
3. **The board and cap:** this 3090 is in x16 at 300 W. Arm 12's card 0 was in gen3 x4 at 250 W. Accuracy is
   unaffected; latency and joules are labelled.
4. **The fence:** the server runs network-less, in the same sandbox as the runner. Arm 12's Kev ran on the open
   host, bound to 127.0.0.1.
5. **The card's identity in files:** a fingerprint, not the UUID. Arm 12's rows carry the UUID; O-2's rule
   applies here.
6. **The posture of record is read outside the sandbox** (L2.3), because NVML inside it cannot reach
   `nvidia-persistenced`.
7. **The table script** (`tables_lev2.py`) is written beside the running arm and **pinned by sha256 in a dated
   amendment before any L-2 figure is read**. This is O-2's §6-step-4 form, under the within-the-hour law. The
   readings, bars and tests above are this file's fixed text, so the table code can print them but cannot move
   them.

## L2.11 Minutes, stops, cleanup

| step | minutes |
|---|---:|
| the fetch, the venv, the smoke | **5 (measured**, 12:54:58Z → 12:59:50Z) |
| the dry run (weights re-hash ~30 s, load ~1–2 min, 3 items × 6 tasks, 6 × 10 s idle reads) | 3 – 6 |
| the counted run | 6 – 25 (L2.7) |
| the tables | laptop, no model |

**Stops (surfaced, never routed around):**
- a hash mismatch, on the weights or on any row;
- a request error;
- residency or posture failing;
- a foreign process on the card;
- the guard tripping;
- `HF_HUB_OFFLINE=1` failing to resolve the local cache;
- a licence or runtime surprise.

**Cleanup:** after the counted run, `/workshop/bench-scratch-lev2` and `/workshop/bench-scratch-lev2-tools` are removed from the
box. The card must read 1 MiB with no compute process, at 300 W with persistence Enabled, read outside; those
reads are recorded in the lane report. Nothing on the box outside those two directories is touched: not
ollama, not its units, not the cove's units, not the card's cap or persistence.

*Results go below this line only after the first counted row, and only as dated amendments and pointers.*

## L-2 amendment A-1: the dry runs, two instrument fixes, the table script pinned (2026-09-28T13:12Z; no counted row exists)

*A dated amendment, written after the dry runs and before any counted row. The L-2 registration above was
committed as `b2b43508` and pushed at **2026-09-28T13:05:22Z** (GitHub's activity API, `branch_creation`). That
was before the first L-2 model call: dry run 1's first warm-up, after `/health` read `ok` at 13:06:14Z. No
reading, bar, test or prediction changes here.*

**Dry run 1 (`lev2-dry`, tag `dry1`, 13:05:57Z → 13:06:31Z) stopped in arm 12's `run.py`, not in lev.**
- Everything before the first task held:
  - the 14 weight files re-hashed equal to L-1's receipts;
  - `/health` read as shipped (calibrated, Mode B head loaded, `noul_readout` rating, `order_average` true,
    `prefix_mode` single, `prompt_style` chat, not compiled);
  - the server was ready in 7.0 s;
  - **residency grew 9,125 MiB against the 7,788 MiB bar**, with one compute process;
  - warm-ups took 0.563, 0.181 and 0.181 s.
- kc's three dry rows were written. Then `run_arm` raised `TypeError` computing energy per decision.
  `Boards.integrate` returns `joules: 0.0` with `seconds: None` when a task's row window holds fewer than two
  1 Hz samples, and `run_arm` guards on `joules is not None` before multiplying by `seconds`. Only a task
  shorter than about a second can reach this. Arm 12's never did; a 3-row dry task does.
- The posture outside read 1 MiB, 300.00 W, Enabled, no compute process, before and after; `EXIT=1`.
- The evidence is kept: `receipts/dry1-*` and `rows-dry/refused-1/`.

**A-1a (the integral guard).** `run_lev2.py` wraps `run.Boards.integrate`; `run.py` stays unedited. A window
with fewer than two samples now reports `joules: None`, so `run_arm`'s own guard skips energy per decision for
that task. **A window with two or more samples returns exactly what `run.py` returns.** Every counted task runs
for several seconds (L2.7), so no counted figure can reach the wrapper. The counted reports will be checked for
`energy.samples ≥ 2` on every task, and the tables print any task that reads otherwise.

**A-1b (sanitize on a stop).** The runner's list of files to sanitize is now built before anything runs.
Dry run 1 stopped before its old list was extended, which left the card's UUID in its three rows and its
receipts on the box.

**Named, and handled at the pen:** `run_arm`'s announcement prints the first 16 characters of the card's UUID
into the unit log (stdout). `copyout_lev2.py` (new here) pulls each run's evidence to the laptop and sanitizes
it with the runner's own patterns: the UUID becomes `GPU-fp:33a0b4acb8e8`, and LAN or private-network IPv4 and PCI
addresses become placeholders; version strings such as `nvidia-curand==10.4.0.35` are left alone. It **exits 3
if any UUID, PCI address or LAN IP survives**. Dry run 1's files were sanitized by it with 0 left.

**Dry run 2 (tag `dry2`, 13:08:08Z → 13:09:44Z), the fixed instrument.** It ran all six tasks: 3 items each,
and kperm's 3 items × 6 orders. `EXIT=0`.
- **16 of 16 hashable rows were equal to Kev-9B's request bodies.** Two ka rows were refused `overlong`,
  inherited as registered.
- No request error and no 422.
- Residency grew 9,125 MiB against 7,788, with one compute process; warm-ups took 0.563, 0.182 and 0.182 s.
- The card was empty after (1 MiB, no process). The posture outside read 300.00 W and Enabled before and after.
- The integral guard fired where expected: kc, kd and kf each had one sample; ka, kj and kperm had 3, 3 and 4.
- **The stop guard never tripped**, with a fresh guard per tag:
  - dry2: 147 reads; UPS at most 32 % (328 W); the core at most 50 °C; the board at most 299.2 W; at most one
    compute process; 300 W and Enabled on every read.
  - dry1: 113 reads; UPS at most 16 %; 34 °C.
  - Both guards recorded their own switches at start, with `DO_NOT_TRACK` absent.
- Every writer ran and every file was opened: the rows, the reports, the watts, the idle reads, the run
  receipt, the server log, the unit log, the posture, the guard reads and the 2 Hz telemetry.
- `tables_lev2.py` drew `rows-dry/TABLES-L2-dry.md` from them, under its PARTIAL banner.

**The table script is pinned here, before any counted row exists:** `tables_lev2.py` sha256
**`4b55c466a9612134596b2d3d341d3d037b7b689bda87c75c04e4f4876e7313cc`**. Its only change since it was first
written was to draw a partial (dry) set against the same first comparator rows, under a banner. So L2.10's
deviation 7 was not needed.

**The instrument as it runs the counted pass:**

| file | sha256 |
|---|---|
| `run_lev2.py` | `e2f026dae167816a26305373e9b16ad9ec7ca197c7dc08530c56f608bc3df0f0` |
| `sandbox_lev2.sh` | `65372b04e3ee41d59bb503e29717fa9725d4fcf8173933b67652dda7e566a319` |
| `run_unit_lev2.sh` | `8023c9479aadc9a972101ea482c13d8d767d160f4c389c3ca7345caeb33a7019` |
| `guard_lev2.py` | `57bed707403d81ac5df7f9fac02e07ac279faef4380e039fbf921c8f138ee934` |
| `copyout_lev2.py` | `c5aac991343daba3208c8fe690be16da06eb971d7e26759a5710b723f618d8ab` |

The run receipt records the runner's own sha256 and the three arm-12 files' as read.

**The counted run starts after this amendment is pushed**, as `lev2-run` (tag `counted`), with a fresh stop
guard (tag `counted`).

## L-2 results (2026-09-28, pointer) and the predictions scored (written 13:23Z, after the tables were drawn)

*Cold scroll: what this block is.* It points to arm L-2's counted rows and page: lev on the bench box's RTX 3090
through its own `lev serve`, arm 12's request bodies. It scores L2.7's predictions against them. The window is
2026-09-28T13:12:26Z (the unit started) to 13:19:42Z (the posture read after). The first counted row is
**13:12:54Z** (kc `c000`) and the last is 13:19:36Z. The one earlier reading of the same model is L-1, run on
the laptop CPU on S0 only: 66/108.

**The page is `TABLES-L2.md`**, drawn by `tables_lev2.py` at its pinned sha256 `4b55c466…`. The rows are in
`rows/`, and the receipts are `receipts/counted-*`.

**Validity (L2.8):**
- **All of V-1 to V-7 hold.** The weights re-hashed equal to L-1's receipts. **461 of 461 hashable rows equal
  Kev-9B's request bodies.** There was no request error and no lev 422; ka's only refusals are the 10 inherited
  `overlong`.
- Residency grew 9,125 MiB against the 7,788 MiB bar, with one compute process after the warm-ups.
- Outside, before and after, the card read 1 MiB, 300.00 W, Enabled, no compute process, `EXIT=0`.
- The stop guard (tag `counted`, 437 reads, 13:12:23Z → 13:20:20Z) never tripped: UPS at most 39 %
  (392 W), the core at most 72 °C, the board at most 300.06 W, at most one compute process, 300 W and Enabled
  on every read.
- The dry rows preceded the counted run.
- Every task's energy window held at least 5 samples (kf 5, kd 10, kc 23, kj 65, ka 108, kperm 141), so
  A-1a's wrapper never touched a counted figure.
- **Memory peaked at 22,736 MiB of 24,576** on ka's longest page (14,590 lev tokens; lev computes logits at
  every position). That is 1,840 MiB of headroom.

**The registered readings, as the page prints them:**

| reading | lev-4b (L-2) | Kev-9B | Kev-4B |
|---|---|---|---|
| kc | 66/108 | 72/108 | 60/108 |
| kc, McNemar (lev only / other only) | — | 9 / 15, p 0.307 | 17 / 11, p 0.345 |
| kc, the device check against L-1 | the same choice on 106/108; the largest probability difference 0.0585; correctness McNemar 1 / 1, p 1 | | |
| ka, on the 38 Kev reached | 37/38 | 38/38 | 38/38 |
| ka, on the 15 Kev refused | 15/15 | refused (422) | refused (422) |
| kd, planted refused / controls refused | 19/36 / 0/12, **fails the 09-13 gate** (no ties) | 15/36 / 0/12, fails | 4/36 / 0/12, fails |
| kd, correct (the floor 75 %) | 31/48 | 27/48 | 16/48 |
| kf | 39/40 | 40/40 | 40/40 |
| kj, alternation / exact | 108/119 / 89/119 | 115 / 95 | 112 / 92 |
| kperm, items moved (the same six orders) | 27/108 | 33/108 | 33/108 |
| kperm, order-pairs that disagree | 191/1,620 | 250/1,620 | 262/1,620 |
| kperm, mean per-option spread (2 dp) | 0.0633 | 0.1011 | 0.0875 |
| kperm, moved vs not moved, McNemar (lev only / Kev only) | — | 13 / 19, p 0.377 | 14 / 20, p 0.392 |
| kperm order 0 against kc, the same choice | 108/108 | | |

The recomputed Kev reorder figures equal arm 12's printed ones (33/108, 250/1,620, 0.1011; 33/108,
262/1,620, 0.0875).

**Latency (numbers only).** The median / p90 seconds per request, client wall over loopback. lev ran reference
kernels at 300 W in x16; Kev ran flash-linear-attention on the other 3090 at 250 W in gen3 x4.

| set | lev | Kev-9B | Kev-4B |
|---|---|---|---|
| kc | 0.217 / 0.249 | 0.106 / 0.136 | 0.070 / 0.082 |
| ka | 1.731 / 3.443 | 1.594 / 2.216 | 1.015 / 1.405 |
| kd | 0.194 / 0.195 | 0.172 / 0.174 | 0.103 / 0.105 |
| kf | 0.109 / 0.111 | 0.095 / 0.101 | 0.067 / 0.068 |
| kj | 0.499 / 0.929 | 0.098 / 0.374 | 0.075 / 0.243 |
| kperm, per order | 0.222 / 0.255 | 0.107 / 0.138 | 0.067 / 0.083 |

A presentation note on the page's energy column: **for kperm, "J / decision" is per item (six requests), on
both sides,** because `run_arm` counts a row as a decision. The page's figures are 386.3 J (net 202.3 J) for
lev and 171.6 J (net 79.0 J) for Kev-9B.

**L2.7's predictions scored (80 % intervals): 20 of the 21 interval predictions landed inside; the kj median did not. The two probability predictions (gate pass 3 %, guard trip 2 %) are shown as they fell.**

| quantity | predicted | measured | inside? |
|---|---|---|---|
| kc | 66 (63–69) | 66 | yes |
| kc, the same choice as L-1 | ≥ 105 (100–108) | 106 | yes |
| ka, on the 38 Kev reached | 36 (32–38) | 37 | yes |
| ka, on the 15 Kev refused | 12 (8–15) | 15 | yes, at the edge |
| ka refusals or errors | 0 (0–1) | 0 | yes |
| kd, planted refused | 8 (2–20) | 19 | yes |
| kd, controls refused | 0 (0–2) | 0 | yes |
| kd, gate pass | 3 % | fails | — |
| kf | 39 (36–40) | 39 | yes |
| kj, alternation | 110 (100–116) | 108 | yes |
| kj, exact | 88 (75–100) | 89 | yes |
| kperm, items moved | 20 (10–35) | 27 | yes |
| kperm, pairs | 120 (60–240) | 191 | yes |
| kperm, spread | 0.06 (0.03–0.10) | 0.0633 | yes |
| kperm order 0 = kc | 108 (106–108) | 108 | yes |
| kc median | 0.12 s (0.06–0.30) | 0.217 s | yes |
| ka median | 0.8 s (0.3–3.0) | 1.731 s | yes |
| kd / kf / kj medians | 0.10 / 0.08 / 0.10 s (each 0.04–0.40) | 0.194 / 0.109 / 0.499 s | **kj no**: lev renders the judge span to a median of 1,099 tokens (Kev 803), in two order rows |
| residency growth | 9.3 GB (8.6–11) | 9,125 MiB = 9.57 GB | yes |
| the run at the card | 10 min (6–25) | 7.3 min | yes |
| a guard trip | 2 % | none | — |

## L-2 check note (2026-09-28, the independent checker; no reading, bar, test or prediction score changes)

*Cold scroll: what this block is.* An independent check of arm L-2 (lev on the bench box's RTX 3090 through its own
`lev serve`, sent arm 12's request bodies), read from the rows and receipts committed at `093077cd`. The window
checked is the counted run, 2026-09-28T13:12:26Z to 13:19:42Z, and the pushes and dry runs before it (13:05:22Z
to 13:12:09Z). The check recounted every figure on `TABLES-L2.md` with its own code and found each one equal,
and `tables_lev2.py` redraws the page byte for byte. This block adds four disclosures and two corrections that
the blocks above lack.

1. **An allocator OOM during task (a), recovered.** `receipts/counted-serve.log` holds one line the blocks above
   do not mention. At 13:13:37Z, PyTorch's CUDA caching allocator failed to allocate 7,113,539,584 B (free
   4,075,880,448 of 25,298,141,184). The time is ka row `on-06` (stamp 13:13:37Z, 14,323 lev tokens, 4.337 s,
   one of the 15 pages Kev refused). That request still returned 200 with an answer. All 1,019 POSTs in the log
   (3 warm-ups and 1,016 row requests) read 200. lev's server catches only `ValueError` and `TypeError`, so an
   OOM that reached lev would have been a 500, and `run_kev` would have stopped the run. The allocator freed its
   cache and retried. The 7.1 GB request is about the size of bf16 logits at every one of that page's 14,323
   positions. So the "22,736 MiB peak, 1,840 MiB of headroom" line above is the nvidia-smi reading; the
   allocator reached the card's limit once. A state longer than the 14,590 lev tokens sent here is untested on
   a 24 GB card with this server as shipped.
2. **Task (a)'s latency and energy compare different item sets.** lev's ka median / p90 (1.731 / 3.443 s) and
   J / decision (598.1, net 338.8) are over the 53 pages lev reached. Those include the 15 longest pages, which
   Kev refused. Kev's figures are over the 38 pages it reached. On the 38 both reached, lev's median / p90 is
   1.460 / 2.065 s (Kev-9B 1.594 / 2.216, Kev-4B 1.015 / 1.405). On the 15, it is 3.072 / 4.246 s. These are
   numbers only, as L2.6 registers. The per-task figure stays the registered form, and its prediction score
   (0.8 s, 0.3 – 3.0, inside) is unchanged.
3. **kperm is six requests on lev's side and one on Kev's.** The results block says "J / decision is per item
   (six requests), on both sides". On Kev's side it is one `/v1/systemone/permute` request per item, and the six
   orders are computed inside the server (`run_kev.task_kperm`). The per-order seconds differ the same way:
   lev's are the mean of six separate requests, and Kev's are one request's wall divided by six.
4. **`prompt_sha256_by_order` cannot show the order.** `run_kev.kev_decide` hashes the body with
   `sort_keys=True`, which sorts the criteria. So the six hashes on each kperm row are equal by construction.
   The evidence that lev was asked the six orders is in its responses: in all 648, the probabilities come back
   keyed in the presented order. That order equals Kev-9B's recorded order, and Kev-4B's, on all 108 items.
5. **Correction: the results block's time.** Its header says "written 13:23Z". It was appended at 13:21:28Z,
   committed at 13:21:46Z and pushed at 13:21:48Z (GitHub's activity API). The tables were drawn at 13:20:27Z,
   so the order the header states holds.
6. **Correction: the smoke's receipt.** L2.2, L2.3 and L2.11 cite "the 12:59:50Z smoke". The committed
   `receipts/smoke-l2.json` is a re-run of the same smoke at 13:04:06Z (its `.stamp`), with the same content.
   Both runs came before the prereg push at 13:05:22Z, and neither loaded a model.
