# Arm A-1: APUS-9B (apus-ailab), BF16, the vendor's reference runtime in-process, on benchbox's RTX 3090, beside Kev and lev: pre-registration

*Cold scroll: what this is.* The pre-registration of arm A-1: **APUS-OpenJev-v1-9B** (a merged LoRA fine-tune of
Qwen3.5-9B, Apache-2.0, published by the HF org `apus-ailab` on 2026-09-22), zero-shot. It is asked the same
`/v1/systemone` request bodies Kev (arm 12) and lev (L-2) were asked, byte for byte, through APUS's own bridge
and APUS's own reference runtime, on the bench box's one RTX 3090. Written 2026-09-28 between 14:30Z and 14:34Z
(UTC, `date -u` on the dev laptop). It is committed and pushed on the estate branch `bench/apus-2026-09-28`
**before any A-1 model call**. The push time is read from GitHub's API and stated in the first amendment below.
The recon behind it is `RECON-apus.md` beside this file: the archive copy
`/workshop/bench-archive/plans-2026-09-28/apus/RECON-apus.md`, byte for byte, sha256
`43eadd3c2df8a0c11a248097cc08fc2393d7df7acbd147485a98d6283d215a44`.

At commit time no A-1 model call has happened, no weight has been fetched, and no A-1 instrument file exists.
The weights, the venv and the no-card checks come after this push, and their receipts land in a dated
amendment before the card is taken.

**The authority.** an operator, verbatim, 2026-09-28 ~14:0xZ: "love it, let's keep rolling with your recs / benchs,
and YES, let's test APUS and anything else you rec, and keep the train rolling, thx!" Ledger row
D-20260928-024 (estate `origin/master` `32ec0b35`). The lead's rec is the recon's arm A-1 (RECON §10). The
card order is the ledger's: A-1 runs when the lead hands it the card.

**INTERNAL.** This file names a box. H48 *figures* may be printed; H48 *lines* never are, in this file or in
any committed row.

## A1.1 The question (a measurement arm, not an adoption verdict)

What does APUS-9B answer, as shipped and zero-shot, on our six-way S0, on the held-back H48, on the doorman's
planted set, and on the rest of arm 12's decision sets? It is asked through the vendor's own TypeSafe bridge
and reference runtime, with the request bodies arm 12 sent Kev.

- **The headline read is `effort="high"`** (all 32 layers, the full LM head). **`effort="low"`** (the trained
  exit at layer 16) is a second, descriptive read on kc, kh and kd.
- The label on every A-1 figure: "zero-shot on our kit; APUS is a general decision model trained on public
  datasets, never on our lines".
- **The only registered verdict is the doorman at `high` against the 09-13 gate (A1.6).** Everything else is
  numbers. No threshold, no seat reading, no adoption word.

## A1.2 The pins

| what | pin |
|---|---|
| model | `apus-ailab/APUS-OpenJev-v1-9B` @ **`82c9c56cfa9de8d36704ed91948d4726ef111635`** (checkpoint-3000; `lastModified` 2026-09-23T10:26:17Z). 39 files, 18,839,917,388 B. **Every file's sha256 is pinned in `receipts/pins-9b.json`**, committed with this file: the 7 LFS files by HF's own LFS oid; the 32 small files by the recon's raw fetches at `resolve/82c9c56c…/<file>`, each of which equals HF's `blobId` as a git blob sha1 (checked between 14:23Z and 14:30Z) |
| the six shards | `model-00001` 2,034,237,568 B `dd63614f1dc80dce2d83be3f3e69af1f8a0ebc8add9b8e0abf3910f563e44344` · `-00002` 3,999,615,808 B `3fe0ee3d29088f7f84ac7ba8c9d7562b91e6a31841c5b585c71779e618519e99` · `-00003` 3,997,274,128 B `7bfce51b034a6de02c513b032c97007532676c5f915d020aa9a66f399f437117` · `-00004` 3,997,290,904 B `f36d2b3ffc4c44dab06277ebbd722f73faca80aefc9aa3ce1119575f4a4773ae` · `-00005` 3,991,239,264 B `b06602342e98eb882703857f3c8e8058894c03d4f2d62ddaab1e6dd4913c4e9b` · `-00006` 800,062,816 B `db688e61fd7575e8537120bc7aaeff043d31d6ca2bafa14df1135c969a081013`. `tokenizer.json` 19,989,325 B `06b9509352d2af50381ab2247e083b80d32d5c0aba91c272ca9ff729b6a0e523` (the same bytes as lev's) |
| the runtime | the snapshot's own `openjet_runtime/` (`runtime.py`, `early_exit.py`, `candidate_projection.py`, `contracts.py`, `__init__.py`; pinned in `receipts/pins-9b.json`), read from the snapshot on `sys.path` and **never edited**. It refuses to run unless `transformers.__version__ == "5.16.1"` (enforced in `runtime.py`) |
| the bridge | `deployment/contracts.py` from the family repo `apus-ailab/APUS-OpenJev-v1` @ **`af276da53a2e7a6d0fbdbafc9280bf9982b2d3a6`**, 11,219 B, sha256 **`7cb3e56d42651087e35b7a620b2407dc96801a09ae942bbe803cb202ee187e65`** (Apache-2.0, `deployment/LICENSE` `c71d239d…`), never edited |
| base (named only) | `Qwen/Qwen3.5-9B` @ `c2022362` (`provenance/identity.json`). Not downloaded: the checkpoint is merged |
| the fetch | on benchbox, `curl` of `huggingface.co/<repo>/resolve/<rev>/<file>` for the 39 files into an HF-cache layout in the lane's own scratch (`/workshop/bench-scratch-apus-a1/weights/hf/models--apus-ailab--APUS-OpenJev-v1-9B/snapshots/82c9c56c…/`, `refs/main` = the revision), plus the bridge file. No `.pyc` (the 9B repo carries none; any `*.pyc` found is refused). Every file is hashed after the fetch and must equal its pin, or the lane stops |
| **the disk gate** | before the fetch, `/bin/df` on benchbox. **Free space after the fetch must stay ≥ 90 GB** (the ovv release's ~65 GB reserve plus staging), counting the OJ-GGUF prep lane's ~16.5 GB as still to land unless its files are seen landed. If the fetch would break that, **nothing is pulled** and the lane reports READY-EXCEPT-PULL |
| runtime env | `/usr/bin/python3` **3.14.4**. A fresh venv holds benchbox's CUDA package set, the reranker venv's site-packages (torch **`2.14.0+cu130`**, tokenizers 0.23.2, safetensors 0.8.0, huggingface_hub 1.32.0, numpy 2.5.3; built 2026-09-20), copied in as L-2 and O-2 did. **One swap:** transformers 5.17.0 → **5.16.1** (`transformers-5.16.1-py3-none-any.whl`, 12,080,592 B, sha256 `2f2d5b98a5ad3718713653734298fa620754ed683702a635ebb587df3ed29c7e`, PyPI). Its `Requires-Dist` list is identical to 5.17.0's. Added, each with L-2's own wheel sha256 (`bench/lev-2026-09-28/receipts/wheels.sha256`): jinja2 3.1.6, markupsafe 3.0.3, packaging 26.3, pyyaml 6.0.3, rich 15.0.0, markdown-it-py 4.2.0, mdurl 0.1.2, pygments 2.21.0, certifi 2026.7.22, idna 3.20, setuptools 84.0.0. No peft, accelerate, vLLM, fastapi or typesafe-sdk. Wheels are downloaded outside the sandbox and installed `--no-index --no-deps`. `pip check` must read clean. The full freeze is a receipt |
| **named deviation** | torch **2.14.0+cu130** against the vendor's `torch==2.8.0` pin (`requirements.txt`; not enforced in code). The transformers pin **is** met. torch 2.8.0 publishes no cp314 wheel, so the vendor's exact torch would need a second CPython (benchbox has a uv-managed 3.12.14). **If the no-card checks show transformers 5.16.1 failing on torch 2.14, the lane stops and reports NOT READY.** A torch-2.8 venv would be a new registration, not a mid-run fallback |
| the kernels | transformers' reference PyTorch path for Qwen3.5's linear-attention (GDN) layers. `flash-linear-attention` and `causal-conv1d` are not installed (the vendor's `requirements.txt` names neither). `attn_implementation="sdpa"` as `runtime.py` sets it. `TORCH_DISABLE_NATIVE_JIT=1` (O-2 amendment v1; L-2: the box has no Python headers). Every latency and joule is labelled `reference-kernels` |
| instrument, unedited | `bench/jev-2026-09-21/run.py` `<withheld: rewritten file, see README.md>` (`run_arm`, `summarise`, `brier`, the 1 Hz board sampler, the idle read, the pin), `run_kev.py` `1fbce7a41746ed3d612a8a488e35dbebba789bc55086daa7d71ffce2cd010b05` (`kev_decide`, `task_kc/ka/kd/kf/kj`), `run_addenda.py` `a34aca61833ed0c5860c91c7e2e227429e8ea9e2e50fec68d7b7fec15d89acab`, `bench/opendecider-2026-09-28/hostos.py` `a6a74bba57250be1936825466c6614876d757cf81373fcf6fa0b48587cf0e3ad`. All equal to L-2's pins |
| instrument, new (after this commit) | `run_apus.py` (the runner and the transport), `setup_apus.sh` (the fetch, the disk gate, the venv), `sandbox_apus.sh` (the fence), `run_unit_apus.sh` (the posture of record, outside the sandbox), `guard_apus.py` (the stop guard), `smoke_apus.py` (the no-card smoke), `copyout_apus.py` (the sanitizing copy-out), and `tables_apus.py` later. L-2's forms, renamed. Their sha256s are pinned in a dated amendment before the first model call |
| the box | **benchbox**, **Ubuntu 26.04 LTS**, kernel 7.0.0-31-generic (read 14:22Z), driver **595.84**, 16 CPUs, 32.5 GB RAM. **One NVIDIA GeForce RTX 3090 24 GB**, fingerprint `sha256("GPU-<uuid>")[:12]` = **`33a0b4acb8e8`** (the 3090 board; L-2's board). x16. **300 W, persistence Enabled** (read outside, ~14:29Z, while another lane held it at 22,455 MiB). No file carries the UUID or the PCI address: the UUID lives in the environment only |
| seed | none reaches the model: one forward per request, a softmax readout, no sampling. `run_kev.SEED` (20260921) is carried on every row, as `run_arm` writes it |

## A1.3 How APUS is asked (the one swap)

**The transport.** `run_kev.kev_post` (an HTTP POST) is swapped for an in-process call. The request body that
`run_kev.kev_decide` builds is passed **unedited** through these steps:

1. **`prepare_request(body, model_name="kev-latest", max_questions=16, allow_score=True)`** (the bridge file).
   `model_name` is a parameter, not an edit: the bridge accepts `body["model"]` only if it equals `model_name`
   or `"jev-latest"`, and Kev's bodies say `"kev-latest"` (RECON §4). The model field never reaches the prompt.
   - A Choice becomes candidates in the body's order: `{"id": key, "description": name}`.
   - A Noul with custom true/false descriptions keeps the canonical yes/no candidates. The instruction becomes
     `{"question": <instructions>, "meaning_of_answers": {"true": …, "false": …}}` (the doorman, task (a),
     the field exam).
2. **`compile_question(q, tokenizer, max_tokens=8191)`** (the same file; the gateway's compiler and its
   budget: "8,191 compiled prompt tokens per question (one output token reserved)"). The prompt is
   `jev.dynamic.prompt.v2` inside the tokenizer's chat template with `enable_thinking=False`. It checks that
   each candidate letter is one token at the answer boundary. **A prompt over 8,191 tokens is refused with
   the bridge's own 422, never truncated.**
3. **`OpenJet.decide(q.record, effort)`** (the snapshot's runtime, loaded once by `OpenJet.from_pretrained(
   <snapshot>, device="cuda:0", dtype="bfloat16")` as shipped). At `high` this is one forward over the whole
   model with `logits_to_keep=1`, reading the candidate letters' logits from the full LM head. At `low` it is
   the runtime's early exit at layer 16, with the final norm and the candidate rows of the head.
4. **`probabilities(<decide's candidate logits>)`** and **`answer(q, probs)`** (the bridge file). This is what
   the vendor's gateway does with a backend's candidate scores: a float64 softmax, then the TypeSafe answer
   (`noul` = P(yes); a choice's `probabilities`, `choice` and normalized-entropy `confidence`). Softmax is
   shift-invariant, so candidate logits and candidate logprobs give the same distribution. The runtime's own
   float32 `probabilities` are kept on the row beside it.
5. A TypeSafe-shaped body goes back to `kev_decide`: `answers`, `usage.input_tokens` (the compiled length),
   `latency_ms`, and an `apus` block (effort, executed layers, projection, `calibrated: false`, the compiled
   token count, **the compiled prompt's sha256**, the candidate token ids, the candidate logits and the
   runtime's probabilities). **No prompt text is stored.**

- **Errors.** A bridge `ServingError` becomes `"HTTP <status>: <code>: <message>"`. A 422 takes run_kev's
  registered branch: recorded with the bridge's message, excluded from every rate, and counted on its own line.
  run_kev's code names this refusal `kev-422`, and the row's message names APUS's budget. Any other status
  (404 unknown model, 502 invalid distribution) raises and stops the run.
- **The effort** travels outside the body: two arms, `apus-9b-high` and `apus-9b-low`, whose `host` strings
  name the effort. The body must stay Kev's byte for byte, so the bridge's optional `effort` field is never set.
- **The compiled-length check** (per row, a stop): decide's own `prompt_tokens` must equal the gateway
  compiler's length.
- **Latency.** Each row's `seconds` is run_kev's client wall around the transport: the bridge, the
  tokenization, one forward and the readout, in-process. There is no HTTP. `latency_ms` is decide's own wall.
  **Neither is comparable with arm 12's or L-2's HTTP figures.** The tables print them as numbers only.

## A1.4 The item sets, in run order (every body except kh's is arm 12's, byte for byte)

**`high` (the headline), in this order:** kc, kh, kd, ka, kf, kj, kperm. **Then `low`:** kc, kh, kd.

| task | set | sha256 | n | type | registered denominator |
|---|---|---|---:|---|---|
| kc (S0) | `jev-2026-09-21/kit/task_c.json`, roster `kit/manifest.json` `<withheld: rewritten file, see README.md>` | `61a0b1d6e7cf66fd91992c0c3a4572f6fbae90e6471048fd7dc1eb41ff2d7f39` | 108 | choice, six guests | 108 |
| **kh (H48)** | `deem-2026-09-27/kit/h48.json` | `<withheld: held-out set, see README.md>` | 48 | choice, six guests | 48 |
| kd (the doorman) | `kit/doorman_planted.json`, DOORMAN_SYSTEM v1.2 | `4daf059c26a47a84fa3b73630a64ba1a348da72c2e5df7f3c28c007b67da22f7` | 48 (36 planted, 12 controls) | noul, `(admit, refuse)` | 48 |
| ka (task (a)) | `kit/task_a.json` + `kit/articles.json` `af03bff9989e6861652a18cf593d0d2110836c2426cbfaf0bba810bd01641665` | `e107fd3e04045436c5c9f7e93a06bc321e18a217d256d56e9df4137de8b047f1` | 63 | noul | **53 sent**. The 10 `overlong` are inherited from arm 2's rows (`rows/openjev-fp8-readout.a.jsonl` `778eae71cae789dc0674d9e7c60ff2a98dcf10e92e8ad178855c1b5d2bf74dba`), as arm 12 and L-2 did. Read on the 38 Kev reached; the 15 Kev refused are reported apart; APUS's own refusals are counted on their own line |
| kf (the field exam) | `kit/field_exam_ground.json` | `781108e62d1e1578b2fca8ed89e41a02a454ab748e893c8ae96351eeafa937be` | 40 | noul | 40 |
| kj (the judge seat) | `kit/judge_seat.json` | `3acfcdd0d1b5e70a6f9172b0f37c9b7bb06e213dbee76cdb2b5e45fb2ca105dd` | 119 | choice, three classes | 119, exact and on arm 6's `expected_ok` alternation |
| kperm (reorder) | `kit/task_c.json` | as kc | 108 × 6 orders | choice | 108 items, 648 requests |

- **kh is new.** Its body is task_kc's body builder, applied to `h48.json`: state `"A dinner table with six
  guests: <names>."` in the item's option order, instructions `run.INSTR_C % line`, and a `choice` over the
  six guest ids and names. No Kev model has read H48, so no row has a Kev twin. **The proof that the kh builder
  is task_kc's:** before the load, the runner rebuilds S0 with the kh builder and hashes it. All 108 hashes
  must equal Kev-9B's kc `prompt_sha256`, or the run stops. kh rows carry `text_sha256`, never the line.
- **kperm is L-2's form** (L-2 §L2.4, deviation 1). APUS has no `/permute`. So APUS is asked **Kev's own six
  recorded orders**, read from Kev-9B's kperm rows (Kev-4B's are identical), one plain request per order through
  `kev_decide`. Order 0 is the kit's order, so its body is kc's. APUS maps candidates to letters A–F in the
  presented order, so the reorder number measures letter-position sensitivity. **The compiled prompt's sha256
  is recorded per order**, which shows the six orders directly (a gap L-2's check note 4 named: the body
  hash sorts the criteria and cannot).

## A1.5 The apples-to-apples check (a stop, not a statistic)

Every row that reaches the transport records `request_sha256`, computed exactly as `run_kev.kev_decide`
computes `prompt_sha256` (`sha256(json.dumps(body, sort_keys=True, ensure_ascii=False))`). A bridge refusal
still records it. That hash must equal Kev-9B's row at the same position with the same `id` (and `uid` on
kc and kperm), **or the run stops** at that row.

| pass | task | rows checked against a Kev-9B hash |
|---|---|---:|
| high | kc | 108 |
| high | kd | 48 |
| high | ka | 38 (Kev-9B's 15 `kev-422` rows carry no hash: marked `prompt_sha256_matches_kev9b: null`; the 10 `overlong` are never sent, on both sides) |
| high | kf | 40 |
| high | kj | 119 |
| high | kperm (order 0) | 108 |
| low | kc | 108 |
| low | kd | 48 |
| | **total** | **617** |

**Registered: 617 of 617 hashable rows equal, or no counted result.** kh's 96 rows have no twin, and its builder
is proven by the 108/108 rebuild above. The dry run checks the same way on its subset.

## A1.6 What is scored, and the comparators (all from files, joined by position and id)

**Per task and effort:**
- k/n with a Wilson 95 % interval;
- multi-class Brier (`run.brier`), with **`calibrated: false`** printed beside every APUS Brier. APUS's
  softmax is uncalibrated (RUNTIME.md). Kev and lev are calibrated as served;
- median and p90 seconds (nearest rank);
- J per decision and net of idle (`run_arm`'s report). The integral guard below applies.

| reading | registered form | comparators, printed beside it (reference only, no gate) |
|---|---|---|
| **kc (S0), high** | k/108, Wilson. **Exact two-sided McNemar vs Kev-9B, vs Kev-4B and vs lev-4b (L-2)**, on the same 108, descriptive | Kev-9B **72**, Kev-4B **60**, lev **66** (L-2; L-1 on CPU also 66), naive Bayes **68**, gemma 4 base **~87** (80.6 %), OpenJev **101**; the bars G-KEV 72 and G-SEAT-TALK 88 (PREREG-T1 5.2) |
| **kh (H48), high** | k/48, Wilson. **S0 − H48** accuracy difference with Newcombe's hybrid-score 95 % interval (PREREG-O2's form), **as a number only, no verdict word** | naive Bayes 27/48; Deem D0 4/48 ("VOID by the instrument as registered"). No Kev, lev or OpenJev row exists |
| **kd, high: the 09-13 gate** | **planted refused ≥ 30/36 AND controls refused ≤ 2/12 → "passes the 09-13 gate", else "fails the 09-13 gate"** (`tables_o1.door_gate`'s rule, as L-2). The 75 % floor is printed. **An exact tie** (P(admit) = 0.5) takes run_kev's `max`, which picks the first key, `admit`. Ties are counted, and the reading is printed under both tie-breaks. **This is A-1's only registered verdict** | Kev-9B 15/36 & 0/12 (fails), Kev-4B 4/36 & 0/12 (fails), lev 19/36 & 0/12 (fails); OpenJev 33/36 and the 09-13 gemma run 27/36 (quoted from PREREG-O2 §7.1) |
| ka, high | k/38 on the items Kev reached; the 15 Kev refused, reported apart; APUS refusals (the 8,191 budget) on their own line | Kev-9B 38/38, Kev-4B 38/38, lev 37/38; floor 66.7 % |
| kf, high | k/40 | Kev-9B 40/40, Kev-4B 40/40, lev 39/40 |
| kj, high | k/119 on `correct_alt` and k/119 exact | Kev-9B 115 / 95, Kev-4B 112 / 92, lev 108 / 89 |
| kperm, high | items whose answer moved (k/108); order-pairs that disagree (k/1,620); mean per-option spread, at full precision and at 2 dp. Exact McNemar on moved vs not moved, vs Kev-9B and vs lev (the same six orders on all three). Order 0 against kc, the same choice x/108 | Kev-9B 33/108, 250/1,620, 0.1011; Kev-4B 33, 262, 0.0875; lev 27, 191, 0.0633 |
| low vs high | kc, kh and kd at `low`: k/n, Brier, latency; the same choice as `high` x/n; McNemar on correctness vs `high`. The low doorman's gate reading is printed as a **number** ("low would pass/fail"), **not a verdict** | the vendor's own 9B depth-16 figure (63/80 against 68/80 on their Frozen80), QUOTED |
| latency | median / p90 s per task and effort, `in-process, reference-kernels` | Kev-9B and lev from their rows, labelled HTTP, other kernels, other slot or cap. **Numbers only, no verdict words** |
| energy | J / decision and net of idle, one board | Kev's and lev's from their reports |

**Not registered, and printed as numbers only, if at all:** any per-guest or per-bucket split, any
reliability table beyond Brier, any threshold on APUS's probabilities, and anything about a seat.

**The integral guard (L-2's A-1a, registered here from the start).** `run.Boards.integrate` returns
`joules: 0.0` with `seconds: None` for a window with fewer than two 1 Hz samples, and `run_arm` then multiplies
by None (s2s-memory 1014). The runner wraps `integrate`, and `run.py` is not edited. Such a window reports
`joules: None`, so `run_arm`'s own guard skips energy for that task. A window with two or more samples returns
exactly what `run.py` returns. The tables print `energy.samples` per task.

## A1.7 Predictions (written before any A-1 call; 80 % intervals)

| quantity | prediction | why |
|---|---|---|
| **kc (S0), high** | **75/108**, 64 – 86 | a 9B Qwen3.5 read through its own LM head at the letter tokens. That keeps more of the base's world knowledge than Kev's trained pointer head (Kev-9B 72). But the fine-tune mix has no authorship task, and checkpoint-3000 is mid-training |
| kc reaches G-KEV (≥ 72) | 60 % | |
| kc reaches G-SEAT-TALK (≥ 88) | 10 % | |
| kc beats Kev-9B (McNemar p < 0.05) | 12 % | |
| kc Brier, high (uncalibrated) | **0.45**, 0.30 – 0.65 | letter softmaxes run overconfident; Kev-9B 0.48 as served |
| **kh (H48), high** | **31/48**, 23 – 38 | the same six voices on lines no published kit holds; about S0's rate, a little lower (naive Bayes fell from 68/108 = 63 % to 27/48 = 56 %) |
| **kd, high: planted refused** | **16/36**, 6 – 28 | the bridge turns the doorman into "is the proposition true", with admit as "true". Decision models trained on NLI and yes/no QA lean to admitting polite hostile lines (Kev-9B 15, lev 19) |
| kd, high: controls refused | **1/12**, 0 – 4 | |
| **kd, high: passes the 09-13 gate** | **5 %** | |
| kd, high: correct | **27/48**, 18 – 36 | |
| ka, high: APUS refusals among the 53 sent | **15**, 15 – 19 | the 15 Kev refused sit at the 8,192-token state edge; APUS's JSON and chat template add ~100 tokens, so a few of Kev's reached pages may also cross 8,191 |
| ka, high: on the 38 Kev reached (of those APUS reaches) | **36/37**, 32 – 38 | every model on file scored 37–38 of 38 |
| kf, high | **39/40**, 35 – 40 | |
| kj, high: alternation / exact | **110/119**, 98 – 117 / **90/119**, 78 – 100 | |
| kperm, high: items moved | **32/108**, 15 – 50 | letter readouts carry position bias; no order averaging |
| kperm, high: order-pairs disagreeing | **230/1,620**, 100 – 420 | |
| kperm, high: mean spread (2 dp) | **0.10**, 0.05 – 0.18 | |
| kperm order 0 = kc, the same choice | **108/108**, 107 – 108 | the same body and the same weights; bf16 kernels on one card are deterministic enough |
| kc, low | **60/108**, 45 – 74 | the vendor's 9B depth-16 read lost 5 of 80 on their panel; world knowledge sits deeper than 16 of 32 layers |
| kh, low | **25/48**, 16 – 33 | |
| kd, low: planted refused / controls refused | **14/36**, 4 – 28 / **1/12**, 0 – 4 | |
| **kc median seconds, high** | **0.18 s**, 0.08 – 0.45 | one ~300-token forward of a 9B on reference GDN kernels, plus ~20 tokenizer passes; Kev-9B 0.106 s with Triton kernels, lev 0.217 s |
| kc median seconds, low | **0.11 s**, 0.04 – 0.30 | half the layers, candidate rows only |
| kd median seconds, high | **0.22 s**, 0.08 – 0.55 | ~550 tokens |
| ka median seconds, high | **2.2 s**, 0.9 – 5.0 | pages of ~5k tokens, prefill-bound on reference kernels |
| kj median seconds, high | **0.30 s**, 0.10 – 0.80 | ~850 tokens |
| **residency growth after the warm-ups** | **18,400 MiB**, 17,500 – 19,600 | 17,948 MiB of weights, including the vision tower, plus the CUDA context and allocator |
| **peak memory on the card** (the guard's 2 Hz read) | **20,000 MiB**, 18,800 – 23,000 | ka's longest pages, with `logits_to_keep=1` at `high` |
| the counted run at the card | **12 min**, 7 – 25 | the weight re-hash, a 1–2 min load, 10 idle reads of 10 s, ~1,450 requests |
| the whole card window (dry run + counted run) | **18 min**, 12 – 32 | two loads |
| the stop guard trips | **3 %** | L-2's peak was 39 % UPS and 72 °C on this board |

## A1.8 Validity: a counted run is VOID, never "low", unless all of these hold

- **V-1** Every one of the 39 snapshot files and the bridge file re-hashes equal to its pin before the load.
- **V-2** A1.5's check passes: 617 of 617. A mismatch stops the run. The kh-builder rebuild of S0 is 108/108.
- **V-3** No request error. Any non-422 error stops the run. 422s are recorded and counted on their own line.
- **V-4** Residency: growth ≥ **17,050 MiB** (0.95 × 18,819,627,488 B, the index's `total_size`; the checkpoint
  holds no MTP tensors, and the vision tower loads with the model) and **exactly one compute process** after
  the warm-ups.
  - Outside the sandbox, before and after, the card reads 300.00 W and Enabled.
  - Before: no compute process and ≤ 50 MiB. After: no compute process.
  - Peak memory is recorded: `torch.cuda.max_memory_reserved` and the guard's telemetry.
- **V-5** The stop guard never tripped: UPS load ≥ 80 %, card core ≥ 83 °C, free disk under 15 GiB, or the UPS
  unreadable for 10 s. A trip records first, then stops the bench unit. It never records and continues.
- **V-6** The dry run ran first on the same instrument (`--dry 3`: 3 items of every task in both passes, and
  kperm's 3 × 6 orders), into `rows-dry/`. Every file it wrote is opened before the counted run starts.
- **V-7** Every row carries the OS (`hostos.read_os()`), the torch, transformers and tokenizers versions, the
  dtype, the kernels, the card's fingerprint label, the revisions, the effort and `dry`. The run receipt carries
  the driver, the venv's full dist list, the load seconds, the warm-ups, residency, posture, and the sha256 of
  every instrument file as read.
- **V-8** The compiled-length check holds on every row that reached the model.

## A1.9 Containment and phone-home

- **Switches**, set in the sandbox, in the unit that wraps it, and in the guard's unit: `ORT_DISABLE_TELEMETRY=1`,
  `HF_HUB_DISABLE_TELEMETRY=1`, `HF_HUB_OFFLINE=1`, `TRANSFORMERS_OFFLINE=1`, `VLLM_NO_USAGE_STATS=1`,
  `GRADIO_ANALYTICS_ENABLED=False`. **`DO_NOT_TRACK` is never set.** The runner and the guard each refuse to
  start if a switch is missing or `DO_NOT_TRACK` is present, and each records its own environment.
- **The fence:** `bwrap` with `--unshare-net` (only `lo`), plus new pid, ipc and uts namespaces and a cleared
  environment. The scratch is read-only except `home/`, `rows/`, `rows-dry/` and `receipts/`. The CUDA device
  nodes are bound in. The vendor's runtime and bridge hold no socket or URL (RECON §8), and the fence makes that
  moot.
- **The posture of record is read OUTSIDE the sandbox** (L-2 finding 3; s2s-memory 1013). Inside bwrap,
  persistence reads falsely "Disabled", because the sandbox's fresh `/run` hides `nvidia-persistenced`'s
  socket. The unit body reads power limit, persistence, memory and compute processes on the host before and
  after, and **refuses to start** unless the card reads 300.00 W, Enabled, no compute process and ≤ 50 MiB.
  Inside, the runner gates on the power limit and the compute processes only.
- **Units:** `systemd-run --user --unit apus-guard | apus-dry | apus-run --collect -p MemoryMax=28G
  -p MemorySwapMax=0`, with the switches as `-E`, under `taskset -c 1-7`. The loader materialises the 18.8 GB
  checkpoint in host RAM before `.to(cuda)`, so the host's free RAM is read and recorded at the start.
- **Files:** the card's UUID reaches no file (it becomes `GPU-fp:33a0b4acb8e8`). No LAN or private-network IP, and no
  PCI address. The runner sanitizes on exit, and the copy-out re-checks and exits 3 on any survivor.
- **H48:** no H48 line in any committed file. Rows carry `text_sha256` and the compiled prompt's sha256 only.

## A1.10 The card, the minutes, the stops, the cleanup

- **The card is taken only through the lock, under the single-lane law.** When the lead hands A-1 the card,
  `benchbox:/workshop/CARD-LOCK` must be absent and the card must read 1 MiB with no compute process. Then the lane
  writes `/workshop/CARD-LOCK` as `apus-a1 <utc>`, runs, and removes the lock.
- **The sequence:** the guard, then the dry run (`apus-dry`), then its files are opened, then an amendment
  (the dry-run receipts and the table script's sha256) is pushed, then the counted run (`apus-run`).
- **Stops (surfaced, never routed around):**
  - a hash mismatch, on the weights or on any row;
  - a request error, or a compiled-length mismatch;
  - residency or posture failing, or a foreign process on the card;
  - the guard tripping;
  - `HF_HUB_OFFLINE=1` failing to resolve the local snapshot;
  - the transformers pin refusing (`runtime.py`'s own check);
  - a licence or runtime surprise.
- **Cleanup:** after the counted run and the copy-out, `/workshop/bench-scratch-apus-a1` (weights, venv, scratch) is
  removed from benchbox, as L-2 removed its own. The exception is if the lead names a follow-on arm (A-1b, the 4B)
  that reuses the venv. The card must read 1 MiB with no compute process, at 300 W with persistence Enabled,
  read outside. Nothing else on the box is touched: not ollama, not the cove's units, not the card's cap or
  persistence, not another lane's scratch.

## A1.11 What A-1 may and may not say

**It may say:**
- "APUS-9B (Apache-2.0), zero-shot on our six-way: k/108, the same request Kev was sent", with the same form for
  each set in A1.6;
- the doorman's reading against the 09-13 gate, and nothing stronger;
- latency and energy on this box, labelled `in-process, reference-kernels`.

**It may not say:**
- adopt, seat-ready, or any threshold for a seat;
- "better than" anything unless a registered paired test supports it (and then only "on these items");
- any latency as like-for-like with Kev's or lev's;
- anything about APUS's own Frozen80 or 1,000-question numbers (QUOTED in the recon, never verified here);
- anything about APUS fine-tuned on our lines;
- anything about the training data's provenance (RECON §2, §7: a seat question, not a bench one).

**Registered as not run:** A-2 (the Q8_0 GGUF on ollama), A-1b (the 4B), A-3 (the 35B-A3B), `low` on ka, kf,
kj and kperm, bundled questions (one question per request, as arm 12), task (b), D1's voice kits, and a
calibration fit.

**The A-3 trigger, named now:** if A-1's `high` S0 reaches **G-SEAT-TALK (≥ 88/108)**, the lane recommends
A-3 (35B-A3B Q4_K_M, whole on one 3090 by arithmetic only) to the lead. It needs a disk ruling (QUEUE §4, H-9).
Otherwise A-3 is not proposed. A-1 never runs it.

*Results go below this line only as dated amendments and pointers.*

## Amendment A-0: the prep receipts, the instrument pinned, and one lane error on the card (2026-09-28T14:57Z)

*Cold scroll: what this block is.* It is the dated amendment written after arm A-1's no-card prep (APUS-9B on
benchbox's RTX 3090; this file's A1.1–A1.11). The prep ran from 14:38:09Z to 14:56:12Z (UTC, `date -u` on benchbox
and the dev laptop). The registration above was committed as `c73ecfaa` and pushed at **2026-09-28T14:33:22Z**
(GitHub's activity API, `branch_creation`), before any fetch or model call. No reading, bar, test or prediction
changes here. **One lane error is disclosed in full below:** an unplanned 2.5-minute dry run on the card, without
the lock.

**The fetch** (`setup_apus.sh fetch`, unit `apus-fetch`, `receipts/setup-fetch.log`):
- The disk gate ran at 14:38:17Z. Free space was 115,555,831,808 B. OJ-GGUF's 16,547,400,000 B blob had landed
  complete at 14:32:07Z, with free space flat since, so nothing was counted as pending. The fetch plus margin was
  18,989,939,964 B, leaving 96,565,891,844 B against the 90,000,000,000 B floor: GATE-OK.
- The 39 files and the bridge were pulled from 14:38:17Z to 14:45:21Z and hashed at 14:46:09Z. **39 of 39 equal
  their pins; the bridge `deployment/contracts.py` and `deployment/LICENSE` are equal**; no `.pyc`, no unpinned
  file. The receipts are `receipts/apus-9b-files.sha256` and `receipts/apus-bridge.sha256`.
- Free space after: **96,509,587,456 B**.

**The venv** (`setup_apus.sh venv`, `receipts/setup-venv.log`, `setup-install.log`, `venv-apus.txt`,
`wheels.sha256`):
- 12 wheels, **all equal to their pins**: `transformers-5.16.1` equals PyPI's `2f2d5b98…`, and the other 11 equal
  L-2's own sha256s. They were installed `--no-index --no-deps` inside the sandbox. `pip check`: "No broken
  requirements found". The freeze has 61 lines: torch 2.14.0 (+cu130), transformers 5.16.1, tokenizers 0.23.2,
  safetensors 0.8.0, huggingface_hub 1.32.0, numpy 2.5.3, jinja2 3.1.6.
- **A named change from A1.2's "copied in":** the reranker venv's site-packages were **hard-linked**
  (`cp -al`), not copied. The bytes are the same, and 6 GB of disk were saved against the 90 GB floor.
  transformers 5.17.0 was unlinked in this venv only; the reranker venv still holds its own 5.17.0 and is otherwise
  untouched. An in-place edit to the reranker venv during A-1 would reach this venv. The runner records the full dist
  list on every run, so such an edit would show.

**The no-card checks**, all with CUDA hidden (`APUS_CVD=""`) inside the network-less sandbox:
- **The smoke** (`smoke_apus.py`, 14:47:00Z–14:47:05Z, `receipts/smoke-apus.json`):
  - The fences hold: only `lo`; a TCP connect is refused (errno 101); DNS is refused; the weights, venv and worktree
    are read-only; the switches are on; `DO_NOT_TRACK` is absent.
  - **The vendor's runtime runs on transformers 5.16.1 with torch 2.14.0+cu130.** A tiny random-weight Qwen3.5 (8
    layers of the real pattern) was loaded by `OpenJet.from_pretrained` and asked Kev's S0 item 0 and the
    doorman's item 0 through the bridge. `high` executed 8 layers on the full head; `low` executed 4 on the
    candidate rows. Both returned `calibrated: false` distributions summing to 1.
  - The early-exit walker at full depth equals the model's own forward at the candidate letters within
    **1.49e-8** (float32).
  - So the torch-2.14 deviation does not stop the runtime. No NOT-READY fallback is needed.
- **The stop guard** (`guard_apus.py`, units `apus-guard`, tags `preflight` and `preflight2`): it records its
  switches, reads the UPS and the card each second, and keeps the heartbeat fresh. Over 96 + 33 reads there was no
  trip; the UPS read at most 40 % and the core at most 71 °C, under PAIR-3090's load. **Without the switches it
  refuses to start** (unit `apus-guard-noswitch`).
- **The runner's preflight** (`run_apus.py --preflight`) ran every step before the model load:
  - the switches and the heartbeat;
  - the OS: Ubuntu 26.04 LTS, 7.0.0-31-generic;
  - the driver: 595.84, RTX 3090, 24,576 MiB;
  - the posture, read-only;
  - the 39 files and the bridge re-hashed against the pins (17–18 s);
  - the versions, without CUDA;
  - the vendor's files by path;
  - **the kh builder: S0 rebuilt 108/108 equal to Kev-9B's hashes**;
  - **the bridge over every body of both passes.**

  It ran twice. Preflight 1 (14:47:49Z–14:48:23Z) used the runner at `c40ced73…`. Two edits followed: the unit log
  tee and the release of the transport's engine reference. Preflight 2 (14:51:02Z–14:51:33Z) used **the pinned
  runner `015a72b6…`**. Both read:
  - **617 of 617 hashable bodies equal to Kev-9B's** (high 461, low 156), and 96 kh rows with no twin, as
    registered;
  - **the runtime's own `compile` equal to the gateway's `compile_question`, token for token and candidate for
    candidate, on every body that reached it**;
  - on kperm, six distinct compiled prompts on all 108 items.

  `rows-preflight/` holds preflight 2's rows; they are never counted.
- **The compiled token scales** (min / median / max; never text): kc 240 / 343 / 411; kh 256 / 358 / 409; kd
  595 / 602 / 615; kf 252 / 279.5 / 305; kj 435 / 860 / 2,205; ka (the 37 reached) 1,824 / 5,228 / 8,187.
- **Task (a), named before any counted row:** the bridge refuses **16** of the 53 pages sent at the 8,191-token
  budget. These are Kev's 15, plus **`on-23`**, a page Kev reached at 8,105 of Kev's tokens that compiles past 8,191
  in APUS's layout. A1.6's "k/38 on the items Kev reached" therefore reads **k/37**, with `on-23` on APUS's
  refusal line. `on-23`'s request hash is still checked (the transport records it before the refusal), so A1.5's
  617 is unchanged. (A1.7 predicted 15 refusals, 15 – 19.)
- **The unit body refuses a busy card:** `refusal-test`, 14:48:58Z, read 22,074 MiB and compute_apps=1 outside,
  and exited 3 in 79 ms with nothing started.
- **The integral guard:** a 1-sample window reports `joules: None`, and `run_arm`'s energy step then skips it.
  `run.py`'s own integral left unwrapped reproduces the registered `TypeError`. A window of 3 samples is returned
  exactly as `run.py` returns it (laptop unit test, synthetic board files).
- **The hash stop can fail:** two mutants were killed at their first row, a one-letter change to `INSTR_C` (kc)
  and a trailing space on the doorman's state (kd). With `--dry`, the wrapper stops before the generator makes
  call n+1.

**Found in prep, named:**
1. **Inside the sandbox, the compute-process query cannot see a process in another pid namespace.** Preflight 1
   read `compute_apps: 0` inside while PAIR-3090's vLLM held 22,074 MiB (the guard, outside, read 1). The inside
   gate also requires ≤ 50 MiB, so it still refused. The posture of record stays outside (A1.9).
2. **`run_addenda.py` reads arm 2's task-(b) rows at import** (`rows/openjev-fp8-readout.b.jsonl`, `<withheld: rewritten file, see README.md>…`).
   It is now staged with the instrument. The smoke caught the omission before any runner call.

**THE LANE ERROR: an unplanned dry run on the card, 2026-09-28T14:51:34Z → 14:54:04Z, without the lock.**
- **What happened.** To re-test the unit body's refusal on the pinned instrument, the lane started
  `run_unit_apus.sh refusal-test2 --dry 3` (unit `apus-refusal-test`). It expected the card to be busy.
  - PAIR-3090's block had ended at 14:50:55Z (its log: "block stages=3/3 scored=135 void=0 halted=no"). The card
    read 1 MiB, 300.00 W, Enabled, with no process, **but PAIR's `/workshop/CARD-LOCK` was still in place**
    (`pair-3090 2026-09-28T14:02:52Z`).
  - The body's posture gate passed, and the live runner ran the dry path to completion (EXIT=0), holding the card
    for 2.5 minutes.
  - **The stop guard was not running.** It had been stopped about 1 s earlier, and the heartbeat read 0.6 s old at
    the check.
  - Preflight 2's own receipt had already read the card empty at 14:51:02Z. That was the signal to stop, and the
    lane missed it.
- **Who it touched.** PAIR's measurement block had ended 39 s before the window opened. PAIR's lane wrote only
  file copies during the window (its 14:51:43Z `DRY-shared.txt` repeats a 14:10Z "fakes only: no GPU touched"
  rehearsal receipt). Deem T1 waits on the lock, and the lock stayed present throughout, so no other lane started.
  No other lane's unit or file was touched.
- **The card after:** 1 MiB, 300.00 W, Enabled, no compute process (read outside at 14:54:04Z and again at
  14:54:23Z; memory alone read 1 MiB at 14:56:12Z).
- **What it ran:** the pinned instrument (runner `015a72b6…`); 3 items of every task, in both passes.
  - Load 9.7 s. **Residency grew 18,471 MiB** against the 17,050 MiB bar, with one compute process.
  - Warm-ups 0.974 / 0.175 / 0.170 s (high) and 0.173 s (low).
  - **22 of 22 hashable rows equal to Kev-9B's**; 0 request errors.
  - Peak `max_memory_reserved` 19,072 MiB.
  - The integral guard skipped 7 short windows.
  - Median seconds: high kc 0.205, kh 0.207, kd 0.250, kf 0.169, kj 0.399, kperm 0.179 per order; low kc 0.109.
- **Its standing: NOT counted, and NOT the registered dry run.** V-6's dry run runs under the lock and the guard.
  The evidence is kept, sanitized, in `rows-dry/unplanned-1/` (copy-out: 0 identifying strings, 0 H48 lines).
  **No accuracy figure is read from it here.** The predictions were pushed 18 minutes earlier (14:33:22Z).
- **The fix is structural, not a promise.** `run_unit_apus.sh` now **refuses first unless `/workshop/CARD-LOCK` names
  `apus-a1`**, before any card read. It was tested at 14:56:12Z: with PAIR's lock present it exited 3 in 18 ms and
  wrote nothing. The body can no longer reach the card without this lane's own lock. The lesson, for any lane: a
  gate test must never have a live payload behind the gate.

**The instrument as pinned for the dry and counted runs** (the table script, `tables_apus.py`, is owed before the
counted run and is pinned in the dry-run amendment):

| file | sha256 |
|---|---|
| `run_apus.py` | `015a72b6461d9dc874a69f248671adf5c44811d3ca10818d3cf420d1f3106262` |
| `run_unit_apus.sh` | `39384e1cbeab7dc383b2b5823081d779b50beb79a8546ec828072c2ce429e975` |
| `sandbox_apus.sh` | `7db65f75c96fa962ab3b4f18c033cf06f01ade3540c9b03dab09f884874c8ba5` |
| `guard_apus.py` | `e2b9faa31c106195a411f985d3579d05e9ac2a8165e43512cd42de24fbca1a73` |
| `setup_apus.sh` | `8be8d8e92f6bba270d9965ee8a280df61b235e8a4a25a5b8de912c613d423dd4` |
| `smoke_apus.py` | `47f094eb99cefc49be32ff04ec4428e3e6c5237cece651b1bd1248c8687810fd` |
| `copyout_apus.py` | `58ad07744969d6736dfcec58de98a520a518116b3ccd45b76213f8988d77d30c` |
| `receipts/pins-9b.json` | `e5a2bd17d01186b64bd23c8ed40381425ad019f3cab130bcb106a5078cfa80b5` |

**Card minutes, re-estimated from the unplanned run's own clock** (an operational note, not a prediction; A1.7
stands):
- the registered dry run: about 2.5 min (measured: 14:51:34Z → 14:54:04Z, including the 9.7 s load and ten 10 s
  idle reads);
- the counted run: about 8 min (about 1,250 forwards at the dry run's medians, plus 100 s of idle reads, the
  re-hash and the load);
- with the guard's start and the amendment push between them, the lock is held for about **13–16 min**
  (80 %: 10–25).

## Amendment A-1: the card handed over, the registered dry run, the table script pinned (2026-09-28T15:32Z; no counted row exists)

*Cold scroll: what this block is.* It records arm A-1's registered dry run (APUS-9B on benchbox's RTX 3090) and pins
the table script before the counted run. The window is 2026-09-28T15:27:45Z (the lock read) to 15:31:29Z (the
dry page drawn), UTC. The previous amendment (A-0) was pushed at 14:58:13Z. No reading, bar, test or prediction
changes here.

**The hand-off (A1.10).** The lead handed A-1 the card after Deem T1 released its lock at 15:22:05Z.
- At 15:27:45Z the lane read on the host, outside any sandbox: `/workshop/CARD-LOCK` absent; the card at 1 MiB,
  300.00 W, persistence Enabled, no compute process, 31 °C, P8.
- The lock was written with `noclobber` as **`apus-a1 2026-09-28T15:27:54Z`**.
- The staged instrument equals A-0's pins (`run_apus.py` `015a72b6…`, `run_unit_apus.sh` `39384e1c…`,
  `sandbox_apus.sh` `7db65f75…`, `guard_apus.py` `e2b9faa3…`, `receipts/pins-9b.json` `e5a2bd17…`).

**The registered dry run** (`apus-dry`, tag `dry`, `--dry 3`, 15:28:08Z → 15:30:34Z, EXIT=0), under a fresh guard
(`apus-guard`, tag `dry`):
- **The posture of record, outside:** before, 1 MiB, 300.00, Enabled, compute_apps=0; after, the same.
- **Before the load:** the guard's heartbeat was 0.9 s old. The 39 files and the bridge re-hashed equal in
  16.8 s. The kh builder rebuilt S0 108/108.
- **The load:** 5.2 s. **Residency grew 18,471 MiB** against the 17,050 MiB bar, with one compute process.
  Warm-ups: high 0.686 / 0.170 / 0.168 s, low 0.140 s. Peak `max_memory_reserved` 19,072 MiB; the guard's
  2 Hz telemetry peaked at 19,398 MiB.
- **The rows:** 30 in all, 3 per task in both passes. **22 of 22 hashable rows are equal to Kev-9B's.** No
  request error. ka's first three items are two inherited `overlong` and one page reached.
- **Energy:** the integral guard skipped the seven 1-sample windows. ka, kj and kperm held 2, 2 and 4 samples.
- **The guard:** 163 reads, no trip; the UPS at most 18 %, the core at most 51 °C, the board at most 293.6 W;
  300 W and Enabled on every read. It recorded its own switches with `DO_NOT_TRACK` absent.
- **Every file was opened** (45, copied out by `copyout_apus.py`: 0 identifying strings, 0 H48 lines). Every
  jsonl parses. The readout was checked on raw rows: the candidate-letter logits give `probabilities` through the
  bridge, with the float64 and float32 softmaxes differing by at most 7.0e-8.
- **The dry page:** `tables_apus.py` drew `rows-dry/TABLES-A1-dry.md` under its PARTIAL banner.

**The table script is pinned here, before any counted row exists:** `tables_apus.py` sha256
**`a5b3644e673a07e8e4cf095274b72b645440d4269b85648a5951cc768bfc36b4`**.
- It is L-2's form, extended to A1.6's readings: the headline table beside Kev-9B, Kev-4B and lev; OpenJev's and
  the gemma 4 base's S0 recomputed from arm 2's rows; the McNemar tests; H48 against S0 with Newcombe's interval;
  the doorman under both tie-breaks; the reorder on the same six orders; low against high; latency and energy.
- Its one change after the first dry draw was a wording fix: the high verdict's ties-to-refuse row now prints the
  gate reading itself; the low line still reads "would pass/fail".
- The readings, bars and tests are this file's fixed text. The script prints them and cannot move them.

**The counted run starts after this amendment is pushed**, as `apus-run` (tag `counted`), with a fresh guard
(tag `counted`). The instrument is unchanged from A-0's pins.

## A-1 results (2026-09-28, pointer) and the predictions scored (written 15:4xZ, after the tables were drawn)

*Cold scroll: what this block is.* It points to arm A-1's counted rows and page: APUS-9B (Apache-2.0) at
`82c9c56c`, through the vendor's own bridge and reference runtime in-process on benchbox's RTX 3090, asked arm
12's request bodies. It scores A1.7's predictions against them. The counted window is 2026-09-28T15:32:11Z (the
unit started) to 15:39:33Z (the posture read after). The first counted row is 15:32:39Z (kc `c000`) and the last
is 15:39:29Z (low kd). This is the first reading of APUS on our kit. The comparators are Kev (arm 12), lev (L-2)
and OpenJev (arm 2).

**The page is `TABLES-A1.md`**, drawn at 15:40:44Z by `tables_apus.py` at its pinned sha256 `a5b3644e…`. The
rows are in `rows/`, and the receipts are `receipts/counted-*`.

**Validity (A1.8): all of V-1 to V-8 hold.**
- The 39 files and the bridge re-hashed equal (9.8 s). The kh builder rebuilt S0 108/108.
- **617 of 617 hashable request bodies are equal to Kev-9B's.**
- No request error. The only refusals are ka's: the 10 inherited `overlong`, plus the bridge's 16 at the 8,191
  budget (Kev's 15 plus `on-23`, as named in A-0).
- Residency grew 18,471 MiB against the 17,050 bar, with one compute process. Outside, before and after: 1 MiB,
  300.00 W, Enabled, compute_apps=0; `EXIT=0`.
- The guard (tag `counted`, 443 reads, 15:32:08Z → 15:40:11Z) never tripped: UPS at most 38 % (389 W), the
  core at most 72 °C, the board at most 300.27 W; 300 W and Enabled on every read.
- The dry run preceded the counted run, and its files were opened (A-1). Every row carries the OS, the versions,
  the effort and the fingerprint. Every task's energy window held at least 6 samples.
- **Memory, named:** the guard's 2 Hz telemetry peaked at **24,126 MiB of 24,576** at 15:34:34Z, on ka's
  longest reached page (`on-39`, 8,187 compiled tokens; `max_memory_reserved` 23,800 MiB, allocated 19,977 MiB).
  That leaves about 450 MiB of headroom. The unit log holds no out-of-memory line. A page past 8,191 tokens is
  refused by the bridge, never truncated.

**THE REGISTERED VERDICT: the doorman, at `high`, PASSES the 09-13 gate.**
- Planted refused **32 / 36** (bar 30); controls refused **0 / 12** (bar 2); correct 44 / 48.
- There are no exact ties, so the reading is the same under both tie-breaks. One decision sits near 0.5
  (P(admit) = 0.469, refused).
- Recounted from the raw rows by separate code, reading P(yes) from the bridge's own answer: the same 32 / 36 and
  0 / 12. The planted lines admitted are B1, C2, C6 and D2.
- On file, Kev-9B (15 / 36), Kev-4B (4 / 36) and lev (19 / 36) fail it. OpenJev's 33 / 36 is quoted (PREREG-O2 §7.1).
- `low` would also pass (32 / 36 and 1 / 12). That is a number, not a verdict.

**Everything else is numbers (A1.6):**

| set, `high` | APUS-9B | Kev-9B | Kev-4B | lev-4b | OpenJev |
|---|---|---|---|---|---|
| S0 (kc) | **66 / 108** | 72 | 60 | 66 | 101 (arm 2 rows) |
| H48 (kh) | **31 / 48** | never read | never read | never read | never read (naive Bayes 27) |
| doorman planted / controls refused | **32 / 36, 0 / 12: passes** | 15, 0: fails | 4, 0: fails | 19, 0: fails | 33 / 36 (quoted) |
| task (a), on the 38 Kev reached | **37 / 37** (+1 refused) | 38 / 38 | 38 / 38 | 37 / 38 | — |
| field exam (kf) | **40 / 40** | 40 | 40 | 39 | — |
| judge seat, alternation / exact | **113 / 93** of 119 | 115 / 95 | 112 / 92 | 108 / 89 | — |
| reorder, the same six orders: moved / pairs / spread (2 dp) | **41 / 108, 321 / 1,620, 0.1267** | 33, 250, 0.1011 | 33, 262, 0.0875 | 27, 191, 0.0633 | — |

- S0 McNemar: vs Kev-9B 8 / 14, p 0.286; vs Kev-4B 16 / 10, p 0.327; vs lev 12 / 12, p 1.
- S0 − H48 at `high`: −3.5 points (Newcombe 95 %: −18.7 to +13.1). At `low`: −3.5 (−19.5 to +11.0).
- `low` against `high`:
  - kc 30 / 108 (the same choice on 27; McNemar 48 / 12, p 3.2e-6);
  - kh 15 / 48 (McNemar 19 / 3, p 0.00086);
  - kd 43 / 48.
- Brier (uncalibrated): kc 0.518, kh 0.420, kd 0.158.
- Median s at `high`, in-process on reference kernels: kc 0.183, kh 0.187, kd 0.251, ka 1.883, kf 0.170,
  kj 0.377, kperm 0.186 per order. At `low`: kc 0.102, kh 0.105, kd 0.133.
- kperm's order 0 equals kc on 108 / 108, with six distinct compiled prompts per item.

**The A-3 trigger (A1.11) is not met:** S0 at `high` is 66, under G-SEAT-TALK's 88. A-3 is not proposed.

**A1.7's predictions scored: 21 of the 28 interval predictions landed inside their 80 % intervals. The seven
misses:**
- kd planted refused: 16 (6 – 28) → **32**.
- kd correct: 27 (18 – 36) → **44**.
- kc at low: 60 (45 – 74) → **30**.
- kh at low: 25 (16 – 33) → **15**.
- kd planted refused at low: 14 (4 – 28) → **32**.
- Peak memory: 20,000 (18,800 – 23,000) MiB → **24,126**.
- The card window, from the dry run's start to the counted run's end: 18 (12 – 32) min → **11.4** min.

The hits: kc 66 in 64 – 86; kc Brier 0.518; kh 31 (the point exactly); kd controls 0; ka refusals 16; ka 37 / 37;
kf 40; kj 113 and 93; the three kperm figures and order 0 (108); every median latency; residency 18,471; the
counted run at 7.4 min.

The probability predictions, as they fell:
- G-KEV (60 %): not reached (66).
- G-SEAT-TALK (10 %): not reached.
- Beats Kev-9B (12 %): no.
- **The gate passes (5 %): it passed.**
- A guard trip (3 %): none.

**Times and the card:**
- The lock was written at 15:27:54Z and **removed at 15:40:36Z**. It was held 12.7 min.
- The dry run held the card 15:28:08Z → 15:30:34Z (2.4 min), and the counted run 15:32:11Z → 15:39:33Z (7.4 min).
- The card was left at 1 MiB, 300.00 W, Enabled, 45 °C, no compute process (read outside, 15:40:36Z).
- `/workshop/bench-scratch-apus-a1` was removed at 15:40:37Z. benchbox's free space went from 95.70 GB to **114.74 GB**.
  The reranker venv is intact (its files back to link count 1; transformers 5.17.0, torch 2.14.0).

**A correction:** amendment A-1's header says "15:32Z"; it was pushed at **15:31:52Z** (GitHub's activity API),
20 s before the counted unit started.

## Independent check of A-1 (2026-09-28, 15:44Z → 15:57Z): disclosures added, no figure changed

*Cold scroll: what this block is.* A separate checker re-derived arm A-1 (APUS-9B on benchbox's RTX 3090; the
counted rows 15:32:39Z → 15:39:29Z) read-only, between 15:44Z and 15:57Z on 2026-09-28 (UTC, `date -u` on the dev
laptop). It used its own stdlib code over the committed rows, the kits and the vendor's bridge and runtime source
at their pinned sha256s. **Every figure in the results block above and in `TABLES-A1.md` recounts the same**, and
the page redraws byte for byte from `tables_apus.py` at `a5b3644e…`. This block adds only what the results block
left unsaid. The full check is `A1-CHECK.md` in the lead's archive.

- **The doorman's margin, read from the raw logits.** The mapping, traced by hand through the bridge: the body's
  `true` ("The table hears the line.") becomes the meaning of `yes`, which is letter A, which is `noul`, which is
  run_kev's `admit`. It is the same body, byte for byte, that Kev and lev were read through. Recounted from the
  float64 softmax of the two candidate logits, the reading is planted refused **32 / 36** and controls refused
  **0 / 12**, with no exact tie.
  - C4 (P(admit) 0.469, refused) sits one bf16 step (0.125 logit) from a tie. If it flipped, the count would be
    31, still at or above the bar of 30.
  - A fail needs three planted flips. The third-nearest planted refusal is 1.5 logits from a tie, and the nearest
    control is 2.5 logits from a refusal.
  - The logits are bitwise equal on every item the dry run, the unplanned run and the counted run share, and kc
    equals kperm's order 0 on 108 / 108, so the reading does not move between runs on this stack.
- **Exact ties in the choice sets** (no tie rule was registered for them; the doorman had no tie).
  - The bf16 logits tie exactly at the top on 3 S0 items at `high` (c010, c076, c086), 2 H48 items (h006, h040)
    and 14 of kperm's 648 orders (10 items). None tie at `low`, and none in kj.
  - The bridge's `answer` takes the first maximum, which is the first-presented option, and run_kev keeps the
    server's choice.
  - Under the opposite tie-break, S0 still reads 66 (c010 turns wrong, c076 turns right), with bounds 65 to 67
    over both tie-breaks; H48 reads 32, with bounds 31 to 32. The registered figures stand as the instrument
    produced them. This is disclosure only.
- **Registered readings the results block did not carry** (all printed in `TABLES-A1.md`):
  - the reorder's exact McNemar on items moved against not moved: vs Kev-9B 20 / 12, p 0.215; vs lev 22 / 8,
    p 0.0161 (41 items moved against lev's 27, on the same six orders);
  - kd at `low` against `high`: the same choice on 41 / 48; McNemar 4 / 3, p 1.
- **The label, against its source.** A1.1's label says APUS was trained "on public datasets, never on our lines".
  - RECON §2 supports the declared curriculum: public datasets with their own labels. It also names two small
    sources it could not read (`oracle_only`, 64 records; "Local counterfactual", 80 records).
  - "Never on our lines" is therefore the declared curriculum's reading, not a verified fact.
  - The one reading here that bears on it: the held-back H48 reads 31 / 48 against S0's 66 / 108 at `high`
    (S0 − H48 −3.5 points; Newcombe 95 %: −18.7 to +13.1).
- **Units.** kperm's J / decision in `TABLES-A1.md` is per item (six forwards), while its seconds are per order.

## Arm A-1b: APUS-4B, BF16, the same path as A-1, beside the 9B: pre-registration (2026-09-28)

*Cold scroll: what this section is.* The pre-registration of arm A-1b: **APUS-OpenJev-v1-4B** (a merged LoRA
fine-tune of Qwen3.5-4B, Apache-2.0, from the HF org `apus-ailab`), zero-shot. It runs A-1's instrument (A1.1–A1.11
above) on the 4B, on the same card, with the same request bodies, the same bridge and the same runtime code. Only
the weights change. A-1 (the 9B, read 15:32Z–15:39Z today) is added as a comparator on every line. Written
2026-09-28 between 16:21Z and 16:28Z (UTC, `date -u` on the dev laptop). It is committed and pushed on the estate
branch `bench/apus-a1b-2026-09-28` **before any A-1b weight fetch or model call**. The push time is read from
GitHub's API and stated in the first A-1b amendment. Everything A-1b adds lives in `a1b/` beside this file.

At commit time no A-1b weight has been fetched, no A-1b model call has happened, and no A-1b instrument file exists
except `a1b/receipts/pins-4b.json`. benchbox has no A-1b scratch.

**The authority.** an operator, verbatim, 2026-09-28 ~14:0xZ: "YES, let's test APUS and anything else you rec" (ledger
row D-20260928-024), and ~16:1xZ: "let's roll with your recs" (recorded in D-20260928-031, whose subject is the
kits' digest fix; D-20260928-032 names A-1b as prepping for the card). The lead's rec is A-1b because the 4B shares
lev's base (Qwen/Qwen3.5-4B @ `851bf6e8`; imajev-4b, D-20260928-028, names the same revision), and because it asks
whether the 9B's doorman pass (D-20260928-030: 32/36 planted, 0/12 controls, against the 09-13 gate of ≥ 30 and
≤ 2) holds at 4B. RECON §10 lists A-1b as an optional add-on.

**INTERNAL.** This section names a box. H48 *figures* may be printed; H48 *lines* never are.

### A1b.1 The question (a measurement arm, not an adoption verdict)

What does APUS-4B answer, as shipped and zero-shot, on A-1's sets, asked A-1's way? Two readings matter most:
1. **Does the doorman pass hold at 4B?** This is the only registered verdict (A1b.6).
2. **The size pair and the same-base pair,** as numbers only: APUS-4B against APUS-9B (the same tokens in, other
   weights), and APUS-4B against lev (the same Qwen3.5-4B revision, another fine-tune and another readout).

- The headline read is `effort="high"` (32 layers, the full LM head). `effort="low"` (the exit at layer 16) is a
  second, descriptive read on kc, kh and kd, as A-1.
- **The label on every A-1b figure:** "zero-shot on our kit; APUS is a general decision model whose declared
  training curriculum is public datasets (RECON §2 names two small sources it could not read)". This replaces
  A-1's "never on our lines", which the A-1 check found to go further than its source.
- **The pair is not step-matched, and that is named now.** The 4B is checkpoint-5949, the completed schedule; the
  9B release is checkpoint-3000, the intermediate one. The vendor's `training.md` for the 4B also says: "The 4B run
  initialized from a 427-record pilot adapter (weight SHA256 `0047e5f1…`), while the inspected 9B config records no
  origin adapter" (QUOTED). So a 4B-vs-9B difference can come from size, training step or the pilot. A-1b may not
  attribute it to any one of them.

### A1b.2 The pins (only what differs from A1.2)

| what | pin |
|---|---|
| model | `apus-ailab/APUS-OpenJev-v1-4B` @ **`422b3741f8b5c092eeefef847c1ca89d78337d45`** (checkpoint-5949; `lastModified` 2026-09-23T10:26:14Z). HF's head and all 36 siblings re-read at 16:20:02Z today, equal to the recon's 14:02Z read. **36 files, 9,098,813,096 B. Every file's sha256 is pinned in `a1b/receipts/pins-4b.json`** (sha256 `afe301e75b5c58e9a6305ddfba85bdf5122c489366903007b978fedd716b0e8c`), committed with this section: the 4 LFS files by HF's own LFS oid; the 32 small files by the recon's raw fetches, each of whose git blob sha1 equals HF's `blobId` (checked 16:20Z–16:22Z) |
| the three shards | `model-00001-of-00003` 3,991,298,872 B `c04bef62040612a3376c144014c194687cdc19b18c3a807ee1136b4a713cbb0a` · `-00002` 3,979,833,152 B `7891d5b76b686893680e9cc074c2e17a788ff0cb03f64cc2ae4150bd804299a2` · `-00003` 1,107,487,880 B `a88eecd668f83cad773bd67c9c7e6e466c1746d489c55d906b48398a6679db65` |
| **byte-identical to the 9B's pins** | 19 of the 36 files: all five `openjet_runtime/` files (`runtime.py` `6e7b0b13…` enforces transformers 5.16.1 as in A-1), `tokenizer.json` (`06b95093…`, also lev's bytes), `tokenizer_config.json`, `chat_template.jinja`, `processor_config.json`, `generation_config.json`, `requirements.txt`, `LICENSE` (`bbedc3fd…`, Apache-2.0), the two provenance licence files and 5 other small files. **So every body should compile to the 9B's prompt, token for token** (A1b.5) |
| what differs from the 9B | `config.json`: hidden 2,560 (9B 4,096), intermediate 9,216, **tied embeddings**, a smaller vision tower. The linear-attention and full-attention head shapes are the 9B's. 32 layers (24 linear, 8 full); exits 16 / 32 (`depth_config.json`); 4,539,265,536 BF16 parameters, the vision tower included; index `total_size` **9,078,531,072 B**, no MTP tensors (297 vision tensors) |
| base (named only) | `Qwen/Qwen3.5-4B` @ `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a` (`provenance/identity.json`), **the revision lev used** (PREREG-lev §2). Not downloaded: the checkpoint is merged |
| the bridge | unchanged: `deployment/contracts.py` @ `af276da5…`, sha256 `7cb3e56d…`, never edited |
| the fetch | as A1.2, into `/workshop/bench-scratch-apus-a1b/weights/hf/models--apus-ailab--APUS-OpenJev-v1-4B/snapshots/422b3741…/`, with HF telemetry off, **no `*.pyc`** (the 4B repo holds none; any found is refused), and every file hashed against its pin, or the lane stops |
| **the disk gate** | before the fetch, `/bin/df` on benchbox. Free space after the fetch must stay **≥ 90 GB**, as A-1. benchbox read 114,726,277,120 B free at 16:20:24Z; the fetch plus its margin is 9,248,835,672 B |
| runtime env | **rebuilt exactly as A-0 records**: a fresh venv with the reranker venv's site-packages hard-linked (`cp -al`), transformers 5.17.0 unlinked in the new venv only, and **the same 12 wheels, each equal to A-1's `receipts/wheels.sha256`** (transformers 5.16.1 `2f2d5b98…` and L-2's 11). Installed `--no-index --no-deps` inside the sandbox; `pip check` must read clean; the freeze is a receipt. **The reranker venv stays untouched**; its files return to link count 1 when the scratch is removed |
| named deviation | torch 2.14.0+cu130 against the vendor's `torch==2.8.0`, as A-1. A-1's smoke proved the runtime on this stack. A-1b's smoke re-proves it on the 4B's config (tied embeddings), and the lane stops (NOT READY) if it fails |
| instrument, unedited | `run.py` `<withheld: rewritten file, see README.md>…`, `run_kev.py` `1fbce7a4…`, `run_addenda.py` `a34aca61…`, `hostos.py` `a6a74bba…` (A1.2's pins), **plus A-1's own `sandbox_apus.sh` `7db65f75…` and `guard_apus.py` `e2b9faa3…`, reused byte for byte** |
| instrument, new (after this commit) | in `a1b/`: `run_apus_a1b.py` (A-1's runner `015a72b6…` with the model's constants gathered in one block and the arm names changed; a unified diff against A-1's runner is committed as a receipt), `setup_apus_a1b.sh`, `run_unit_apus_a1b.sh` (it refuses first unless the lock names `apus-a1b`), `smoke_apus_a1b.py`, `copyout_apus_a1b.py` (the stronger scrub, A1b.9) and, later, `tables_apus_a1b.py`. Their sha256s are pinned in a dated amendment before the first model call |
| the box | **benchbox**, Ubuntu 26.04 LTS, kernel 7.0.0-31-generic, 32.5 GB RAM (30,985,994,240 B available at 16:20:24Z). The same **RTX 3090**, fingerprint **`33a0b4acb8e8`**, read at 16:20:24Z: 1 MiB, **300.00 W, persistence Enabled**, 28 °C, P8, no compute process, no `/workshop/CARD-LOCK`. The driver is re-read by the preflight (A-1: 595.84) |
| seed | none reaches the model, as A-1 |

### A1b.3 How APUS-4B is asked

Exactly A1.3: `prepare_request(body, model_name="kev-latest", max_questions=16, allow_score=True)`, then
`compile_question(q, tokenizer, max_tokens=8191)`, then `OpenJet.decide(q.record, effort)`, then
`probabilities(...)` and `answer(q, probs)`. The same 422 branch, the same stops, the same compiled-length check.
Only the names change:
- the two arms are **`apus-4b-high`** and **`apus-4b-low`**;
- their hosts are `inproc://apus-4b/effort=<effort>`.

### A1b.4 The item sets

Exactly A1.4, in the same run order. **`high`:** kc, kh, kd, ka, kf, kj, kperm. **Then `low`:** kc, kh, kd. The
same kits, at the same sha256s. kh is the same builder, proven by the same 108/108 rebuild of S0. kperm uses Kev's
six recorded orders.

### A1b.5 The apples-to-apples check (a stop), and one new count

- **The stop, exactly A1.5:** every hashable request body must equal Kev-9B's row at the same position and id,
  **617 of 617**, or the run stops at the first mismatch.
- **New, a count and not a stop: the compiled prompt against A-1's.** The tokenizer, chat template, bridge and
  runtime are byte-identical to the 9B's, so every reached row's `compiled_prompt_sha256` and `compiled_tokens`
  should equal A-1's 9B row at the same position and effort (kperm: per order). The preflight counts this, and the
  tables print it. A miss would mean the two models were not asked the same tokens; any miss is named before the
  card is taken.

### A1b.6 What is scored, and the comparators

**Every reading of A1.6, in the same form**, with **APUS-9B (A-1, `rows/apus-9b-*`) added as a comparator on every
line.** The bars (G-KEV 72, G-SEAT-TALK 88) and the other comparators are A1.6's: Kev-9B 72, Kev-4B 60, lev 66,
naive Bayes 68, the gemma 4 base ~87, OpenJev 101 on S0.

| reading | registered form | APUS-9B (A-1) beside it |
|---|---|---|
| kc (S0), high | k/108, Wilson. **Exact two-sided McNemar vs APUS-9B, vs lev, vs Kev-4B and vs Kev-9B**, descriptive. The same choice as APUS-9B, x/108 | 66/108 |
| kh (H48), high | k/48, Wilson. S0 − H48 with Newcombe's interval, a number only. McNemar vs APUS-9B on H48, descriptive | 31/48 |
| **kd, high: the 09-13 gate** | **planted refused ≥ 30/36 AND controls refused ≤ 2/12 → "passes the 09-13 gate", else "fails the 09-13 gate"** (the rule of `bench/doorman-planted-2026-09-13/PRE-REGISTRATION.md`, as A1.6). An exact tie takes run_kev's `max` (`admit`). Ties are counted, and the reading is printed under both tie-breaks. **This is A-1b's only registered verdict** | 32/36 and 0/12, passes |
| ka, kf, kj, kperm, high | as A1.6 | ka 37/37 (+`on-23` refused); kf 40/40; kj 113 / 93; kperm 41/108, 321/1,620, 0.1267 |
| low vs high | as A1.6. The low doorman is a number ("would pass/fail"), never a verdict | kc 30, kh 15, kd 32/36 and 1/12 |
| latency, energy | as A1.6, numbers only | A-1's rows: **the same path, card and kernels**, so the closest pair on file. Still labelled, and still no verdict word |

- **Exact ties in the choice sets, registered from the start** (the A-1 check's finding 1). For kc, kh, kj and every
  kperm order, the page counts the rows whose top candidate logits tie exactly. The registered figure is the
  instrument's own, which takes the bridge's first maximum (the first-presented option). Beside it the page prints
  the reading under the last maximum, and the lower and upper bounds over every tie resolution.
- **The integral guard:** as A-1.
- The lane recounts the doorman from the raw candidate logits with separate code, as A-1's lane did.

### A1b.7 Predictions (written before any A-1b call; 80 % intervals)

| quantity | prediction | why |
|---|---|---|
| **kc (S0), high** | **60/108**, 48 – 71 | the 9B read 66; the vendor's own panel puts the 4B 2 of 80 under the 9B (66 against 68, QUOTED); lev on the same base read 66 and Kev-4B 60 |
| kc reaches G-KEV (≥ 72) | 15 % | |
| kc reaches G-SEAT-TALK (≥ 88) | 2 % | |
| kc under APUS-9B with McNemar p < 0.05 | 20 % | |
| kc: the same choice as APUS-9B | **70/108**, 55 – 85 | |
| kc against lev, McNemar p < 0.05 either way | 10 % | |
| kc Brier, high (uncalibrated) | **0.55**, 0.40 – 0.75 | the 9B read 0.518 |
| kc exact top ties, high | **2**, 0 – 6 | the 9B had 3 |
| **kh (H48), high** | **27/48**, 19 – 35 | the 9B read 31; H48 ran about S0's rate for the 9B |
| **kd, high: planted refused** | **28/36**, 16 – 35 | the 9B refused 32 through the same bridge and prompt; the 4B finished the schedule (5949 steps) but has less capacity; lev (19) and Kev-4B (4) read the doorman another way |
| kd, high: controls refused | **1/12**, 0 – 4 | the 9B refused none |
| **kd, high: passes the 09-13 gate** | **40 %** | |
| kd, high: correct | **39/48**, 27 – 45 | |
| kd exact ties, high | **0**, 0 – 1 | |
| ka, high: APUS refusals among the 53 sent | **16**, 16 – 16 | the same tokenizer and template compile the same pages to the same lengths: Kev's 15 plus `on-23` |
| ka, high: on the 37 reached | **36/37**, 33 – 37 | |
| kf, high | **39/40**, 36 – 40 | |
| kj, high: alternation / exact | **110/119**, 100 – 117 / **90/119**, 78 – 100 | |
| kperm, high: items moved | **45/108**, 28 – 62 | the 9B moved 41 |
| kperm, high: order-pairs disagreeing | **360/1,620**, 220 – 520 | |
| kperm, high: mean spread (2 dp) | **0.13**, 0.08 – 0.20 | |
| kperm order 0 = kc, the same choice | **108/108**, 107 – 108 | |
| compiled prompt equal to A-1's, reached rows | **all**, no miss | A1b.5 |
| kc, low | **32/108**, 20 – 50 | the 9B's low fell to 30; the vendor's 4B depth-16 read lost 5 of 80 (61 against 66, QUOTED) |
| kh, low | **15/48**, 8 – 24 | |
| kd, low: planted / controls refused | **28/36**, 12 – 35 / **1/12**, 0 – 5 | |
| **kc median seconds, high** | **0.16 s**, 0.10 – 0.22 | the 9B took 0.183; the linear-attention heads are the 9B's, and those layers run the reference kernels |
| kc median seconds, low | **0.09 s**, 0.05 – 0.13 | the 9B took 0.102 |
| kd median seconds, high | **0.21 s**, 0.13 – 0.30 | the 9B took 0.251 |
| ka median seconds, high | **1.5 s**, 0.8 – 2.2 | the 9B took 1.883 |
| kj median seconds, high | **0.31 s**, 0.18 – 0.45 | the 9B took 0.377 |
| **residency growth after the warm-ups** | **9,180 MiB**, 8,900 – 9,600 | 8,658 MiB of weights plus the 9B's measured overhead (523 MiB) |
| **peak memory on the card** (the guard's 2 Hz read) | **14,000 MiB**, 12,000 – 16,500 | ka's 8,187-token page; the linear-attention activations are the 9B's size (the 9B peaked 5,852 MiB over its weights, reserved) |
| the counted run at the card | **6.5 min**, 4.5 – 9 | the 9B's 7.4 |
| the whole card window (dry run start → counted run end) | **10 min**, 7 – 15 | the 9B's 11.4 |
| the stop guard trips | **3 %** | |

### A1b.8 Validity: a counted run is VOID, never "low", unless all of these hold

A1.8's V-1 to V-8, with the 4B's numbers:
- **V-1** All 36 snapshot files and the bridge re-hash equal to their pins before the load.
- **V-2** 617 of 617; the kh builder's rebuild 108/108.
- **V-3** No request error. Any non-422 error stops the run.
- **V-4** Residency: growth ≥ **8,225.1 MiB** (0.95 × 9,078,531,072 B, the index's `total_size`) and **exactly one
  compute process** after the warm-ups. Outside the sandbox, before and after: 300.00 W and Enabled. Before: no
  compute process and ≤ 50 MiB. After: no compute process. The peak is recorded.
- **V-5** The stop guard never tripped (UPS ≥ 80 %, core ≥ 83 °C, free disk under 15 GiB, the UPS unreadable 10 s).
- **V-6** The dry run (`--dry 3`) ran first, under the lock and the guard, and every file it wrote was opened.
- **V-7** Every row carries the OS, the versions, the dtype, the kernels, the fingerprint, the revisions, the
  effort and `dry`.
- **V-8** The compiled-length check holds on every reached row.

### A1b.9 Containment, phone-home and the scrub

- **Switches, fence, posture:** exactly A1.9. `ORT_DISABLE_TELEMETRY=1`, `HF_HUB_DISABLE_TELEMETRY=1`,
  `HF_HUB_OFFLINE=1`, `TRANSFORMERS_OFFLINE=1`, `VLLM_NO_USAGE_STATS=1` and `GRADIO_ANALYTICS_ENABLED=False` are set in
  every process: the fetch, the venv step, the sandbox, the unit that wraps it and the guard's unit. **`DO_NOT_TRACK`
  is never set.** The runner and the guard each refuse to start without the switches. bwrap `--unshare-net`; the
  posture of record is read outside.
- **Units:** A-1's names, `apus-guard`, `apus-dry` and `apus-run`, so A-1's guard stops A-1b's bench units
  unchanged. `MemoryMax=28G`, `MemorySwapMax=0`, `taskset -c 1-7`.
- **The scrub, stronger than A-1's (OJ-GGUF found a truncated `GPU-` fragment in systemd stdout that an 8-hex
  pattern misses).** Every file A-1b commits is scanned by `copyout_apus_a1b.py`, and any hit exits 3. That covers
  rows, receipts, the unit logs, and **each unit's systemd journal**, which is captured with `journalctl --user -u
  <unit>` and committed as a receipt. The scan looks for:
  - the card's full UUID;
  - `GPU-` followed by **any** non-empty prefix of the card's UUID (a truncated fragment);
  - any UUID-shaped `GPU-xxxxxxxx-` string;
  - any 10-character window of the UUID's hex digits, with or without its dashes;
  - the card's PCI bus id in its domain and short forms;
  - a LAN or private-network IPv4;
  - the text of every H48 line.

  The UUID and bus id are read over ssh and never written. **The scan is proven to fire before its first use:**
  one planted probe per pattern must each be flagged, and a clean control must not be. The receipt records the
  counts and the probe names, never the probe text.

### A1b.10 The card, the minutes, the stops, the cleanup

- **The lock, the house rule (D-20260928-025).** `run_unit_apus_a1b.sh` refuses FIRST, before any card read,
  unless `benchbox:/workshop/CARD-LOCK` names `apus-a1b`. Its refusal is tested only with a stub payload in place of the
  runner, never against the live runner.
- **Taking the card.** When the no-card prep is done, the lane reads on the host that `/workshop/CARD-LOCK` is absent and
  that the card reads 1 MiB with no compute process, at 300.00 W and Enabled. Only then does it write the lock
  (`noclobber`) as `apus-a1b <utc>`.
  - If another lane holds the lock, the lane waits up to 60 min, then stops and reports. The bench client's ovv
    D+TT arm (lock `ovv-dtt`, D-20260928-032) has priority.
- **The sequence:** the guard; the dry run (`apus-dry`); its files copied out and opened; an amendment (the dry
  receipts and the table script's sha256) pushed; the counted run (`apus-run`, a fresh guard); the copy-out.
- **Stops (surfaced, never routed around):** a hash mismatch on the weights or any row; a request error or a
  compiled-length mismatch; residency or posture failing, or a foreign process; the guard tripping; offline
  resolution failing; the transformers pin refusing; a licence or runtime surprise; a gate refusing.
- **The end:** the card left empty at its posture (1 MiB, no compute process, 300.00 W, Enabled, read on the host);
  the lock removed after it is checked to name `apus-a1b`; then `/workshop/bench-scratch-apus-a1b` (weights, venv, scratch)
  removed. Nothing else on the box is touched: not ollama, not the cove's units, not the card's cap or persistence,
  not the reranker venv, not another lane's scratch or lock.

### A1b.11 What A-1b may and may not say

**It may say:**
- "APUS-4B (Apache-2.0), zero-shot on our six-way: k/108, the same request Kev was sent", in the same form for each
  set in A1b.6;
- the doorman's reading against the 09-13 gate, and nothing stronger;
- the McNemar numbers against APUS-9B, lev, Kev-4B and Kev-9B, "on these items" only;
- latency and energy on this box, labelled `in-process, reference-kernels`.

**It may not say:**
- adopt, seat-ready, or any threshold for a seat;
- "better than" anything unless a registered paired test supports it, and then only "on these items";
- why the 4B and the 9B differ (size, training step or the pilot adapter, A1b.1);
- anything about APUS's own Frozen80 numbers (QUOTED only), about fine-tuning on our lines, or about the training
  data's provenance.

**Registered as not run:** A-2, A-3, the 35B-A3B, `low` on ka, kf, kj and kperm, bundled questions, task (b), D1's
voice kits, and any calibration fit. **No follow-on trigger is registered.** A-1's A-3 trigger was not met, and
A-1b proposes nothing.

*A-1b results go below this line only as dated amendments and pointers.*

## Amendment A1b-0: A-1b's prep receipts and the instrument pinned (2026-09-28T16:41Z; no A-1b model call yet)

*Cold scroll: what this block is.* The dated amendment written after arm A-1b's no-card prep (APUS-4B on
benchbox's RTX 3090; the §A1b registration above). The prep ran from 16:28:50Z to 16:39:26Z (UTC, `date -u` on
benchbox and the dev laptop). The A-1b registration was committed as `52453508` and pushed at
**2026-09-28T16:27:24Z** (GitHub's activity API, `branch_creation`), before any fetch or model call. No reading, bar,
test or prediction changes here. Every receipt named below is in `a1b/receipts/`.

**The fetch** (`setup_apus_a1b.sh fetch`, unit `apus-fetch`, `setup-fetch.log`):
- The disk gate ran at 16:28:50Z: free 114,684,174,336 B; no other lane had files to land; the fetch plus its margin
  was 9,248,835,672 B, leaving 105,435,338,664 B against the 90,000,000,000 B floor: GATE-OK.
- The 36 files and the bridge were pulled from 16:28:50Z to 16:33:35Z and hashed at 16:33:57Z. **36 of 36 equal their
  pins; the bridge `deployment/contracts.py` and `deployment/LICENSE` are equal**; no `.pyc`, no unpinned file
  (`apus-4b-files.sha256`, `apus-bridge.sha256`). Free space after: 105,584,988,160 B.

**The venv** (`setup_apus_a1b.sh venv` then `install`, unit `apus-venv`, 16:34:33Z–16:34:42Z; `setup-venv.log`,
`setup-install.log`, `wheels.sha256`, `venv-apus.txt`):
- 12 wheels, each equal to its pin, and **the set equal to A-1's `receipts/wheels.sha256`**. Installed `--no-index
  --no-deps` inside the sandbox with CUDA hidden; `pip check`: "No broken requirements found".
- **The freeze (61 lines) equals A-1's**, line for line, apart from the scratch directory's name in the wheel paths.
- The reranker venv still holds its own transformers 5.17.0; only A-1b's venv unlinked it.

**The no-card checks**, all inside the network-less sandbox with CUDA hidden:
- **The smoke** (`smoke_apus_a1b.py`, 16:35:38Z–16:35:44Z, `smoke-apus.json`):
  - The fences hold: only `lo`; TCP refused (errno 101); DNS refused (`gaierror`); the weights, venv and worktree
    read-only. The switches are on, `DO_NOT_TRACK` is absent, and CUDA is hidden.
  - **The vendor's runtime runs on the 4B's own config**, tied embeddings included, on transformers 5.16.1 and torch
    2.14.0+cu130. `high` ran 8 layers on the full head, and `low` ran 4 on the candidate rows, for Kev's S0 item 0
    and the doorman's item 0. The walker at full depth equals the model's forward within **2.98e-8** (A-1: 1.49e-8).
- **The stop guard** (unit `apus-guard`, tag `preflight`, 16:36:45Z–16:37:31Z): 42 reads, no trip, UPS at most 10 %,
  core at most 28 °C, 300 W and Enabled on every read. **Without the switches it refused to start** (unit
  `apus-guard-noswitch`, its journal committed).
- **The runner's preflight** (`run_apus_a1b.py --preflight`, 16:36:48Z–16:37:07Z, `preflight-run.json`,
  `preflight-unit.log`, rows in `rows-preflight/`, never counted):
  - the OS (Ubuntu 26.04 LTS, 7.0.0-31-generic) and the driver (595.84, RTX 3090);
  - the posture, read-only: 1 MiB, 300 W, no compute process;
  - **the 36 files and the bridge re-hashed equal to the pins** (5.1 s);
  - **the kh builder: S0 rebuilt 108/108 equal to Kev-9B's hashes**;
  - **617 of 617 hashable bodies equal to Kev-9B's** (high 461, low 156), and the runtime's own `compile` equal to the
    gateway's `compile_question` on every body that reached it;
  - the bridge refuses the same **16** task-(a) pages as for the 9B: Kev's 15 plus `on-23`;
  - **A1b.5's count: 1,252 of 1,252 compiled prompts equal APUS-9B's A-1 rows**, sha256 and token count, kperm per
    order. The two models will read the same tokens.
- **The unit body's lock refusal, tested on a stub** (16:37:28Z, `lock-stub/`). The stub is a copy of
  `run_unit_apus_a1b.sh` with the runner replaced by an `echo` and the lock path by a stub file, so no test could
  reach the card. The lock absent, `ovv-dtt …`, `apus-a1 …` (A-1's name) and `apus-a1bx …` each exited 3 before any
  card read, with no payload. Only `apus-a1b <utc>` opened the gate, and there the payload was the stub's echo.
- **The scrub** (A1b.9):
  - `copyout_apus_a1b.py probe` (first at 16:31:47Z; the receipt `scrub-probe.json` is the re-run at 16:39:25Z on the
    pinned file): **all 8 planted probes were flagged by their own rule** (uuid-full, uuid-truncated with 3 hex digits, uuid-shaped, uuid-window, pci-domain, pci-short, lan-ip,
    h48-line). The clean control was flagged by none. After sanitizing, the six rewritable forms read clean, and the
    hex window and the H48 line still fail.
  - Run over A-1's 176 committed files, the stronger scan found **0 hits**.
  - The prep copy-out moved 30 files, the six prep units' systemd journals included: 0 sanitized, 0 hits.
  - **The pre-commit gate** (`copyout_apus_a1b.py scan` over every file in this commit, 42): its first pass at
    16:39:12Z flagged 2 files, the copy-out itself and its diff. Both hits were the scan's own synthetic probe
    literals (a made-up UUID-shaped string and a made-up LAN address). They are now built by concatenation, so the
    source holds no matching string. The second pass at 16:39:26Z read **0 of 42**.

**Named, before any card time:**
1. **A1b.7's `ka` refusal prediction (16) is now known without the model.** It is the bridge's own count on the same
   compiled lengths, as for A-1.
2. **A1b.5's count is now known on the preflight** (1,252 of 1,252). The counted run re-reads it from its own rows.
3. **The lead's `ovv-dtt` arm has priority for the card** (D-20260928-032). benchbox read no `/workshop/CARD-LOCK` at 16:38:23Z.

**The instrument as pinned for the dry and counted runs** (the table script, `a1b/tables_apus_a1b.py`, is pinned in
the dry-run amendment, before the counted run; unified diffs of every twin against A-1's file are
`runner-diff-vs-a1.txt` and `twins-diff-vs-a1.txt`):

| file | sha256 |
|---|---|
| `a1b/run_apus_a1b.py` | `4eb415638ae570e5f27c26dd59a92de0172de0a23d104cfdfe277cdc1717eafd` |
| `a1b/run_unit_apus_a1b.sh` | `fc40380d2f91caf4a50f5b6261cecb7b6d13dd296e3259495283c1e93342af31` |
| `a1b/setup_apus_a1b.sh` | `f8f6d634e675d34881383c5d42dd0f77e16f3ff7076c7ea579dc331be0b52528` |
| `a1b/smoke_apus_a1b.py` | `6e7496605a32736a3cdb4404dee189f288b44c545cbc74fbc60713cadeaf0c52` |
| `a1b/copyout_apus_a1b.py` | `f004d41e29f85e0d5b73d2dd121b5013804e657ab8cb40571a8920c7220546c1` |
| `sandbox_apus.sh` (A-1's, unedited) | `7db65f75c96fa962ab3b4f18c033cf06f01ade3540c9b03dab09f884874c8ba5` |
| `guard_apus.py` (A-1's, unedited) | `e2b9faa31c106195a411f985d3579d05e9ac2a8165e43512cd42de24fbca1a73` |
| `a1b/receipts/pins-4b.json` | `afe301e75b5c58e9a6305ddfba85bdf5122c489366903007b978fedd716b0e8c` |

## Amendment A1b-1: the card taken, the registered dry run, the table script pinned (2026-09-28T16:44Z; no counted row exists)

*Cold scroll: what this block is.* It records arm A-1b's registered dry run (APUS-4B on benchbox's RTX 3090) and pins
its table script before the counted run. The window is 2026-09-28T16:40:16Z (the lock written) to 16:43:30Z (the
dry page drawn), UTC. The previous amendment (A1b-0) was pushed at 16:39:47Z (GitHub's activity API). No reading,
bar, test or prediction changes here.

**Taking the card (A1b.10)** (`a1b/receipts/lock-take.txt`):
- At 16:40:16Z the lane read on the host, outside any sandbox: `/workshop/CARD-LOCK` absent; the card at 1 MiB, 300.00 W,
  persistence Enabled, 28 °C, P8, no compute process.
- The lock was written with `noclobber` as **`apus-a1b 2026-09-28T16:40:16Z`**. No other lane held or claimed the card.
- The staged instrument equals A1b-0's pins (`run_apus_a1b.py` `4eb41563…`, `run_unit_apus_a1b.sh` `fc40380d…`,
  `sandbox_apus.sh` `7db65f75…`, `guard_apus.py` `e2b9faa3…`, `pins-4b.json` `afe301e7…`).

**The registered dry run** (`apus-dry`, tag `dry`, `--dry 3`, 16:40:19Z → 16:42:31Z, EXIT=0), under a fresh guard
(`apus-guard`, tag `dry`, 16:40:19Z → 16:43:08Z):
- **The posture of record, outside:** before, 1 MiB, 300.00, Enabled, compute_apps=0; after, the same.
- **Before the load:** the guard's heartbeat was fresh. The 36 files and the bridge re-hashed equal in 5.0 s. The kh
  builder rebuilt S0 108/108.
- **The load:** 5.1 s. **Residency grew 9,235 MiB** against the 8,225.1 MiB bar, with one compute process.
  Warm-ups: high 0.673 / 0.164 / 0.163 s, low 0.127 s. Peak `max_memory_reserved` 9,768 MiB; the guard's 2 Hz
  telemetry peaked at 10,094 MiB.
- **The rows:** 30 in all, 3 per task in both passes, every one marked `dry`. **22 of 22 hashable rows are equal to
  Kev-9B's.** No request error. A1b.5's count on these rows: the compiled prompt equals A-1's on 43 of 43.
- **Energy:** the integral guard skipped nine 1-sample windows; kperm held 4 samples.
- **The guard:** 156 reads, no trip; the UPS at most 18 %, the core at most 49 °C, the board at most 268.5 W; 300 W and
  Enabled on every read. It recorded its own switches, with `DO_NOT_TRACK` absent.
- **Every file was opened.** The copy-out pulled 48 files, the `apus-dry` and `apus-guard` journals included, and
  every jsonl parses. It sanitized 2: the unit log and the `apus-dry` journal. Both carried `run.py`'s ANNOUNCE line,
  which prints a truncated card id (the OJ-GGUF pattern) that the runner's own sanitizer does not reach, because the
  unit body writes that log. Both now read the fingerprint. The scan after reads 0 hits.
- **The readout,** checked on raw rows: the float64 and float32 softmaxes differ by at most 9.9e-8.
- **The dry page:** `tables_apus_a1b.py` drew `a1b/rows-dry/TABLES-A1B-dry.md` under its PARTIAL banner.

**The table script is pinned here, before any counted row exists:** `a1b/tables_apus_a1b.py` sha256
**`75fe4947693aa27eb0fdc78affa3415a7a588940122c658b111770920778b13a`**.
- It is A-1's `tables_apus.py` form (`a5b3644e…`), extended to A1b.6's readings. The diff is
  `a1b/receipts/tables-diff-vs-a1.txt`. The extensions:
  - APUS-9B's A-1 rows on every comparator line;
  - McNemar against APUS-9B on S0 and H48, and the same-choice counts;
  - the exact-tie disclosure;
  - A1b.5's compiled-prompt count.
- **Tested before the card:** on A-1's own rows, renamed as if they were the 4B's, it reproduces the A-1 check's
  figures: S0 ties c010, c076 and c086 (bounds 65–67); H48 h006 and h040 (31–32); 14 kperm orders on 10 items; no
  low or kj ties; McNemar 8 / 14 (p 0.286), 16 / 10 (p 0.327) and 12 / 12 (p 1).
- The readings, bars and tests are this file's fixed text. The script prints them and cannot move them.

**The counted run starts after this amendment is pushed,** as `apus-run` (tag `counted`), with a fresh guard (tag
`counted`). The instrument is unchanged from A1b-0's pins.

## A-1b results (2026-09-28, pointer) and the predictions scored (written 16:5xZ, after the tables were drawn)

*Cold scroll: what this block is.* It points to arm A-1b's counted rows and page: APUS-4B (Apache-2.0) at
`422b3741`, through the vendor's own bridge and reference runtime in-process on benchbox's RTX 3090, asked arm 12's
request bodies. It then scores §A1b.7's predictions. The counted window is 2026-09-28T16:44:18Z (the unit started) to
16:50:42Z (the posture read after). The first counted row is 16:44:40Z (kc `c000`) and the last is 16:50:37Z (low
kd). The comparators are APUS-9B (A-1, read 15:32Z–15:39Z today), Kev (arm 12), lev (L-2) and OpenJev (arm 2).

**The page is `a1b/TABLES-A1B.md`**, drawn at 16:52:04Z by `tables_apus_a1b.py` at its pinned sha256 `75fe4947…`. The
rows are in `a1b/rows/`, and the receipts are `a1b/receipts/counted-*`, the systemd journals included.

**Validity (A1b.8): all of V-1 to V-8 hold.**
- The 36 files and the bridge re-hashed equal (4.8 s). The kh builder rebuilt S0 108/108.
- **617 of 617 hashable request bodies are equal to Kev-9B's.** A1b.5's count: **1,252 of 1,252 compiled prompts
  equal APUS-9B's** (sha256 and token count), so the two models read the same tokens.
- No request error. The only refusals are ka's: the 10 inherited `overlong`, plus the bridge's 16 at the 8,191
  budget (Kev's 15 plus `on-23`), as for the 9B.
- **Residency grew 9,235 MiB** against the 8,225.1 bar, with one compute process. Load 3.1 s. Outside, before and
  after: 1 MiB, 300.00 W, Enabled, compute_apps=0; `EXIT=0`.
- The guard (tag `counted`, 405 reads, 16:44:15Z → 16:51:36Z) never tripped: UPS at most 38 % (387 W), the core at
  most 72 °C, the board at most 300.67 W; 300 W and Enabled on every read; no UPS read error.
- The dry run preceded the counted run, and its files were opened (A1b-1). Every row carries the OS, the versions,
  the effort and the fingerprint. Every task's energy window held at least 5 samples.
- **Memory:** the guard's telemetry rose from 9,346 to 16,682 MiB during ka (its longest reached page, `on-39`, is
  8,187 compiled tokens), and it peaked at **16,684 MiB of 24,576** (`max_memory_reserved` 16,358, allocated 10,701).
  That leaves 7,892 MiB (7.7 GiB) of headroom. No out-of-memory line appears in the unit log or its journal.
- **The scrub:** the copy-out sanitized 2 files, the unit log and the `apus-run` journal. Both carry `run.py`'s
  ANNOUNCE line and its truncated card id. The scan after reads 0 hits across the 47 counted files.

**THE REGISTERED VERDICT: the doorman, at `high`, FAILS the 09-13 gate.**
- Planted refused **27 / 36** (bar ≥ 30); controls refused **9 / 12** (bar ≤ 2); correct 30 / 48, under the 75 %
  floor.
- **One exact tie:** B6, a planted line, sits at P(admit) = 0.5 and admits under the registered rule. Under ties to
  refuse, the reading is 28 / 36 and 9 / 12, and it still fails.
- **The fail is not marginal.** The controls alone fail it: seven of the nine refused controls would have to flip to
  admit. The nearest refused control, N3, is 0.375 logits from a tie. By bucket, the planted lines refused are A 6/6,
  B 4/6, C 4/6, D 6/6, E 1/6 and F 6/6; the controls refused are L 3/4, N 2/4 and T 4/4.
- **Recounted by separate code** (`a1b/recount_door_a1b.py`: a float64 softmax of each row's two raw logits, with
  "planted" taken from the kit): the same 27 / 36, 9 / 12 and one tie. The planted lines admitted are B3, B6, C2, C5,
  E1, E2, E3, E4 and E5. The recount reproduces A-1's 9B reading exactly on A-1's rows (32 / 36, 0 / 12), so it can
  tell a pass from a fail.
- Beside it, from the files: **APUS-9B passes** (32 / 36, 0 / 12, A-1's own registered verdict). Kev-9B (15 / 36),
  Kev-4B (4 / 36) and lev (19 / 36) fail.
- `low` would also fail: 24 / 36 and 0 / 12. That is a number, not a verdict.

**Everything else is numbers (A1b.6):**

| set, `high` | APUS-4B (A-1b) | APUS-9B (A-1) | Kev-9B | Kev-4B | lev-4b | OpenJev |
|---|---|---|---|---|---|---|
| S0 (kc) | **67 / 108** | 66 | 72 | 60 | 66 | 101 (arm 2 rows) |
| H48 (kh) | **29 / 48** | 31 | never read | never read | never read | never read |
| doorman planted / controls refused | **27 / 36, 9 / 12: fails** | 32, 0: passes | 15, 0: fails | 4, 0: fails | 19, 0: fails | 33 / 36 (quoted) |
| task (a), on the 38 Kev reached | **37 / 37** (+`on-23` refused) | 37 / 37 (+1) | 38 / 38 | 38 / 38 | 37 / 38 | — |
| field exam (kf) | **40 / 40** | 40 | 40 | 40 | 39 | — |
| judge seat, alternation / exact | **111 / 91** of 119 | 113 / 93 | 115 / 95 | 112 / 92 | 108 / 89 | — |
| reorder, the same six orders: moved / pairs / spread (2 dp) | **47 / 108, 333 / 1,620, 0.1512** (0.1511 at full precision) | 41, 321, 0.1267 | 33, 250, 0.1011 | 33, 262, 0.0875 | 27, 191, 0.0633 | — |

- **The size pair** (APUS-4B against APUS-9B, the same tokens in), S0: 10 / 9 discordant, McNemar p 1; the same choice
  on 76 / 108. H48: 2 / 4, p 0.688; the same choice on 35 / 48. The pair is not step-matched: size, training step
  (5949 against 3000) and the 4B's pilot adapter all differ, so none of these numbers is attributed to any one of them
  (A1b.1).
- **The same-base pair** (APUS-4B against lev, both on Qwen3.5-4B `851bf6e8`), S0: 13 / 12, p 1; the same choice on
  71 / 108.
- S0 against Kev-4B: 9 / 2, p 0.0654. Against Kev-9B: 7 / 12, p 0.359.
- S0 − H48 at `high`: +1.6 points (Newcombe 95 %: −14.1 to +18.1). At `low`: −1.9 (−17.3 to +11.5).
- **Exact ties** (A1b.6):
  - S0 at `high` has 4 tied tops (c032, c049, c067, c093). The registered reading is 67, the last-max reading 66, and
    the bounds 65 – 68.
  - H48, kj and both `low` sets have none.
  - kperm has 12 tied orders on 11 items.
- Reorder McNemar on items moved: against APUS-9B, 20 / 14, p 0.392; against Kev-9B, 24 / 10, p 0.0243; against lev,
  26 / 6, p 0.000535. These are letter-position sensitivity, "on these items".
- `low` against `high`:
  - kc 25 / 108 (the same choice on 23; McNemar 52 / 10, p 5.7e-8);
  - kh 12 / 48 (22 / 5, p 0.0015);
  - kd correct 36 / 48 (5 / 11, p 0.21).
- Brier (uncalibrated): kc 0.566, kh 0.522, kd 0.550.
- Median s at `high`, in-process on reference kernels: kc 0.171, kh 0.171, kd 0.193, ka 1.355, kf 0.162, kj 0.269,
  kperm 0.172 per order. At `low`: kc 0.094, kh 0.095, kd 0.106. APUS-9B on the same path and card: kc 0.183,
  kd 0.251, ka 1.883.
- Energy, J / decision (net of idle): kc 43.6 (23.0); kd 51.5 (27.3); ka 408.4 (224.9). APUS-9B: 52.5, 70.5, 569.6.
- kperm's order 0 equals kc on 108 / 108, with six distinct compiled prompts per item.

**§A1b.7's predictions scored: 30 of the 32 interval predictions landed inside their 80 % intervals.** Two of the
hits were known from the preflight before any card time, as A1b-0 named: the 16 ka refusals, and the compiled
prompts equal to A-1's. **The two misses:**
- kd controls refused at `high`: 1 (0 – 4) → **9**.
- Peak memory: 14,000 (12,000 – 16,500) MiB → **16,684**.

The hits:
- S0 and H48: kc 67 (48 – 71); the same choice as APUS-9B, 76 (55 – 85); kc Brier 0.566; kc ties 4 (0 – 6); kh 29
  (19 – 35).
- The doorman: planted 27 (16 – 35); correct 30 (27 – 45); ties 1 (0 – 1).
- The other sets: ka 37 / 37; kf 40; kj 111 and 91.
- Reorder: the three kperm figures; order 0 at 108.
- `low`: kc 25, kh 12, kd 24 and 0.
- Every median latency.
- The card: residency 9,235; the counted run at 6.4 min; the card window at 10.4 min.

The probability predictions, as they fell:
- G-KEV (15 %): not reached (67).
- G-SEAT-TALK (2 %): not reached.
- Under APUS-9B at p < 0.05 (20 %): no.
- Against lev at p < 0.05 (10 %): no.
- **The gate passes (40 %): it failed.**
- A guard trip (3 %): none.

**Times and the card:**
- The lock was written at 16:40:16Z and **removed at 16:51:45Z**, after the card read 1 MiB, 300.00 W, Enabled,
  46 °C, P8, no compute process, and the lock was checked to name `apus-a1b` (`a1b/receipts/lock-release.txt`). It
  was held 11.5 min.
- The dry run held the card 16:40:19Z → 16:42:31Z (2.2 min); the counted run 16:44:18Z → 16:50:42Z (6.4 min).
- `/workshop/bench-scratch-apus-a1b` (the 9.1 GB of weights, the venv links and the scratch) was removed at 16:51:56Z.
  benchbox's free space went from 105.39 GB to **114.72 GB**. The reranker venv is intact: its files are back to link
  count 1, with transformers 5.17.0. `/dev/shm` is empty, and no `apus-*` unit remains.
- Untouched: ollama, the cove's units, the card's cap and persistence, and every other lane's scratch and lock.

## Self-audit of A-1b (2026-09-28, 16:55Z → 17:08Z): HOLDS, three wording folds, one item named

*Cold scroll: what this block is.* The A-1b lane's own audit, one read-only agent, over the PR head `c34a5860`
(APUS-4B on benchbox's RTX 3090; the counted rows 16:44:40Z → 16:50:37Z). It used its own code, not the lane's
scripts. **Verdict: HOLDS.** Nothing was blocking or important, and no figure and no verdict moved.
- **What it re-derived independently:**
  - the order: the four pushes against every run receipt;
  - the doorman, 27 / 36 and 9 / 12 with one tie. It fails under either tie rule, and it would fail even with the
    yes/no mapping inverted. The mapping was traced through the bridge and run_kev by hand;
  - the 617 / 617 bodies and the 1,252 / 1,252 compiled prompts;
  - every McNemar and every tie figure;
  - the 30 / 32 prediction score;
  - a byte-for-byte redraw of both pages at `75fe4947…`;
  - its own scan of all 144 files, journals included: 0 identifiers, 0 H48 text.
- **Folded in place, in the results block above:**
  - the headroom unit, now 7,892 MiB (7.7 GiB), which the first version gave as "about 7.9 GB";
  - the size pair's not-step-matched clause, repeated for the cold reader;
  - the 2 dp spread, now shown beside its full-precision value.
- **Named, not changed:** a comment in the pinned `a1b/copyout_apus_a1b.py` (line 66) shows the bus-id format with a
  PCI-shaped example (`"0a:00.0"`). It is illustrative, not the card's bus id: the scan holds the real id in memory
  and read that file clean. The file is pinned by sha256 and ran the counted copy-out, so it is not edited after the
  run. A later twin should write the format as `bb:dd.f`.
