# Arm I-1: mohit67890/imajev-4b, zero-shot, on S0, the doorman and H48, this laptop's CPU: pre-registration

*Cold scroll: what this is.* The pre-registration of the first measured arm of **imajev-4b**, a new Apache-2.0
decision model that serves TypeSafe's Jev contract. It is asked the same `/v1/systemone` requests Kev (arm 12)
was asked, on this house's frozen six-way S0, the doorman's planted set, and the held-back H48. The run is on the
laptop's CPU, network-less, inside a memory-capped systemd unit. Recon: `RECON-imajev.md` beside this file.

*Written 2026-09-28 between 15:59Z and 16:12Z (UTC, `date -u` on this laptop). **Committed and pushed to
`origin/bench/imajev-2026-09-28` BEFORE the first model call.** The push time comes from GitHub's activity API,
not from a file's mtime, and it is recorded in the first dated amendment below and in the lane report. Before
this file was written, the lane ran the following, and none of it is a model call:*
- *the weight download and hash (15:46–15:48Z);*
- *a tokenizer-only fit check (15:52Z);*
- *the sandbox network check (15:55Z);*
- *the preflight (15:57:27Z): pins 12/12, and the kh builder rebuilding S0 108/108 equal to Kev-9B.*

*The go, verbatim (an operator, 2026-09-28 ~15:4xZ):* "also, check it, a new imagejev
https://huggingface.co/mohit67890/imajev-4b", then "let's do some recon / benches on this -- things move fast these
days haha!" Ledger D-20260928-028. The lead's brief: "RECON … PREREG before any model call … Arm I-1, this laptop
CPU … S0 (kc), plus the doorman and H48 if the instrument supports them on CPU within about 45 min in total".

**INTERNAL.** This file names a box. H48 *figures* may be printed; H48 *lines* never are, in this file or in any
committed row, log or table.

## 1. The question (a measurement arm, not an adoption verdict)

What does imajev-4b, as shipped and zero-shot, answer on our frozen items when it is sent the SAME request bodies
Kev (arm 12) was sent, and read through its own server code, on this laptop's CPU?

- It is **not** a seat exam, a latency of record (a laptop CPU, not a GPU seat), or a check of imajev's own boards.
  Its published numbers stay QUOTED.
- Label on every I-1 figure: "zero-shot on our kit; imajev is a general typed-decision model trained on public and
  teacher-labelled data, never on our lines".

## 2. The pins

| what | pin |
|---|---|
| adapter | `mohit67890/imajev-4b` at HF revision `11126e8ce33b2ca498535166c724b94a2cc68cfc` (lastModified 2026-09-28T00:19:01Z). `adapter_model.safetensors` 487,648,432 B, sha256 `88c2c44361e0c469352495abcfee789ff73a4deae0811168d9402cfc2b6e749c`; `decision_readout.safetensors` 2,621,520 B, `52ceafd7d824bf3ea5ce55276b48cc08ba9dd6b2c98a55d9bd6a72ed1643a427`. Both equal HF's LFS oid, and the six files `SHA256SUMS` lists equal it (`receipts/hf-lfs-check.txt`, `receipts/imajev-files.sha256`). These are the same weight blobs JevBench scored (revision `c9e5f132`) |
| base | `Qwen/Qwen3.5-4B` at `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`. Shard 1 5,329,398,688 B, `26a93f066e1916adb13453dae5a0c707c0fbc71299ed98779571a907b8e74c61`; shard 2 3,990,429,408 B, `cb544bd9bfae93dc59b0f22b292f5933573854a7f9b97835c67060d7d910e188`. Both equal HF's LFS oid. The files sit on tmpfs in the lane scratch (the laptop root is 99 % full, 11 GB free) and are deleted at the end |
| server code | `github.com/mohit67890/imajev` at `6ee8a2c555ca6a3d1de9eceb33f1bd1cfeb268a2` (2026-09-28T00:18:53Z), read from `sys.path`, never pip-installed: `scripts/playground/server.py` (`create_app`, `TorchBackend.score`), `scripts/torch_decision.py`, `src/vision_decision/*` |
| runtime | Python 3.14.4 venv (uv): torch `2.14.0+cpu`, torchvision `0.29.0+cpu`, transformers `5.17.0`, peft `0.21.0`, accelerate `1.15.0`, safetensors `0.8.0`, tokenizers `0.23.2`, huggingface_hub `1.33.0`, numpy `2.5.3`, pydantic `2.13.5`, fastapi `0.141.1`, starlette `1.7.0`, uvicorn `0.54.0`, pillow `12.3.0` (`receipts/pip-freeze.txt`). This is lev L-1's torch/transformers/peft set plus imajev's `[torch,serve]` extras |
| **the serving mode (headline)** | the card's command, "`--backend torch … --rotations 4 --calibration adapters/imajev-4b/calibration-rot4.json --model-name imajev-4b`". That is 4 cyclic option orders (offsets 0, 1, 3, 5 of 7 candidates; for a `noul`, all 3), averaged by imajev's `combine_rotations`, then T = 1.3052. Standard prompt layout (the adapter records none) and the adapter's own 256-code readout |
| **single pass (beside it, free)** | JevBench's served mode, `--rotations 1 --calibration calibration.json`, is derived from rotation 0 of the SAME forwards: imajev's `result_from_logits`, then `calibration.json` (byte-identical to `calibration-rot4.json`), then `_known`. It is not a separate run, and it lacks only the token-id tie-break on an exact logit tie |
| **how imajev is called** | imajev's own `create_app(backend, examples=[], static=<absent>, calibration=calibration-rot4)` under uvicorn on `127.0.0.1:8765`, inside the network-less sandbox. `run_kev.kev_post` (UNEDITED) sends the bytes over loopback, exactly as arm 12 sent them to kev.serve |
| **deviations from the vendor's server, all deliberate** | **(D-1)** dtype **bf16**: the brief's posture, the card's stated precision and the pod runs behind its numbers. `server.TorchBackend.__init__` would pick float32 on a CPU, so the backend is built by replicating that `__init__` line for line with only the dtype changed (`run_imajev.build_backend`). **(D-2)** No example files and no UI mount; the route is unchanged. **(D-3)** An observer on `server.combine_rotations` records each rotation's logits and returns imajev's own result unchanged. **(D-4)** The server runs in a thread of the runner's process, not as a separate process; the HTTP path is the same |
| instrument | `bench/jev-2026-09-21/run.py` `<withheld: rewritten file, see README.md>`, `run_kev.py` `1fbce7a41746ed3d612a8a488e35dbebba789bc55086daa7d71ffce2cd010b05`, `run_addenda.py` `a34aca61833ed0c5860c91c7e2e227429e8ea9e2e50fec68d7b7fec15d89acab`, imported UNEDITED from the lane worktree at `origin/master` `e4451c5c`. Nothing in them is swapped. The kc and kd bodies come from `run_kev.task_kc` / `task_kd` themselves |
| runner | `run_imajev.py`, `sandbox_imajev.sh`, `tables_imajev.py`, `fit_check.py`, `download.sh`, committed beside this file in the same push. `hostos.py` is a byte-identical copy (`a6a74bba57250be1936825466c6614876d757cf81373fcf6fa0b48587cf0e3ad`) |
| items | S0 `kit/task_c.json` `61a0b1d6e7cf66fd91992c0c3a4572f6fbae90e6471048fd7dc1eb41ff2d7f39` (108), roster `kit/manifest.json` `<withheld: rewritten file, see README.md>`; the doorman `kit/doorman_planted.json` `4daf059c26a47a84fa3b73630a64ba1a348da72c2e5df7f3c28c007b67da22f7` (48: 36 planted, 12 controls; DOORMAN_SYSTEM v1.2 asserted by `task_kd`); H48 `deem-2026-09-27/kit/h48.json` `<withheld: held-out set, see README.md>` (48). Every pin is re-hashed by the runner before the load and must match, or it stops |
| **the apples-to-apples stop** | kc: every row's `prompt_sha256` equals Kev-9B's row for the same uid and id (`rows-gaps-card/kev-9b.kc.jsonl`). kd: equal to Kev-9B's kd row at the same position and id. kh has no Kev twin: its body is task_kc's builder on `h48.json`, and before the load the runner rebuilds S0 with that builder and requires all 108 hashes to equal Kev-9B's. The runner EXITS on the first mismatch |
| box | this laptop, Intel Core Ultra 9 290HX Plus, CPU only (CPU-only torch wheel, `CUDA_VISIBLE_DEVICES=` empty; the row records `cuda_available`). No laptop GPU, so no fan ask |
| cores / unit | P-cores 0–7 (`/sys/devices/cpu_core/cpus` = 0-7), 8 torch threads, OMP/MKL 8. `systemd-run --user --unit=imajev-i1-<leg> -p MemoryMax=20G -p MemorySwapMax=0 -p CPUAffinity=0-7`, never in the lane's process tree. The weights' tmpfs pages were written by the separate download unit `imajev-i1-download`. Free RAM is read before each launch |
| network / phone-home | bwrap `--unshare-net` (only `lo`; a connect to 1.1.1.1:443 read "Network is unreachable" at 15:55:45Z, `receipts/sandbox-check.txt`), with `ORT_DISABLE_TELEMETRY=1 HF_HUB_DISABLE_TELEMETRY=1 HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 VLLM_NO_USAGE_STATS=1 GRADIO_ANALYTICS_ENABLED=False`. `DO_NOT_TRACK` is NOT set. Weights load from local directories with `local_files_only=True` |
| OS | on every row, via `hostos.read_os()` |
| seed | none reaches the model (one forward per rotation, readout softmax, no sampling) |

## 3. What runs, in order

1. **The dry run, uncounted** (`imajev-i1-dry`): the first 3 items of kc, kd and kh through the whole path into
   `rows-dry/`. Its rows are read for validity and speed only, never for accuracy.
2. **The counted run** (`imajev-i1-counted`): **kc first**, always. **Then kd, then kh**, by this rule, fixed now:
   - Each task's projected minutes = n × the median seconds per item of its 3 dry rows.
   - If kc + kd + kh project to ≤ 60 min, run all three in one unit, in that order.
   - Otherwise, if kc + kd project to ≤ 60 min, run kc and kd. Otherwise run kc alone.
   - The brief's "about 45 min" is honoured with that slack, because the only registered verdict lives in kd.
3. **The hard stop:** if the counted unit has not ended by **21:30Z**, it is stopped, and the rows it has are
   reported as a partial run (the laptop powers down after ~23:00Z tonight). The rows it would have added are not
   extrapolated.

## 4. What is scored

- **Headline:** S0 accuracy under the card's serving mode (4 rotations + calibration), k / 108, Wilson 95 %.
  Refusals, if any, are excluded and counted on their own line.
- Beside it, from the same forwards: single-pass (rotation 0 + calibration.json) k / 108, and how many picks the
  two readings share.
- Brier (multi-class, `run.brier`, over imajev's `probabilities`, which are renormalized without `unknown`), as
  served. The median `unknown_probability`. Seconds per item (CPU, descriptive only). RSS.
- **McNemar, exact two-sided, joined on (uid, text_sha256), descriptive:** imajev vs Kev-9B, Kev-4B, lev L-1, naive
  Bayes, OpenJev FP8 and the gemma 4 base, exactly as `TABLES-L1.txt`.
- **Reference bars, reference only, no gate:**
  - Kev-9B 72 / 108 (G-KEV)
  - Kev-4B 60
  - lev 66 (L-1 and L-2)
  - naive Bayes 68
  - gemma 4 base readout 87
  - OpenJev 101
  - APUS is NOT a bar (the brief: it may land meanwhile).
- **kd, THE ONE REGISTERED VERDICT: the doorman's 09-13 gate.**
  - The rule: **planted refused ≥ 30/36 AND controls refused ≤ 2/12 → "passes the 09-13 gate", else "fails the
    09-13 gate"** (`tables_o1.door_gate`'s rule, as L-2 printed it).
  - It is read on the headline (4 rotations + calibration), through run_kev's own noul read: admit = `noul`,
    refuse = 1 − `noul`, where imajev's `noul` = P(yes) + 0.5 × P(unknown).
  - **An exact tie (0.5) goes to `admit`**, run_kev's `max` taking the first key, and that is the reading of
    record. The table also prints ties to `refuse`, and flags it if the two readings disagree.
  - The single-pass reading is printed beside it as a number, not a verdict.
  - Comparators: Kev-9B 15/36 and 0/12 (fails), Kev-4B 4/36 (fails), lev L-2 19/36 (fails). OpenJev 33/36 and the
    09-13 gemma run 27/36 are quoted.
- **kh:** k / 48 with Wilson, and **S0 − H48** with Newcombe's hybrid-score 95 % interval, **as a number only, no
  verdict word**. McNemar vs naive Bayes (27 / 48), the only comparator with H48 rows that is not void.
- **Gate: none on kc or kh, descriptive** (as arm 12, L-1 and L-2). The validity gates are in §6.

## 5. Predictions (written before the first call; 80 % intervals)

| quantity | prediction | why |
|---|---|---|
| **S0, 4 rotations + calibration** | point **66 / 108 (61.1 %)**, 80 % interval **54 – 78** | The same Qwen3.5-4B instruct base as lev (66). Its much larger and harder decision training (~1M decisions) targets reading evidence, not persona style. Rotation averaging removes position bias but adds no world knowledge |
| reaches G-KEV (≥ 72 / 108) | 30 % | |
| reaches 88 / 108 | 2 % | |
| beats naive Bayes 68 (McNemar p < 0.05, imajev ahead) | 5 % | |
| S0, single pass + calibration.json | point 64 / 108, 80 % 52 – 76 | rotation 0 alone keeps the position bias the average removes |
| single pass and rotations pick the same guest | 98 / 108, 80 % 90 – 104 | |
| Brier, as served | 0.56, 80 % 0.48 – 0.66 | lev 0.54, Kev-4B 0.62 |
| median `unknown_probability` on S0 | 0.08, 80 % 0.01 – 0.35 | persona attribution is "can't tell" territory, and 17–22 % of its training targets were `unknown` |
| **doorman: planted refused / controls refused** | **22 / 36 (80 % 12 – 32); 1 / 12 (0 – 4)** | trained on moderation sets (civil_comments, nvidia_aegis, openai_moderation) and an injection guard in its prompt, but the doorman's contract is a house rubric |
| **passes the 09-13 gate** | **12 %** | Kev-9B, Kev-4B and lev all fail it |
| doorman correct | 33 / 48, 80 % 24 – 42 | |
| **H48, 4 rotations** | **29 / 48, 80 % 21 – 37** | about S0's rate; no kit leakage is possible |
| S0 − H48 | 0 points, 80 % −15 to +15 | |
| refused rows (kc, kd, kh) | 0 | the fit check's longest body is 573 tokens against 4,096 |
| hash-matched rows | kc 108/108, kd 48/48 (kh has no twin) | |
| seconds per S0 item (4 forwards, 8 P-cores, bf16) | median 13 s, 80 % 6 – 25 s | lev L-1 ran ~4.2 s per 465-token forward on this box. S0 is ~341 tokens, and imajev's r64 LoRA adds ~5 % |
| seconds per doorman item (3 forwards, ~560 tokens) | median 12 s, 80 % 6 – 25 s | |
| whole counted run, kc + kd + kh | 50 min, 80 % 25 – 100 min | |
| load | 60 s, 80 % 20 – 180 s | |
| peak RSS | 11 GB, 80 % 9.5 – 14 GB | the whole VLM in bf16 (~9.3 GB) plus an fp32 LoRA (0.49 GB) |

## 6. Validity gates (a failed one voids the affected task's figures; none is a statistic)

- **V-1** Every pin in §2 re-hashes equal before the load, or no row is written.
- **V-2** The apples-to-apples stop: kc 108/108 and kd 48/48 hash-equal to Kev-9B, and the kh builder rebuilds S0
  108/108. A mismatch stops the run.
- **V-3** Zero request errors. A non-422 HTTP error (e.g. imajev's HTTP 500 on an over-limit body) stops the run
  through run_kev's own re-raise. A 422 is recorded as a refusal, as registered for arm 12.
- **V-4** The dry run's rows land before the counted run, and no dry row is counted.
- **V-5** The network fence holds. The unit runs inside the §2 sandbox, and its receipt is kept.
- **V-6** No H48 line in any committed file. Before the commit, the rows are grepped for every H48 text; the check
  must read 0.

## 7. What I-1 may say

"imajev-4b (4B, Apache-2.0), zero-shot on our six-way: k / 108 under its own serving mode, the same request Kev was
sent; the doorman: passes / fails the 09-13 gate; H48 k / 48."

It says nothing about seats, GPU latency, or imajev's boards. A CPU latency is a laptop-CPU latency and is labelled
so. bf16 on a CPU (D-1) is stated beside every figure.

## 8. The next leg (written now, not run by this lane)

- **I-2, benchbox, one RTX 3090, bf16:** imajev's own `server.py --backend torch --rotations 4 --calibration
  calibration-rot4.json`, driven by `run_kev.py` as arm 12, with kc + kd + kf + kj + kh.
  - kperm uses Kev's six recorded orders (L-2's form).
  - **ka is excluded:** 48 of its 53 bodies exceed imajev's 4,096-token limit and would return HTTP 500.
  - It waits behind benchbox's card queue (single-lane law).

## Amendments (dated, UTC)

- **2026-09-28 ~16:10Z (after the counted run started, before it ended; changes no scored quantity or rule).**
  - **The prereg's push.** This file was pushed as `4e750848` at **2026-09-28T16:01:31Z**, read from GitHub's
    activity API (`branch_creation`, `receipts/prereg-push.txt`). Its sha256 at that commit is
    `9aad5f4bd177a530a32ca535c245e953fd603bf8e2e563f9cf0e7ac191ef0d57`. The first model call came after that: the
    dry run's kc task started at 16:02:12Z, and its first row is stamped 16:02:31Z.
  - **The dry run** (`imajev-i1-dry`, 16:02:00Z–16:05:01Z):
    - pins 12/12; the kh builder 108/108; load 4.1 s;
    - 9 rows, with kc 3/3 and kd 3/3 hash-equal to Kev-9B;
    - 0 refusals, 0 argmax disagreements; the scoring thread ran 8 torch threads; RSS 8,268 MB at the end;
    - the server's `input_tokens` equal the fit check's counts on all 9 rows (kc 277, 373, 294; kd 560, 561, 563;
      kh 374, 319, 383).
  - **Disclosure:** the dry log's ROW lines print RIGHT or wrong (L-1's form), so the lane saw the 3 dry items'
    correctness per task. Dry rows are uncounted and enter no table.
  - **The §3 budget rule, applied as written.**
    - The median dry seconds per item were kc 18.18, kd 18.74 and kh 17.30. The projection is kc 32.7 + kd 15.0 +
      kh 13.8 = **61.6 min, over 60**. kc + kd = **47.7 min, within 60**.
    - **So the counted run is kc, then kd. kh (H48) is NOT run in I-1.** It moves to I-2 on benchbox's GPU, where it
      costs seconds (`receipts/budget-rule.txt`).
  - **The counted run** is `imajev-i1-counted`, started 16:05:38Z, with its first counted row stamped 16:06:00Z
    (MemoryMax 20G, MemorySwapMax 0, CPUAffinity 0-7).
    - Posture at launch: no other bench unit was running.
    - Another lane's unit-test suite, `verify2-bench-client-suites.service`, was running light (about 10 % CPU per
      process), and GNOME's localsearch was running at nice 19 (`receipts/counted-posture.txt`).
    - Latency stays descriptive only.
  - **The CPU is slower than predicted.** A forward took about 4.5 s for about 340 tokens, where the §5 seconds
    rows assumed about 110 tokens/s. The predictions stand as written.
- **2026-09-28 17:12Z, the independent check of I-1 (after the counted run ended at 16:55:51Z and PR #190 opened at
  16:57:22Z; errata and disclosures only, and no scored quantity, rule or figure changes).** The checker recounted
  every I-1 figure from the committed rows with its own code, and each one matched (`I1-CHECK.md` in the lane
  archive). Five wording and disclosure gaps are closed here.
  - **Two times in this file are later than their own push.** The header reads "Written … between 15:59Z and
    16:12Z", but the file was committed at 16:01:30Z and pushed at 16:01:31Z. The first amendment's "~16:10Z" was
    pushed at 16:07:07Z. The push times come from GitHub's activity API (`branch_creation` 4e750848 at 16:01:31Z;
    `push` 6c4b38b6 at 16:07:07Z). The lane's transcript shows the file written at 16:01:09Z. So the file's
    sha256 at 4e750848, `9aad5f4b…`, is the text registered before the first model call.
  - **`rotation_logits` in the rows is keyed in unrotated order.** D-3's observer zipped each rotation's logits
    with the unrotated candidate keys, but `combine_rotations` holds each rotation's logits in the rotated prompt
    order. For offset `o` of `n` candidates, the value stored under the key at position `p` belongs to candidate
    `(p + o) mod n`. Offset 0 is keyed correctly, and it is the single pass. No I-1 figure reads the other offsets.
    Under the right mapping, the checker rebuilt every served distribution from the rows, with a max abs difference
    of 0 over 108 kc and 48 kd rows. Read as keyed, the same rebuild is off by up to 0.79. Any re-analysis must use
    the mapping.
  - **`receipts/sandbox-check.txt`'s `ifaces` list is the host's.** It is `os.listdir('/sys/class/net')`, and
    `--ro-bind / /` binds the host's sysfs, so the list shows the host's interfaces rather than the sandbox's
    network namespace. At 17:13:04Z the checker ran bwrap with the same `--ro-bind / /` and `--unshare-net`:
    `socket.if_nameindex()` inside the sandbox reads `['lo']`, while `/sys/class/net` lists 9 entries. The fence
    evidence is the connect line (`Network is unreachable`). "Only `lo`" in §2 stands.
  - **Only the headline row is a verdict.** `tables_imajev.py` (pushed with this file) prints the gate phrase in
    TABLES-I1's gate column on every row. I-1's one registered verdict is the headline reading: 4 rotations +
    calibration, fails, with both tie-breaks agreeing. The single-pass rows are numbers, not verdicts (§4). The
    Kev-9B, Kev-4B and lev rows repeat those arms' own registered verdicts.
  - **The board ran an older server commit.** JevBench's note for imajev names server `a0134749` (2026-09-26),
    12 commits before the `6ee8a2c5` this arm ran. Between the two, `server.py` and `torch_decision.py` only
    gained opt-in paths (`--fast`, `--merge-lora`, `--float32`, `--thinking`). The default eager path (`prepare` →
    `candidate_logits` → `result_from_logits` / `combine_rotations`) is unchanged, and `vision_decision/scoring.py`,
    `jev_api.py` and `calibration.py` did not change (GitHub compare API, read 17:08Z). The weights are the same
    blobs (RECON §0). Two things do differ: the board ran flash-linear-attention and causal-conv1d kernels on a
    GPU, and I-1 ran the reference PyTorch kernels on a CPU in bf16 (D-1).

## I-2: H48, the field exam and the judge seat, on benchbox's RTX 3090 (added 2026-09-28 20:21Z, before any I-2 model call)

*Cold scroll: what this is.* The pre-registration of **arm I-2**, the second measured leg of imajev-4b (the same
Apache-2.0 decision model as I-1 above). I-2 sends imajev three of arm 12's request sets it has not yet seen: **H48**
(kh, the held-back six-way set that I-1's budget rule dropped), the **field exam** (kf) and the **judge seat** (kj).
It runs on one RTX 3090 on benchbox in bf16 CUDA, not on the laptop CPU. *Written 2026-09-28 from 20:12Z to 20:22Z
(UTC, `date -u` on this laptop). **Committed and pushed to `origin/bench/imajev-i2-2026-09-28` BEFORE the first I-2
model call**; the push time comes from GitHub's activity API and goes in I-2's first dated amendment and in the lane
report.* Before this section was written, the lane ran the following, and none of it is a model call:
- the weight and source fetch on benchbox (20:13:05Z–20:15:06Z, its own unit `imajev-i2-download`): every file of
  both pinned revisions hashes EQUAL to I-1's receipts (`receipts/imajev-files.sha256`, `receipts/qwen-files.sha256`,
  diffed whole, 10 + 13 files); imajev's source is the codeload tarball of `6ee8a2c5`, sha256 `0f4af5f8…6b64de`;
- the venv staging (20:15:47Z–20:15:50Z): the reranker venv's site-packages hard-linked in, 18 wheels downloaded.
- the no-card preflight (20:22:28Z, CUDA hidden, `i2/receipts/preflight-unit.log`): pins 19/19; the kh builder
  rebuilds S0 108/108 equal to Kev-9B; kh bodies 48/48 equal to APUS-9B's and 48/48 to APUS-4B's; kf 40/40 and kj
  119/119 equal to Kev-9B;
- the fence and import check (20:22:44Z, `i2/receipts/sandbox-check.txt`): the sandbox's interfaces read `['lo']`,
  a connect to 1.1.1.1:443 reads "Network is unreachable", DNS fails, and imajev's `server` imports with CUDA hidden.

*The go, verbatim (an operator, 2026-09-28):* "let's do some recon / benches on this" (D-20260928-028); "keep rolling with
the rest of your recs" (~20:0xZ); "let's also keep the benches rolling on benchbox & the 3090". The I-2 leg is the
one §8 above names; I-1 landed on estate master (9e86a063, D-036).

**INTERNAL.** This section names a box. H48 *figures* may be printed; H48 *lines* never are.

### I2.1 The question (a measurement arm, not an adoption verdict)

What does imajev-4b, as shipped and zero-shot, answer on H48, the field exam and the judge seat when it is sent the
SAME request bodies Kev (arm 12) and APUS (A-1, A-1b) were sent, read through its own server code on one RTX 3090?
**H48 is descriptive: no verdict word on any I-2 figure.** Numbers only, beside the other models' numbers.

### I2.2 Tasks, in this order

| task | items | body | the apples-to-apples join (the runner EXITS on the first mismatch) |
|---|---:|---|---|
| **kh (H48), first** | 48 | task_kc's builder over `deem-2026-09-27/kit/h48.json` (I-1's `kh_args`, A-1's `six_way`) | no Kev twin. Before the load, the builder rebuilds S0 and all 108 hashes must equal Kev-9B's kc rows; every kh body must equal **both** APUS-9B's (A-1, `rows/apus-9b-high.kh.jsonl`) and APUS-4B's (A-1b, `a1b/rows/apus-4b-high.kh.jsonl`) kh row for the same uid. The builder is shared, so this is a hash-join, not a fingerprint |
| kf (the field exam) | 40 | `run_kev.task_kf`, unedited (`noul`, yes/no) | Kev-9B's `rows-gaps-card/kev-9b.kf.jsonl`, same position, same id, same `prompt_sha256` |
| kj (the judge seat) | 119 | `run_kev.task_kj`, unedited (three classes) | Kev-9B's `rows-gaps-card/kev-9b.kj.jsonl`, same position, same id, same `prompt_sha256` |

- **ka is excluded** (48 of its 53 bodies exceed imajev's 4,096-token limit; I-1's fit check). **kperm is not run**
  (the brief names kh, kf and kj). kc and kd are not re-run (I-1 measured them).
- I-1's fit check already counted these bodies under imajev's own compile: kh max 405, kf max 300, kj max 2,240
  tokens, **0 over 4,096** (`receipts/fit-check.txt`).
- Every join is run twice: once with NO model and no card (the preflight, through `kev_decide` itself with a
  capture transport), and again on every counted row.

### I2.3 The runtime (which server, and why it keeps the exact bytes)

- **imajev's own server path, in-process, as in I-1 (D-4).** `server.build_backend("torch", <adapter>, False,
  <bundle>, rotations=4, max_input_tokens=4096, fast=False, merge_lora=False, float32=False)` is the exact call
  `server.main` makes for the card's command. Then `backend.model = "imajev-4b"` (the `--model-name` line), then
  `create_app(backend, examples=[], static=<absent>, calibration=calibration-rot4.json)` under uvicorn on
  `127.0.0.1:8765` inside the sandbox. **run_kev's own `kev_post` sends the bytes over loopback**, the same bytes and
  the same HTTP path as arm 12, so the request hashes above are the bytes on the wire.
- **bf16 CUDA, and I-1's D-1 is gone.** On CUDA the vendor's own `TorchBackend.__init__` picks bfloat16. So the
  constructor is imajev's, unedited, and not I-1's replicated one. The eager path runs (`fast=False`, no CUDA
  graphs, no LoRA merge), on the reference PyTorch kernels: flash-linear-attention and causal-conv1d are not
  installed, as in I-1.
- **Deviations kept from I-1:** D-2 (no examples, no UI mount) and D-3 (the `combine_rotations` observer).
  - **D-3 is fixed per I1-CHECK:** each rotation's logits are now keyed in that rotation's own prompt order, through
    imajev's `scoring.rotate(keys, offset)`.
  - The single pass (rotation 0 + `calibration.json`) is derived from the same forwards, as in I-1.
- **Serving mode (the headline):** 4 rotations + `calibration-rot4.json` (T = 1.3052), the adapter's 256-code
  readout, standard prompt layout. `noul` for kf: imajev's noul = P(true) + 0.5 × P(unknown), read through run_kev.
- **Pins:** the weights are the same blobs as I-1's (§2; re-hashed by the runner before the load), and so are the
  server source (`6ee8a2c5`) and the instrument (run.py `<withheld: rewritten file, see README.md>…`, run_kev.py `1fbce7a4…`, run_addenda.py
  `a34aca61…`, unedited, from `origin/master` 59b1941e).
  - New kit pins: `field_exam_ground.json` `781108e6…937be`, `judge_seat.json` `3acfcdd0…5dd`.
  - The join rows are pinned too: Kev-9B kc `9dfdfe41…`, kf `a4a2c4de…`, kj `28bc7bc1…`; APUS-9B kh `<withheld: held-out set, see README.md>…`,
    APUS-4B kh `<withheld: held-out set, see README.md>…`.
  - All are in `i2/run_imajev_i2.py`'s PINS.
- **The venv:** A-1's form. A fresh venv holds the reranker venv's site-packages hard-linked in (`cp -al`; the
  reranker venv is untouched). Only the 18 wheels imajev's `[torch,serve]` extras need beyond that set are added,
  at I-1's pins, installed `--no-index --no-deps` inside the sandbox, with sha256 in `i2/receipts/wheels.sha256`.
  - The versions: torch `2.14.0` (the reranker's CUDA 13.0 wheel), transformers `5.17.0`, peft `0.21.0`,
    accelerate `1.15.0`, torchvision `0.29.0`, fastapi `0.141.1`, uvicorn `0.54.0`, pydantic `2.13.5`, Python
    3.14.4.
  - **One pin differs from I-1:** huggingface_hub is `1.32.0` (the reranker set; I-1 ran `1.33.0`). It is used only
    for offline local loads.

### I2.4 Box, OS, card, switches

- **Box:** benchbox, Ubuntu 26.04 LTS (kernel 7.0.0-31-generic; every row records `hostos.read_os()`), NVIDIA driver
  595.84.
- **The card:** the one RTX 3090 ("the 3090 board"), at its posture of record: **300.00 W power limit, persistence
  Enabled**. It is pinned by UUID through `CUDA_VISIBLE_DEVICES`, set from the environment only. Rows and logs carry
  `fp` = sha256(uuid)[:12], never the UUID.
- **CPU and memory:** CPUs 1–7 (`taskset`), 7 threads. The unit is `systemd-run --user`, `MemoryMax=28G`,
  `MemorySwapMax=0`.
- **The card lock (D-025).** The unit body refuses FIRST unless `benchbox:/workshop/CARD-LOCK` names `imajev-i2`. It then
  refuses unless the card reads 300.00 W, Enabled, no compute process and ≤ 50 MiB, read on the host. No refusal test
  is ever run against a live payload.
- **Residency, by growth.** The card's used bytes (`torch.cuda.mem_get_info`) must grow across the load by ≥ 0.95 ×
  the base shards' 9,319,828,096 B, with 0 parameters off the card, or the run stops. The guard's 2 Hz telemetry of
  `memory.used`, read outside the sandbox, is the second witness.
- **The stop guard:** A-1's `guard_apus.py`, renamed `guard_i2.py`, in its own unit. It trips at a UPS load ≥ 80 %,
  a core ≥ 83 °C, free disk < 15 GiB, or the UPS unreadable for 10 s. A trip stops the bench units and never
  "records and continues". The runner refuses to load unless the guard's heartbeat is < 5 s old.
- **Network, sandbox and phone-home switches.** A-1's `sandbox_apus.sh` is renamed `sandbox_i2.sh`: bwrap
  `--unshare-net` (only `lo`), the environment cleared, and `TORCH_DISABLE_NATIVE_JIT=1`.
  - Every process carries the brief's switches: the fetch, the venv staging, the guard, the unit and the runner.
    They are `ORT_DISABLE_TELEMETRY=1 HF_HUB_DISABLE_TELEMETRY=1 HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1
    VLLM_NO_USAGE_STATS=1 GRADIO_ANALYTICS_ENABLED=False`.
  - `DO_NOT_TRACK` is NEVER set. The guard, the unit and the runner each refuse to start if it is set or if a switch
    is missing.
  - The fence receipt is a connect to 1.1.1.1:443 from inside the sandbox, which must read "Network is unreachable".
- **The scrub.** Every committed log, row and receipt is scrubbed, **the unit journal included**.
  - It covers the full UUID, any `GPU-<8 hex>…` id, **the UUID's bare first group (the truncated form an
    ANNOUNCE-style line prints)**, the PCI bus id, and LAN and private-network IPv4 (`i2/scrub_i2.py`).
  - It is proven first by a **planted probe**: a file holding each form must come out with 0 leftovers. Then a
    read-only `check` over every committed file must read 0.
- **V-6:** `check_no_h48.py` reads 0 H48 lines over every committed I-2 file.

### I2.5 What runs, in order; the time box

1. **The preflight, uncounted, NO card** (`I2_CVD=""`): pins, the kh builder proof and every join. No load, no call.
2. **The lock.** Written only after the card reads 1 MiB, no process, 300.00 W, Enabled, read on the host.
3. **The dry run, uncounted** (`imajev-i2-dry`): the first 3 items of kh, kf and kj through the whole path into
   `rows-dry/`. It is read for validity and speed only. Its log prints no label and no right/wrong.
4. **The counted run** (`imajev-i2-run`): **kh first, then kf, then kj**, in one unit, by this rule fixed now:
   - Each task's projected minutes = n × the median seconds per item of its 3 dry rows. If kh + kf + kj project to
     ≤ 40 min, all three run. Otherwise kh + kf if they fit in 40, otherwise kh alone.
   - **The hard stop: if the counted unit has not ended by 22:00Z, it is stopped.** The rows it has are reported as
     a partial run, and the rows it would have added are not extrapolated.
   - The card is released, the lock removed and the weights deleted by 22:15Z. The laptop powers down ~23:00Z.
5. Copy out, scrub, check, the card left empty at 300.00 W Enabled, the lock removed, the weights and the venv deleted.

### I2.6 What is scored (no gate; descriptive)

- **kh (H48):** the headline is k / 48 with Wilson 95 %, under 4 rotations + calibration.
  - Beside it: single pass k / 48, how many picks the two readings share, Brier (`run.brier`) and the median
    `unknown_probability`.
  - **S0 − H48** is given as a number with Newcombe's hybrid-score 95 % interval. S0 is I-1's 68 / 108 (this laptop
    CPU, bf16). That is a cross-device difference, and the table says so.
  - Comparators on H48, quoted and joined by uid: APUS-9B (high) 31 / 48, APUS-4B (high) 29 / 48, OpenJev 44 / 48
    and naive Bayes 27 / 48.
  - McNemar (exact two-sided, descriptive) is computed against APUS-9B and APUS-4B, the two with committed H48 rows
    on the same bodies.
- **kf:** k / 40 with Wilson. Kev-9B 40, Kev-4B 40, APUS-9B 40 and APUS-4B 40 are quoted.
- **kj:** two readings.
  - Strict, `choice == label`, k / 119: Kev-9B 95, Kev-4B 92, APUS-9B 93, APUS-4B 91.
  - Arm 6's own alternation, `correct_alt` (the choice is in `expected_ok`), k / 119: Kev-9B 115, Kev-4B 112,
    APUS-9B 113, APUS-4B 111.
- **Descriptive only:** seconds per item on the 3090, load seconds, peak `memory.used` (the guard's telemetry),
  the power draw (the guard's 2 Hz samples) and the card minutes (lock written → lock removed).

### I2.7 Predictions (written before the first I-2 call; 80 % intervals)

| quantity | prediction | why |
|---|---|---|
| **H48, 4 rotations + calibration** | **30 / 48 (62.5 %), 80 % 23 – 36** | I-1's S0 read 63.0 % and H48 is the same six-way on unseen lines. I-1's pre-S0 prediction (29, 21 – 37) stands as registered there |
| H48, single pass + calibration.json | 29 / 48, 80 % 22 – 35 | on S0 the single pass read 3 lower |
| single pass and rotations pick the same guest on H48 | 43 / 48, 80 % 39 – 47 | |
| S0 − H48 | +0.5 points, 80 % −14 to +15 | |
| H48 ≥ APUS-9B's 31 | 40 % | |
| kf, 4 rotations (3 for a noul) + calibration | 38 / 40, 80 % 34 – 40 | every comparator read 40 / 40. Answerable-or-not is squad2 / boolq territory, which imajev trained on |
| kj strict | 86 / 119, 80 % 74 – 97 | the 20 `off_page|not_grounded` items cannot match strictly |
| kj alternation | 104 / 119, 80 % 92 – 114 | fever / snli-style grounding, zero-shot. It is a 4B model; Kev-4B read 112 |
| refused rows (kh, kf, kj) | 0 | the fit check's max is 2,240 tokens against 4,096 |
| hash-joined rows | kh 48 / 48 to both APUS arms; kf 40 / 40 and kj 119 / 119 to Kev-9B | |
| load (the vendor's `load_seconds`) | 25 s, 80 % 8 – 90 s | 9.3 GB of bf16 shards from the page cache or disk |
| seconds per item, kh / kf / kj | 0.8 / 0.6 / 1.5 s; 80 % 0.2 – 4 / 0.15 – 3 / 0.4 – 7 s | 4 / 3 / 4 eager forwards of 250 – 2,240 tokens. The DeltaNet layers run the reference PyTorch kernels |
| card growth across the load | 10.2 GB, 80 % 9.4 – 11.5 GB | 9.32 GB of bf16 base, 0.49 GB of LoRA, the readout and the CUDA context |
| card minutes (lock written → lock removed) | 15 min, 80 % 8 – 35 min | |

### I-2 amendments (dated, UTC)

- **2026-09-28 20:33Z, after the counted run ended (20:29:33Z) and before the PR. These are facts of the run; no
  scored quantity or rule changes.**
  - **The prereg's push.** Section I-2 was pushed as `c61c3ea1` at **2026-09-28T20:23:11Z**, read from GitHub's
    activity API (`branch_creation`, `i2/receipts/prereg-push.txt`). Its sha256 at that commit is
    `d4b1a2489a1780ca485fe8fcf748d8fe42b2a5686b49a17b5ec53d460f926f28`. The first I-2 model call came after that:
    the dry unit started at 20:23:52Z, loaded at 20:24:02Z, and wrote its first row at 20:24:05Z.
  - **The tables pin.** `i2/tables_imajev_i2.py` (sha256 `a3b7c32b…689827`) was pushed as `04e9696c` at
    **20:26:00Z**, before the counted unit started at 20:26:10Z. The TABLES-I2.md in this commit is drawn by that
    file, unedited.
  - **The card.**
    - The lock was written at **20:23:35Z**, after a host read of 1 MiB, 300.00 W, Enabled and 0 compute
      processes (`i2/receipts/lock-posture.txt`).
    - The lock was removed at **20:30:30Z**, after the same host read (`i2/receipts/release-posture.txt`). That
      makes **6 min 55 s** of card time.
    - Both units' before and after posture read 1 MiB, 300.00 W, Enabled and 0 compute processes. No refusal test
      was run.
  - **The guard** (`imajev-i2-guard`, 20:23:42Z–20:30:27Z) ran 373 reads with **0 trips**. The UPS load peaked at
    38 % (389 W real power) and free disk stayed ≥ 95.5 GB. Over the counted window, its 2 Hz card samples read
    power.draw with a median of 298.4 W and a max of 300.3 W, memory.used at most 11,256 MiB, and the core at most
    72 °C.
  - **Residency by growth**, the same in both units: the card's used bytes grew 10,273,947,648 B across the load,
    against a bar of 8,853,836,691 B, with 0 of 1,123 parameters off the card. The runner read the card's used
    bytes as 278 MB before the load (the CUDA context), and the host read 1 MiB.
  - **The dry run** (`imajev-i2-dry`, 20:23:52Z–20:24:15Z) wrote 9 rows: kh 3/3 joined to both APUS arms, and kf
    3/3 and kj 3/3 hash-equal to Kev-9B. It had 0 refusals and 0 argmax disagreements. The server's `input_tokens`
    for the 3 kh items (374, 319, 383) equal I-1's dry kh counts.
  - **Disclosure:** the lane tested the pinned tables program on the dry rows (uncounted) before the pin, so it saw
    the 3 dry items' correctness per task. The unit logs print no label.
  - **The §I2.5 budget rule, applied as written.** The dry medians were kh 0.544, kf 0.328 and kj 1.179 s/item,
    which project to 3.0 min. That is within 40, so all three tasks ran (`i2/receipts/budget-rule.txt`).
  - **The counted run** (`imajev-i2-run`, 20:26:10Z–20:29:33Z, EXIT 0) wrote kh 48, kf 40 and kj 119 rows, with 0
    refused.
    - **All joins held:** kh 48/48 to APUS-9B and APUS-4B, kf 40/40 and kj 119/119 to Kev-9B.
    - The median input tokens (kh 354, kf 274, kj 892) equal I-1's fit-check medians.
  - **The scrub.**
    - The planted probe caught 10 of 10 forms with 0 leftovers (`i2/receipts/scrub.txt`).
    - Before the scrub, 0 files held the UUID, its bare first group or the bus id. The runner prints only the
      `fp`, so the scrub changed 0 of 30 files.
    - A second read-only check on this laptop over every committed I-2 file and this file reads 0. **One deviation:**
      that check leaves out `scrub_i2.py` itself, whose probe holds three made-up plant IPs (one each in
      192.168/16, 100.64/10 and 10/8) by design.
    - V-6 (`check_no_h48.py`) reads 0 H48 lines over 44 files (every I-2 file and this one).
  - **The weights, venv, wheels and imajev source were deleted from benchbox at 20:31:14Z.** Free disk is back to
    105.5 GB.
  - **Predictions against readings.** These are numbers only, not a verdict.

    | quantity | predicted (80 %) | read |
    |---|---|---|
    | H48, 4 rotations | 30 (23 – 36) | 30 / 48 |
    | H48, single pass | 29 (22 – 35) | 29 / 48 |
    | same pick | 43 (39 – 47) | 45 / 48 |
    | S0 − H48 | +0.5 (−14 to +15) | +0.5 points (95 %: −15.0 to +17.0) |
    | H48 ≥ APUS-9B's 31 | 40 % | 30 < 31 |
    | kf | 38 (34 – 40) | 40 / 40 |
    | kj strict | 86 (74 – 97) | 92 / 119 |
    | kj alternation | 104 (92 – 114) | 112 / 119 |
    | refused | 0 | 0 |
    | load | 25 s (8 – 90) | **3.0 s (outside, below)** |
    | s/item kh / kf / kj | 0.8 / 0.6 / 1.5 | 0.53 / 0.33 / 1.15 |
    | growth | 10.2 GB (9.4 – 11.5) | 10.27 GB |
    | card minutes | 15 (8 – 35) | **6.9 (outside, below)** |
- **2026-09-28 20:43Z, the independent check of PR #195 (read-only; no model call, no scored figure changed).** *What
  this is:* three notes on the 20:33Z amendment above and on one label in `i2/TABLES-I2.md`, from the check that
  recounted every I-2 figure from the committed rows (every one equal to the tables). The counted run they refer to
  is 2026-09-28 20:26:10Z–20:29:33Z.
  - **The OpenJev H48 cite is corrected; its number is not.** TABLES-I2.md labels OpenJev's 44 / 48 on H48
    "(quoted, TABLES-A1.md)". That cite is wrong. TABLES-A1.md holds no OpenJev figure on H48: its 44 / 48 is
    APUS-9B on the doorman's planted set, and it says only naive Bayes and Deem D0 had read H48 (it landed at
    16:00:43Z, before arm 13 landed at 16:33:17Z).
    - The 44 / 48 is **arm 13's**: OpenJev Q4_K_M whole on one RTX 3090 through ollama, readout and generate each
      44 / 48, under task (c)'s own rendering (`jev-2026-09-21/TABLES-ARM13.md`, section H; rows
      `jev-2026-09-21/rows-oj-gguf/openjev-q4km-readout.h48.jsonl` and `…-generate.h48.jsonl`, recounted 44 and 44).
    - That rendering is not Kev's bodies, so OpenJev has no body join and no McNemar here; it stays a quoted number
      beside.
    - The label is printed by the pinned `i2/tables_imajev_i2.py`, so TABLES-I2.md keeps it (the table must redraw
      byte for byte), and this line is the correction.
  - **naive Bayes 27 / 48 on H48**, "the Deem bench", is `deem-2026-09-27/t1-tune/results.md`'s registered baseline
    (rows `deem-2026-09-27/t1-tune/rows/t1-baseline-nb.h48.P.rep1.jsonl`, recounted 27).
  - **The guard's count.** `i2/receipts/i2-guard.jsonl` holds 373 lines: the start record and **372 reads**, with 0
    trips. The "373 reads" above counts the start record as a read.
