# Arm 13 of the Jev bench: OpenJev Q4_K_M whole on one RTX 3090, through arm 1's ollama route (pre-registration, 2026-09-28)

*Cold scroll: what this file is.* The Jev bench (`README.md` beside this file, benchbox, 2026-09-21) measured
OpenJev, a decision model that answers by reading the scores of option letters, against the estate's own
decisions. OpenJev scored **101 of 108** on the six-way voice set S0 in FP8 across two RTX 3090s (arm 2) and in
BF16 and FP8 on one RTX PRO 6000 (arm 3). This file registers **arm 13**: the OpenJev org's own **Q4_K_M GGUF**
of the same weights, whole on **one** 24 GB RTX 3090, asked through the ollama path arm 1 used. It was written
2026-09-28 between 14:19Z and 15:1xZ (UTC, fresh `date -u` reads on the dev laptop and on benchbox), and it is
committed and pushed on the estate branch `bench/oj-gguf-2026-09-28` **before any arm-13 model call**. The push
time is read from GitHub and stated in amendment A-1. Only dated amendments are appended; no registered text
changes.

**Before this commit, and not a model call:** the pull of the GGUF into the bench store through a CPU-only
ollama instance (14:29:31Z → 14:32:28Z), the blob's sha256 recomputed (14:34:25Z → 14:35:07Z), one `/api/show`
(it reads the file's header and manifest, and loads nothing), a read of the GGUF header by a stdlib script,
`run_oj_gguf.py --check` (every prompt built and joined with a stub in place of the server), and
`mock_oj_gguf.py` on the dev laptop (the runner's whole card path against a fake `nvidia-smi` and a fake
`ollama serve`: a dry leg and a counted leg DONE, 1,056 counted rows, the recommendations restart exercised,
no UUID or PCI address left in any written file; and four negative runs, each refused where it must be: no
lock, another lane's lock, a render offset, a thinking marker at position 0; `receipts/oj-gguf/mock-*.json`). The card was never
touched: another lane held it behind `benchbox:/workshop/CARD-LOCK` throughout, and the pull instance was started with
CUDA hidden (its own log: `inference compute id=cpu library=cpu`, the only device line).

**The authority.** an operator, verbatim, 2026-09-28 ~14:0xZ: *"love it, let's keep rolling with your recs / benchs,
and YES, let's test APUS and anything else you rec, and keep the train rolling, thx!"* DECISIONS.md
D-20260928-024 (origin/master `32ec0b35`) names OJ-GGUF in the lead's card order. The recipe is
`/workshop/bench-archive/plans-2026-09-28/benchbox-day-2026-09-29/QUEUE-3090-2026-09-28.md`, rank 2 and §3 note 2.

**INTERNAL.** This file names the box and the board. The board's house name in §2 never goes on a page.

---

## 1. The question (a measurement arm, not an adoption verdict)

**Can the top decider serve whole on ONE house-class 24 GB card, and how fast?** The FP8 file (30.41 GB) cannot
load on a 24 GB card: on 2026-09-21 it needed both of benchbox's 3090s, tensor-parallel over the host bridge
(arm 2, 0.350 s a decision). The OpenJev org published `openjev/openjev-GGUF` on 2026-09-24, built from the
revision this bench measured (`5ec9e5fd`), and its card says Q4_K_M "fits 24 GB cards (RTX 3090 / 4090)".

It is **not** a seat exam, an adoption verdict, or a claim about any serving path but this one. **OpenJev stays
BENCH-ONLY** (section 3): these weights never become a seat.

## 2. The pins

| what | pin |
|---|---|
| model file | `openjev/openjev-GGUF` at revision **`208220bc40401670015640fc7216b9a583515abc`** (committed 2026-09-24T00:26:04Z; repo `main` read equal to it at 14:24:10Z and again at 14:35:14Z, after the pull), file `OpenJev-Q4_K_M.gguf`, **16,547,400,000 B**, sha256 **`7baa5501dfeb7d2b80d7bfa3fe85373304557e30efcea8fb350b5f3f1252ef21`**: HF's LFS oid at that revision, the repo's own `SHA256SUMS` and `MANIFEST.json`, ollama's layer digest, and the sha256 recomputed on benchbox at 14:35:07Z all agree (`receipts/oj-gguf/blob.sha256`) |
| what it was built from | the org's `MANIFEST.json`: source `openjev/openjev` revision **`5ec9e5fd2f80a6fff386779b1e5ac7e389971889`** (the revision arm 3 served in BF16), `llama.cpp b11147 (prebuilt ubuntu cuda-12.8)`, `convert_hf_to_gguf.py b11147, --no-nextn, mtp_num_hidden_layers 0`. Text only: no vision tower, no projector. No imatrix is named |
| the file's own header | `general.architecture` **qwen35**, 64 blocks (full attention every 4th: 16 full-attention layers, 48 linear-attention layers), hidden 5,120, 24 query heads and 4 KV heads of 256, vocab 248,320, `general.parameter_count` 26,895,998,464, `file_type` 15 (Q4_K_M). 851 tensors, 16,536,406,016 B: 433 Q4_K (12,076,646,400 B), 65 Q6_K (4,449,177,600 B), 353 F32 (10,582,016 B). `token_embd.weight` 715,161,600 B (Q4_K), `output.weight` 1,042,944,000 B (Q6_K) (`receipts/oj-gguf/gguf-header.json`) |
| the chat template | the GGUF's `tokenizer.chat_template`, 8,952 chars, sha256 **`c3cf9e34abf4f9e36c2d72165aa9c132d3e2a725b6c2586aaa3a8af9d7a81041`**: **byte-identical** to `chat_template.jinja` of `openjev/openjev@5ec9e5fd` (arm 3's BF16) and of `openjev/openjev-FP8@4ec320f2` (arms 2, 3 FP8, 4 and gap 4). With `enable_thinking` false it ends the prompt `<|im_start|>assistant\n<think>\n\n</think>\n\n` |
| the ollama manifest | `/workshop/bench-store/ollama/manifests/hf.co/openjev/openjev-GGUF/Q4_K_M`, 416 B, sha256 `a16c4708a12d7ab413ce942622d70d04ea33d3f6bf95b072e329c1ec038b920e`: ONE layer (the model, digest = the blob's sha256), config `9aa66413…3c22` (family `qwen35`). **No template layer, no params layer, no licence layer.** `/api/show` reports the GGUF's own jinja as the template and capabilities `tools, thinking, completion` (`receipts/oj-gguf/api-show-summary.json`) |
| the model's own parameters | none in the manifest. The GGUF carries `general.sampling.temp` 1.0, `top_k` 20, `top_p` 0.95. Every request sets `temperature` 0, `seed` 20260921, `num_predict` 1 (readout) or 8 (generate), `num_ctx` 16,384, `think` false, `logprobs` true, `top_logprobs` 20 (`run.py` `call_ollama`, unedited). The readout softmaxes six letter logprobs; any renormalisation over a top-k set adds one constant to all six and cancels in that softmax, so the GGUF's sampling keys cannot move a readout unless a letter falls out of the top 20, which is counted (floored rows) |
| runtime | **ollama 0.32.13**, `/usr/local/bin/ollama` (sha256 `39c8dea520be728599c4b8fda8f3466b4caca0c50fff6e5c065d0368c74d551c`, installed 2026-08-14, the binary arm 1 and arm S used), started by the runner as its own child with the store `/workshop/bench-store/ollama` on **127.0.0.1:11436** (arm 1's port). **One slot** (`<the runtime's parallel-requests variable>=1`), one loaded model, `OLLAMA_KEEP_ALIVE=-1`, `OLLAMA_CONTEXT_LENGTH=16384`, **`num_ctx` 16,384** on every request (the model card's `-c 16384`; arms 2 and 3's `--max-model-len 16384`). Flash attention requested as arm 1 did (`OLLAMA_FLASH_ATTENTION=1`); what ollama resolves is read back from its runner's argv |
| the box | **benchbox**, **Ubuntu 26.04 LTS** (`/etc/os-release`), kernel `7.0.0-31-generic`, NVIDIA driver **595.84** (open module), AMD Ryzen 7 3700X (16 threads), 30,990 MiB RAM. Every row carries `hostos.read_os()` (`bench/opendecider-2026-09-28/hostos.py`, sha256 `a6a74bba…e3ad`, the canonical's byte-twin) |
| the card | **one NVIDIA GeForce RTX 3090 24 GB** (EVGA XC3 Ultra), fingerprint **`33a0b4acb8e8`** = `sha256("GPU-<uuid>")[:12]` (the board L-2 and O-2 read). Internal only: the house calls it the 3090 board, UUID `GPU-fde82ed0…`. In benchbox's **x16** slot (PCIe gen 4 max), the x4 slot empty, **power limit 300 W** (default 350 W), **persistence Enabled**, 24,576 MiB. The UUID lives only in the run's environment (`OJG_CUDA`); every file the run writes is rewritten with the fingerprint before it leaves the box |
| the UPS | NUT `pr1500@localhost`; 40 % load read at 14:4xZ with the other lane's two vLLM seats resident |
| instrument, unedited | `run.py` `<withheld: rewritten file, see README.md>`, `run_addenda.py` `a34aca61…89acab`, `run_gaps_card.py` `3ea0a585…43bae66` (full values in `run_oj_gguf.py` `INSTRUMENT_SHA256`; the runner refuses a moved file) |
| instrument, new in this commit | `run_oj_gguf.py` (the runner and the Arm rows), `guard_oj_gguf.py` (L-2's guard `57bed707…e934` with the names changed), `mock_oj_gguf.py` (the no-card whole-path mock; never run on the box). Their sha256s are this commit's and every run receipt records the runner's as read |
| the kits | S0 `kit/task_c.json` `61a0b1d6…2d7f39` (108); the doorman `kit/doorman_planted.json` `4daf059c…22f7` (48); H48 `../deem-2026-09-27/kit/h48.json` `<withheld: held-out set, see README.md>` (48, a LIVE control: rows carry each line's sha256, never its text); the roster `kit/manifest.json` `<withheld: rewritten file, see README.md>` |
| seed | 20260921 on every request (`Arm.seed`); gap 4's orders are `random.Random(20260921 + k)` (`run_gaps_card.permutations_for`, unedited) |

### The Arm rows (`run_oj_gguf.register_arms`)

| arm | model string | mode | host | cards | num_ctx | quant | box | cap |
|---|---|---|---|---|---:|---|---|---:|
| `openjev-q4km-readout` | `hf.co/openjev/openjev-GGUF:Q4_K_M` | readout | 127.0.0.1:11436 | the one card, by UUID from the environment | 16,384 | Q4_K_M GGUF, `openjev/openjev-GGUF@208220bc`, blob `7baa5501…`, llama.cpp b11147 from `openjev/openjev@5ec9e5fd` | benchbox | the limit read at start (300 W) |
| `openjev-q4km-generate` | the same | generate | the same | the same | 16,384 | the same | benchbox | the same |

They differ in the readout and in nothing else, as arms 1 and 2's pairs did.

## 3. The licence, and a discrepancy recorded rather than resolved

| repository (HF API, read 2026-09-28T14:24:36Z) | card `license` | a LICENSE file in the repo |
|---|---|---|
| `openjev/openjev` (the base; `main` `0b6bb6e5`) | `cc-by-nc-4.0` | `LICENSE` (the CC BY-NC 4.0 text), `LICENSE-APACHE-2.0` (for `helper/` and `serve/`) |
| `openjev/openjev-FP8` | `cc-by-nc-4.0` | none |
| `openjev/openjev-MLX`, `openjev/openjev-MLX-4bit` | `cc-by-nc-4.0` | the 4-bit carries the same two files as the base |
| `m3kro/openjev-NVFP4-W4A4` (community) | `cc-by-nc-4.0` | the same two files |
| **`openjev/openjev-GGUF` @`208220bc`** | **`apache-2.0`** | **none**: 13 files, none of them a licence |

The base repo's own README, section "Licence": *"OpenJev weights are released under **CC BY-NC 4.0**: free for
research and other non-commercial use, with attribution. For commercial use, open a discussion on this
repository. The files in `helper/` and `serve/` are Apache 2.0."* The GGUF is a quantisation of those weights
(its own `MANIFEST.json` names the source revision). **The GGUF card's `apache-2.0` is a discrepancy, not a
grant**: a format conversion does not relicense weights, and the org's statement on the base repo governs.
So arm 13 is **BENCH-ONLY**, like arms 2, 3, 4 to 10 and the gaps: nothing here serves, and nothing here is
published. Whether to ask about the discrepancy in discussion #1 (the house's open licence enquiry,
`huggingface.co/openjev/openjev/discussions/1`, no reply yet) is an operator's call, from his account.

## 4. How it runs (the lane's commands; no operator paste)

### 4.1 The server

`run_oj_gguf.py` starts `/usr/local/bin/ollama serve` as its own child, in its own unit's cgroup, with **arm 1's
unit environment** (`receipts/2026-09-21-substrate.md`: `OLLAMA_HOST=127.0.0.1:11436`,
`OLLAMA_MODELS=/workshop/bench-store/ollama`, `OLLAMA_VULKAN=0`, `GGML_VK_VISIBLE_DEVICES=-1`, `HIP_VISIBLE_DEVICES=`,
`ROCR_VISIBLE_DEVICES=`, `OLLAMA_KEEP_ALIVE=-1`, `OLLAMA_FLASH_ATTENTION=1`) and these deltas, each named:

| delta | why |
|---|---|
| `CUDA_VISIBLE_DEVICES=<this card's UUID>` | the card (arm 1's was the other 3090 in the x4 slot) |
| `OLLAMA_CONTEXT_LENGTH=16384` (arm 1: 131,072) | arm 1 needed the docent's 82 k-token pages; every arm-13 prompt is under 600 tokens, and 16,384 is arms 2 and 3's window and the model card's. Set on the server too, so no request reloads the model at a different window |
| `<the runtime's parallel-requests variable>=1`, `OLLAMA_MAX_LOADED_MODELS=1` | explicit; the build's own default is already one slot (its `server config` line at 14:29:30Z) |
| `OLLAMA_NO_CLOUD=1` | the house phone-home switch (D-20260928-008 / -016), in the instance's OWN environment |
| `OLLAMA_NOPRUNE=1` | the store is shared: an ollama start otherwise deletes blobs no manifest names (none today, 55 of 55 referenced at 14:2xZ; the switch keeps it so) |

**benchbox's system ollama is not touched**: `ollama.service` and `ollama-11435.service` read inactive and
stay so. It has not had its phone-home switch yet (D-20260928-023: benchbox's kit2 step waits until the day's
card lanes are done); arm 13's instance carries the switch itself.

### 4.2 The switches, read back rather than assumed

- **The server's own words.** Its `server config` line must read `OLLAMA_NO_CLOUD:true` and
  `OLLAMA_VULKAN:false`, and its device lines must show exactly one CUDA device, filtered to this card's UUID,
  and no Vulkan device (D-20260920-133).
- **Its environ, after that line.** `/proc/<pid>/environ` of the server must carry `OLLAMA_NO_CLOUD=1` and this
  card's UUID, and no `DO_NOT_TRACK`. It is read after the startup line, never the launcher's (the 09-28
  false FAIL, D-20260928-016).
- **The runner and the guard** run in units launched with `OLLAMA_NO_CLOUD=1 ORT_DISABLE_TELEMETRY=1
  HF_HUB_DISABLE_TELEMETRY=1 HF_HUB_OFFLINE=1 VLLM_NO_USAGE_STATS=1`. The runner is stdlib only; it refuses to
  start if any is missing. **`DO_NOT_TRACK` is never set** anywhere in this arm, and the runner refuses to
  start if it is present.

### 4.3 A fetcher the switch may not cover, fenced by time

ollama 0.32.13 carries a **model-recommendations fetcher**. Its cache, `/workshop/.ollama/cache/model-recommendations.json`
(1,156 B, a list of `:cloud` models), was last written **2026-09-26T17:26:41Z**, before the house switch-offs,
and was **not** rewritten when the pull instance started with `OLLAMA_NO_CLOUD=1` at 14:29:30Z. That instance
logged `model recommendations cache sleep scheduled wait=3h47m29.9s`; arm S's instance at 11:32:03Z drew
`wait=4h38m32.2s`, so the wait is drawn per start. Whether `OLLAMA_NO_CLOUD` silences the fetch is **not
known**. So the runner reads the server's own `wait=`, and if the next attempt would fall **within 2 h** (the
arm's window is under 45 min) it **restarts the server before any model is loaded**, at most three starts in
all, then stops; every start's draw and log are kept. It records the cache file's mtime before and after the
run. This is flagged to the lead for the phone-home train. It is not fixed here.

### 4.4 The stop guard, the lock, the units

- **The guard** (`guard_oj_gguf.py`, its own unit `oj-gguf-guard`, launched with the switches): L-2's guard with
  the names changed. Each second it reads the UPS, free disk, the card's compute processes, power limit and
  persistence; it samples the card at 2 Hz. It **trips on UPS load ≥ 80 %, card core ≥ 83 °C, free disk under
  15 GiB, or the UPS unreadable for 10 s**. A trip records first, then stops `oj-gguf-run` and `oj-gguf-dry`,
  which ends the ollama child with them. It never records and continues. The runner refuses to start unless the
  guard's heartbeat is under 10 s old.
- **The lock** (the lead's protocol): the lane reads `/workshop/CARD-LOCK` absent and the card at **1 MiB with no
  compute process**, outside any sandbox, then writes `/workshop/CARD-LOCK` as `oj-gguf <utc>`. The lane removes the
  lock when the arm is done.
- **The lock is checked FIRST, structurally** (D-20260928-025, the lead's rule from 14:51Z): the first statement
  of the runner's card path and of the guard's `main` is `require_lock()`, which refuses unless the lock's first
  word is `oj-gguf`, **before any posture read, any `nvidia-smi`, any server start and any model call**. A
  posture check is not a lock check. Proved with no card: `mock_oj_gguf.py --negative nolock` and
  `--negative otherlock` (the lock reading `deem-t1 2026-09-28T14:56:35Z`) both refuse in the runner and in the
  guard (exit 4), with **zero** calls in the fake `nvidia-smi`'s call log and the fake server never started.
  No refusal test is ever run against the live runner on the box.
- **The units:**
  `systemd-run --user --unit oj-gguf-guard --collect -p RuntimeMaxSec=5400 -E OJG_CUDA=<uuid> -E <the switches> python3 -B <stage>/bench/jev-2026-09-21/guard_oj_gguf.py --tag <tag>`,
  then `systemd-run --user --unit oj-gguf-dry|oj-gguf-run --collect -p MemoryMax=20G -p MemorySwapMax=0 -E OJG_CUDA=<uuid> -E <the switches> python3 -B <stage>/bench/jev-2026-09-21/run_oj_gguf.py --tag <tag> [--dry 3]`.
  `<stage>` is `benchbox:/workshop/bench-oj-gguf/wt`, a `git archive` of this branch's pushed head (A-1 names it).
  Nothing is pinned to cores, as in arm 1.

## 5. The item sets, and what each task sends (every prompt is an earlier arm's, byte for byte)

| task | set | rows | the prompt | joined to (by position; a differing prompt sha256 STOPS the run) | set fingerprint |
|---|---|---:|---|---|---|
| `c` | S0, the kit's option order | 108 per mode | `run.task_c` | arm 3 BF16 `rows-arm3/openjev-bf16-largecard.c.jsonl` (the same 108 hashes as arms 1, 2 and 3 FP8, and gap 4's k = 0) | `7ca6bcedfd7e190a` |
| `h48` | H48 | 48 per mode | `run.task_c`'s own shape (`c_shape_rows`), over H48's items | none exists: **no arm has read H48 in this layout.** `--check` proves the builder is task (c)'s twin: over S0's items it rebuilds arm 3's 108 hashes | `<withheld: held-out set, see README.md>` |
| `d` | the doorman's planted set | 48 per mode | `run_addenda.task_d` (DOORMAN_SYSTEM v1.2, `"Host: " + line`, [A] admit / [B] refuse) | arm 4 `rows-addenda/openjev-fp8-readout.d.jsonl` | `e2f8caccd6d4e8a2` |
| `creorder` | S0 under six option orders (k = 0 is the kit's) | 648, readout | `run_gaps_card.task_creorder` | gap 4 `rows-gaps-card/openjev-fp8-readout.creorder.jsonl`, same (k, uid) at every position | `6f99313d4e314957` |

The fingerprint is `sha256` over the concatenated sorted distinct prompt sha256s, first 16 hex
(`run_oj_gguf.fingerprint`). `--check` at this commit: every join 108 / 48 / 648 equal; H48's builder 108 of 108;
H48's text hashes all equal their kit's; 1,920 stubbed calls, no server, no card
(`receipts/oj-gguf/check.json`).

**Render parity rides every row that has a reference.** `render_offset` = ollama's own `prompt_eval_count` minus
the token count vLLM reported for the same prompt under the byte-identical template. **Expected 0**: on this
same ollama build, arm 1's `gemma4:26b` counts equalled vLLM's (arm 10) on all 108 S0 prompts. A missing
`<think>\n\n</think>\n\n` closure would read −2.

## 6. The run, in order

### 6.1 Before the server (each one a stop)
First, before anything else, the lock names `oj-gguf` (4.4); then the switches are present and `DO_NOT_TRACK` absent; the guard's heartbeat is fresh;
the three imported instrument files equal their pins; the OS and the driver are read; the card reads **300 W,
persistence Enabled, no compute process, ≤ 50 MiB**; the manifest's sha256 equals `a16c4708…`; **the blob is
re-hashed** (16.5 GB, ~40 s) and must equal `7baa5501…` at 16,547,400,000 B.

### 6.2 The server and residency (each one a stop)
`/api/version` = `0.32.13`; the switch and device reads of 4.2; a recommendations wait of at least 2 h (4.3:
a shorter draw restarts the server, three starts at most);
`/api/show`'s template sha256 = `c3cf9e34…` (so the prompt is rendered from the pinned jinja); **three
discarded warm-ups** (S0 item 0, readout, never rows). Then **residency by growth** (D-20260921-001): the card's
`memory.used` after the warm-ups minus the reading before the server must be **≥ 14,334 MiB**, which is 0.95 ×
(16,536,406,016 B of tensors − 715,161,600 B of `token_embd`, which llama.cpp keeps in host memory) / 2^20.
**Exactly one compute process** on the card, and ollama's own `offloaded N/65 layers to GPU` line must read 65
of 65 when it is printed. `/api/ps` and the runner's argv are recorded, never the proof.

### 6.3 The dry leg (`--tag dry --dry 3`, rows-dry/, never counted)
The first 3 items of every (mode, task) in the order of 6.4: 21 rows. **No counted row exists unless every
one of these holds on the dry rows**, else the arm stops and reports, and any fix is a dated amendment before
the counted run:
1. **Render parity: every render offset is 0** (S0, the doorman, the reorder rows).
2. **No readout row's position-0 token is a thinking marker** (`<think>`, `</think>`, `<|im_start|>`,
   `<|im_end|>`). A letter is not required at position 0: arm 1's base model put `[` there on 88 of 108 rows
   and read its letters from the top 20.
3. No error, no unparsed generation, no refusal.

Gates 1 and 2 are proved able to fire: `mock_oj_gguf.py --negative offset` (the fake adds 2 prompt tokens)
stops at `dry generate.c: render parity broken`, and `--negative think` (the fake puts `<think>` at position 0)
stops at `dry readout.c: position 0 is a thinking marker` (`receipts/oj-gguf/mock-negative-*.json`).

### 6.4 The counted run (`--tag counted`, rows/), in this order
The headline first, so a stopped run still holds it.

| # | arm · task | rows | expected minutes |
|---:|---|---:|---:|
| 1 | `openjev-q4km-generate` · c (S0, **the headline**) | 108 | 2.5 |
| 2 | `openjev-q4km-readout` · c | 108 | 2.5 |
| 3 | `openjev-q4km-readout` · h48 | 48 | 1.2 |
| 4 | `openjev-q4km-generate` · h48 | 48 | 1.2 |
| 5 | `openjev-q4km-readout` · d | 48 | 1.6 |
| 6 | `openjev-q4km-generate` · d | 48 | 1.6 |
| 7 | `openjev-q4km-readout` · creorder | 648 | 13.7 |
| | **total** | **1,056** | **~24** |

Each task goes through `run.run_arm` unedited: its idle read (10 s), its announcement, the 1 Hz board sampler
by UUID, the per-task pin and report, `summarise`. The model stays resident across all seven (one model string,
`KEEP_ALIVE=-1`).

### 6.5 After
The server is stopped; the card must read no compute process and ≤ 50 MiB. Every file the run wrote is
rewritten: the UUID becomes `GPU-fp:33a0b4acb8e8`, any PCI address `<pci>`, any LAN or private-network IPv4
`<lan-ip>`. The laptop copies the rows and receipts into `rows-oj-gguf/`, `rows-oj-gguf-dry/` and
`receipts/oj-gguf/`, and greps them for the UUID, PCI addresses and LAN IPs before any commit. The tables
script (`tables_oj_gguf.py`) computes only the figures registered in sections 7 and 8, from the rows and the
comparators' own files; it is pinned by a dated amendment before the first counted row.

## 7. The readings and the bars, written before any call

| id | reading | the bar, exactly | printed beside it |
|---|---|---|---|
| **W** | whole on one 24 GB card | **"WHOLE"** iff 6.2's residency holds (growth ≥ 14,334 MiB, one compute process, 65/65 when printed) **and** the peak `memory.used` over every task's 1 Hz sampler stays under the card's 24,576 MiB. Otherwise the arm stops at 6.2 and prints **"NOT WHOLE"** with the reading | growth, peak MiB and headroom; `/api/ps` size and size_vram; the runner's argv |
| **C-1** | **the carry bar (the headline)**: does OpenJev's six-way result carry to Q4_K_M, whole on one 3090, through ollama? | on **S0 generate** (kit order, strict: an unparsed answer is wrong): **"CARRIES"** iff **correct ≥ 97 / 108** **and** the chosen guest equals **arm 3 BF16's** on **≥ 100 / 108** items. Otherwise **"DOES NOT CARRY"**. 97 is BF16's 101 less 4, and equals the FP8 file's own worst option order in gap 4 (97 / 108 at k = 1 and k = 2); 100 allows 8 changed answers, over twice the 3.4 % the org measured between its Q4_K_M and 16-bit answers (96.6 % agreement on 1,789 rows) | the Wilson 95 % interval; an exact two-sided McNemar on correctness against arm 3 BF16 (descriptive); agreement with arm 3 FP8 and arm 2 FP8 generate |
| **C-2** | the readout's carry (secondary) | the same bar on **S0 readout**, read only when **no row is floored** (every option letter inside ollama's top 20). With any floored row it prints **"qualified: N rows had a letter outside the top 20"** instead of a verdict (ollama gives 20; the org's own GGUF validation read llama.cpp's top 64) | floored rows and letters; Brier; the 10-bin reliability table; readout vs generate same choice |
| **D** | the doorman | **the 09-13 gate: planted refused ≥ 30 / 36 AND controls refused ≤ 2 / 12**, on the readout rows: "passes / fails the 09-13 gate". The generate rows are printed under the same gate | arm 4 FP8 readout 33 / 36 and 0 / 12 (passes); the v1.2 contract split (the 24 inside v1.2, C + D 12) |
| **R** | reorder, descriptive (no bar) | items whose chosen guest is not the same under all six orders (of 108); order-pairs that disagree (of 1,620); correct per order; the letter-position table | gap 4 FP8 readout: 15 / 108 items, 92 / 1,620 pairs (5.68 %), orders 101 / 97 / 97 / 98 / 99 / 100. **The k = 0 control:** k = 0's prompts are S0's, so its choices against task 2's are ollama's own determinism, printed as a count |
| **H** | H48 | **S0 − H48 on generate**, strict, with Newcombe's hybrid-score 95 % interval (independent sets): **"CONTAMINATION SIGNAL"** iff the lower bound is above 0; otherwise "no signal" (the control cannot prove absence). The same on readout, printed | H48 by chair (mistral 28 / gemma 20) and S0 by chair (57 / 51); the timeline below; the other readers' H48 counts from their own files, labelled as other renderings |
| **L** | how fast, and at what energy (descriptive) | median and p95 seconds per decision per task; ollama's own `prompt_eval` ms; joules per decision on ONE board, raw and net of the 10 s idle read; the model's load-to-warm seconds | the comparison table of section 10 |

**H48's timeline, stated rather than inferred.** `openjev/openjev@5ec9e5fd` is dated 2026-09-21T10:06:02Z. The
wall's lines have been public since 2026-09-19 (so the COI's rule R-1 flags every S0 and H48 cell as
contamination-possible for wall leakage); the S0 kit went public with exhibit fifty-eight at 2026-09-23T13:17:23Z,
after those weights, and the GGUF (2026-09-24T00:24Z) is a post-training quantisation of them. So kit leakage
is excluded by the timeline, and H48 here reads mostly as **generalisation to 48 lines never in any kit**.

## 8. Predictions, written before the first call (80 % intervals)

| figure | prediction | 80 % interval |
|---|---|---|
| W: whole on the card | yes (p ≈ 0.95) | |
| growth after the warm-ups | 17,600 MiB | 15,800 – 19,800 |
| peak `memory.used` | 18,300 MiB | 16,500 – 21,000 |
| S0 generate, strict (the headline) | **100 / 108** | 97 – 103 |
| S0 readout | 100 / 108 | 97 – 103 |
| S0 generate agrees with arm 3 BF16 | 105 / 108 | 102 – 108 |
| S0 readout and generate choose the same | 107 / 108 | 105 – 108 |
| C-1 | **CARRIES** (p ≈ 0.80) | |
| floored S0 readout rows | 0 | 0 – 3 |
| render offset 0 on S0 | 108 / 108 | 106 – 108 |
| H48 generate, strict | 43 / 48 | 39 – 46 |
| H | no signal (p ≈ 0.9) | |
| doorman, planted refused (readout) | 33 / 36 | 30 – 35 |
| doorman, controls refused (readout) | 0 / 12 | 0 – 1 |
| D | passes the 09-13 gate (p ≈ 0.85) | |
| reorder: items whose guest moved | 16 / 108 | 10 – 24 |
| reorder: order-pairs that disagree | 100 / 1,620 | 60 – 150 |
| reorder k = 0: same choice as S0 readout | 108 / 108 | 106 – 108 |
| S0 readout, median s per decision | **1.25 s** | 0.95 – 1.90 |
| S0 generate, median s | 1.30 s | 1.00 – 1.95 |
| doorman readout, median s | 1.70 s | 1.20 – 2.60 |
| S0 readout, ollama's own `prompt_eval` median | 300 ms | 150 – 600 |
| S0 readout, J per decision (one board, raw) | 250 J | 150 – 450 |
| load to warm (server start → third warm-up) | 35 s | 15 – 90 |
| **card minutes, the dry leg and the counted run together** | **~29 min** | 22 – 40 |

*Why ~1.25 s and not arm 3's 0.09 s:* on this ollama build every request has carried ~0.8–0.9 s that is not the
model's (arm 1 §R2: ~870 ms on 286-token prompts; arm S 09-28: first token 1,018 ms against vLLM's 226 ms), and
a dense 27B at Q4_K_M prefills a ~292-token prompt in a few hundred ms on one 3090. If ollama's overhead is not
there for this model, the latency falls well under the interval, and that is a finding, not a failure.

The predictions are scored after the tables are drawn: each observed value inside or outside its interval, and
the count of hits.

## 9. Validity: a counted reading is VOID, never "low", unless all of these hold

- **V-1** Every task answered once: 108 / 48 / 48 / 648 rows, positional, no duplicate (task, position).
- **V-2** Every S0, doorman and reorder row's prompt sha256 equals its reference row's (enforced per row; a
  mismatch stops the run), and H48's set fingerprint equals `<withheld: held-out set, see README.md>`.
- **V-3** The blob and the manifest equal their pins (6.1); the template equals its pin (6.2).
- **V-4** Residency holds (6.2) and the card's peak stays under 24,576 MiB (reading W).
- **V-5** **Render parity on at least 106 of 108 S0 rows in each mode.** Otherwise the S0 readings print
  **"VOID: render parity"** and C-1 and C-2 are not read. Every non-zero offset is listed with its row.
- **V-6** No readout row's position 0 is a thinking marker (6.3's list).
- **V-7** Zero errors and timeouts; the posture reads 300 W and Enabled before and after; the guard never
  tripped; the lock named `oj-gguf` at start (the runner's and the guard's receipts both record it); the switches were read back (4.2) and the recommendations wait
  was ≥ 2 h (4.3).

## 10. The comparison plan

**The decisions are like-for-like where the prompt is:** the same prompt sha256 per item, the same chat template
bytes, the same letters, the same calibration constants (`READOUT_T` 0.85), the same seed. **The runtime, the
precision and the board are not**, and every table says which.

| comparator | what it is | S0 | latency (S0 median) | file |
|---|---|---|---|---|
| **arm 3 BF16** | the same weights at 16-bit, vLLM, one RTX PRO 6000 (the inference box's 96 GB card), 420 W | 101 / 108 | 0.092 s | `rows-arm3/openjev-bf16-largecard.c.jsonl` |
| arm 3 FP8 | the FP8 file, native FP8, the same card | 101 / 108 | 0.087 s | `rows-arm3/openjev-fp8-largecard.c.jsonl` |
| **arm 2 FP8** | the FP8 file, vLLM 0.29.0, **two** 3090s tensor-parallel over PHB, 250 W each | 101 / 108 (readout and generate) | 0.350 s (readout), 0.392 s (generate) | `rows/openjev-fp8-*.c.jsonl` |
| arm 1 | a different model (the jevify gemma4-26B-A4B adapter, Q4_K_M) on **the same ollama build**, one 3090 in gen3 x4 at 250 W | 90 / 108 | 1.017 s | `rows/jev-readout.c.jsonl` |
| arm 4, gap 4 | the FP8 file on arm 2's server: the doorman, the six orders | see section 7 | 0.58 s (doorman) | `rows-addenda/`, `rows-gaps-card/` |

- **Accuracy:** arm 13 against arm 3 BF16 is the carry question (C-1). Against arm 3 FP8 and arm 2 it is the
  same question at another precision. Per-item agreement and an exact McNemar are printed for each; the
  confusion table (six guests) is printed beside arm 2's.
- **Latency:** **cross-runtime and cross-board in every row**, and the table's header says so. The nearest
  runtime twin is arm 1 (the same ollama binary), with a different model and board. The nearest model twin is
  arm 3, on a different runtime and a 96 GB card. Nothing here is a ratio against a vLLM row without that
  label.
- **Energy:** arm 13's joules are **one** board's; arm 2's are the sum of two; arm 3's are a different card.
  A joules-per-decision figure here is never subtracted from theirs.
- **H48:** no OpenJev reading exists to compare with. Deem D0 (4 / 48), OpenDecider nano (5 / 48), small
  (25 / 48), its base (26 / 48) and naive Bayes (27 / 48) are printed from their own files as **other readers
  under other renderings**, counts only.

## 11. What arm 13 may and may not say

**It may say:** whether OpenJev's org-built Q4_K_M, as pulled (`208220bc`, `7baa5501…`), served by ollama
0.32.13 on one RTX 3090 at 300 W in x16, sits whole on the card and with what headroom; its S0, H48, doorman and
reorder figures under this protocol, beside arms 2 and 3 on the same prompts; its seconds and joules per
decision through this server.

**It may not say:**
- anything about serving OpenJev: the weights are CC BY-NC by the base repo's statement (section 3), so the
  arm licenses a **finding**, never a seat;
- that this is llama.cpp's speed (ollama's per-request overhead is inside every number) or vLLM-comparable;
- anything about the other GGUF quantisations, the 10,000-question set, the targeted readout, or long pages
  (every arm-13 prompt is under 600 tokens);
- "beats" or "matches" arm 3 on a count difference alone: C-1 is the registered reading, and McNemar is printed
  beside it as description.

## 12. Deviations, each named

1. **The queue's shorthand "arm 3 (FP8, TP=2, 0.35 s)"** is this bench's **arm 2**. Arm 3 is the BF16 and FP8 pair
   on one RTX PRO 6000 (0.092 s and 0.087 s). Both are comparators, under their README names.
2. **Not in a sandbox.** L-2 and O-2 ran their servers in a network-less `bwrap`; arm 1 and arm S ran ollama
   bare with its switches. Arm 13 follows arm 1 and arm S, with the switch read back (4.2) and the fetcher
   fenced by time (4.3).
3. **Doorman and H48 in both modes; reorder in readout only.** The lead's list is "S0 generate + readout, the
   doorman, reorder × 6, H48". The generate rows of the doorman and H48 are added because C-1's headline is a
   generate count and H must subtract like from like (96 rows, ~3 min). Reorder's generate twin (648 more rows,
   ~14 min) is not run.
4. **`num_ctx` 16,384, not arm 1's 131,072** (4.1).
5. **The run's order puts S0 generate first** (6.4), so the headline survives a stopped run.

## 13. Minutes, stops, cleanup

- **Card minutes:** ~29 (22 – 40), the dry leg and the counted run together.
- **Disk:** the blob stays in `/workshop/bench-store/ollama` (the handoff law: no bench's store is deleted), as the
  queue's disk ledger (§4) budgets. benchbox's `/` read 132,109,893,632 B free before the pull and
  115,557,040,128 B after (14:35:14Z), against the §4.3 line of 65 GB for the follow-up plus ~17 GB for OVV-D's
  staging: 82 GB, cleared by 33.5 GB. The run writes megabytes of rows and receipts.
- **Stops:** every stop of 6.1, 6.2 and 6.3 prints its reason and leaves the card empty (the server is the
  runner's child; a guard trip stops both). The lane removes `/workshop/CARD-LOCK` when it is done, including after a
  stop.

## 14. Sources

- The queue: `/workshop/bench-archive/plans-2026-09-28/benchbox-day-2026-09-29/QUEUE-3090-2026-09-28.md` rank 2, §3 note 2,
  §4; `DAY-PLAN.md` §4.3.
- This bench: `README.md` §4, §7, §8, §R2, §R6, §Arm 2, §Arm 3; `TABLES-ADDENDA.md` arm 4; `TABLES-GAPS-CARD.md`
  gap 4; `receipts/2026-09-21-substrate.md` (arm 1's unit).
- `bench/lev-2026-09-28/PREREG-lev.md` L2.2 and L2.3 and `guard_lev2.py` (the posture, guard and sanitiser
  pattern); `bench/deem-2026-09-27/PREREG-v0.md` §7 (H48) and `kit/README.md`;
  `bench/ollama-vs-vllm-2026-09-28/s/receipts/bringup.md` (ollama 0.32.13 on this card).
- DECISIONS.md D-20260921-001, D-20260920-133, D-20260928-008, -016, -023, -024, and the lead's lock rule
  D-20260928-025 (relayed 2026-09-28 ~15:0xZ).
- Hugging Face API, read 2026-09-28T14:24Z–14:35Z: `openjev/openjev-GGUF` (the tree with LFS oids at `208220bc`,
  `MANIFEST.json`, `SHA256SUMS`, `README.md`, the commits), `openjev/openjev` (`chat_template.jinja` at
  `5ec9e5fd`, `LICENSE` and README at `0b6bb6e5`, the commits), `openjev/openjev-FP8@4ec320f2`
  (`chat_template.jinja`), the org listing and `?search=openjev` (licence tags). Their own claims, never verified.
- The box: benchbox, read-only reads 14:19Z–14:45Z, and the pull (`receipts/oj-gguf/`).

---

## A-1 (2026-09-28, written 15:08Z–15:1xZ): the push time, the tables and the launcher pinned, the staging

*Cold scroll: what this amendment is.* The registration above was committed as `22e53775` and pushed at
**2026-09-28T15:05:00Z** (GitHub's activity API for `refs/heads/bench/oj-gguf-2026-09-28`: `branch_creation`,
after `22e53775b2d084f1be71507017e30404108efda6`). **No arm-13 model call exists**, dry or counted; the card has
not been read by this arm. This amendment pins three files and records the staging. No registered text above
changes.

- **The tables script, pinned before any row:** `tables_oj_gguf.py`, sha256
  `550620f4d51dc6e8b5f210f2368c77aee51672e82281a93edf1568c86b84f619`. It computes the readings of §7 (W, C-1,
  C-2, D, R, H, L), the validity lines of §9, and §8's predictions scored, from the rows, the run receipt and the
  comparators' own files; nothing else. It was drawn once over the mock's fake rows to prove it runs end to end;
  those numbers are the reference arms' echoed back by the fake server and mean nothing.
- **The launcher:** `launch_oj_gguf.sh`, sha256 `9128892fd06b9535e23f57e5953c9115e569c0971b9b24f319797cc97475d2e5`.
  It is §4.4's unit lines as a script. Its first act is the lock check (D-20260928-025): it refuses unless
  `/workshop/CARD-LOCK` names `oj-gguf`, before it reads the card's UUID. Then it requires exactly one card with
  fingerprint `33a0b4acb8e8`, starts the guard (tag `arm13`, one guard for both legs) if it is not running, waits
  for a heartbeat at most 5 s old, and starts `oj-gguf-dry` (`--tag dry --dry 3`) or `oj-gguf-run`
  (`--tag counted`). `stop-guard` stops the guard. It is never run to test a refusal on the box.
- **The mock** (`mock_oj_gguf.py`, never run on the box) now writes its fake lock's time in UTC; sha256
  `f6ee62020d2d2cd5b6106944398cb1e2a843714fcc1602226c265e3fd97600a6`. `run_oj_gguf.py` (`6263e3fd…0e52`) and
  `guard_oj_gguf.py` (`8e626613…6e91`) are unchanged since `22e53775`.
- **The staging:** `benchbox:/workshop/bench-oj-gguf/wt` is a `git archive` of this amendment's commit, holding the
  runner, the guard, the launcher, the tables, the three imported instrument files, the kits and the reference
  rows (`/workshop/bench-oj-gguf/STAGED.json` names the commit and a tree hash). `run_oj_gguf.py --check` on benchbox's
  own Python 3.14.4 read every join equal (108 / 48 / 648), H48 `<withheld: held-out set, see README.md>`, 1,920 stubbed calls, no
  server, no card.
- **The disk** read 94,963,687,424 B free at 15:05:55Z (other lanes' writes since the pull's 115,557,040,128 B
  at 14:35:14Z), still 13 GB above §13's 82 GB line.
- **The card** is not ours: `/workshop/CARD-LOCK` read `deem-t1 2026-09-28T14:56:35Z` at 15:05:55Z. Arm 13 waits for the
  lead's hand-off.

---

## A-2 (2026-09-28, written 16:15Z–16:2xZ): one label fix in the pinned tables script, after the first draw

*Cold scroll: what this amendment is.* After the counted run, the first draw of `TABLES-ARM13.md` printed the
paired discordant counts against arm 3 BF16 as "0 / 2" (only BF16 right / only this arm right). A direct recount
from the two row files reads the other way round: **BF16 right and this arm wrong on 2 items (positions 37 and
102), the reverse on 0.** `tables_oj_gguf.py`'s `paired()` returned the two counts under each other's names.
**No registered reading moves:** C-1 and C-2 read the correct count and the same-guest count, which were right;
McNemar's p is symmetric in the two counts. The fix renames the counts to match their arithmetic; nothing else
in the file changed. The fixed file's sha256 is `2f83ac0982d42df62324a92a80b0784398ea858f2ea6dfd6126f91e91ba6cf6d`
(A-1 pinned `550620f4…f619`). The mock could not have caught it: its fake server echoed the reference arm's own
answers, so both discordant counts were 0.

## Results (2026-09-28, pointer) — drawn in `TABLES-ARM13.md`

*Cold scroll: what was measured.* OpenJev's Q4_K_M GGUF (`208220bc`, blob `7baa5501…`), served by ollama 0.32.13
on one RTX 3090 (fingerprint `33a0b4acb8e8`, 300 W, x16) on benchbox, asked 1,056 counted decisions between
**2026-09-28T15:47:53Z and 16:12:27Z**, after a 21-row dry leg (15:44:38Z–15:46:56Z) that passed every gate
(render offset 0 on all 15 referenced rows, a letter at position 0 on all 12 readout rows). The lock was held
15:44:31Z → 16:13:04Z (28.6 min); the guard read 1,562 times and never tripped (UPS load at most 38 %, card
core at most 66 °C). Zero errors, zero refusals, zero unparsed answers, zero floored rows.

| reading | result |
|---|---|
| **W** | **WHOLE**: grew 16,727 MiB against the 14,334 MiB bar, 65 / 65 layers offloaded, one compute process; peak 16,730 MiB of 24,576 (7,846 MiB headroom); ollama's own buffer line: CUDA0 15,088.32 MiB of weights, 682.03 MiB in host memory |
| **C-1 (the headline)** | **CARRIES**: S0 generate **99 / 108** (91.7 %, Wilson 84.9–95.6 %), the same guest as arm 3 BF16 on **106 / 108**; the two changed answers are items BF16 had right (McNemar p 0.5) |
| **C-2** | **CARRIES**: S0 readout 99 / 108, 0 floored rows, render parity 108 / 108; readout and generate choose the same guest on 108 / 108 |
| **D** | **passes the 09-13 gate**: planted refused 32 / 36, controls refused 0 / 12 (arm 4's FP8: 33 / 36, 0 / 12) |
| **R** | 14 / 108 items moved, 101 / 1,620 order-pairs disagree (6.2 %); correct per order 99 / 94 / 100 / 100 / 97 / 99 (FP8 gap 4: 15 / 108, 92 / 1,620, 101 / 97 / 97 / 98 / 99 / 100); the k = 0 control matched S0 readout on 108 / 108 |
| **H** | **no signal**: H48 44 / 48 in both modes; S0 − H48 +0.0 pts (−8.4 to +11.9) |
| **L** | S0 median **1.14 s** (generate) and **1.13 s** (readout), p95 1.22 / 1.20 s; ollama's own prompt_eval 312 ms; doorman 1.54–1.57 s; about **211–214 J** a decision on one board (net of idle 63–84 J). Beside it: arm 2 (FP8, two 3090s, vLLM) 0.350 s; arm 3 (vLLM, one RTX PRO 6000) 0.092 s BF16 and 0.087 s FP8; arm 1 (a different model on this ollama build) 1.017 s |
| predictions | **18 of 19** numeric predictions inside their 80 % intervals; the miss is load-to-warm (11.6 s against 15–90 s: the blob re-hash had just put the file in the page cache) |

**The recommendations fetcher (4.3):** each start drew a wait of 3 h 13 m (dry) and 4 h 02 m (counted), no
restart was needed, and the cache file's mtime did not move across either leg.

## Independent check (2026-09-28, written 16:30Z–16:3xZ): the readings hold; three things disclosed

*Cold scroll: what this block is.* A checker who is not the lane recounted arm 13's registered readings from the
committed rows with its own stdlib code. It also read the receipts, the push times (GitHub's activity API:
`22e53775` at 15:05:00Z and `aedf918e` at 15:08:54Z, both before the dry leg's first model call at about
15:45:00Z) and benchbox's raw copies (64 files byte-equal to the committed ones; the two runner logs are equal
once their 14 ANNOUNCE lines carry the fingerprint). Every figure above and in `TABLES-ARM13.md` recounts the
same. `tables_oj_gguf.py` at this head redraws `TABLES-ARM13.md` byte for byte, and A-1's pinned script redraws it
with only the discordant-count cells swapped ("0 / 2"), as A-2 says. **No figure changes here.** The text above
leaves out three things:

- **H48's other readers.** Section 10 names five other readers of H48, and `TABLES-ARM13.md` prints four. Naive
  Bayes (27 / 48, section 10's own figure) is missing because A-1's script names no file for it. Its rows
  (`bench/deem-2026-09-27/t1-tune/rows/t1-baseline-nb.h48.P.rep1.jsonl`) came into master with `1b4f1e14`
  (committed 15:43Z), after the script was pinned, and are not on this branch. That file reads 27 of 48
  correct. It is a count only, under another rendering, as section 10 says.
- **V-7's "posture before / after" line** prints the receipt's `card_empty_after` as its second value, not the
  posture after the run. That posture (`posture_after` in `run/counted-run.json`, 16:12:31Z) reads 300 W,
  persistence Enabled, 0 compute processes and 1 MiB, so V-7 holds as registered.
- **The dry leg asked 28 items, not 21.** `run_oj_gguf.joined` stops a dry task only after its generator has
  built the fourth row. So each of the seven dry tasks made one extra model call that became no row: 31
  `/api/chat` calls in `run/dry-serve.log` (3 warm-ups, 21 rows and 7 discarded). The counted run has no limit
  and made 1,059 calls (3 warm-ups and 1,056 rows). The guard's one read of 2 compute processes (15:44:58Z) falls
  inside the dry server's device discovery (15:44:55.9Z–15:44:58.8Z), before any load. The residency read
  (15:45:07Z) and every read under load show 1.
