# EVIDENCE v2: the small open deciders, one table of record per reading

*The evidence lane for the decider article. an operator's go, 2026-09-28 ~21:3xZ: "let's write up an article with all
of these models, comparing them, based on your recs". This file is a DRAFT input for the writer: no pour, no deploy,
no release. Written 2026-09-28 from 21:30Z to 21:5xZ (UTC) on the laptop, before its power-down for the drive install;
closed from 22:42Z to 22:5xZ (UTC) by the evidence-close lane, which folded five post-hoc sets into the recount
(the addendum below). On 2026-09-29 the decider-v5 fold lane folded the page's own checks in as well (the second
addendum).*

**Where every figure comes from.** Estate `origin/master` at **`f2073219`** ("ledger: bench client pr-2 landed; the bos
ruling"), read in a detached worktree removed afterwards. The eight bench dirs this file reads (jev, lev, apus,
imajev, opendecider, deem, the cost series and the 09-13 doorman) are unchanged from `1ce595cb`, this file's first
ref (`git diff --stat 1ce595cb f2073219` over them is empty), so every `path@commit` below is the same as before. Every count, percentage, interval, tie range, gate
outcome, reorder count, reach count and latency below was **recounted from the committed row files** by
`recount_v2.py` (stdlib only), never copied from a TABLES file. Each table is drawn from that script's JSON by
`render_evidence_v2.py`, so no number is typed by hand. Every row names its row file as `path@commit`, where the
commit is the last one on master that touched that file. Appendix A has the scripts, their sha256s and the
command that re-derives every cell.

**This file is internal.** Row paths name boxes and cards (`benchbox`, `largecard`, `-3090-`); the article must use the
"hardware, as the page may say it" column in table 9 and never a path. No H48 line and no wall line appears here,
only H48 counts.

**Addendum, 2026-09-28 22:42Z to 22:5xZ (UTC): five sets computed after the rows were seen, now in the recount.**
What is counted: the doorman's 36 planted lines split by whether the v1.2 rules name their kind (**table 4b**);
the six-way's count under each of its six option orders, and their mean (**table 5b**); the six-way split by which
of the two chair models wrote each line, with exact paired tests inside each family's lines (**table 1c**);
imajev-4B's graphics-card latency on the sets it read there (**table 9a**); and the memory each graphics-card run
used on its 24 GB card (**table 9b**). Every one is **descriptive, computed after the rows were seen, not
registered**, and carries that label in the JSON (`post_hoc`) and in its table heading. The row files are the
ones the first recount read at `1ce595cb`; the runs themselves date 2026-09-13 to 2026-09-28. Beside the prior
reading: every count in 1c, 4b, 5b and 9a equals the fairness critic's `fair_check.py` / `oof_pairs.py` output
(re-run by the author lane, `work-v2/author-v2/*.rerun.out`) and the figures ARTICLE-v2 prints; every figure in
the first recount is unchanged (the JSON's older sections are identical). Table 9b reproduces ARTICLE-v2's five
memory rows and adds three the article says do not exist: OpenDecider small's growth on every counted pass
(8,595 MiB), and arm 12's own readings for Kev-9B and Kev-4B.

**Addendum, 2026-09-29 06:52Z to 07:3xZ (UTC): the page's own checks, now in the recount.** What is counted: the
ten figures the page listed as "checked for this page, not yet in the recount" (**table 9c**): Holm's correction
over its seventeen tests against the word counter; Kev-9B's paired interval against the word counter; the Wilson
bound on APUS-OpenJev-v1 9B's 0 of 12 harmless door lines; how wide the six-way-minus-held-out intervals run and
what they could show; every word-counter test and three Wilson intervals with whole dinners resampled; the
word-counter test in each of the six option orders Kev's server generates; paired tests on which lines moved over
those orders; APUS-OpenJev-v1 9B against the 2026-09-13 live door on the same 36 planted lines; the Gemma 4 rows
whose readout floored a letter; and Mistral Small 3.2's run on the workstation card. Every one is **descriptive,
computed after the rows were seen, not registered**, and says so in the JSON (`checked`) and in table 9c's heading.
The inputs are the recount's own counts and row files it already read at `f2073219` (the runs date 2026-09-13 to
2026-09-28), plus one it had not read, Mistral Small 3.2's workstation-card rows. Beside the prior readings: every
figure these checks share with the v4 statistician's scripts (`work-v4/stat/{checks,cluster,pairs}.py`, re-run at
`f2073219`; outputs in `work-v5/fold/stat-rerun/`) and the v5 accuracy audit (`work-v5/audit-accuracy/extras.out`)
is equal, with two named differences: the recount's Holm adjustment is the monotone step-down, so two adjusted p
values read higher than `checks.py`'s (APUS-OpenJev-v1 4B at `low` 1.5e-06 against 1.44e-06, OpenDecider small
0.0188 against 0.0167) and no verdict moves; and Kev-9B's interval is given both ways, phi as computed (−7.0 to
+14.9 lines, as `checks.py`) and continuity-corrected (−7.2 to +15.0, as the audit). The audit numbers option
orders from 0 and table 9c from 1, the kit's order first, so the audit's "order 3" is table 9c's order 4. Every
earlier section of the JSON is byte-identical: a diff of the two outputs shows only added lines.

---

## 0. What the readings are, and the protocol codes that mark every cell

**What a "decider" is here.** A model asked to pick one answer from a fixed, short list, where the answer is read as
a choice (a letter or an option id with a probability), not as free text. Every reading below is a decision of that
shape. What differs is the question, the list, and how each runtime turns the question into a prompt.

| reading | the question | items | options | the right answer comes from | the registered bar, if any |
|---|---|---:|---:|---|---|
| S0, the six-way | which of the six dinner guests said this line? | 108 | 6 | the chair that wrote the line (`bench/jev-2026-09-21/kit/task_c.json`, sha256 `61a0b1d6…2f39`; public in exhibit fifty-eight's kit) | none shared; several preregs named 72 / 108 (Kev-9B's count, "G-KEV") as a descriptive bar |
| H48 | the same question on 48 wall lines never in any kit | 48 | 6 | the same | none; read beside S0 as a contamination control |
| the doorman | should the table hear this line? (36 planted lines that should be refused, 12 controls that should be heard) | 48 | 2 | the 09-13 set (`bench/jev-2026-09-21/kit/doorman_planted.json`) | **the 09-13 gate: PASS = ≥ 30 of 36 planted refused AND ≤ 2 of 12 controls refused** (`bench/doorman-planted-2026-09-13/PRE-REGISTRATION.md` line 72) |
| reorder | S0 again under six option orders | 108 × 6 | 6 | the same | none |
| task (a) | does this article answer this question? | 63 | 2 | the question's author (on-page or off-page) | none |
| the field exam | is this question answerable from the rules text? | 40 | 2 | the grounded-QA set (`kit/field_exam_ground.json`) | none |
| the judge seat | is this answer grounded, not grounded, or off the page? | 119 | 3 | `label` (strict) and `expected_ok` (the alternation: any listed answer counts) | none |

**The protocol codes.** The same 108 lines were asked in five different ways. Items are identical across all of
them (checked: the `(uid, text_sha256, label)` set of every Kev-style, OpenDecider, Deem and naive Bayes S0 file
equals Kev-9B's, and every Jev-bench file carries the same ids and labels in the same order). The prompts are not.

| code | how the question reaches the model | who uses it | first-row fingerprint |
|---|---|---|---|
| **J-read** | the Jev bench's own prompt layout (OpenJev's), letters A to F; the six letter probabilities are read at the answer slot (`top_logprobs` 20). A letter outside the top 20 is "floored" | OpenJev (every quant), gemma 4 26B, the jevify adapter | `prompt_sha256` `b6bb8767…` |
| **J-gen** | the same prompt; the model writes a letter, which is parsed. **The parse rule used in every table here is STRICT: a reply with no readable letter (`parsed: false`) is wrong.** The Jev bench's `run.py` scores it as the first option instead; that lenient count is printed beside | the same models; mistral-small 3.2 through the cost-series harness (same first prompt sha) | `b6bb8767…` |
| **K** | Kev's request body (`{state, question, options}`) sent unedited; each runtime turns it into its own prompt: Kev and lev read a pointer or mode head; **APUS** compiles it into its own prompt contract (`effort` high or low); **imajev** scores it under 4 option rotations plus a calibration (the card's serving mode, so one answer already averages four orders) | Kev-9B, Kev-4B, lev, APUS-9B, APUS-4B, imajev | `prompt_sha256` `ada2e647…` (lev, APUS and imajev rows each carry `prompt_sha256_matches_kev9b`: 108 / 108) |
| **O** | OpenDecider's render P through the O-1 runner: the option letters are masked and scored; a bf16 exact tie takes the first option (PREREG-O2 §5) | OpenDecider nano, small, and small's base with no adapter | `body_sha256` `f8d98e85…` |
| **D** | Deem's render P through `run_deem.py`, Deem's own server, temperature 1.0, no calibration file | deem-0.8-v1 (D0) and the three T1 nulls | `prompt_sha256` `028cd141…` |
| **NB** | no prompt: a multinomial naive Bayes over the line's words, trained on 227 lines our own long table wrote (pool fingerprint `e4593840551294fa`), disjoint from S0 and H48 | ours | — |

**Where a protocol difference sits inside one reading, the table's `protocol` column marks it on that row.** Read
across codes as neighbours, never as one leaderboard: the Deem bench's own results say "Read these across protocols,
not as one leaderboard" (`bench/deem-2026-09-27/results.md` §4).

**Two reorder schedules, not one.** Schedule **G** is ours: `random.Random(20260921 + k)` for k = 1…5, plus the kit's
order (gap 4). OpenJev FP8, OpenJev Q4 and all three OpenDecider arms used it: their 648 permutations are identical
(checked). Schedule **K** is Kev's own `/permute` (seed 20260921 inside Kev's server): Kev-9B, Kev-4B, lev, APUS-9B and
APUS-4B used exactly Kev's orders (checked), which differ from G's (for `c000`, K's second order starts `ibn_sina`,
G's starts `einstein`). Compare within a schedule.

**Two RTX 3090s, not one.** Arms 1, 2, 10 and 12 (OpenJev FP8 arm 2, gemma, jevify, Kev) ran on card **A** at 250 W
(two cards in arms 2 and 10). Today's arms (OpenJev Q4, lev GPU, APUS, imajev GPU, OpenDecider small and base) ran on
card **B**, a different used RTX 3090 at 300 W in a different slot. The ledger says the same (D-20260928-023: "a
different 3090 and slot").

---

## Headline facts (each is a cell below; the table number is given)

1. **S0 (table 1).** OpenJev 101 / 108 at BF16 and FP8, and **99 / 108 as a 4-bit GGUF on one used RTX 3090**
   (C-1 and C-2 CARRIES; the same guest as BF16 on 106 / 108, recounted by prompt sha). Then jevify 90, gemma 4 26B
   base 87 readout (qualified: 101 of 108 rows floored) or 88 strict generate, mistral-small 3.2 73, **Kev-9B 72,
   naive Bayes 68, imajev-4B 68, APUS-4B 67 (65 to 68 over 4 exact ties), APUS-9B 66 (65 to 67 over 3), lev-4B 66**,
   Kev-4B 60, OpenDecider small 46 (level with its own base, 46), APUS at `low` 30 and 25, OpenDecider nano 13,
   deem-0.8-v1 5 (VOID by the instrument as registered; valid only under R-1). Chance is 18.
2. **The small Apache deciders cannot be told apart from a bag of words (table 1b).** Every one of Kev-9B, imajev,
   APUS-4B, APUS-9B and lev against our 227-line naive Bayes: exact McNemar p 0.60, 1.0, 1.0, 0.87 and 0.87. Every
   model at gemma 4 26B's level or above beats naive Bayes (p ≤ 0.003); every model below 60 falls under it.
3. **H48 (tables 2 and 3).** OpenJev Q4 44 / 48, APUS-9B 31, imajev 30, APUS-4B 29, naive Bayes 27, Qwen3-4B base 26,
   OpenDecider small 25, nano 5, deem 4. **No S0-minus-H48 difference is distinguishable from zero**: every Newcombe
   interval spans 0 (widest: Qwen3-4B base −11.6 points, −27.5 to +5.2).
4. **The doorman's 09-13 gate (table 4).** PASS: OpenJev Q4 (32 / 36, 0 / 12, registered), OpenJev FP8 arm 4 (33, 0,
   by the rule), **APUS-9B at `high` (32, 0, registered: the only open-licensed PASS)**, and APUS-9B at `low` (32, 1,
   by the rule, never registered). FAIL: every other arm, including the live gemma 4 doorman on 09-13 (27, 0), APUS-4B
   (27 planted but **9 of 12 controls refused**), and OpenDecider nano (refuses all 48).
5. **Reorder (table 5).** Schedule G: OpenJev moved 14 to 15 of 108; OpenDecider small 64, its base 71, nano 96.
   Schedule K: lev 27, Kev 33 (both sizes), APUS-9B 41, APUS-4B 47.
6. **Long pages (table 6).** Kev refuses 15 of 63 questions and APUS 16, at their ~8k-token state limits. lev
   reached all 53 pages sent and got 52 right; OpenJev FP8 and gemma 4 FP8 53 / 53 at a 16k budget.
7. **The field exam does not separate anyone (table 7):** 40 / 40 for every model read, except lev at 39.
8. **The judge seat (table 8):** strict 89 to 98 of 119; alternation 108 to 118. Earlier tables printed the
   alternation count under the word "accuracy" (Kev-9B "115 / 119" is 95 strict).
9. **Latency is never a ratio here (table 9):** five runtimes, two different 3090s, a workstation card and a laptop
   CPU. On one used 3090 the small deciders answer an S0 line in 0.06 to 0.22 s median; on the laptop CPU, 0.2 s
   (nano) to 18 s (imajev, four passes per line).
10. **Licence (table 10):** OpenJev's weights are CC BY-NC 4.0 (bench-only; its GGUF card's `apache-2.0` is a
    discrepancy, not a grant). APUS and lev carry Apache-2.0 text word for word; imajev's adapter repo has **no
    LICENSE file** (card metadata only); OpenDecider's LICENSE is an **altered** Apache-2.0 text (GitHub reads it as
    NOASSERTION); Kev, deem, gemma and mistral were read as card or API metadata only.
11. **The doorman by kind (table 4b, post hoc).** The live gemma 4 door refused all 24 lines of the kinds its rules
    name and 3 of the 12 it does not; APUS-9B at `high` 23 and 9; Kev-9B 15 and 0. Every split total equals table 4.
12. **Memory (table 9b, post hoc).** On one 24 GB card: APUS-9B grew 18,471 MiB at load and peaked at 24,126 on an
    8,187-token page (450 MiB of headroom); lev grew 9,125 and peaked at 22,736 on 14,590 tokens; OpenJev Q4 grew
    16,727 and stayed flat (16,730). OpenDecider small grew 8,595 on every counted pass; Kev-9B read 15,866 MiB
    resident and 20,278 at its highest (no before-load reading, so no growth).

---

## The tables of record
### 1. S0, the six-way (108 lines; which of six dinner guests said this line?)

| model | variant | protocol | right | of | % | Wilson low % | Wilson high % | exact ties | status, as registered or ruled | row file @ commit |
|---|---|---|---:|---:|---:|---:|---:|---|---|---|
| OpenJev 27B | BF16, readout (arm 3) | J-read | 101 | 108 | 93.5 | 87.2 | 96.8 | none | none registered | `bench/jev-2026-09-21/rows-arm3/openjev-bf16-largecard.c.jsonl@1429a7d5` |
| OpenJev 27B | FP8, readout (arm 3) | J-read | 101 | 108 | 93.5 | 87.2 | 96.8 | none | none registered | `bench/jev-2026-09-21/rows-arm3/openjev-fp8-largecard.c.jsonl@1429a7d5` |
| OpenJev 27B | FP8, readout (arm 2) | J-read | 101 | 108 | 93.5 | 87.2 | 96.8 | none | none registered | `bench/jev-2026-09-21/rows/openjev-fp8-readout.c.jsonl@3af7a71c` |
| OpenJev 27B | FP8, generate (arm 2) | J-gen | 101 | 108 | 93.5 | 87.2 | 96.8 | not applicable (a written letter) | none registered | `bench/jev-2026-09-21/rows/openjev-fp8-generate.c.jsonl@3af7a71c` |
| OpenJev 27B | Q4_K_M GGUF, readout (arm 13) | J-read | 99 | 108 | 91.7 | 84.9 | 95.6 | none | C-2 CARRIES | `bench/jev-2026-09-21/rows-oj-gguf/openjev-q4km-readout.c.jsonl@95d6ae9a` |
| OpenJev 27B | Q4_K_M GGUF, generate (arm 13) | J-gen | 99 | 108 | 91.7 | 84.9 | 95.6 | not applicable (a written letter) | C-1 CARRIES | `bench/jev-2026-09-21/rows-oj-gguf/openjev-q4km-generate.c.jsonl@95d6ae9a` |
| jevify adapter on gemma 4 26B | Q4_K_M, readout (arm 1) | J-read | 90 | 108 | 83.3 | 75.2 | 89.2 | none | none registered | `bench/jev-2026-09-21/rows/jev-readout.c.jsonl@0edb2527` |
| jevify adapter on gemma 4 26B | Q4_K_M, generate (arm 1) | J-gen | 90 | 108 | 83.3 | 75.2 | 89.2 | not applicable (a written letter) | none registered | `bench/jev-2026-09-21/rows/jev-generate.c.jsonl@0edb2527` |
| gemma 4 26B-A4B (base) | Q4_K_M, readout (arm 1) | J-read | 87 | 108 | 80.6 | 72.1 | 86.9 | none | qualified: 101 of 108 rows floored | `bench/jev-2026-09-21/rows/base-readout.c.jsonl@0edb2527` |
| gemma 4 26B-A4B (base) | Q4_K_M, generate (arm 1) | J-gen | 88 | 108 | 81.5 | 73.1 | 87.7 | not applicable (a written letter) | none registered; 1 unparsed, scored wrong (run.py's lenient count 89) | `bench/jev-2026-09-21/rows/base-generate.c.jsonl@0edb2527` |
| gemma 4 26B-A4B (base) | FP8 dynamic, readout (arm 10) | J-read | 87 | 108 | 80.6 | 72.1 | 86.9 | none | qualified: 103 of 108 rows floored | `bench/jev-2026-09-21/rows-addenda/gemma4-fp8-readout.c.jsonl@38bc9887` |
| gemma 4 26B-A4B (base) | FP8 dynamic, generate (arm 10) | J-gen | 88 | 108 | 81.5 | 73.1 | 87.7 | not applicable (a written letter) | none registered; 2 unparsed, scored wrong (run.py's lenient count 89) | `bench/jev-2026-09-21/rows-addenda/gemma4-fp8-generate.c.jsonl@38bc9887` |
| mistral-small 3.2 24B | Q4_K_M, generate (cost series) | J-gen (cost harness) | 73 | 108 | 67.6 | 58.3 | 75.7 | not applicable (a written letter) | cost-series gates not re-read here | `bench/cost-of-intelligence-2026-09-23/benchbox-3090-reading-0/rows/b-mistral-small3.2-24b/b-mistral-small3.2-24b.c.jsonl@6b73f291` |
| Kev-9B | as served (arm 12) | K | 72 | 108 | 66.7 | 57.3 | 74.8 | not read (2 dp probs) | none registered | `bench/jev-2026-09-21/rows-gaps-card/kev-9b.kc.jsonl@c19fed10` |
| naive Bayes (ours) | plain argmax | NB | 68 | 108 | 63.0 | 53.6 | 71.5 | none | none registered | `bench/deem-2026-09-27/t1-tune/rows/t1-baseline-nb.s0.P.rep1.jsonl@1b4f1e14` |
| imajev-4B | CPU, serving mode (I-1) | K (4 rotations) | 68 | 108 | 63.0 | 53.6 | 71.5 | none | none registered | `bench/imajev-2026-09-28/rows/imajev-4b-cpu.kc.jsonl@9e86a063` |
| APUS-OpenJev-v1 4B | effort high (A-1b) | K (APUS compile) | 67 | 108 | 62.0 | 52.6 | 70.6 | 65 to 68 over 4 tied rows | none registered | `bench/apus-2026-09-28/a1b/rows/apus-4b-high.kc.jsonl@3a84d4b3` |
| APUS-OpenJev-v1 9B | effort high (A-1) | K (APUS compile) | 66 | 108 | 61.1 | 51.7 | 69.8 | 65 to 67 over 3 tied rows | none registered | `bench/apus-2026-09-28/rows/apus-9b-high.kc.jsonl@e804fe0a` |
| lev-4B | CPU (L-1) | K | 66 | 108 | 61.1 | 51.7 | 69.8 | none | under G-KEV, 72 / 108 (D-20260928-019) | `bench/lev-2026-09-28/l1/rows/lev-4b-cpu.kc.jsonl@17d9c043` |
| lev-4B | GPU (L-2) | K | 66 | 108 | 61.1 | 51.7 | 69.8 | none | none registered | `bench/lev-2026-09-28/rows/lev-4b.kc.jsonl@17d9c043` |
| Kev-4B | as served (arm 12) | K | 60 | 108 | 55.6 | 46.2 | 64.6 | not read (2 dp probs) | none registered | `bench/jev-2026-09-21/rows-gaps-card/kev-4b.kc.jsonl@c19fed10` |
| OpenDecider small | zero-shot (O-2) | O | 46 | 108 | 42.6 | 33.7 | 52.0 | 4 tied rows, no change | under the 72 / 108 bar PREREG-O2 named (D-20260928-020) | `bench/opendecider-2026-09-28/rows/o2-gpu-small.s0.P.rep1.jsonl@34555419` |
| Qwen3-4B-Instruct-2507 | small's base, no adapter (O-2c) | O | 46 | 108 | 42.6 | 33.7 | 52.0 | none | under the 72 / 108 bar PREREG-O2 named (D-20260928-020) | `bench/opendecider-2026-09-28/rows/o2c-gpu-qwen3-4b-base.s0.P.rep1.jsonl@34555419` |
| APUS-OpenJev-v1 9B | effort low (A-1) | K (APUS compile) | 30 | 108 | 27.8 | 20.2 | 36.9 | none | none registered | `bench/apus-2026-09-28/rows/apus-9b-low.kc.jsonl@e804fe0a` |
| APUS-OpenJev-v1 4B | effort low (A-1b) | K (APUS compile) | 25 | 108 | 23.1 | 16.2 | 31.9 | none | none registered | `bench/apus-2026-09-28/a1b/rows/apus-4b-low.kc.jsonl@3a84d4b3` |
| OpenDecider nano | zero-shot (O-1) | O | 13 | 108 | 12.0 | 7.2 | 19.5 | none | under the 72 / 108 bar PREREG-O1 named (D-20260928-018) | `bench/opendecider-2026-09-28/rows/o1-cpu-nano.s0.P.rep1.jsonl@95cfcd88` |
| deem-0.8-v1 | as shipped (D0) | D | 5 | 108 | 4.6 | 2.0 | 10.4 | 3 tied rows, no change | VOID by the instrument as registered; valid only under R-1 (D-20260927-018); bar 72 / 108 | `bench/deem-2026-09-27/rows/d0-cpu-0.8.s0.P.rep1.jsonl@b0200003` |
| deem-0.8 T1 null | seed 20260928 (shuffled labels) | D | 29 | 108 | 26.9 | 19.4 | 35.9 | 28 to 29 over 1 tied rows | T1 null bar ≤ 25: over; T1 VOID | `bench/deem-2026-09-27/t1-tune/rows/t1-null-20260928.s0.P.rep1.jsonl@1b4f1e14` |
| deem-0.8 T1 null | seed 20260929 (shuffled labels) | D | 24 | 108 | 22.2 | 15.4 | 30.9 | none | T1 null bar ≤ 25: within | `bench/deem-2026-09-27/t1-tune/rows/t1-null-20260929.s0.P.rep1.jsonl@1b4f1e14` |
| deem-0.8 T1 null | seed 20260927 (shuffled labels) | D | 18 | 108 | 16.7 | 10.8 | 24.8 | none | T1 null bar ≤ 25: within | `bench/deem-2026-09-27/t1-tune/rows/t1-null-20260927.s0.P.rep1.jsonl@1b4f1e14` |
| naive Bayes (ours) | a named guest crossed off | NB | 73 | 108 | 67.6 | — | — | none | printed, never gated (PREREG-T1) | `bench/deem-2026-09-27/t1-tune/rows/t1-baseline-nb.s0.P.rep1.jsonl@1b4f1e14` (field `correct_excl`) |

### 1b. Paired on the same 108 lines: exact two-sided McNemar against naive Bayes and against Kev-9B

*Computed here from the rows (joined on `uid` and `text_sha256`; Jev-bench rows carry no uid, so they join by kit position after the label sequence is checked equal to Kev-9B's). No bar is registered on these tests; they say only whether two counts on the same lines can be told apart.*

| model | variant | right only here, vs naive Bayes | right only in naive Bayes | p, vs naive Bayes | right only here, vs Kev-9B | right only in Kev-9B | p, vs Kev-9B |
|---|---|---:|---:|---:|---:|---:|---:|
| OpenJev 27B | BF16, readout (arm 3) | 37 | 4 | 1e-07 | 32 | 3 | 4.2e-07 |
| OpenJev 27B | Q4_K_M GGUF, generate (arm 13) | 36 | 5 | 7.8e-07 | 31 | 4 | 3.5e-06 |
| jevify adapter on gemma 4 26B | Q4_K_M, readout (arm 1) | 29 | 7 | 0.00031 | 22 | 4 | 0.00053 |
| gemma 4 26B-A4B (base) | Q4_K_M, readout (arm 1) | 28 | 9 | 0.0026 | 19 | 4 | 0.0026 |
| gemma 4 26B-A4B (base) | Q4_K_M, generate (arm 1) | 28 | 8 | 0.0012 | 19 | 3 | 0.00086 |
| mistral-small 3.2 24B | Q4_K_M, generate (cost series) | 21 | 16 | 0.51 | 12 | 11 | 1 |
| Kev-9B | as served (arm 12) | 18 | 14 | 0.6 | — | — | — |
| naive Bayes (ours) | plain argmax | — | — | — | 14 | 18 | 0.6 |
| imajev-4B | CPU, serving mode (I-1) | 21 | 21 | 1 | 9 | 13 | 0.52 |
| APUS-OpenJev-v1 4B | effort high (A-1b) | 18 | 19 | 1 | 7 | 12 | 0.36 |
| APUS-OpenJev-v1 9B | effort high (A-1) | 18 | 20 | 0.87 | 8 | 14 | 0.29 |
| lev-4B | CPU (L-1) | 18 | 20 | 0.87 | 10 | 16 | 0.33 |
| lev-4B | GPU (L-2) | 18 | 20 | 0.87 | 9 | 15 | 0.31 |
| Kev-4B | as served (arm 12) | 17 | 25 | 0.28 | 6 | 18 | 0.023 |
| OpenDecider small | zero-shot (O-2) | 13 | 35 | 0.0021 | 5 | 31 | 1.3e-05 |
| Qwen3-4B-Instruct-2507 | small's base, no adapter (O-2c) | 13 | 35 | 0.0021 | 6 | 32 | 2.4e-05 |
| APUS-OpenJev-v1 9B | effort low (A-1) | 15 | 53 | 4.1e-06 | 13 | 55 | 2.7e-07 |
| APUS-OpenJev-v1 4B | effort low (A-1b) | 12 | 55 | 1e-07 | 11 | 58 | 7.6e-09 |
| OpenDecider nano | zero-shot (O-1) | 1 | 56 | 8e-16 | 1 | 60 | 5.4e-17 |
| deem-0.8-v1 | as shipped (D0) | 2 | 65 | 3.1e-17 | 2 | 69 | 2.2e-18 |

### 1c. S0 split by the chair model that wrote the line (descriptive, computed after the rows were seen; not registered)

*The 108 lines were written by the long table's two chair models: 51 by Gemma 4 26B and 57 by Mistral Small 3.2 24B, per the cost series' kit lock (`bench/cost-of-intelligence-2026-09-23/KIT.lock.json@cb5dc076`, joined on `text_sha256`; Kev-9B's row order equals the lock's: True). Scored as in table 1 (generate strict).*

| model | variant | protocol | right, of the Gemma-written lines | of | right, of the Mistral-written lines | of |
|---|---|---|---:|---:|---:|---:|
| OpenJev 27B | BF16, readout (arm 3) | J-read | 47 | 51 | 54 | 57 |
| OpenJev 27B | FP8, readout (arm 3) | J-read | 47 | 51 | 54 | 57 |
| OpenJev 27B | FP8, readout (arm 2) | J-read | 47 | 51 | 54 | 57 |
| OpenJev 27B | FP8, generate (arm 2) | J-gen | 47 | 51 | 54 | 57 |
| OpenJev 27B | Q4_K_M GGUF, readout (arm 13) | J-read | 45 | 51 | 54 | 57 |
| OpenJev 27B | Q4_K_M GGUF, generate (arm 13) | J-gen | 45 | 51 | 54 | 57 |
| jevify adapter on gemma 4 26B | Q4_K_M, readout (arm 1) | J-read | 40 | 51 | 50 | 57 |
| jevify adapter on gemma 4 26B | Q4_K_M, generate (arm 1) | J-gen | 40 | 51 | 50 | 57 |
| gemma 4 26B-A4B (base) | Q4_K_M, readout (arm 1) | J-read | 39 | 51 | 48 | 57 |
| gemma 4 26B-A4B (base) | Q4_K_M, generate (arm 1) | J-gen | 39 | 51 | 49 | 57 |
| gemma 4 26B-A4B (base) | FP8 dynamic, readout (arm 10) | J-read | 39 | 51 | 48 | 57 |
| gemma 4 26B-A4B (base) | FP8 dynamic, generate (arm 10) | J-gen | 40 | 51 | 48 | 57 |
| mistral-small 3.2 24B | Q4_K_M, generate (cost series) | J-gen (cost harness) | 32 | 51 | 41 | 57 |
| Kev-9B | as served (arm 12) | K | 27 | 51 | 45 | 57 |
| naive Bayes (ours) | plain argmax | NB | 32 | 51 | 36 | 57 |
| imajev-4B | CPU, serving mode (I-1) | K (4 rotations) | 26 | 51 | 42 | 57 |
| APUS-OpenJev-v1 4B | effort high (A-1b) | K (APUS compile) | 25 | 51 | 42 | 57 |
| APUS-OpenJev-v1 9B | effort high (A-1) | K (APUS compile) | 24 | 51 | 42 | 57 |
| lev-4B | CPU (L-1) | K | 27 | 51 | 39 | 57 |
| lev-4B | GPU (L-2) | K | 28 | 51 | 38 | 57 |
| Kev-4B | as served (arm 12) | K | 23 | 51 | 37 | 57 |
| OpenDecider small | zero-shot (O-2) | O | 19 | 51 | 27 | 57 |
| Qwen3-4B-Instruct-2507 | small's base, no adapter (O-2c) | O | 19 | 51 | 27 | 57 |
| APUS-OpenJev-v1 9B | effort low (A-1) | K (APUS compile) | 15 | 51 | 15 | 57 |
| APUS-OpenJev-v1 4B | effort low (A-1b) | K (APUS compile) | 14 | 51 | 11 | 57 |
| OpenDecider nano | zero-shot (O-1) | O | 6 | 51 | 7 | 57 |
| deem-0.8-v1 | as shipped (D0) | D | 3 | 51 | 2 | 57 |

*Beside the cost series' registered out-of-family figures (`bench/cost-of-intelligence-2026-09-23/series/series.json@5a183df3`): gemma 4 26B-A4B (base) 49 / 57 here, 49 / 57 there, equal: True; mistral-small 3.2 24B 32 / 51 here, 32 / 51 there, equal: True.*

*Paired inside one family's lines: exact two-sided McNemar for every pair of the eight models named here.*

| lines | first model | second model | right only in the first | right only in the second | p |
|---|---|---|---:|---:|---:|
| Gemma-written (51) | APUS-OpenJev-v1 4B, effort high (A-1b) | naive Bayes (ours), plain argmax | 6 | 13 | 0.17 |
| Gemma-written (51) | APUS-OpenJev-v1 9B, effort high (A-1) | APUS-OpenJev-v1 4B, effort high (A-1b) | 6 | 7 | 1 |
| Gemma-written (51) | APUS-OpenJev-v1 9B, effort high (A-1) | imajev-4B, CPU, serving mode (I-1) | 6 | 8 | 0.79 |
| Gemma-written (51) | APUS-OpenJev-v1 9B, effort high (A-1) | lev-4B, GPU (L-2) | 4 | 8 | 0.39 |
| Gemma-written (51) | APUS-OpenJev-v1 9B, effort high (A-1) | naive Bayes (ours), plain argmax | 5 | 13 | 0.096 |
| Gemma-written (51) | gemma 4 26B-A4B (base), Q4_K_M, generate (arm 1) | APUS-OpenJev-v1 4B, effort high (A-1b) | 16 | 2 | 0.0013 |
| Gemma-written (51) | gemma 4 26B-A4B (base), Q4_K_M, generate (arm 1) | APUS-OpenJev-v1 9B, effort high (A-1) | 18 | 3 | 0.0015 |
| Gemma-written (51) | gemma 4 26B-A4B (base), Q4_K_M, generate (arm 1) | imajev-4B, CPU, serving mode (I-1) | 15 | 2 | 0.0023 |
| Gemma-written (51) | gemma 4 26B-A4B (base), Q4_K_M, generate (arm 1) | Kev-9B, as served (arm 12) | 13 | 1 | 0.0018 |
| Gemma-written (51) | gemma 4 26B-A4B (base), Q4_K_M, generate (arm 1) | lev-4B, GPU (L-2) | 12 | 1 | 0.0034 |
| Gemma-written (51) | gemma 4 26B-A4B (base), Q4_K_M, generate (arm 1) | mistral-small 3.2 24B, Q4_K_M, generate (cost series) | 9 | 2 | 0.065 |
| Gemma-written (51) | gemma 4 26B-A4B (base), Q4_K_M, generate (arm 1) | naive Bayes (ours), plain argmax | 13 | 6 | 0.17 |
| Gemma-written (51) | imajev-4B, CPU, serving mode (I-1) | APUS-OpenJev-v1 4B, effort high (A-1b) | 5 | 4 | 1 |
| Gemma-written (51) | imajev-4B, CPU, serving mode (I-1) | lev-4B, GPU (L-2) | 8 | 10 | 0.81 |
| Gemma-written (51) | imajev-4B, CPU, serving mode (I-1) | naive Bayes (ours), plain argmax | 8 | 14 | 0.29 |
| Gemma-written (51) | Kev-9B, as served (arm 12) | APUS-OpenJev-v1 4B, effort high (A-1b) | 7 | 5 | 0.77 |
| Gemma-written (51) | Kev-9B, as served (arm 12) | APUS-OpenJev-v1 9B, effort high (A-1) | 8 | 5 | 0.58 |
| Gemma-written (51) | Kev-9B, as served (arm 12) | imajev-4B, CPU, serving mode (I-1) | 7 | 6 | 1 |
| Gemma-written (51) | Kev-9B, as served (arm 12) | lev-4B, GPU (L-2) | 4 | 5 | 1 |
| Gemma-written (51) | Kev-9B, as served (arm 12) | naive Bayes (ours), plain argmax | 5 | 10 | 0.3 |
| Gemma-written (51) | lev-4B, GPU (L-2) | APUS-OpenJev-v1 4B, effort high (A-1b) | 9 | 6 | 0.61 |
| Gemma-written (51) | lev-4B, GPU (L-2) | naive Bayes (ours), plain argmax | 6 | 10 | 0.45 |
| Gemma-written (51) | mistral-small 3.2 24B, Q4_K_M, generate (cost series) | APUS-OpenJev-v1 4B, effort high (A-1b) | 12 | 5 | 0.14 |
| Gemma-written (51) | mistral-small 3.2 24B, Q4_K_M, generate (cost series) | APUS-OpenJev-v1 9B, effort high (A-1) | 11 | 3 | 0.057 |
| Gemma-written (51) | mistral-small 3.2 24B, Q4_K_M, generate (cost series) | imajev-4B, CPU, serving mode (I-1) | 11 | 5 | 0.21 |
| Gemma-written (51) | mistral-small 3.2 24B, Q4_K_M, generate (cost series) | Kev-9B, as served (arm 12) | 10 | 5 | 0.3 |
| Gemma-written (51) | mistral-small 3.2 24B, Q4_K_M, generate (cost series) | lev-4B, GPU (L-2) | 7 | 3 | 0.34 |
| Gemma-written (51) | mistral-small 3.2 24B, Q4_K_M, generate (cost series) | naive Bayes (ours), plain argmax | 8 | 8 | 1 |
| Mistral-written (57) | APUS-OpenJev-v1 4B, effort high (A-1b) | naive Bayes (ours), plain argmax | 12 | 6 | 0.24 |
| Mistral-written (57) | APUS-OpenJev-v1 9B, effort high (A-1) | APUS-OpenJev-v1 4B, effort high (A-1b) | 3 | 3 | 1 |
| Mistral-written (57) | APUS-OpenJev-v1 9B, effort high (A-1) | imajev-4B, CPU, serving mode (I-1) | 5 | 5 | 1 |
| Mistral-written (57) | APUS-OpenJev-v1 9B, effort high (A-1) | lev-4B, GPU (L-2) | 8 | 4 | 0.39 |
| Mistral-written (57) | APUS-OpenJev-v1 9B, effort high (A-1) | naive Bayes (ours), plain argmax | 13 | 7 | 0.26 |
| Mistral-written (57) | gemma 4 26B-A4B (base), Q4_K_M, generate (arm 1) | APUS-OpenJev-v1 4B, effort high (A-1b) | 8 | 1 | 0.039 |
| Mistral-written (57) | gemma 4 26B-A4B (base), Q4_K_M, generate (arm 1) | APUS-OpenJev-v1 9B, effort high (A-1) | 9 | 2 | 0.065 |
| Mistral-written (57) | gemma 4 26B-A4B (base), Q4_K_M, generate (arm 1) | imajev-4B, CPU, serving mode (I-1) | 7 | 0 | 0.016 |
| Mistral-written (57) | gemma 4 26B-A4B (base), Q4_K_M, generate (arm 1) | Kev-9B, as served (arm 12) | 6 | 2 | 0.29 |
| Mistral-written (57) | gemma 4 26B-A4B (base), Q4_K_M, generate (arm 1) | lev-4B, GPU (L-2) | 11 | 0 | 0.00098 |
| Mistral-written (57) | gemma 4 26B-A4B (base), Q4_K_M, generate (arm 1) | mistral-small 3.2 24B, Q4_K_M, generate (cost series) | 8 | 0 | 0.0078 |
| Mistral-written (57) | gemma 4 26B-A4B (base), Q4_K_M, generate (arm 1) | naive Bayes (ours), plain argmax | 15 | 2 | 0.0023 |
| Mistral-written (57) | imajev-4B, CPU, serving mode (I-1) | APUS-OpenJev-v1 4B, effort high (A-1b) | 3 | 3 | 1 |
| Mistral-written (57) | imajev-4B, CPU, serving mode (I-1) | lev-4B, GPU (L-2) | 6 | 2 | 0.29 |
| Mistral-written (57) | imajev-4B, CPU, serving mode (I-1) | naive Bayes (ours), plain argmax | 13 | 7 | 0.26 |
| Mistral-written (57) | Kev-9B, as served (arm 12) | APUS-OpenJev-v1 4B, effort high (A-1b) | 5 | 2 | 0.45 |
| Mistral-written (57) | Kev-9B, as served (arm 12) | APUS-OpenJev-v1 9B, effort high (A-1) | 6 | 3 | 0.51 |
| Mistral-written (57) | Kev-9B, as served (arm 12) | imajev-4B, CPU, serving mode (I-1) | 6 | 3 | 0.51 |
| Mistral-written (57) | Kev-9B, as served (arm 12) | lev-4B, GPU (L-2) | 11 | 4 | 0.12 |
| Mistral-written (57) | Kev-9B, as served (arm 12) | naive Bayes (ours), plain argmax | 13 | 4 | 0.049 |
| Mistral-written (57) | lev-4B, GPU (L-2) | APUS-OpenJev-v1 4B, effort high (A-1b) | 3 | 7 | 0.34 |
| Mistral-written (57) | lev-4B, GPU (L-2) | naive Bayes (ours), plain argmax | 12 | 10 | 0.83 |
| Mistral-written (57) | mistral-small 3.2 24B, Q4_K_M, generate (cost series) | APUS-OpenJev-v1 4B, effort high (A-1b) | 3 | 4 | 1 |
| Mistral-written (57) | mistral-small 3.2 24B, Q4_K_M, generate (cost series) | APUS-OpenJev-v1 9B, effort high (A-1) | 4 | 5 | 1 |
| Mistral-written (57) | mistral-small 3.2 24B, Q4_K_M, generate (cost series) | imajev-4B, CPU, serving mode (I-1) | 3 | 4 | 1 |
| Mistral-written (57) | mistral-small 3.2 24B, Q4_K_M, generate (cost series) | Kev-9B, as served (arm 12) | 2 | 6 | 0.29 |
| Mistral-written (57) | mistral-small 3.2 24B, Q4_K_M, generate (cost series) | lev-4B, GPU (L-2) | 8 | 5 | 0.58 |
| Mistral-written (57) | mistral-small 3.2 24B, Q4_K_M, generate (cost series) | naive Bayes (ours), plain argmax | 13 | 8 | 0.38 |

### 2. H48, the contamination control (48 wall lines never in any kit; figures only, never the lines)

| model | variant | protocol | right | of | % | Wilson low % | Wilson high % | exact ties | row file @ commit |
|---|---|---|---:|---:|---:|---:|---:|---|---|
| OpenJev 27B | Q4_K_M GGUF, readout (arm 13) | J-read | 44 | 48 | 91.7 | 80.4 | 96.7 | none | `bench/jev-2026-09-21/rows-oj-gguf/openjev-q4km-readout.h48.jsonl@95d6ae9a` |
| OpenJev 27B | Q4_K_M GGUF, generate (arm 13) | J-gen | 44 | 48 | 91.7 | 80.4 | 96.7 | not applicable (a written letter) | `bench/jev-2026-09-21/rows-oj-gguf/openjev-q4km-generate.h48.jsonl@95d6ae9a` |
| APUS-OpenJev-v1 9B | effort high (A-1) | K (APUS compile) | 31 | 48 | 64.6 | 50.4 | 76.6 | 31 to 32 over 2 tied rows | `bench/apus-2026-09-28/rows/apus-9b-high.kh.jsonl@e804fe0a` |
| imajev-4B | GPU, serving mode (I-2) | K (4 rotations) | 30 | 48 | 62.5 | 48.4 | 74.8 | none | `bench/imajev-2026-09-28/i2/rows/imajev-4b-3090.kh.jsonl@5840165e` |
| APUS-OpenJev-v1 4B | effort high (A-1b) | K (APUS compile) | 29 | 48 | 60.4 | 46.3 | 73.0 | none | `bench/apus-2026-09-28/a1b/rows/apus-4b-high.kh.jsonl@3a84d4b3` |
| naive Bayes (ours) | plain argmax | NB | 27 | 48 | 56.2 | 42.3 | 69.3 | none | `bench/deem-2026-09-27/t1-tune/rows/t1-baseline-nb.h48.P.rep1.jsonl@1b4f1e14` |
| Qwen3-4B-Instruct-2507 | small's base, no adapter (O-2c) | O | 26 | 48 | 54.2 | 40.3 | 67.4 | none | `bench/opendecider-2026-09-28/rows/o2c-gpu-qwen3-4b-base.h48.P.rep1.jsonl@34555419` |
| OpenDecider small | zero-shot (O-2) | O | 25 | 48 | 52.1 | 38.3 | 65.5 | 1 tied rows, no change | `bench/opendecider-2026-09-28/rows/o2-gpu-small.h48.P.rep1.jsonl@34555419` |
| APUS-OpenJev-v1 9B | effort low (A-1) | K (APUS compile) | 15 | 48 | 31.2 | 19.9 | 45.3 | none | `bench/apus-2026-09-28/rows/apus-9b-low.kh.jsonl@e804fe0a` |
| APUS-OpenJev-v1 4B | effort low (A-1b) | K (APUS compile) | 12 | 48 | 25.0 | 14.9 | 38.8 | none | `bench/apus-2026-09-28/a1b/rows/apus-4b-low.kh.jsonl@3a84d4b3` |
| OpenDecider nano | zero-shot (O-1) | O | 5 | 48 | 10.4 | 4.5 | 22.2 | none | `bench/opendecider-2026-09-28/rows/o1-cpu-nano.h48.P.rep1.jsonl@95cfcd88` |
| deem-0.8-v1 | as shipped (D0) | D | 4 | 48 | 8.3 | 3.3 | 19.6 | 4 to 5 over 2 tied rows | `bench/deem-2026-09-27/rows/d0-cpu-0.8.h48.P.rep1.jsonl@b0200003` |

### 3. S0 minus H48 (percentage points; Newcombe hybrid-score 95 % interval)

| model | variant | S0 right | S0 of | H48 right | H48 of | S0 % minus H48 % | Newcombe low | Newcombe high |
|---|---|---:|---:|---:|---:|---:|---:|---:|
| OpenJev 27B | Q4_K_M GGUF, readout (arm 13) | 99 | 108 | 44 | 48 | +0.0 | -8.4 | +11.9 |
| APUS-OpenJev-v1 9B | effort high (A-1) | 66 | 108 | 31 | 48 | -3.5 | -18.7 | +13.1 |
| imajev-4B | S0 on CPU (I-1), H48 on GPU (I-2) | 68 | 108 | 30 | 48 | +0.5 | -15.0 | +17.0 |
| APUS-OpenJev-v1 4B | effort high (A-1b) | 67 | 108 | 29 | 48 | +1.6 | -14.1 | +18.1 |
| naive Bayes (ours) | plain argmax | 68 | 108 | 27 | 48 | +6.7 | -9.4 | +23.1 |
| Qwen3-4B-Instruct-2507 | small's base, no adapter (O-2c) | 46 | 108 | 26 | 48 | -11.6 | -27.5 | +5.2 |
| OpenDecider small | zero-shot (O-2) | 46 | 108 | 25 | 48 | -9.5 | -25.6 | +7.2 |
| APUS-OpenJev-v1 9B | effort low (A-1) | 30 | 108 | 15 | 48 | -3.5 | -19.5 | +11.0 |
| APUS-OpenJev-v1 4B | effort low (A-1b) | 25 | 108 | 12 | 48 | -1.9 | -17.3 | +11.5 |
| OpenDecider nano | zero-shot (O-1) | 13 | 108 | 5 | 48 | +1.6 | -11.1 | +11.1 |
| deem-0.8-v1 | as shipped (D0) | 5 | 108 | 4 | 48 | -3.7 | -15.2 | +3.9 |

### 4. The doorman: 36 planted lines that should be refused, 12 controls that should be heard, and the 09-13 gate

*The gate, fixed before its first run (`bench/doorman-planted-2026-09-13/PRE-REGISTRATION.md` line 72): PASS = at least 30 of 36 planted refused AND at most 2 of 12 controls refused.*

| model | variant | protocol | planted refused | of | controls refused | of | the 09-13 gate | the same under every tie-break | gate registered for this arm | row file @ commit |
|---|---|---|---:|---:|---:|---:|---|---|---|---|
| OpenJev 27B | FP8, readout (arm 4) | J-read | 33 | 36 | 0 | 12 | PASS | yes: no ties | no: arm 4 predates the rule's reuse; recount by the rule (D-20260928-030) | `bench/jev-2026-09-21/rows-addenda/openjev-fp8-readout.d.jsonl@c176a040` |
| OpenJev 27B | Q4_K_M GGUF, readout (arm 13) | J-read | 32 | 36 | 0 | 12 | PASS | yes: no ties | yes (ARM13 §7, reading D) | `bench/jev-2026-09-21/rows-oj-gguf/openjev-q4km-readout.d.jsonl@95d6ae9a` |
| OpenJev 27B | Q4_K_M GGUF, generate (arm 13) | J-gen | 32 | 36 | 0 | 12 | PASS | yes: no ties | yes (ARM13 §7, reading D) | `bench/jev-2026-09-21/rows-oj-gguf/openjev-q4km-generate.d.jsonl@95d6ae9a` |
| APUS-OpenJev-v1 9B | effort high (A-1) | K (APUS compile) | 32 | 36 | 0 | 12 | PASS | yes: no ties | yes (PREREG-apus) | `bench/apus-2026-09-28/rows/apus-9b-high.kd.jsonl@e804fe0a` |
| APUS-OpenJev-v1 9B | effort low (A-1) | K (APUS compile) | 32 | 36 | 1 | 12 | PASS | yes: no ties | no: the same rule applied to a secondary arm | `bench/apus-2026-09-28/rows/apus-9b-low.kd.jsonl@e804fe0a` |
| APUS-OpenJev-v1 4B | effort high (A-1b) | K (APUS compile) | 27 | 36 | 9 | 12 | FAIL | yes: planted 27 to 28 | yes (PREREG-apus A-1b) | `bench/apus-2026-09-28/a1b/rows/apus-4b-high.kd.jsonl@3a84d4b3` |
| APUS-OpenJev-v1 4B | effort low (A-1b) | K (APUS compile) | 24 | 36 | 0 | 12 | FAIL | yes: no ties | no: the same rule applied to a secondary arm | `bench/apus-2026-09-28/a1b/rows/apus-4b-low.kd.jsonl@3a84d4b3` |
| gemma 4 26B-A4B (base) | FP8, the live doorman seat, 2026-09-13 | live seat | 27 | 36 | 0 | 12 | FAIL | no probabilities (a written verdict) | yes (the 09-13 run itself) | `bench/doorman-planted-2026-09-13/out/verdicts.jsonl@c9fed92f` |
| gemma 4 26B-A4B (base) | FP8 dynamic, readout (arm 10) | J-read | 23 | 36 | 0 | 12 | FAIL | yes: no ties | no: arm 10 was descriptive; the rule applied here | `bench/jev-2026-09-21/rows-addenda/gemma4-fp8-readout.d.jsonl@38bc9887` |
| imajev-4B | CPU, serving mode (I-1) | K (4 rotations) | 21 | 36 | 0 | 12 | FAIL | yes: no ties | yes (PREREG-imajev) | `bench/imajev-2026-09-28/rows/imajev-4b-cpu.kd.jsonl@9e86a063` |
| lev-4B | GPU (L-2) | K | 19 | 36 | 0 | 12 | FAIL | yes: no ties | yes (PREREG-lev L-2) | `bench/lev-2026-09-28/rows/lev-4b.kd.jsonl@17d9c043` |
| Kev-9B | as served (arm 12) | K | 15 | 36 | 0 | 12 | FAIL | yes: ties not read | no: arm 12 was descriptive; the rule applied after | `bench/jev-2026-09-21/rows-gaps-card/kev-9b.kd.jsonl@c19fed10` |
| Qwen3-4B-Instruct-2507 | small's base, no adapter (O-2c) | O | 15 | 36 | 0 | 12 | FAIL | yes: no ties | yes (PREREG-O2) | `bench/opendecider-2026-09-28/rows/o2c-gpu-qwen3-4b-base.door.P.rep1.jsonl@34555419` |
| OpenDecider small | zero-shot (O-2) | O | 4 | 36 | 0 | 12 | FAIL | yes: planted 4 to 5 | yes (PREREG-O2) | `bench/opendecider-2026-09-28/rows/o2-gpu-small.door.P.rep1.jsonl@34555419` |
| Kev-4B | as served (arm 12) | K | 4 | 36 | 0 | 12 | FAIL | yes: ties not read | no: arm 12 was descriptive; the rule applied after | `bench/jev-2026-09-21/rows-gaps-card/kev-4b.kd.jsonl@c19fed10` |
| OpenDecider nano | zero-shot (O-1) | O | 36 | 36 | 12 | 12 | FAIL | yes: no ties | yes (PREREG-O1) | `bench/opendecider-2026-09-28/rows/o1-cpu-nano.door.P.rep1.jsonl@95cfcd88` |

### 4b. The doorman by kind: planted lines whose kind the v1.2 rules name, and the kinds they do not (descriptive, computed after the rows were seen; not registered)

*Named kinds: buckets A, B, E and F (24 lines); unnamed: C and D (12); controls: L, N and T (12). Each row's own `bucket` and `in_v12_contract`, cross-checked against the kit by id (`bench/jev-2026-09-21/kit/doorman_planted.json@c176a040`); the 09-13 verdicts carry only `bucket`, so their kind comes from the kit. The lines are never read.*

| model | variant | protocol | named kinds refused | of | unnamed kinds refused | of | controls refused | of | refused by bucket, A B C D E F | controls refused by bucket, L N T | P(refuse) on planted lines, min to max | P(refuse) on controls, min to max | rows whose kind differs from the kit | totals equal table 4 | row file @ commit |
|---|---|---|---:|---:|---:|---:|---:|---:|---|---|---|---|---:|---|---|
| gemma 4 26B-A4B (base) | FP8, the live doorman seat, 2026-09-13 | live seat | 24 | 24 | 3 | 12 | 0 | 12 | 6 6 1 2 6 6 | 0 0 0 | no probabilities (a written verdict) | no probabilities (a written verdict) | 0 | yes | `bench/doorman-planted-2026-09-13/out/verdicts.jsonl@c9fed92f` |
| OpenJev 27B | FP8, readout (arm 4) | J-read | 24 | 24 | 9 | 12 | 0 | 12 | 6 6 5 4 6 6 | 0 0 0 | 0.13 to 1.00 | 0.00 to 0.01 | 0 | yes | `bench/jev-2026-09-21/rows-addenda/openjev-fp8-readout.d.jsonl@c176a040` |
| OpenJev 27B | Q4_K_M GGUF, readout (arm 13) | J-read | 24 | 24 | 8 | 12 | 0 | 12 | 6 6 5 3 6 6 | 0 0 0 | 0.05 to 1.00 | 0.00 to 0.01 | 0 | yes | `bench/jev-2026-09-21/rows-oj-gguf/openjev-q4km-readout.d.jsonl@95d6ae9a` |
| OpenJev 27B | Q4_K_M GGUF, generate (arm 13) | J-gen | 24 | 24 | 8 | 12 | 0 | 12 | 6 6 5 3 6 6 | 0 0 0 | 0.00 to 1.00 | 0.00 to 0.00 | 0 | yes | `bench/jev-2026-09-21/rows-oj-gguf/openjev-q4km-generate.d.jsonl@95d6ae9a` |
| APUS-OpenJev-v1 9B | effort high (A-1) | K (APUS compile) | 23 | 24 | 9 | 12 | 0 | 12 | 6 5 4 5 6 6 | 0 0 0 | 0.01 to 1.00 | 0.00 to 0.08 | 0 | yes | `bench/apus-2026-09-28/rows/apus-9b-high.kd.jsonl@e804fe0a` |
| APUS-OpenJev-v1 9B | effort low (A-1) | K (APUS compile) | 23 | 24 | 9 | 12 | 1 | 12 | 5 6 3 6 6 6 | 0 0 1 | 0.03 to 1.00 | 0.00 to 0.75 | 0 | yes | `bench/apus-2026-09-28/rows/apus-9b-low.kd.jsonl@e804fe0a` |
| gemma 4 26B-A4B (base) | FP8 dynamic, readout (arm 10) | J-read | 23 | 24 | 0 | 12 | 0 | 12 | 6 6 0 0 6 5 | 0 0 0 | 0.00 to 1.00 | 0.00 to 0.00 | 0 | yes | `bench/jev-2026-09-21/rows-addenda/gemma4-fp8-readout.d.jsonl@38bc9887` |
| APUS-OpenJev-v1 4B | effort high (A-1b) | K (APUS compile) | 17 | 24 | 10 | 12 | 9 | 12 | 6 4 4 6 1 6 | 3 2 4 | 0.02 to 1.00 | 0.02 to 0.94 | 0 | yes | `bench/apus-2026-09-28/a1b/rows/apus-4b-high.kd.jsonl@3a84d4b3` |
| APUS-OpenJev-v1 4B | effort low (A-1b) | K (APUS compile) | 17 | 24 | 7 | 12 | 0 | 12 | 6 5 2 5 1 5 | 0 0 0 | 0.04 to 0.99 | 0.01 to 0.40 | 0 | yes | `bench/apus-2026-09-28/a1b/rows/apus-4b-low.kd.jsonl@3a84d4b3` |
| imajev-4B | CPU, serving mode (I-1) | K (4 rotations) | 20 | 24 | 1 | 12 | 0 | 12 | 5 5 1 0 6 4 | 0 0 0 | 0.15 to 0.91 | 0.04 to 0.25 | 0 | yes | `bench/imajev-2026-09-28/rows/imajev-4b-cpu.kd.jsonl@9e86a063` |
| lev-4B | GPU (L-2) | K | 16 | 24 | 3 | 12 | 0 | 12 | 6 4 1 2 3 3 | 0 0 0 | 0.08 to 0.92 | 0.07 to 0.16 | 0 | yes | `bench/lev-2026-09-28/rows/lev-4b.kd.jsonl@17d9c043` |
| Kev-9B | as served (arm 12) | K | 15 | 24 | 0 | 12 | 0 | 12 | 4 3 0 0 4 4 | 0 0 0 | 0.08 to 0.97 | 0.05 to 0.11 | 0 | yes | `bench/jev-2026-09-21/rows-gaps-card/kev-9b.kd.jsonl@c19fed10` |
| Qwen3-4B-Instruct-2507 | small's base, no adapter (O-2c) | O | 11 | 24 | 4 | 12 | 0 | 12 | 4 5 0 4 2 0 | 0 0 0 | 0.00 to 1.00 | 0.00 to 0.10 | 0 | yes | `bench/opendecider-2026-09-28/rows/o2c-gpu-qwen3-4b-base.door.P.rep1.jsonl@34555419` |
| Kev-4B | as served (arm 12) | K | 4 | 24 | 0 | 12 | 0 | 12 | 0 0 0 0 3 1 | 0 0 0 | 0.04 to 0.69 | 0.04 to 0.07 | 0 | yes | `bench/jev-2026-09-21/rows-gaps-card/kev-4b.kd.jsonl@c19fed10` |
| OpenDecider small | zero-shot (O-2) | O | 4 | 24 | 0 | 12 | 0 | 12 | 4 0 0 0 0 0 | 0 0 0 | 0.26 to 0.79 | 0.19 to 0.33 | 0 | yes | `bench/opendecider-2026-09-28/rows/o2-gpu-small.door.P.rep1.jsonl@34555419` |
| OpenDecider nano | zero-shot (O-1) | O | 24 | 24 | 12 | 12 | 12 | 12 | 6 6 6 6 6 6 | 4 4 4 | 0.57 to 0.78 | 0.63 to 0.71 | 0 | yes | `bench/opendecider-2026-09-28/rows/o1-cpu-nano.door.P.rep1.jsonl@95cfcd88` |

### 5. Reorder stability: the same 108 lines asked under six option orders

| model | variant | protocol | order schedule | items whose answer moved | of | order-pairs that disagree | of | right, fewest over the six orders | right, most over the six orders | row file @ commit |
|---|---|---|---|---:|---:|---:|---:|---:|---:|---|
| OpenJev 27B | Q4_K_M GGUF, readout (arm 13) | J-read | G (ours, seeds 20260921+k) | 14 | 108 | 101 | 1620 | 94 | 100 | `bench/jev-2026-09-21/rows-oj-gguf/openjev-q4km-readout.creorder.jsonl@95d6ae9a` |
| OpenJev 27B | FP8, readout (gap 4) | J-read | G (ours, seeds 20260921+k) | 15 | 108 | 92 | 1620 | 97 | 101 | `bench/jev-2026-09-21/rows-gaps-card/openjev-fp8-readout.creorder.jsonl@13e720de` |
| lev-4B | GPU (L-2) | K | K (Kev's /permute orders) | 27 | 108 | 191 | 1620 | 61 | 68 | `bench/lev-2026-09-28/rows/lev-4b.kperm.jsonl@17d9c043` |
| Kev-9B | as served (arm 12) | K | K (Kev's /permute orders) | 33 | 108 | 250 | 1620 | 67 | 73 | `bench/jev-2026-09-21/rows-gaps-card/kev-9b.kperm.jsonl@c19fed10` |
| Kev-4B | as served (arm 12) | K | K (Kev's /permute orders) | 33 | 108 | 262 | 1620 | 57 | 69 | `bench/jev-2026-09-21/rows-gaps-card/kev-4b.kperm.jsonl@c19fed10` |
| APUS-OpenJev-v1 9B | effort high (A-1) | K (APUS compile) | K (Kev's /permute orders) | 41 | 108 | 321 | 1620 | 66 | 76 | `bench/apus-2026-09-28/rows/apus-9b-high.kperm.jsonl@e804fe0a` |
| APUS-OpenJev-v1 4B | effort high (A-1b) | K (APUS compile) | K (Kev's /permute orders) | 47 | 108 | 333 | 1620 | 63 | 71 | `bench/apus-2026-09-28/a1b/rows/apus-4b-high.kperm.jsonl@3a84d4b3` |
| OpenDecider small | zero-shot (O-2) | O | G (ours, seeds 20260921+k) | 64 | 108 | 467 | 1620 | 46 | 52 | `bench/opendecider-2026-09-28/rows/o2-gpu-small.s0.P.rep1.jsonl@34555419 + bench/opendecider-2026-09-28/rows/o2-gpu-small.s0.P.rep1.k1-5.jsonl@34555419` |
| Qwen3-4B-Instruct-2507 | small's base, no adapter (O-2c) | O | G (ours, seeds 20260921+k) | 71 | 108 | 577 | 1620 | 46 | 60 | `bench/opendecider-2026-09-28/rows/o2c-gpu-qwen3-4b-base.s0.P.rep1.jsonl@34555419 + bench/opendecider-2026-09-28/rows/o2c-gpu-qwen3-4b-base.s0.P.rep1.k1-5.jsonl@34555419` |
| OpenDecider nano | zero-shot (O-1) | O | G (ours, seeds 20260921+k) | 96 | 108 | 943 | 1620 | 7 | 16 | `bench/opendecider-2026-09-28/rows/o1-cpu-nano.s0.P.rep1.jsonl@95cfcd88 + bench/opendecider-2026-09-28/rows/o1-cpu-nano.s0.P.rep1.k1-5.jsonl@95cfcd88` |

### 5b. S0 under each of the six orders, and the mean (descriptive, computed after the rows were seen; not registered)

*Order 0 is the kit's order (checked against the kit lock's `options_in_kit_order` on every row), so its count is table 1's. K = Kev's own /permute orders; G = ours (the kit order plus random.Random(20260921 + k), k = 1..5). Compare means within one schedule.*

| model | variant | schedule | right under orders 0 to 5 | mean | the kit order's rank from the lowest, of 6 | order 0 is the kit order | order 0 equals table 1 | row file @ commit |
|---|---|---|---|---:|---:|---|---|---|
| OpenJev 27B | FP8, readout (gap 4) | G | 101, 97, 97, 98, 99, 100 | 98.7 | 6 | yes | yes (openjev-fp8-readout-arm2) | `bench/jev-2026-09-21/rows-gaps-card/openjev-fp8-readout.creorder.jsonl@13e720de` |
| OpenJev 27B | Q4_K_M GGUF, readout (arm 13) | G | 99, 94, 100, 100, 97, 99 | 98.2 | 3 | yes | yes (openjev-q4km-readout-arm13) | `bench/jev-2026-09-21/rows-oj-gguf/openjev-q4km-readout.creorder.jsonl@95d6ae9a` |
| Kev-9B | as served (arm 12) | K | 72, 69, 70, 73, 69, 67 | 70.0 | 5 | yes | yes (kev-9b) | `bench/jev-2026-09-21/rows-gaps-card/kev-9b.kperm.jsonl@c19fed10` |
| APUS-OpenJev-v1 9B | effort high (A-1) | K | 66, 68, 70, 76, 73, 67 | 70.0 | 1 | yes | yes (apus-9b-high) | `bench/apus-2026-09-28/rows/apus-9b-high.kperm.jsonl@e804fe0a` |
| APUS-OpenJev-v1 4B | effort high (A-1b) | K | 67, 64, 63, 71, 66, 63 | 65.7 | 5 | yes | yes (apus-4b-high) | `bench/apus-2026-09-28/a1b/rows/apus-4b-high.kperm.jsonl@3a84d4b3` |
| lev-4B | GPU (L-2) | K | 66, 67, 65, 65, 68, 61 | 65.3 | 4 | yes | yes (lev-4b-gpu) | `bench/lev-2026-09-28/rows/lev-4b.kperm.jsonl@17d9c043` |
| Kev-4B | as served (arm 12) | K | 60, 58, 59, 63, 57, 69 | 61.0 | 4 | yes | yes (kev-4b) | `bench/jev-2026-09-21/rows-gaps-card/kev-4b.kperm.jsonl@c19fed10` |
| Qwen3-4B-Instruct-2507 | small's base, no adapter (O-2c) | G | 46, 52, 57, 55, 60, 53 | 53.8 | 1 | yes | yes (opendecider-base-qwen3-4b) | `bench/opendecider-2026-09-28/rows/o2c-gpu-qwen3-4b-base.s0.P.rep1.jsonl@34555419 + bench/opendecider-2026-09-28/rows/o2c-gpu-qwen3-4b-base.s0.P.rep1.k1-5.jsonl@34555419` |
| OpenDecider small | zero-shot (O-2) | G | 46, 46, 50, 52, 50, 48 | 48.7 | 1 | yes | yes (opendecider-small) | `bench/opendecider-2026-09-28/rows/o2-gpu-small.s0.P.rep1.jsonl@34555419 + bench/opendecider-2026-09-28/rows/o2-gpu-small.s0.P.rep1.k1-5.jsonl@34555419` |
| OpenDecider nano | zero-shot (O-1) | G | 13, 11, 12, 16, 8, 7 | 11.2 | 5 | yes | yes (opendecider-nano) | `bench/opendecider-2026-09-28/rows/o1-cpu-nano.s0.P.rep1.jsonl@95cfcd88 + bench/opendecider-2026-09-28/rows/o1-cpu-nano.s0.P.rep1.k1-5.jsonl@95cfcd88` |

*The rank counts ties low: a kit-order count level with another order's takes the lower rank.*

### 6. Task (a) reach: does this article answer this question? (63 questions over long pages)

| model | variant | protocol | questions | reached | right of those reached | refused by the model's server as over its limit | not sent: over the context budget (arm 2's 16k; 32k in the 32k row) | median s, reached | row file @ commit |
|---|---|---|---:|---:|---:|---:|---:|---:|---|
| OpenJev 27B | FP8, readout, 16k context (arm 2) | J-read | 63 | 53 | 53 | 0 | 10 | 6.46 | `bench/jev-2026-09-21/rows/openjev-fp8-readout.a.jsonl@3af7a71c` |
| OpenJev 27B | FP8, readout, 32k context, the 10 long pages only | J-read | 10 | 6 | 6 | 0 | 4 | 32.68 | `bench/jev-2026-09-21/rows-addenda/openjev-fp8-readout-32k.along.jsonl@9879cab6` |
| gemma 4 26B-A4B (base) | FP8 dynamic, readout (arm 10) | J-read | 63 | 53 | 53 | 0 | 10 | 2.39 | `bench/jev-2026-09-21/rows-addenda/gemma4-fp8-readout.a.jsonl@38bc9887` |
| lev-4B | GPU (L-2) | K | 63 | 53 | 52 | 0 | 10 | 1.73 | `bench/lev-2026-09-28/rows/lev-4b.ka.jsonl@17d9c043` |
| Kev-9B | as served (arm 12) | K | 63 | 38 | 38 | 15 | 10 | 1.59 | `bench/jev-2026-09-21/rows-gaps-card/kev-9b.ka.jsonl@c19fed10` |
| Kev-4B | as served (arm 12) | K | 63 | 38 | 38 | 15 | 10 | 1.02 | `bench/jev-2026-09-21/rows-gaps-card/kev-4b.ka.jsonl@c19fed10` |
| APUS-OpenJev-v1 9B | effort high (A-1) | K (APUS compile) | 63 | 37 | 37 | 16 | 10 | 1.88 | `bench/apus-2026-09-28/rows/apus-9b-high.ka.jsonl@e804fe0a` |
| APUS-OpenJev-v1 4B | effort high (A-1b) | K (APUS compile) | 63 | 37 | 37 | 16 | 10 | 1.36 | `bench/apus-2026-09-28/a1b/rows/apus-4b-high.ka.jsonl@3a84d4b3` |

### 7. The field exam: is this question answerable from the rules text? (40 items, yes or no)

| model | variant | protocol | right | of | row file @ commit |
|---|---|---|---:|---:|---|
| OpenJev 27B | FP8, readout (arm 5) | J-read | 40 | 40 | `bench/jev-2026-09-21/rows-addenda/openjev-fp8-readout.f.jsonl@c176a040` |
| gemma 4 26B-A4B (base) | FP8 dynamic, readout (arm 10) | J-read | 40 | 40 | `bench/jev-2026-09-21/rows-addenda/gemma4-fp8-readout.f.jsonl@38bc9887` |
| Kev-9B | as served (arm 12) | K | 40 | 40 | `bench/jev-2026-09-21/rows-gaps-card/kev-9b.kf.jsonl@c19fed10` |
| Kev-4B | as served (arm 12) | K | 40 | 40 | `bench/jev-2026-09-21/rows-gaps-card/kev-4b.kf.jsonl@c19fed10` |
| APUS-OpenJev-v1 9B | effort high (A-1) | K (APUS compile) | 40 | 40 | `bench/apus-2026-09-28/rows/apus-9b-high.kf.jsonl@e804fe0a` |
| APUS-OpenJev-v1 4B | effort high (A-1b) | K (APUS compile) | 40 | 40 | `bench/apus-2026-09-28/a1b/rows/apus-4b-high.kf.jsonl@3a84d4b3` |
| imajev-4B | GPU, serving mode (I-2) | K (4 rotations) | 40 | 40 | `bench/imajev-2026-09-28/i2/rows/imajev-4b-3090.kf.jsonl@5840165e` |
| lev-4B | GPU (L-2) | K | 39 | 40 | `bench/lev-2026-09-28/rows/lev-4b.kf.jsonl@17d9c043` |

### 8. The judge seat: grounded, not grounded, or off the page? (119 items, 3 options)

| model | variant | protocol | right, strict (the one label) | right, alternation (any answer in `expected_ok`) | of | row file @ commit |
|---|---|---|---:|---:|---:|---|
| OpenJev 27B | FP8, readout (arm 6) | J-read | 98 | 118 | 119 | `bench/jev-2026-09-21/rows-addenda/openjev-fp8-readout.j.jsonl@c176a040` |
| Kev-9B | as served (arm 12) | K | 95 | 115 | 119 | `bench/jev-2026-09-21/rows-gaps-card/kev-9b.kj.jsonl@c19fed10` |
| APUS-OpenJev-v1 9B | effort high (A-1) | K (APUS compile) | 93 | 113 | 119 | `bench/apus-2026-09-28/rows/apus-9b-high.kj.jsonl@e804fe0a` |
| imajev-4B | GPU, serving mode (I-2) | K (4 rotations) | 92 | 112 | 119 | `bench/imajev-2026-09-28/i2/rows/imajev-4b-3090.kj.jsonl@5840165e` |
| Kev-4B | as served (arm 12) | K | 92 | 112 | 119 | `bench/jev-2026-09-21/rows-gaps-card/kev-4b.kj.jsonl@c19fed10` |
| APUS-OpenJev-v1 4B | effort high (A-1b) | K (APUS compile) | 91 | 111 | 119 | `bench/apus-2026-09-28/a1b/rows/apus-4b-high.kj.jsonl@3a84d4b3` |
| lev-4B | GPU (L-2) | K | 89 | 108 | 119 | `bench/lev-2026-09-28/rows/lev-4b.kj.jsonl@17d9c043` |

### 9. Latency per S0 item (seconds; one request at a time)

| model | variant | protocol | median s | p95 s (nearest rank) | timed rows | runtime | hardware, as the page may say it | what is NOT like-for-like | row file @ commit |
|---|---|---|---:|---:|---:|---|---|---|---|
| OpenJev 27B | FP8, readout (arm 3) | J-read | 0.087 | 0.088 | 108 | vLLM 0.29.0, native FP8 | a workstation card: RTX PRO 6000 Blackwell (96 GB), 420 W | the same workstation card | `bench/jev-2026-09-21/rows-arm3/openjev-fp8-largecard.c.jsonl@1429a7d5` |
| OpenJev 27B | BF16, readout (arm 3) | J-read | 0.092 | 0.096 | 108 | vLLM 0.29.0 | a workstation card: RTX PRO 6000 Blackwell (96 GB), 420 W | a 96 GB workstation card; one letter read, no generation | `bench/jev-2026-09-21/rows-arm3/openjev-bf16-largecard.c.jsonl@1429a7d5` |
| OpenJev 27B | FP8, readout (arm 2) | J-read | 0.350 | 0.417 | 108 | vLLM 0.29.0, tensor-parallel 2 | two used RTX 3090s, 250 W each, over an x4 link | two cards and an all-reduce over a slow link on every token | `bench/jev-2026-09-21/rows/openjev-fp8-readout.c.jsonl@3af7a71c` |
| OpenJev 27B | Q4_K_M GGUF, readout (arm 13) | J-read | 1.131 | 1.196 | 108 | ollama 0.32.13 | one used RTX 3090 (card B), 300 W | a 4-bit quant through ollama; a different 3090 from arm 12's | `bench/jev-2026-09-21/rows-oj-gguf/openjev-q4km-readout.c.jsonl@95d6ae9a` |
| OpenJev 27B | Q4_K_M GGUF, generate (arm 13) | J-gen | 1.141 | 1.215 | 108 | ollama 0.32.13 | one used RTX 3090 (card B), 300 W | writes a letter; the same card as arm 13 readout | `bench/jev-2026-09-21/rows-oj-gguf/openjev-q4km-generate.c.jsonl@95d6ae9a` |
| jevify adapter on gemma 4 26B | Q4_K_M, readout (arm 1) | J-read | 1.017 | 1.093 | 108 | ollama | one used RTX 3090 (card A), 250 W | ollama, 131k context allocated | `bench/jev-2026-09-21/rows/jev-readout.c.jsonl@0edb2527` |
| gemma 4 26B-A4B (base) | Q4_K_M, readout (arm 1) | J-read | 1.025 | 1.116 | 108 | ollama | one used RTX 3090 (card A), 250 W | ollama, 131k context allocated | `bench/jev-2026-09-21/rows/base-readout.c.jsonl@0edb2527` |
| gemma 4 26B-A4B (base) | FP8 dynamic, readout (arm 10) | J-read | 0.120 | 0.151 | 108 | vLLM 0.29.0, tensor-parallel 2 | two used RTX 3090s, 250 W each | two cards, tensor-parallel | `bench/jev-2026-09-21/rows-addenda/gemma4-fp8-readout.c.jsonl@38bc9887` |
| mistral-small 3.2 24B | Q4_K_M, generate (cost series) | J-gen (cost harness) | 0.619 | 0.752 | 108 | ollama (cost-series harness) | one used RTX 3090, 300 W | a different harness; 4k context | `bench/cost-of-intelligence-2026-09-23/benchbox-3090-reading-0/rows/b-mistral-small3.2-24b/b-mistral-small3.2-24b.c.jsonl@6b73f291` |
| Kev-4B | as served (arm 12) | K | 0.070 | 0.083 | 108 | kev.serve (transformers + FastAPI), bf16 | one used RTX 3090 (card A), 250 W | flash-linear-attention kernels; card A at 250 W | `bench/jev-2026-09-21/rows-gaps-card/kev-4b.kc.jsonl@c19fed10` |
| Kev-9B | as served (arm 12) | K | 0.106 | 0.137 | 108 | kev.serve (transformers + FastAPI), bf16 | one used RTX 3090 (card A), 250 W | flash-linear-attention kernels; card A at 250 W | `bench/jev-2026-09-21/rows-gaps-card/kev-9b.kc.jsonl@c19fed10` |
| OpenDecider small | zero-shot (O-2) | O | 0.063 | 0.067 | 108 | transformers, bf16 | one used RTX 3090 (card B), 300 W | card B; masked-letter scoring | `bench/opendecider-2026-09-28/rows/o2-gpu-small.s0.P.rep1.jsonl@34555419` |
| APUS-OpenJev-v1 4B | effort high (A-1b) | K (APUS compile) | 0.171 | 0.178 | 108 | the vendor runtime in-process, BF16 | one used RTX 3090 (card B), 300 W | its own compiled prompt; card B | `bench/apus-2026-09-28/a1b/rows/apus-4b-high.kc.jsonl@3a84d4b3` |
| APUS-OpenJev-v1 9B | effort high (A-1) | K (APUS compile) | 0.183 | 0.208 | 108 | the vendor runtime in-process, BF16 | one used RTX 3090 (card B), 300 W | its own compiled prompt; card B | `bench/apus-2026-09-28/rows/apus-9b-high.kc.jsonl@e804fe0a` |
| lev-4B | GPU (L-2) | K | 0.217 | 0.250 | 108 | lev serve 0.1.1 (reference kernels), bf16 | one used RTX 3090 (card B), 300 W | reference kernels, not Kev's; card B at 300 W | `bench/lev-2026-09-28/rows/lev-4b.kc.jsonl@17d9c043` |
| OpenDecider nano | zero-shot (O-1) | O | 0.199 | 0.246 | 108 | opendecider package, CPU, sandboxed | a laptop CPU | CPU; a 400M encoder | `bench/opendecider-2026-09-28/rows/o1-cpu-nano.s0.P.rep1.jsonl@95cfcd88` |
| deem-0.8-v1 | as shipped (D0), rep 2, uncontended rows | D | 0.421 | 0.542 | 104 | deem server, TorchBackend, bf16, torch 2.14.0+cpu | a laptop CPU, 8 performance cores | CPU; a 0.8B model | `bench/deem-2026-09-27/rows/d0-cpu-0.8.s0.P.rep2.jsonl@b0200003` |
| lev-4B | CPU (L-1) | K | 8.465 | 10.696 | 108 | lev 0.1.1 in-process, bf16, torch 2.14.0+cpu | a laptop CPU, 8 performance cores | CPU, not a GPU | `bench/lev-2026-09-28/l1/rows/lev-4b-cpu.kc.jsonl@17d9c043` |
| imajev-4B | CPU, serving mode (I-1) | K (4 rotations) | 18.134 | 25.395 | 108 | imajev server (uvicorn), bf16, 4 rotations per item | a laptop CPU, 8 performance cores | CPU, and four forward passes per item | `bench/imajev-2026-09-28/rows/imajev-4b-cpu.kc.jsonl@9e86a063` |

### 9a. imajev-4B on the graphics card (I-2), per item, one request at a time (descriptive, computed after the rows were seen; not registered)

*imajev-4B on one used RTX 3090 (card B, 300 W), its own server in bf16, four option orders inside every answer, one request at a time. imajev's S0 ran only on the laptop CPU (I-1, table 9); these are the other sets.*

| set | timed rows | median s | p95 s (nearest rank) | row file @ commit |
|---|---:|---:|---:|---|
| H48, the held-out lines (the same question as S0) | 48 | 0.516 | 0.581 | `bench/imajev-2026-09-28/i2/rows/imajev-4b-3090.kh.jsonl@5840165e` |
| the field exam | 40 | 0.331 | 0.336 | `bench/imajev-2026-09-28/i2/rows/imajev-4b-3090.kf.jsonl@5840165e` |
| the judge seat | 119 | 1.149 | 2.151 | `bench/imajev-2026-09-28/i2/rows/imajev-4b-3090.kj.jsonl@5840165e` |

### 9b. Memory on one 24 GB card: what each graphics-card run added at load, and its highest reading (descriptive, computed after the rows were seen; not registered)

*growth = memory.used after the warm-ups minus the reading before the load, as each run's own residency check recorded it; highest = the highest memory.used sample in the run's telemetry (2 Hz) or, for arms 12 and 13, its per-task nvidia-smi CSVs (1 Hz); longest input = the most tokens among rows whose request had started by the first highest sample. The card's total, 24,576 MiB, is cited from bench/jev-2026-09-21/ARM13-OJ-GGUF.md, the card row (24,576 MiB).*

| model | precision and runtime | grew at load, MiB | resident before the first task, MiB | highest, MiB | the highest first seen | headroom, MiB | longest input by the highest, tokens | longest input in the run, tokens | growth from |
|---|---|---:|---:|---:|---|---:|---|---|---|
| OpenJev 27B, Q4_K_M GGUF (arm 13, card B) | Q4_K_M, ollama 0.32.13 | 16,727 | — | 16,730 | `bench/jev-2026-09-21/rows-oj-gguf/openjev-q4km-generate.c.watts.csv` | 7,846 | the per-task CSVs are not timed against the rows here, so only the run's longest is given | 524 (E4, d) | `bench/jev-2026-09-21/receipts/oj-gguf/run/counted-run.json@95d6ae9a` |
| APUS-OpenJev-v1 9B (A-1, card B) | bf16, the vendor runtime in-process (one load serves high and low) | 18,471 | — | 24,126 | 2026-09-28T15:34:34Z | 450 | 8,187 (on-39, ka) | 8,187 (on-39, ka) | `bench/apus-2026-09-28/receipts/counted-run.json@e804fe0a` |
| APUS-OpenJev-v1 4B (A-1b, card B) | bf16, the vendor runtime in-process (one load serves high and low) | 9,235 | — | 16,684 | 2026-09-28T16:47:38Z | 7,892 | 8,187 (on-39, ka) | 8,187 (on-39, ka) | `bench/apus-2026-09-28/a1b/receipts/counted-run.json@3a84d4b3` |
| lev-4B (L-2, card B) | bf16, lev serve 0.1.1, reference kernels | 9,125 | — | 22,736 | 2026-09-28T13:14:15Z | 1,840 | 14,590 (on-27, ka) | 14,590 (on-27, ka) | `bench/lev-2026-09-28/receipts/counted-run.json@17d9c043` |
| imajev-4B (I-2, card B) | bf16, imajev's own server, four orders per answer | 9,798 | — | 11,256 | 2026-09-28T20:29:13Z | 13,320 | 2,240 (j-108, kj) | 2,240 (j-108, kj) | `bench/imajev-2026-09-28/i2/receipts/counted-unit.log@5840165e` |
| OpenDecider small (O-2, card B) | bf16, transformers | 8,595 | — | 8,596 | 2026-09-28T12:04:43Z | 15,980 | no request had started: the highest came at load | 521 (E4, door) | `bench/opendecider-2026-09-28/rows/o2-gpu-small.*.pass.json (residency_after_warmup.growth_bytes, 12 counted passes)` |
| Kev-9B (arm 12, card A) | bf16, kev.serve (transformers + FastAPI) | not recorded | 15,866 | 20,278 | `bench/jev-2026-09-21/rows-gaps-card/kev-9b.ka.watts.csv` | 4,298 | the per-task CSVs are not timed against the rows here, so only the run's longest is given | 8,105 (on-23, ka) | `none: arm 12's files record no reading from before the load` |
| Kev-4B (arm 12, card A) | bf16, kev.serve (transformers + FastAPI) | not recorded | 8,776 | 13,358 | `bench/jev-2026-09-21/rows-gaps-card/kev-4b.ka.watts.csv` | 11,218 | the per-task CSVs are not timed against the rows here, so only the run's longest is given | 8,105 (on-23, ka) | `none: arm 12's files record no reading from before the load` |

*Kev-9B's record also states "~19 GB of GPU memory in bf16" (bench/jev-2026-09-21/GAPS-CARD.md line 351; TABLES-GAPS-CARD.md line 186): a statement, not one of these readings. Arm 12 recorded no reading from before the load, so no growth is printed for Kev; the resident column is the first idle sample before its first task, with the server loaded and no work in flight. No census of other tenants was recorded; Kev-4B's lowest reading (8,776 MiB), taken between Kev-9B's tasks, shows the two servers were not resident together. OpenDecider small's highest is read from the telemetry inside its own pass windows; the same file also covers its base model's passes.*

### 9c. Checked for the page: the correction for seventeen tests, Kev-9B's paired interval, the bound on 0 of 12, the held-out widths, the resample by dinner, the tests in each option order, the paired shuffle counts, the 9B against the live door, Gemma 4's floored lines and Mistral Small 3.2's other run (descriptive, computed after the rows were seen; not registered)

*None of these checks was registered before its rows were seen, and no bar applies to any of them. Three read sets whose rows a public kit holds back: bound_on_0_of_12 (the doorman), heldout_interval_widths (the held-out lines) and door_9b_against_the_live_door (the doorman and the 2026-09-13 live door's verdicts). Every figure below is in `recount-v2.json` under `checked`, drawn from the recount's own counts or read from the row files it names.*

**The correction for seventeen tests.** Holm's step-down correction over the seventeen exact McNemar tests against the word counter: the page's word-counter table row by row, and deem-0.8-v1 in its caption, at 0.05.

| model | variant | right | of | right only here | right only in the word counter | side | exact p | Holm-adjusted p | told apart after Holm |
|---|---|---:|---:|---:|---:|---|---:|---:|---|
| OpenJev 27B | BF16, readout (arm 3) | 101 | 108 | 37 | 4 | ahead | 1e-07 | 1.5e-06 | yes |
| OpenJev 27B | Q4_K_M GGUF, generate (arm 13) | 99 | 108 | 36 | 5 | ahead | 7.8e-07 | 1e-05 | yes |
| jevify adapter on gemma 4 26B | Q4_K_M, readout (arm 1) | 90 | 108 | 29 | 7 | ahead | 0.00031 | 0.0034 | yes |
| gemma 4 26B-A4B (base) | Q4_K_M, generate (arm 1) | 88 | 108 | 28 | 8 | ahead | 0.0012 | 0.012 | yes |
| mistral-small 3.2 24B | Q4_K_M, generate (cost series) | 73 | 108 | 21 | 16 | ahead | 0.51 | 1 | no |
| Kev-9B | as served (arm 12) | 72 | 108 | 18 | 14 | ahead | 0.6 | 1 | no |
| imajev-4B | CPU, serving mode (I-1) | 68 | 108 | 21 | 21 | level | 1 | 1 | no |
| APUS-OpenJev-v1 4B | effort high (A-1b) | 67 | 108 | 18 | 19 | behind | 1 | 1 | no |
| APUS-OpenJev-v1 9B | effort high (A-1) | 66 | 108 | 18 | 20 | behind | 0.87 | 1 | no |
| lev-4B | GPU (L-2) | 66 | 108 | 18 | 20 | behind | 0.87 | 1 | no |
| Kev-4B | as served (arm 12) | 60 | 108 | 17 | 25 | behind | 0.28 | 1 | no |
| OpenDecider small | zero-shot (O-2) | 46 | 108 | 13 | 35 | behind | 0.0021 | 0.019 | yes |
| Qwen3-4B-Instruct-2507 | small's base, no adapter (O-2c) | 46 | 108 | 13 | 35 | behind | 0.0021 | 0.019 | yes |
| APUS-OpenJev-v1 9B | effort low (A-1) | 30 | 108 | 15 | 53 | behind | 4.1e-06 | 4.9e-05 | yes |
| APUS-OpenJev-v1 4B | effort low (A-1b) | 25 | 108 | 12 | 55 | behind | 1e-07 | 1.5e-06 | yes |
| OpenDecider nano | zero-shot (O-1) | 13 | 108 | 1 | 56 | behind | 8e-16 | 1.3e-14 | yes |
| deem-0.8-v1 | as shipped (D0) | 5 | 108 | 2 | 65 | behind | 3.1e-17 | 5.3e-16 | yes |

10 of 17 are told apart after the correction (10 before it): 4 ahead of the word counter and 6 behind. The largest adjusted p among them is 0.019; the smallest among the rest 1. Every row at or above Gemma 4's 88 ahead after the correction: yes. Every row below 60 behind: yes. No row between told apart: yes.

**Kev-9B's paired interval.** Newcombe's (1998) 95 per cent interval for Kev-9B's six-way share minus the word counter's, paired line by line: each share's Wilson limits combined through the phi coefficient of the 2 x 2 table, given with phi as computed and with its numerator continuity-corrected.

| both right | only Kev-9B right | only the word counter right | neither right | Kev-9B's edge, lines (points) | phi | 95 % interval, points | 95 % interval, lines |
|---:|---:|---:|---:|---|---|---|---|
| 54 | 18 | 14 | 22 | +4 (+3.7) | 0.3525, as computed | -6.5 to +13.8 | -7.0 to +14.9 |
| 54 | 18 | 14 | 22 | +4 (+3.7) | 0.3322, continuity-corrected | -6.6 to +13.9 | -7.2 to +15.0 |

**The bound on 0 of 12.** The Wilson 95 per cent interval on APUS-OpenJev-v1 9B at high refusing the door's harmless lines (the page: "0 of 12 still allows about one harmless line in four"): 0 of 12, Wilson 0.0 to 24.2 per cent, the upper limit one line in 4.1 (from `door.apus-9b-high: controls_refused of controls_n`). The page's other bound, 40 of 40 on the field exam, is field.<model>.wilson (table 7's row files; its Wilson limits are in the JSON).

**The held-out intervals' widths.** The width of each six-way-minus-held-out Newcombe interval (s0_minus_h48), its lower arm (the observed difference less the lower limit), and, holding the six-way count, the smallest kit advantage whose interval would clear zero (percentage points; table 3's readings).

| model | variant | S0 | H48 | S0 % minus H48 % | Newcombe 95 % | width | lower arm | the largest H48 count that would clear zero | the kit advantage that would clear zero |
|---|---|---|---|---:|---|---:|---:|---:|---:|
| OpenDecider small | zero-shot (O-2) | 46 / 108 | 25 / 48 | -9.5 | -25.6 to +7.2 | 32.8 | 16.1 | 12 | 17.6 |
| Qwen3-4B-Instruct-2507 | small's base, no adapter (O-2c) | 46 / 108 | 26 / 48 | -11.6 | -27.5 to +5.2 | 32.7 | 16.0 | 12 | 17.6 |
| naive Bayes (ours) | plain argmax | 68 / 108 | 27 / 48 | +6.7 | -9.4 to +23.1 | 32.5 | 16.1 | 22 | 17.1 |
| APUS-OpenJev-v1 4B | effort high (A-1b) | 67 / 108 | 29 / 48 | +1.6 | -14.1 to +18.1 | 32.2 | 15.7 | 21 | 18.3 |
| imajev-4B | GPU, serving mode (I-2) | 68 / 108 | 30 / 48 | +0.5 | -15.0 to +17.0 | 32.0 | 15.5 | 22 | 17.1 |
| APUS-OpenJev-v1 9B | effort high (A-1) | 66 / 108 | 31 / 48 | -3.5 | -18.7 to +13.1 | 31.8 | 15.2 | 21 | 17.4 |
| APUS-OpenJev-v1 9B | effort low (A-1) | 30 / 108 | 15 / 48 | -3.5 | -19.5 to +11.0 | 30.5 | 16.0 | 6 | 15.3 |
| APUS-OpenJev-v1 4B | effort low (A-1b) | 25 / 108 | 12 / 48 | -1.9 | -17.3 to +11.5 | 28.8 | 15.4 | 4 | 14.8 |
| OpenDecider nano | zero-shot (O-1) | 13 / 108 | 5 / 48 | +1.6 | -11.1 to +11.1 | 22.2 | 12.7 | 0 | 12.0 |
| OpenJev 27B | Q4_K_M GGUF, readout (arm 13) | 99 / 108 | 44 / 48 | +0.0 | -8.4 to +11.9 | 20.3 | 8.4 | 38 | 12.5 |
| deem-0.8-v1 | as shipped (D0) | 5 / 108 | 4 / 48 | -3.7 | -15.2 to +3.9 | 19.2 | 11.5 | none | none |

Near the word counter's level (small deciders with a held-out reading whose six-way count is within 5 lines of the word counter's 68: apus-4b-high, apus-9b-high, imajev-4b-gpu): widths 31.8 to 32.2, lower arms 15.2 to 15.7, and the kit advantage that would clear zero, the S0 count held, 17.1 to 18.3 points.

**The resample by dinner.** The six-way's 108 lines come from 16 dinners; each paired test against the word counter, and the three single-count Wilson intervals the page prints, with whole dinners resampled (20000 resamples, seed 20260929, percentile interval) and with a dinner-robust standard error; a verdict is ahead or behind at p < 0.05 by the exact test, and by dinner when the resampled 95 per cent interval excludes zero. 16 dinners, 5 to 9 lines each; each line's dinner from `bench/cost-of-intelligence-2026-09-23/KIT.lock.json@cb5dc076`.

| model | variant | split | exact p | verdict | difference, points | SE, lines independent | SE, by dinner | design effect | by-dinner z p | by-dinner resample 95 %, points | resample p | verdict by dinner | the same |
|---|---|---|---:|---|---:|---:|---:|---:|---:|---|---:|---|---|
| OpenJev 27B | BF16, readout (arm 3) | 37 to 4 | 1e-07 | ahead | +30.6 | 5.17 | 5.49 | 1.13 | 2.7e-08 | +19.8 to +40.8 | 0 | ahead | yes |
| OpenJev 27B | Q4_K_M GGUF, generate (arm 13) | 36 to 5 | 7.8e-07 | ahead | +28.7 | 5.27 | 5.27 | 1.00 | 5.1e-08 | +18.4 to +38.7 | 0 | ahead | yes |
| jevify adapter on gemma 4 26B | Q4_K_M, readout (arm 1) | 29 to 7 | 0.00031 | ahead | +20.4 | 5.22 | 5.64 | 1.17 | 0.0003 | +9.7 to +31.1 | 0.0002 | ahead | yes |
| gemma 4 26B-A4B (base) | Q4_K_M, generate (arm 1) | 28 to 8 | 0.0012 | ahead | +18.5 | 5.29 | 5.89 | 1.24 | 0.0017 | +6.8 to +28.8 | 0.0035 | ahead | yes |
| mistral-small 3.2 24B | Q4_K_M, generate (cost series) | 21 to 16 | 0.51 | not told apart | +4.6 | 5.64 | 6.47 | 1.32 | 0.47 | -7.5 to +16.5 | 0.51 | not told apart | yes |
| Kev-9B | as served (arm 12) | 18 to 14 | 0.6 | not told apart | +3.7 | 5.25 | 5.53 | 1.11 | 0.5 | -6.7 to +14.3 | 0.55 | not told apart | yes |
| imajev-4B | CPU, serving mode (I-1) | 21 to 21 | 1 | not told apart | +0.0 | 6.03 | 5.74 | 0.91 | 1 | -10.7 to +11.1 | 1 | not told apart | yes |
| APUS-OpenJev-v1 4B | effort high (A-1b) | 18 to 19 | 1 | not told apart | -0.9 | 5.66 | 5.64 | 0.99 | 0.87 | -11.5 to +10.0 | 0.93 | not told apart | yes |
| APUS-OpenJev-v1 9B | effort high (A-1) | 18 to 20 | 0.87 | not told apart | -1.9 | 5.73 | 8.08 | 1.99 | 0.82 | -17.0 to +13.8 | 0.85 | not told apart | yes |
| lev-4B | GPU (L-2) | 18 to 20 | 0.87 | not told apart | -1.9 | 5.73 | 6.87 | 1.44 | 0.79 | -15.1 to +10.9 | 0.84 | not told apart | yes |
| Kev-4B | as served (arm 12) | 17 to 25 | 0.28 | not told apart | -7.4 | 5.99 | 6.86 | 1.31 | 0.28 | -20.6 to +5.6 | 0.3 | not told apart | yes |
| OpenDecider small | zero-shot (O-2) | 13 to 35 | 0.0021 | behind | -20.4 | 6.14 | 5.88 | 0.92 | 0.00053 | -31.3 to -9.1 | 0.0007 | behind | yes |
| Qwen3-4B-Instruct-2507 | small's base, no adapter (O-2c) | 13 to 35 | 0.0021 | behind | -20.4 | 6.14 | 5.26 | 0.73 | 0.00011 | -30.1 to -10.3 | 0.0001 | behind | yes |
| APUS-OpenJev-v1 9B | effort low (A-1) | 15 to 53 | 4.1e-06 | behind | -35.2 | 6.88 | 7.74 | 1.27 | 5.5e-06 | -49.5 to -20.2 | 0 | behind | yes |
| APUS-OpenJev-v1 4B | effort low (A-1b) | 12 to 55 | 1e-07 | behind | -39.8 | 6.57 | 7.74 | 1.39 | 2.7e-07 | -54.2 to -24.8 | 0 | behind | yes |
| OpenDecider nano | zero-shot (O-1) | 1 to 56 | 8e-16 | behind | -50.9 | 5.01 | 5.79 | 1.34 | 1.4e-18 | -61.5 to -39.6 | 0 | behind | yes |
| deem-0.8-v1 | as shipped (D0) | 2 to 65 | 3.1e-17 | behind | -58.3 | 5.12 | 5.81 | 1.29 | 9.5e-24 | -69.0 to -47.1 | 0 | behind | yes |

Verdicts changed by resampling dinners: 0 of 17. Paired standard errors wider by dinner (design effect over 1): 12 of 17.

| model | variant | right | of | Wilson 95 %, % | by-dinner resample 95 %, % | SE, independent | SE, by dinner | design effect | wider by dinner |
|---|---|---:|---:|---|---|---:|---:|---:|---|
| OpenJev 27B | BF16, readout (arm 3) | 101 | 108 | 87.2 to 96.8 | 89.8 to 97.2 | 2.4 | 1.9 | 0.64 | no |
| Kev-9B | as served (arm 12) | 72 | 108 | 57.3 to 74.8 | 58.6 to 74.5 | 4.5 | 4.3 | 0.88 | no |
| naive Bayes (ours) | plain argmax | 68 | 108 | 53.6 to 71.5 | 53.4 to 72.2 | 4.6 | 5.0 | 1.14 | yes |

**The tests in each option order.** The exact McNemar test against the word counter in each of the six option orders Kev's server generates, joined on each line's uid; order 1 is the kit's order (post_hoc.six_order.<model>.kit_order_is_order_0). Each cell: right of 108 (right only here to right only in the word counter, exact p).

| model | variant | order 1, the kit's | order 2 | order 3 | order 4 | order 5 | order 6 | closest | row file @ commit |
|---|---|---|---|---|---|---|---|---|---|
| Kev-9B | as served (arm 12) | 72 (18 to 14, p 0.60) | 69 (18 to 17, p 1.00) | 70 (20 to 18, p 0.87) | 73 (21 to 16, p 0.51) | 69 (20 to 19, p 1.00) | 67 (16 to 17, p 1.00) | order 4, p 0.51 | `bench/jev-2026-09-21/rows-gaps-card/kev-9b.kperm.jsonl@c19fed10` |
| APUS-OpenJev-v1 4B | effort high (A-1b) | 67 (18 to 19, p 1.00) | 64 (20 to 24, p 0.65) | 63 (18 to 23, p 0.53) | 71 (22 to 19, p 0.76) | 66 (21 to 23, p 0.88) | 63 (20 to 25, p 0.55) | order 3, p 0.53 | `bench/apus-2026-09-28/a1b/rows/apus-4b-high.kperm.jsonl@3a84d4b3` |
| APUS-OpenJev-v1 9B | effort high (A-1) | 66 (18 to 20, p 0.87) | 68 (16 to 16, p 1.00) | 70 (19 to 17, p 0.87) | 76 (24 to 16, p 0.27) | 73 (17 to 12, p 0.46) | 67 (19 to 20, p 1.00) | order 4, p 0.27 | `bench/apus-2026-09-28/rows/apus-9b-high.kperm.jsonl@e804fe0a` |
| lev-4B | GPU (L-2) | 66 (18 to 20, p 0.87) | 67 (19 to 20, p 1.00) | 65 (18 to 21, p 0.75) | 65 (18 to 21, p 0.75) | 68 (19 to 19, p 1.00) | 61 (15 to 22, p 0.32) | order 6, p 0.32 | `bench/lev-2026-09-28/rows/lev-4b.kperm.jsonl@17d9c043` |
| Kev-4B | as served (arm 12) | 60 (17 to 25, p 0.28) | 58 (14 to 24, p 0.14) | 59 (17 to 26, p 0.22) | 63 (19 to 24, p 0.54) | 57 (17 to 28, p 0.14) | 69 (20 to 19, p 1.00) | order 5, p 0.14 | `bench/jev-2026-09-21/rows-gaps-card/kev-4b.kperm.jsonl@c19fed10` |

Among the four the page shuffled (Kev-9B as served (arm 12), APUS-OpenJev-v1 4B effort high (A-1b), APUS-OpenJev-v1 9B effort high (A-1), lev-4B GPU (L-2)), the closest is APUS-OpenJev-v1 9B effort high (A-1) at order 4, p 0.268; told apart in any order: no.

**The paired shuffle counts.** Exact McNemar tests on which lines' answers moved over the six orders Kev's server generates, joined on each line's uid.

| first | second | moved, first | moved, second | moved only in the first | moved only in the second | exact p | told apart |
|---|---|---:|---:|---:|---:|---:|---|
| Kev-9B as served (arm 12) | APUS-OpenJev-v1 9B effort high (A-1) | 33 | 41 | 12 | 20 | 0.215 | no |
| lev-4B GPU (L-2) | APUS-OpenJev-v1 9B effort high (A-1) | 27 | 41 | 8 | 22 | 0.016 | yes |
| lev-4B GPU (L-2) | Kev-9B as served (arm 12) | 27 | 33 | 13 | 19 | 0.377 | no |

**The 9B against the live door.** APUS-OpenJev-v1 9B at high against the 2026-09-13 live door on the same 36 planted lines, joined by id; a kind is named when the door's v1.2 rules name it. Ids, counts and kinds only: no line is read.

| planted lines | the 9B refused | the live door refused | only the 9B refused (named, unnamed) | only the live door refused (named, unnamed) | exact p | the 9B's rows record the live verdict on every line | each row's kind equals the kit's |
|---:|---:|---:|---|---|---:|---|---|
| 36 | 32 | 27 | 6 (0, 6) | 1 (1, 0) | 0.125 | yes | yes |

Row files: `bench/apus-2026-09-28/rows/apus-9b-high.kd.jsonl@e804fe0a`, `bench/doorman-planted-2026-09-13/out/verdicts.jsonl@c9fed92f`, `bench/jev-2026-09-21/kit/doorman_planted.json@c176a040`.

**Gemma 4's floored lines.** Rows whose `floored` list is not empty: at least one option letter fell outside the log-probabilities the server returned, so its score was floored.

| model | variant | rows | floored rows | right | row file @ commit |
|---|---|---:|---:|---:|---|
| gemma 4 26B-A4B (base) | Q4_K_M, readout (arm 1) | 108 | 101 | 87 | `bench/jev-2026-09-21/rows/base-readout.c.jsonl@0edb2527` |
| gemma 4 26B-A4B (base) | FP8 dynamic, readout (arm 10) | 108 | 103 | 87 | `bench/jev-2026-09-21/rows-addenda/gemma4-fp8-readout.c.jsonl@38bc9887` |

**Mistral Small 3.2's other run.** Mistral Small 3.2 24B's six-way run on the workstation card, scored as the recount scores the card-A run (s0.mistral-small3.2-24b-generate-coi): strict, a reply with no readable letter wrong.

| run | right, strict | of | unparsed | right, lenient | the first and last stamp, UTC | the same prompts in the same order as the card-A run | row file @ commit |
|---|---:|---:|---:|---:|---|---|---|
| the workstation card | 74 | 108 | 0 | 74 | 2026-09-23T18:30:40Z to 2026-09-23T18:31:49Z | yes | `bench/cost-of-intelligence-2026-09-23/largecard-reading-0/rows/b1-mistral-small3.2-cd/b1-mistral-small3.2-cd.c.jsonl@6b73f291` |
| card A (table 1) | 73 | 108 | 0 | 73 | 2026-09-23T18:56:57Z to 2026-09-23T18:58:05Z | — | `bench/cost-of-intelligence-2026-09-23/benchbox-3090-reading-0/rows/b-mistral-small3.2-24b/b-mistral-small3.2-24b.c.jsonl@6b73f291` |

*Notes on tables 1 to 9b, each tied to its cell.*

- **Table 1, "exact ties".** A tied row has two or more options at exactly the same top probability. The low end
  counts every tied row wrong; the high end counts it right when the label is among the tied options. Kev's server
  rounds probabilities to 2 dp, so ties are not read for Kev. Generate rows are one-hot, so ties do not apply.
- **Table 1, the T1 nulls.** PREREG-T1 5.4 bars each shuffled-label null at ≤ 25 / 108. Seed 20260928 reads 29
  (28 under the other tie-break), over the bar either way, so **every tuned T1 figure is VOID** and none was read
  (D-20260928-027, -029). The nulls are printed because they are the reason; they are not a model anyone would run.
- **Table 3, imajev.** Its S0 is the CPU arm (I-1) and its H48 the GPU arm (I-2): the same request bodies (I-2's H48
  bodies are hash-equal to APUS's, 48 / 48), two runtimes.
- **Table 4, "gate registered for this arm".** "yes" means the arm's own prereg named the 09-13 gate before its first
  call. "no" means this file applied the same rule to rows that were not registered against it; print those as "by
  the 09-13 rule", never as a registered PASS or FAIL.
- **Table 5.** "Right, fewest / most over the six orders" is the S0 count under each of the six orders. For
  OpenDecider's base the kit order is its lowest (46) and one shuffle reaches 60.
- **Table 6.** The 10 pages "not sent" were over arm 2's 16k budget and every later arm inherited that refusal so the
  denominators match (GAPS-CARD, arm 12). At a 32k budget OpenJev FP8 reached 6 of those 10 and got all 6 right.
  Kev refuses at `INFER_MAX_STATE = 8192` (HTTP 422, "branch too long"); APUS at its 8,191-token context. lev
  reached every page sent; its check disclosed a recovered allocator OOM on the 14,323-token page, so a state over
  14,590 tokens is untested on a 24 GB card with its server (D-20260928-023). imajev was not sent task (a): 48 of the
  53 article bodies exceed its 4,096-token limit (RECON-imajev).
- **Table 9.** Card A and card B are two different used RTX 3090s (A at 250 W, B at 300 W, different slots). The
  deem row is rep 2's 104 uncontended rows: rep 1 was not instrumented, and rep 2 chose the same guest on 108 / 108.
  Every other row times all 108. Arm 2 and arm 10 sum two cards; arm 3 is the 96 GB workstation card.
- **Tables 1c, 4b, 5b, 9a and 9b** are descriptive, computed after the rows were seen, and not registered; no bar
  applies, and none may be printed as a registered result. 1c's paired tests are many (56 pairs) and uncorrected:
  read one as a description of those lines, never as a finding. 4b's split totals equal table 4 on every row.
  5b's order 0 is the kit's order on every row and equals table 1's count.
- **Table 9b.** Growth is each run's own residency check (memory.used after the warm-ups minus the reading before
  the load; imajev's is recorded in bytes, 10,273,947,648 B = 9,798 MiB; OpenDecider small's is the same
  9,012,510,720 B on all 12 counted passes). The highest is the telemetry's highest sample (2 Hz) or, for arms 12
  and 13, the highest in their per-task nvidia-smi CSVs (1 Hz). "Longest input by the highest" counts rows whose
  request had started (end stamp minus seconds) by the first highest sample: APUS-9B's peak came during on-39
  itself, while APUS-4B's came after its long pages, during the judge seat, as the allocator kept its cache. The
  card total is nvidia-smi's 24,576 MiB (torch reports 25,298,141,184 B, 24,126 MiB).

### 10. Licence: the text each layer actually carries, as read, with the read date

| model | layer | what was read | text class, as read | read (UTC) | source file |
|---|---|---|---|---|---|
| OpenJev 27B (BF16 and FP8) | weights | the `LICENSE` file in `openjev/openjev` @`0b6bb6e5`; card `license` of both repos | the full CC BY-NC 4.0 text (19,347 B); card `cc-by-nc-4.0`. A second file, `LICENSE-APACHE-2.0`, covers only `helper/` and `serve/` code | 2026-09-28 14:02–14:05 and 14:24:36 | `bench/apus-2026-09-28/RECON-apus.md` §7; `bench/jev-2026-09-21/ARM13-OJ-GGUF.md` §3 |
| OpenJev 27B, Q4_K_M GGUF | weights (a quantisation) | card `license` of `openjev/openjev-GGUF` @`208220bc`; its 13 files | card says `apache-2.0`; no LICENSE file. Recorded as "a discrepancy, not a grant": a format conversion does not relicense weights | 2026-09-28 14:24:36 | `bench/jev-2026-09-21/ARM13-OJ-GGUF.md` §3 |
| gemma 4 26B-A4B | weights | HF card metadata of `google/gemma-4-26B-A4B-it` | `apache-2.0` (metadata; the LICENSE text was not compared in these files) | 2026-09-28 07:54 | `bench-archive/plans-2026-09-28/opendecider/RECON-house-fit.md` line 332 |
| jevify adapter | adapter and its GGUF | HF card metadata | `license: gemma` on both, over an `apache-2.0` base; not resolved | 2026-09-28 08:02–08:05 | `EVIDENCE-TABLES.md` §1a row 5, Appendix B |
| mistral-small 3.2 24B | weights | HF card metadata | `apache-2.0` (metadata) | 2026-09-28 07:54 | `RECON-house-fit.md` line 333 |
| Kev-9B, Kev-4B | code, adapters and heads, bases | GitHub API `spdx_id`; HF card metadata | `Apache-2.0` (metadata; the LICENSE text was not compared word for word in these files) | 2026-09-28 08:02–08:05 (and arm 12's prereg, 2026-09-22) | `EVIDENCE-TABLES.md` §1a row 14; `bench/jev-2026-09-21/GAPS-CARD.md` arm 12 |
| naive Bayes | ours | our own stdlib code and our own long table's lines | no third-party licence | — | `bench/deem-2026-09-27/t1-tune/baseline_nb.py`, PREREG-T1 |
| lev-4B | code | `LICENSE` of the GitHub repo | apache.org's LICENSE-2.0.txt word for word after whitespace normalisation (difflib 1.0; `[yyyy]` left unfilled) | 2026-09-28, the lev recon (~11:3x–11:45) | `bench/lev-2026-09-28/RECON-lev.md` §4 |
| lev-4B | adapter | HF card metadata and README | `apache-2.0`; "released under Apache-2.0, the same license as the base model". Four of its 26 training sets carry CC BY-NC tags (a seat caveat, not a licence on the weights) | 2026-09-28 | the same |
| APUS-OpenJev-v1 9B and 4B | weights | `LICENSE`, 11,544 B, sha256 `bbedc3fd…`, in each repo | Apache-2.0 word for word; the appendix reads "2026 Alibaba Cloud". Byte-identical to the Qwen3.5 bases' LICENSE | 2026-09-28 14:02–14:05 | `bench/apus-2026-09-28/RECON-apus.md` §7 |
| APUS-OpenJev-v1 | runtime code; family repo | `deployment/LICENSE` and two copies; the root `LICENSE` | runtime: the canonical Apache-2.0 text; root: MIT for the family repo's own docs and code, scoped by `LICENSE_NOTICES.md` ("Qwen-derived model weights: Apache License 2.0") | 2026-09-28 14:02–14:05 | the same |
| imajev-4B | code | `LICENSE` of `mohit67890/imajev` @`6ee8a2c5` | sections 1–9 word for word apache.org's text; the appendix replaced by a filled notice | 2026-09-28, the imajev recon (~15:4x) | `bench/imajev-2026-09-28/RECON-imajev.md` §2 |
| imajev-4B | adapter | HF repo `mohit67890/imajev-4b` @`11126e8c` | **no LICENSE file**; card metadata `apache-2.0` only | 2026-09-28 (~15:4x) | the same |
| OpenDecider nano and small | code and all six model repos | `LICENSE` (the same bytes, md5 `4bdc91ed…`, in all six model repos) | **an altered Apache-2.0 text** (word similarity 0.9917): §6 drops "reasonable and customary use in"; §9 drops "choose to offer, and charge a fee for, acceptance of". GitHub's detector reads `NOASSERTION`. Cards say `apache-2.0` | 2026-09-28 07:48–07:58 | `bench-archive/plans-2026-09-28/opendecider/RECON-sources.md` L1, L2, §2.1 |
| OpenDecider nano, small | bases | HF card metadata | nano on `jhu-clsp/ettin-encoder-400m`: `mit`; small on `Qwen/Qwen3-4B-Instruct-2507`: `apache-2.0` with a LICENSE file | 2026-09-28 07:48–07:58 | the same, L3, L5 |
| deem-0.8-v1 | weights and code | HF card metadata; the GitHub repo's `LICENSE` | cards `apache-2.0`; the repo LICENSE is the Apache 2.0 text; base Qwen3.5 `apache-2.0` | 2026-09-27 | `bench-archive/plans-2026-09-26/RECON-deem-decision-models.md` lines 381–399 |

*Publish only these facts as sourced. OpenJev is bench-only (CC BY-NC); the page must not speak of serving it: arm
13 says it "sits WHOLE" on one 3090, and nothing more (ARM13 §11, D-20260928-034). The house enquiry (discussion #1,
opened 2026-09-22T01:19:18Z) had no reply at the last read (RECON-apus §7).*

---

## 11. What we did not measure

- **H48 for Kev-9B, Kev-4B and lev.** H48 was never sent to Kev (the imajev rows say so); lev's legs had no H48 set.
- **H48 for OpenJev at BF16 or FP8.** Only the Q4 GGUF (arm 13) read H48.
- **gemma 4 26B on H48 and on reorder.** Its reorder rate "is owed, and was not measured" (TABLES-GAPS-CARD, gap 4).
- **gemma 4 on the judge seat.** No arm-10 judge-seat row file exists on master.
- **mistral-small 3.2 on anything but S0** in this file (the cost series' other sets were not re-read here).
- **imajev:** reorder (its serving mode already averages four rotations; no six-order pass), task (a) (over its
  4,096-token limit), S0 and the doorman on a GPU (I-1 was CPU; I-2 read H48, the field exam and the judge seat).
- **OpenDecider:** task (a), the field exam and the judge seat (not in the O-1 or O-2 sets).
- **deem-0.8-v1:** the doorman, reorder, task (a), the field exam and the judge seat (D0 read S0, H48 and the D1
  voice sets only). **deem-9b-v1 (D2) was never run.**
- **Any tuned deem checkpoint.** T1 closed VOID by its own null bar; no final was read (D-20260928-027).
- **APUS at `low`** on reorder, task (a), the field exam and the judge seat (it read S0, the doorman and H48).
  **APUS-OpenJev-v1 35B-A3B** was not benched. The APUS 9B and 4B checkpoints are not step-matched (9B ckpt-3000,
  4B ckpt-5949 from a pilot adapter), so nothing here says why the sizes differ (D-20260928-035).
- **The other Kev sizes** (0.5b, 0.6b, 0.8b, 8b, 27b on HF): only 9B and 4B were measured.
- **The hosted Jev.** It was never run. Every Jev-family figure here is an open-weights reproduction; never write
  "we tested Jev".
- **Brier and calibration across protocols.** Not recomputed here: readout, generate, pointer-head and masked-letter
  probabilities are different kinds, and Kev's are rounded to 2 dp.
- **Energy per decision.** Watts files exist for several arms; none was recounted here.
- **Concurrency.** Every latency is one request at a time.
- **Memory growth at load for Kev** (arm 12 recorded no reading before the load; table 9b gives its resident and
  highest readings), and **the smallest card each model fits**: every graphics-card run used a 24 GB card.
- **The OS of the older arms.** Jev-bench rows (arms 1, 2, 3, 10, 12) do not record it; print "OS not recorded".
- **A rulesage-shaped test, and APUS-9B's training-data provenance review**: the two steps D-20260928-044 names next
  toward any seat.
- **The D1 voice route and nudge menu** exist for deem D0 and OpenDecider only (`bench/deem-2026-09-27/d1-voice/`,
  `bench/opendecider-2026-09-28/results*.md`); they are outside this file's readings.

---

## 12. Flags for the writer (found while recounting; each with its cell)

1. **The judge seat has two scorings, and earlier tables printed one of them as "accuracy".** TABLES-GAPS-CARD's
   Kev-9B "96.6 %, 115 / 119" and OpenJev "118 / 119" are the alternation counts; strict is 95 and 98 (table 8).
   Name the scoring every time.
2. **S0's `id` is not unique** (50 distinct ids over 108 rows). Join on `prompt_sha256` or `(uid, text_sha256)`.
   A join on `id` silently drops rows: this lane's first C-1 recount read 99 same-guest items that way, against the
   registered 106 (correct by prompt sha, and by position).
3. **gemma 4 FP8 (arm 10) generate has three counts by parse rule:** 89 (`run.py`, lenient; TABLES-ADDENDA), 88
   (strict by the rows' `parsed` field; this file), 87 (the cost series' strict parser, which also counts three
   name-only replies as format failures; EVIDENCE-TABLES row 10). Name the rule wherever a generate count appears.
   The Q4 base (arm 1) is 88 strict, 89 lenient (D-20260923-020).
4. **mistral-small 3.2 has two runs:** 73 / 108 here (the cost series' RTX 3090 reading at 300 W) and 74 / 108 in
   EVIDENCE-TABLES (the workstation-card reading). Cite the run.
5. **The ledger's day board (D-20260928-044) puts two arms in one cell.** "OpenJev FP8/BF16 101 and Q4 … 99 / 44 /
   33·0": the H48 44 and C-1 99 are the Q4's, the doorman 33 · 0 is FP8 arm 4's. The Q4's own doorman reads 32 · 0
   (table 4). Print them apart.
6. **Two PASSes in table 4 are by the rule, not registered:** OpenJev FP8 arm 4 (33, 0) and APUS-9B at `low` (32, 1).
7. **APUS-4B's S0 is 65 to 68 over four exact ties** (the ledger prints 67 only). APUS-9B's is 65 to 67 (as the
   ledger says).
8. **Contamination, carefully.** S0 went public on the wall 2026-09-19 and in the kit 2026-09-23. Every model
   released or updated after 2026-09-19 could have seen it; H48 is the control, and no S0-minus-H48 gap is
   distinguishable from zero at n = 48 (table 3). That is "no signal at this n", not "clean". Separately, imajev's
   "#1" on JevBench is on that board's public split (86.1 % public against 37.0 % sealed; D-20260928-035).
9. **Four unrelated projects are called "openjev"** (RECON-apus §2): the CC BY-NC `openjev/openjev`, the Apache
   framework `S1LV3RJ1NX/openjev`, the `openjev/openjev-GGUF` builds, and APUS-OpenJev-v1. Name the repo.
10. **Two different RTX 3090s** carry the GPU rows (card A at 250 W for arms 1, 2, 10 and 12; card B at 300 W for
    today's arms). Never compare a card-A latency with a card-B one as if one card.

---

## Appendix A: how every cell is re-derived

Both scripts sit beside this file and use the standard library only. They read row files (and, for table 9b, the
runs' own receipts and telemetry) from an estate checkout at `f2073219` and never call a model or the network.

```
git -C /workshop/estate worktree add --detach <dir> f2073219
python3 recount_v2.py <dir> recount-v2.json          # every figure, from the rows
python3 render_evidence_v2.py recount-v2.json        # tables 1 to 9, 1b, 1c, 4b, 5b, 9a, 9b, 9c, from that JSON
```

| file | sha256 |
|---|---|
| `recount_v2.py` | `<withheld: rewritten file, see README.md>` |
| `render_evidence_v2.py` | `f87b25543b7777c3733ff7e37e2a95b1354ef5967bbf0015bc515e56108bf765` |
| `recount-v2.json` (the output of record) | `<withheld: rewritten file, see README.md>` |

*The first recount's three files (`<withheld: rewritten file, see README.md>…`, `5f2baa19…`, `<withheld: rewritten file, see README.md>…`) are kept in `work-v2/before-close/`. The
new `recount_v2.py` writes every older section of the JSON identically (checked key by key); the new
`render_evidence_v2.py` draws table 1b directly after table 1, so its output from the old JSON equals the old
tables byte for byte.*

*The fold of 2026-09-29 added the JSON's `checked` section and table 9c, and keeps the three files it replaced in
`work-v5/fold/before-fold/`. At `f2073219` the new `recount_v2.py` writes every other section of the JSON byte for
byte as before (a diff of the two outputs: 1,235 lines added, none removed or changed; `work-v5/fold/`). From the
pre-fold JSON the new `render_evidence_v2.py` writes the pre-fold tables byte for byte; from the new JSON it writes
the same tables with table 9c after them.*

The rules the recount applies (also in the script's docstring): correct = the row's `choice` equals its `label`
(rows that carry a boolean `correct` are cross-checked, and every such file agreed: 0 mismatches); generate =
strict (`parsed` false is wrong); ties by exact float equality at the top; Wilson 95 % with z = 1.959963984540054;
Newcombe's hybrid score for S0 minus H48; exact two-sided McNemar on discordant items; doorman planted = `planted`
true, refused = `choice` "refuse"; latency = median and nearest-rank p95 of `seconds`.

Every recount above that the ledger also states agrees with it: S0 13, 46, 46, 66, 66, 67, 68, 72, 60, 5, 99, 101;
H48 44, 31, 29, 30, 27, 5, 4; the doorman 32·0, 27·9, 21·0, 19·0, 15·0, 4·0, 4–5·0, 36·12, 33·0 and the 09-13
live 27·0; reorder 41, 47, 27, 33, 15; lev CPU and GPU the same guest on 106 / 108; McNemar Kev-9B vs naive Bayes
p 0.597, lev vs naive Bayes 0.871, APUS-9B vs Kev-9B 0.286, nano vs Kev-9B 5.4e-17; the judge seat 92 / 112 for
imajev; the T1 nulls 18, 29 (28), 24.

## Appendix B: the ledger rows this file leans on (DECISIONS.md at `1ce595cb`)

| row | what it records |
|---|---|
| D-20260921-225 | Kev zero-shot is not a seat; OpenJev's 93.5 is the best of six option orders |
| D-20260927-018 | deem D0: R-1 (bf16 exact ties) ratified, labelled as after the fact |
| D-20260927-019 | deem T1: the recipe fixed a priori |
| D-20260928-018 | OpenDecider O-1 (nano) landed |
| D-20260928-019 | lev L-1 read, 66 / 108 |
| D-20260928-020 | OpenDecider O-2 (small and its base) landed; amendment v2 upheld |
| D-20260928-023 | lev L-2 landed: the doorman fails the 09-13 gate; long pages read |
| D-20260928-026 | PAIR-3090 landed (not a decider; cited only for the card's day) |
| D-20260928-027, -029 | deem T1 closes VOID by its own null bar; landed |
| D-20260928-030 | APUS-9B passes the 09-13 gate at `high` |
| D-20260928-034 | OJ-GGUF (arm 13) landed: C-1 and C-2 CARRIES; "sits WHOLE" |
| D-20260928-035, -036 | APUS-4B fails the gate; imajev I-1 landed |
| D-20260928-044 | imajev I-2 landed; the day's decider board |
| D-20260928-045 | the go for the article (quoted at the top) |
| D-20260928-047 | the go for the recs train, "decison model write up / comparissons" first (on master from `31b942a3`; read at `f2073219`) |
