# RECON B — the published gemma4:26b decode timeline (SHIPPED data only)

> SANITIZED AT PUBLICATION (2026-08-21): estate-internal box names, home
> paths, and private overlay addresses in this working document are replaced
> with public aliases ([workspace], the-gpu-box, the-dev-laptop, the-vps,
> [private-peer], [agent-memory]); two receipt filenames beginning with a box
> name are renamed to speed-forensic-2026-08-16.md / speed-probe-2026-08-16.json.
> Quoted clock strings are re-expressed in UTC (same instants) and no longer
> byte-match their sources; every load-bearing timestamp is stated in UTC.
> Measurements, counts, shas, and quoted third-party text are untouched.


Recon agent, written 2026-08-21. Every row below carries a receipt: a file path
+ line, or a verbatim command + output excerpt. Where a number that circulates
in briefs could **not** be found in shipped data, it is marked **UNVERIFIED**
rather than narrated.

Constraint honoured: no builds, no benches, no model calls. File reads,
`/bin/grep`, and the `scribe` CLI only.

---

## 0 · How the sweep was run (reproducible)

```
$ find [workspace]/s2s-research-hub/site -type d -name data | sort
[workspace]/s2s-research-hub/site/august-arrivals/data
[workspace]/s2s-research-hub/site/chair-trials/data
[workspace]/s2s-research-hub/site/cove-voice-head-to-head/data
[workspace]/s2s-research-hub/site/data
[workspace]/s2s-research-hub/site/llms-txt/data
[workspace]/s2s-research-hub/site/outside-judges/data
[workspace]/s2s-research-hub/site/reading-is-fast/data
[workspace]/s2s-research-hub/site/seat-trials/data
[workspace]/s2s-research-hub/site/the-compressed-photograph/data
[workspace]/s2s-research-hub/site/the-move/data
[workspace]/s2s-research-hub/site/the-open-call/data
[workspace]/s2s-research-hub/site/three-at-the-table/data
[workspace]/s2s-research-hub/site/voice-trials/data
```

```
$ cd [workspace]/s2s-research-hub/site && /bin/grep -arlo "gemma4:26b" . | sort -u
```
(50 files; the ones carrying a **decode rate** for that tag are the five kits in
§1. `august-arrivals/data/residency.json`, `seat-rows.json`,
`the-new-kid/index.html`, `outside-judges/*`, `llms-txt/*` and
`the-open-call/*` mention the tag but publish **no tok/s for it** — residency
bytes, seat verdicts, gate pass/fail and model digests only. Verified by
grepping each for `tok`/`decode`; e.g. `august-arrivals/data/seat-rows.json:624`
is a seat row whose only timing field is `"warm_p50_ms": 2310`.)

**Models present in `reading-is-fast/data/matrix-runs.json`:** `qwen3.8:27b`
only —
```
$ /bin/grep -ano "\"model\": \"[^\"]*\"" matrix-runs.json | sort -u -t: -k3
64:"model": "qwen3.8:27b"
```
So the draft-head matrix kit is **not** a gemma4:26b source. Its C5 rows
(115.69 / 133.54 / 126.995) are `qwen3.8:27b`, per
`reading-is-fast/data/MATRIX-PROTOCOL.md:5` ("**Target:** `qwen3.8:27b`") and
`:18-24`. Do not mistake them for the MoE arm.

**Models in `the-compressed-photograph/data/ladder-runs.json`:** the four 12B
builds only (`gemma4:12b`, `-it-bf16`, `-it-q8_0`, `-it-qat`) — no 26B rows.

---

## 1 · The chronological table — every SHIPPED gemma4:26b decode reading

| # | date (UTC) | decode tok/s | n | runtime version | posture | kit + file path | comparable? |
|---|---|---:|---:|---|---|---|---|
| R1 | 2026-08-12 22:46–23:10Z | **76.98** median (primary) / **80.37** (contention-reruns substituted); single flagged-sample rerun **144.26** | 13 prompts | **UNVERIFIED** (kit pins no version; sibling kit of the same publication date pins 0.32.9) | `think:false`, `stream:false`, `num_ctx 32768`, `format:json` for dialogue; **three generation lanes concurrent + a live cove bench** | cove-voice-head-to-head · `site/cove-voice-head-to-head/data/rows.json:530-531`, sample at `:2032,:2042` | ❌ **NON-COMPARABLE** — heavy contention by design; real cove prompts; JSON-constrained decode; one sample included a 23.1 s cold load |
| R2 | 2026-08-12 18:53:24–18:53:36Z | **124.975** median (103.18–125.40) | 10 | **ollama 0.32.9** | C5 posture, `num_ctx`/`context_length` 32768, kv `q8_0`, flash on, `think:false`; **32k prompt tier**; ran beside live app traffic | chair-trials · `site/chair-trials/data/c5-stopwatch.json:4621-4648` (arm header `:4270`) | ⚠️ **TIER MISMATCH** — this is the **32k** tier. The 1k and 8k cells for this arm returned **0/10 usable counters** (`:4472-4477`, `:4552-4557`), so there is no 08-12 figure at the tiers everything later uses |
| R3 | 2026-08-13 20:02–20:06Z | journal-line rates **133.50 – 154.06** (per-request, live production) | ~13 lines | **ollama 0.32.9** (stated in the record's own `pattern_status`) | live rulesage-live asks, prompts 2.7k–4.8k tokens, `n_ctx_slot = 32768`; co-residents `muse-glimmer:30b-q8_0-dflash` + `nomic-embed-text` | the-move · `site/the-move/data/c6-vps-rerun.json:1293,494,881,1873,253,2305,6568,9579,8676,6885,1080,1675,9160,2505,1478` | ❌ **NON-COMPARABLE** — live serving, variable prompts, no repeat-set; several lines are 3–9-token generations. Attribution rests on the record's own `journal.pattern: "gemma4:26b"`, and the record itself stamps `pattern_status: "UNCONFIRMED"` |
| R4 | 2026-08-15 | **138.600** — derived p50 over **n = 1,982 live rulings** | 1,982 | **UNVERIFIED** for that date (the kit does not pin the runtime on this row) | production observation, explicitly "**not a bench cell**"; same `eval_count/eval_duration` definition | three-at-the-table · `site/three-at-the-table/data/active-weights-derivation.md:358` | ⚠️ **PRODUCTION, NOT A BENCH** — but it is the only shipped figure for 08-15 and it is a large-n p50 |
| R5 | 2026-08-19 22:19:57–22:20:10Z (Q1) / 22:21:32–22:21:45Z (Q2) | **213.07** (211.82–213.23) Q1 · **210.72** (207.27–211.95) Q2 | 5 + 5 | **ollama 0.32.13** | temp 0, seed 0, `think:false`, `num_ctx 32768`, `num_predict` 220 (Q1) / 320 (Q2), serial, 2 s pauses, 1 unscored warm-up, `keep_alive -1` (resident); **beside live traffic, disclosed** | three-at-the-table · `site/three-at-the-table/data/robber-asks.json:9-19` (medians) ← `site/the-compressed-photograph/data/robber-runs.json:70-130,~1100+` (raw) | ✅ **COMPARABLE to the A/B NEW arm** — same model, same prompt, same runtime, byte-identical reply (see §3) |
| R6 | 2026-08-21 07:24:30–07:25:04Z (~1k) / 07:27:52–07:28:27Z (~8k) | **208.04** (206.63–209.31) @~1k · **201.72** (200.35–203.86) @~8k | 10 + 10 | **ollama 0.32.13** | temp 0, seed 0, `think:false`, `stream:false`, `num_ctx 32768`, `num_predict 256`, serial 2 s, `keep_alive -1`; **verified zero-contention window**; co-residents `qwen3.8:27b`, `gemma4:12b`, later the a3b visitor (4 seated, 63.80 GiB) | three-at-the-table · `site/three-at-the-table/data/STOPWATCH-WINDOW-RUN-2026-08-21.md:118,127` ← raw `stopwatch-window-runs.json` | ✅ **CLEANEST BENCH ROW ON THE TIMELINE** — but the prompt is a repeated-sentence builder, not the robber prompt |
| R7 | 2026-08-21 07:44:45–07:45:15Z (OLD) / 07:45:31–07:45:59Z (NEW) | **154.09** (153.82–154.22) on 0.32.9 · **203.00** (201.37–203.63) on 0.32.13 | 10 + 10 | **0.32.9 vs 0.32.13, same hour, same card** | temp 0, seed 0, `think:false`, `num_ctx 32768`, `num_predict 220`, robber Q1 (68 prompt tokens), 1 warm-up, 2 s pauses, empty card | `[workspace]/rs-runtime-dividend-2026-08-19/ab-runs.json` + `results-window/AB-RUN-2026-08-21.md` | ✅ **THE ONLY CONTROLLED CROSS-VERSION PAIR** — see §3 for the one caveat (the two arms did not generate the same text) |

### Medians recomputed by hand from the shipped values (not transcribed)

- **R2** `[103.18, 124.19, 124.28, 124.51, 124.88, 125.07, 125.28, 125.33, 125.34, 125.40]` → (124.88 + 125.07)/2 = **124.975** ✅ matches `c5-stopwatch.json:4635`.
- **R5 Q1** `[211.82, 212.54, 213.07, 213.20, 213.23]` → **213.07** ✅ matches `robber-asks.json:12`.
- **R5 Q2** `[207.27, 208.55, 210.72, 210.88, 211.95]` → **210.72** ✅.
- **R6 ~1k** `[206.63, 207.64, 207.82, 207.87, 207.95, 208.14, 208.54, 208.62, 209.00, 209.31]` → (207.95 + 208.14)/2 = **208.045 → 208.04** ✅.
- **R6 ~8k** `[200.35, 201.25, 201.36, 201.58, 201.68, 201.75, 201.80, 201.93, 203.34, 203.86]` → (201.68 + 201.75)/2 = **201.715 → 201.72** ✅.
- **R7 OLD** `[153.82, 153.84, 153.99, 153.99, 154.08, 154.09, 154.12, 154.15, 154.18, 154.22]` → (154.08 + 154.09)/2 = **154.085 → 154.09** ✅.
- **R7 NEW** `[201.37, 202.01, 202.05, 202.83, 202.99, 203.02, 203.20, 203.36, 203.38, 203.63]` → (202.99 + 203.02)/2 = **203.005 → 203.00** ✅.
- **R7 ratio** 203.005 / 154.085 = **1.3175 → +31.7%** ✅ matches the run note's claim.

---

## 2 · Receipts, reading by reading

### R2 — chair-trials C5 stopwatch, 2026-08-12, ollama 0.32.9

`site/chair-trials/data/c5-stopwatch.json` — arm header at line 4270:
```
4270	      "arm": "gemma4:26b",
4278	        "not_run_reason": "prod resident; unloading the live seat is out of contract."
```
The `warm-32000` block at line 4621:
```
4630	          "first_at": "2026-08-12T18:53:24Z",
4631	          "last_at": "2026-08-12T18:53:36Z",
4632	          "decode_tok_s": {
4634	            "min": 103.18,
4635	            "median": 124.975,
4636	            "max": 125.4,
```
The 1k tier is empty — `"warm-1000"` at line 4461 carries
`"decode_tok_s": { "n": 0, "min": null, "median": null, "max": null }`
(lines 4472-4477), and the arm's own `primary_tier` block says so plainly:
```
4285	        "usable": [ "0/10", "0/10" ],
4289	        "median_decode_tok_s": [ null, null ],
4293	        "measurable": false,
4294	        "note": "the primary metric is decode tok/s at the 1k tier; an arm with 0 usable counters here has no primary figure whatever its outcome state says about its other tiers"
```
The 8k tier (line 4541) is likewise `n: 0`.

Runtime pin — `site/chair-trials/data/provenance.json:264-269`:
```
264	  "runtime": {
265	    "daemon": "ollama 0.32.9",
266	    "kv_cache": "q8_0",
267	    "flash_attention": "on",
268	    "context_length": 32768,
```
Same pin in `site/chair-trials/data/roster.json:28-32` (`"ollama_version": "0.32.9"`).

Contention posture — `roster.json:8`: *"One 96G VRAM workstation, which also
answered live requests for three of our apps throughout these trials."*

`num_predict` for the C5 warm blocks: **UNVERIFIED** — not stated in the kit.
`think` is false for this arm (`counting-rules.json:74`: *"Thinking postures
were run for two arms only; every other row is think:false"*).

### R3 — the-move C6 serving rerun, 2026-08-13, ollama 0.32.9

Header, `site/the-move/data/c6-vps-rerun.json:1-6`:
```
 "leg": "C6",
 "at": "2026-08-13T20:02:36Z",
 "base_url": "https://rulesage-live.strata2signal.com",
```
Residents at send (`:37-45`): `gemma4:26b` (ctx 32768, `size_vram
17474319810`), `muse-glimmer:30b-q8_0-dflash`, `nomic-embed-text:latest`.

The journal timing lines, all quoted verbatim from the file:
```
:1293  eval time = 3782.84 ms /  505 tokens (7.49 ms per token, 133.50 tokens per second)
:494   eval time = 7395.47 ms / 1024 tokens (7.22 ms per token, 138.46 tokens per second)
:881   eval time = 2616.36 ms /  363 tokens (7.21 ms per token, 138.74 tokens per second)
:1873  eval time = 2817.72 ms /  391 tokens (7.21 ms per token, 138.76 tokens per second)
:253   eval time = 6542.10 ms /  909 tokens (7.20 ms per token, 138.95 tokens per second)
:2305  eval time = 6490.54 ms /  903 tokens (7.19 ms per token, 139.13 tokens per second)
:6568  eval time = 3025.37 ms /  423 tokens (7.15 ms per token, 139.82 tokens per second)
:9579  eval time = 2694.61 ms /  377 tokens (7.15 ms per token, 139.91 tokens per second)
:8676  eval time = 2722.68 ms /  381 tokens (7.15 ms per token, 139.94 tokens per second)
:6885  eval time = 4137.80 ms /  584 tokens (7.09 ms per token, 141.14 tokens per second)
:1080  eval time = 3305.21 ms /  496 tokens (6.66 ms per token, 150.07 tokens per second)
:1675  eval time = 3438.89 ms /  517 tokens (6.65 ms per token, 150.34 tokens per second)
:9160  eval time = 3727.91 ms /  562 tokens (6.63 ms per token, 150.75 tokens per second)
:2505  eval time = 3390.06 ms /  512 tokens (6.62 ms per token, 151.03 tokens per second)
:1478  eval time =   58.42 ms /    9 tokens (6.49 ms per token, 154.06 tokens per second)
```
**This is where a "154" lives in shipped data** — `:1478`, 2026-08-13T20:03:57Z,
and it is a **9-token generation**. The substantial-length lines (363–1024
tokens) cluster at **133.5 – 151.0 tok/s** on 0.32.9. Three further lines in the
same file (2233.90, 273.25, 286.78 tok/s at 3–139 tokens) are **not** plausible
26B decode rates and are almost certainly the drafter/embedder co-residents —
they are excluded above.

Attribution caveat, quoted from the file itself (`:1462`):
> `"pattern_status": "UNCONFIRMED — the guide's journal signature on ollama 0.32.9 has not been read first-hand; the raw window is stored so a corrected pattern can be re-applied offline without re-running an ask"`

and `:1465`:
> `"pattern_status": "UNCONFIRMED — ollama 0.32.9's runner does not necessarily print eval counters to the journal at this log level; an empty match is NOT evidence that nothing generated"`

Timestamp caveat — `site/the-move/data/CORRECTIONS.md` (2026-08-20): 958
offset-form clock readings were re-expressed as UTC on 2026-08-20; measurements
unchanged, quoted journal lines no longer byte-match their source.

### R4 — the 08-15 production p50, 138.600

`site/three-at-the-table/data/active-weights-derivation.md:358`, verbatim:
```
| gemma4:26b | active path | 2.594 | 138.600 — a production observation, not a bench cell: derived p50 over n=1,982 live rulings, 2026-08-15, same eval_count/eval_duration definition | 360 GB/s = 0.360 TB/s |
```

### R5 — the robber mini-bench, 2026-08-19, ollama 0.32.13

`site/three-at-the-table/data/robber-asks.json:6`:
```
 "measured_utc": "2026-08-19 ~22:3X UTC, runtime ollama 0.32.13",
```
`:9-19`:
```
   "question": "Q1-simple",
   "model": "gemma4:26b",
   "decode_tok_s_median": 213.07,
   "decode_tok_s_min": 211.82,
   "decode_tok_s_max": 213.23,
   "n": 5,
   "replies_byte_identical": true,
   "reply_sha256": "d2db508f5f499e2c4a4b66115e8fc97ad79cc51d89d4728bbb42e9c01a856a25",
   "prompt_eval_count": 68,
   "eval_count_first": 130
```
Posture from the pre-registration
(`site/the-compressed-photograph/data/ROBBER-PREREG.md:21-26`):
> "n=5 scored per arm per question after 1 unscored warm-up per arm; serial, 2s
> pause; temperature 0, seed 0, think false, num_ctx 32768, num_predict 220 (Q1)
> / 320 (Q2); runtime counters only. Residents called with keep_alive -1"

and `:3-5`: *"Runs beside live traffic (serial, paced, warm-resident-first);
contention disclosed."*

Raw rows: `site/the-compressed-photograph/data/robber-runs.json` — the five Q1
calls at 2026-08-19T22:19:57Z → 22:20:10Z, all `eval_count: 130`, all
`reply_sha256: d2db508f…`.

### R6 — the zero-contention window stopwatch, 2026-08-21, ollama 0.32.13

`site/three-at-the-table/data/STOPWATCH-WINDOW-RUN-2026-08-21.md:118` and `:127`:
```
| gemma4:26b | MoE 25.8B, 8 of 128 experts |  996 |  8,886 (8,623–9,420) ⚠ | **208.0** (206.6–209.3) | 573 ⚠ | 107 |
| gemma4:26b | MoE 25.8B, 8 of 128 experts | 8007 | 160,137 (158,470–161,656) ⚠ | **201.7** (200.3–203.9) | 578 ⚠ | 112 |
```
Runtime + card state, `:90-92`:
> "Runtime `0.32.13`, `OLLAMA_MAX_LOADED_MODELS=5`, `OLLAMA_KV_CACHE_TYPE=q8_0`,
> `OLLAMA_FLASH_ATTENTION=1`."

Posture, `:101-104`:
> "temperature 0, seed 0, `think` false, `num_ctx` 32768, `num_predict` 256,
> `stream` false. Strictly serial, 2 s between calls, no concurrency anywhere.
> **Zero-contention window** … every figure below was taken with no other caller
> on the card."

Contention proof, `:68`: `| foreign /api/generate | **0** | — |`.

Warm-up vs scored control, `:170` and `:174`:
```
| ~1000 | gemma4:26b | 205.63 | 208.04 | 1.012× |
| ~8000 | gemma4:26b | 199.39 | 201.72 | 1.012× |
```

⚠️ **Do not quote this run's prefill or first-token columns.** The run found and
published its own defect (`:180-186`): a byte-identical prompt at temp 0 makes
the ten scored calls prompt-cache hits, inflating prefill up to 21×. *"the decode
column stands."*

**Prompt-shape difference vs R5/R7:** R6 uses a 996-token / 8,007-token
repetition-built prompt; R5 and R7 use the 68-token robber Q1. Decode rate is
`eval_count ÷ eval_duration` and excludes prefill, but KV depth differs — the
run itself measures the ~1k → ~8k penalty at 208.04 → 201.72 (−3.0%).

---

## 3 · The A/B leg — arms quoted verbatim

`[workspace]/rs-runtime-dividend-2026-08-19/ab-runs.json`, header lines 1-6:
```json
{
 "prereg": "AB-LEG-PREREG.md (sha pinned in PREREG-INDEX.txt)",
 "deviation": "no docker/podman on the box; official v0.32.9 release binary as an isolated second instance (own port+store) — substance preserved, vehicle changed, disclosed",
 "new_version": {
  "version": "0.32.13"
 },
```

**Arm OLD-0.32.9** — 10 rows, `2026-08-21T07:44:45Z` → `07:45:15Z`:
| i | utc | eval_count | eval_duration_ns | decode_tok_s |
|---:|---|---:|---:|---:|
| 0 | 07:44:45Z | 128 | 831,244,000 | 153.99 |
| 1 | 07:44:48Z | 128 | 830,373,000 | 154.15 |
| 2 | 07:44:51Z | 128 | 831,249,000 | 153.99 |
| 3 | 07:44:55Z | 128 | 829,981,000 | 154.22 |
| 4 | 07:44:58Z | 128 | 830,722,000 | 154.08 |
| 5 | 07:45:01Z | 128 | 832,145,000 | 153.82 |
| 6 | 07:45:05Z | 128 | 830,527,000 | 154.12 |
| 7 | 07:45:08Z | 128 | 830,224,000 | 154.18 |
| 8 | 07:45:11Z | 128 | 830,700,000 | 154.09 |
| 9 | 07:45:15Z | 128 | 832,057,000 | 153.84 |

All ten `reply_sha256: 5ad07c931cbffda072d56a238df805cc2d74d842d396e5808dd542548fa971be`,
all `prompt_eval_count: 68`. **Median 154.09, range 153.82–154.22, 10/10 byte-identical.**

**Arm NEW-0.32.13** — 10 rows, `07:45:31Z` → `07:45:59Z`:
| i | utc | eval_count | eval_duration_ns | decode_tok_s |
|---:|---|---:|---:|---:|
| 0 | 07:45:31Z | 130 | 645,565,000 | 201.37 |
| 1 | 07:45:34Z | 130 | 640,319,000 | 203.02 |
| 2 | 07:45:37Z | 130 | 639,185,000 | 203.38 |
| 3 | 07:45:40Z | 130 | 639,762,000 | 203.20 |
| 4 | 07:45:44Z | 130 | 639,251,000 | 203.36 |
| 5 | 07:45:47Z | 130 | 638,415,000 | 203.63 |
| 6 | 07:45:50Z | 130 | 640,927,000 | 202.83 |
| 7 | 07:45:53Z | 130 | 643,520,000 | 202.01 |
| 8 | 07:45:56Z | 130 | 640,414,000 | 202.99 |
| 9 | 07:45:59Z | 130 | 643,403,000 | 202.05 |

All ten `reply_sha256: d2db508f5f499e2c4a4b66115e8fc97ad79cc51d89d4728bbb42e9c01a856a25`,
all `prompt_eval_count: 68`. **Median 203.00, range 201.37–203.63, 10/10 byte-identical.**

Run note, `results-window/AB-RUN-2026-08-21.md:12-23`, verbatim:
> "## Result (gemma4:26b, robber Q1, temp 0/seed 0/ctx 32768/npred 220, n=10 +
> warm-up per arm, empty card) … **P1 CONFIRMED: +31.7% from the runtime alone**
> (>25% registered threshold)."

Deviation, `:4-10`: no docker/podman on the box; OLD ran the official v0.32.9
release **binary** on port 11436 with a blob-for-blob scratch store (18.0 GB);
the first start attempt hit a foreign listener on 11435 and sidestepped rather
than pid-guessing.

Posture read from the harness itself
(`[workspace]/rs-runtime-dividend-2026-08-19/leg2b_ab.py:55-58`):
```python
def call(host, npred=220, keep="5m"):
    body = json.dumps({"model": MODEL, "prompt": Q1, "stream": False, "think": False,
        "options": {"temperature": 0, "seed": 0, "num_ctx": 32768, "num_predict": npred},
        "keep_alive": keep}).encode()
```
Note `keep_alive: "5m"` — the A/B ran gemma4:26b as a **visitor**, not as the
pinned `keep_alive -1` resident that R5 and R6 used.

### 🔗 The strongest cross-link on the timeline

The NEW arm's reply hash **`d2db508f5f499e2c4a4b66115e8fc97ad79cc51d89d4728bbb42e9c01a856a25`**
with `eval_count 130` and `prompt_eval_count 68` is **byte-identical** to the
published 08-19 robber Q1 cell:
```
$ /bin/grep -a -A9 '"model": "gemma4:26b"' site/the-compressed-photograph/data/robber-runs.json \
  | /bin/grep -a '"eval_count"\|"reply_sha256"\|"question"' | head -3
   "question": "Q1-simple",
   "eval_count": 130,
   "reply_sha256": "d2db508f5f499e2c4a4b66115e8fc97ad79cc51d89d4728bbb42e9c01a856a25",
```
So R5 (published, 08-19) and R7-NEW (this morning) are **the same model doing the
same work on the same runtime** — the cleanest same-model-same-card pair the
archive owns.

### ⚠️ Three flags on the A/B that the article must carry

1. **The two arms did not generate the same text.** OLD produced
   `5ad07c93…` at `eval_count 128`; NEW produced `d2db508f…` at `eval_count 130`.
   Decode tok/s is a per-token rate so the comparison survives, but "same weights,
   same prompt, same answer, just faster" is **false** — the runtimes disagree on
   the output. That is worth a sentence, not a hidden footnote.
2. **R5 (213.07, beside live traffic, resident) is FASTER than R7-NEW (203.00,
   empty card, visitor) on identical work** — −4.7% in the *quieter* condition.
   Whatever explains that (visitor vs `keep_alive -1` seating, the OLD instance
   having just been SIGTERMed 16 s earlier at 07:45:15Z, clock/thermal state, the
   scratch-store blob copy), it runs **against** the contention narrative and is
   unexplained by anything in the shipped record. **Open question.**
3. **The "~145 with a 70B co-resident" premise is UNVERIFIED** — see §4.

---

## 4 · UNVERIFIED — claims that do not have a shipped receipt

**(a) The "~145" archive baseline.** No shipped kit publishes a gemma4:26b decode
median near 145. Exhaustive grep of every `site/*/data/` file for `14x.y`
alongside `tok`/`decode` returns only: `cove-voice-head-to-head/data/rows.json`
`144.26` (a **single contention-rerun sample**, not a median, from the most
contended run on the timeline) and the-move journal lines at 141.14–154.06 (live
serving, several of them 3–9-token generations). The nearest shipped **median-like**
figure for the era is **R4's 138.600 p50 over n=1,982 live rulings, 2026-08-15**.

`scribe` corroborates that the 145 was always a soft in-session recollection, and
that its provenance was itself under investigation:
```
$ scribe similar "the 145 tok/s gemma reading had a 70B model co-resident on the card" --k 5
[0.6619] 2026-08-16T21:27:03Z  claude-opus-5  Benchmark Glimmer model against golden test sets
    I'm finding the baseline ~143 tok/s figure for gemma4:26b from 08-15, plus today's 205 observation…
```
```
$ scribe search "gemma4:26b 145 tok"
[0.1449] 2026-08-16T21:30:08Z  claude-opus-5
    145 baseline from the 08-14 and 08-15 logs, and figuring out where that number was actually measured…
```
**Recommendation:** the article should cite **138.600 (08-15, n=1,982, production
p50)** as the shipped pre-upgrade production figure and stop using ~145, or state
plainly that ~145 is an unreceipted recollection.

**(b) The "70B co-resident" explanation in `AB-LEG-PREREG.md:11` and
`AB-RUN-2026-08-21.md:19-22`.** No 70B model appears anywhere in the shipped
residency record. The full census roster:
```
$ /bin/grep -aho "\"model\": \"[^\"]*\"" site/august-arrivals/data/residency.json | sort -u
"model": "gemma4:12b"          "model": "muse-glimmer:30b-q4_K_M-dflash"
"model": "gemma4:12b-it-q8_0"  "model": "muse-glimmer:30b-q8_0-dflash"
"model": "gemma4:26b"          "model": "nemotron3:33b"
"model": "granite4.1:30b-q8_0" "model": "nemotron-3.5-lightning:30b-a3b"
"model": "muse-glimmer:30b"    "model": "nomic-embed-text:latest"
                               "model": "olmo-3.1:32b-think-q4_K_M"
                               "model": "qwen3.5:27b"
                               "model": "qwen3.6:27b"
```
Largest tag on the card in that era is **33B**. And the residents actually
recorded during the 08-13 serving leg:
```
$ /bin/grep -aho "\"name\": \"[a-z0-9:._-]*\"" site/the-move/data/c6-vps-rerun.json | sort | uniq -c
     47 "name": "gemma4:26b"
     39 "name": "muse-glimmer:30b-q8_0-dflash"
     19 "name": "nomic-embed-text:latest"
      2 "name": "gemma4:12b"
```
`scribe search "70B co-resident gemma4 decode contention"` → `no matches`.
**The 70B is not in the record.** The AB-RUN note's decomposition ("~32 points
runtime, the remainder contention") rests on it and should not be published as
written.

**(c) The cove-voice kit's runtime version (R1).** The kit pins posture but no
`ollama_version`; grep for `ollama|0\.32|runtime` in `rows.json` and `gate.json`
returns nothing. Its sibling kits published the same week pin 0.32.9, but that is
inference, not a receipt.

**(d) `num_predict` for the chair-trials C5 warm blocks (R2).** Not stated in the kit.

---

## 5 · The honest timeline, stated once

On **0.32.9** the shipped gemma4:26b readings run **124.975** (32k tier, quiet-ish
bench, 08-12), **133.5–151.0** (live serving journal, 08-13), **138.600**
(production p50 over 1,982 rulings, 08-15), and — measured this morning under
proper control — **154.09** (robber Q1, empty card, 68-token prompt).

On **0.32.13** they run **213.07** (robber Q1 beside traffic, 08-19), **208.04 /
201.72** (~1k/~8k, verified zero-contention window, 08-21), and **203.00**
(robber Q1, empty card, A/B NEW arm, 08-21).

The **only** apples-to-apples cross-version pair in existence is R7:
**154.09 → 203.00, +31.7%, same hour, same card, same prompt, same posture,
runtime the only variable.** Everything before it differs in tier, prompt,
contention, or measurement kind, and the table above says which.

The **one thing the timeline cannot yet explain** is R5 (213.07, beside traffic)
sitting 4.7% *above* R7-NEW (203.00, empty card) on byte-identical work.
