# RECON-C — the estate legend, and the NOT-claims

> SANITIZED AT PUBLICATION (2026-08-21): estate-internal box names, home
> paths, and private overlay addresses in this working document are replaced
> with public aliases ([workspace], the-gpu-box, the-dev-laptop, the-vps,
> [private-peer], [agent-memory]); two receipt filenames beginning with a box
> name are renamed to speed-forensic-2026-08-16.md / speed-probe-2026-08-16.json.
> Quoted clock strings are re-expressed in UTC (same instants) and no longer
> byte-match their sources; every load-bearing timestamp is stated in UTC.
> Measurements, counts, shas, and quoted third-party text are untouched.


Recon agent, 2026-08-21. Sources: the local conversation archive (`scribe` CLI +
the raw transcript `.jsonl` on disk), estate memory files under
`[agent-memory]/`, and the bench/forensic
artifacts those sessions wrote. Read-only throughout; no build, bench or model
was run for this report.

**Reading rule for this document.** Every claim below carries one of: a verbatim
command + output, a file path (+ line or key) with quoted content, or an archive
message with its UTC timestamp and speaker. Anything I could not verify is
labelled **UNVERIFIED** and is not smoothed over.

---

## 0. Executive answer to the three jobs

1. **The legend is TRUE, and the load-bearing distinction is: it is an UPSTREAM
   ollama default change, not an estate-side config change.** Ollama v0.32.10
   stopped supplying a default `repeat_penalty` of 1.1 to models whose Modelfile
   sets none. The estate never set the parameter — before or after. The archive
   supports "sampling silently off" in *both* senses: the sampler's `penalties`
   stage was silently dropped (the speed win) **and** every product on that
   daemon silently began generating with no repetition penalty (the quality
   hole, still unmeasured as of the last archive record).
   **But three corrections to the legend as memory currently states it** — see §2.4.

2. **The NOT-claims are real, and the mission's framing of (a) is wrong.** The
   Muse Glimmer seat was BLOCKED by the **`think:false` drops-the-schema trap**,
   not by anything numbered 412. The `412` receipt belongs to a *different*
   model and a *different* week: `qwen3.8:27b-nvfp4` answered its pull with
   `412: this model requires macOS`. Both are documented below with receipts.
   Glimmer also carries a *third*, earlier NOT-claim: on release day it was
   **MLX / Apple-Silicon only** and could not run on the estate's CUDA box at all.

3. **Nine dated decode readings for `gemma4:26b` between 2026-07-24 and
   2026-08-21 corroborate the ~145 → ~205 arc** — with the honest caveat that
   they were taken at four different prompt lengths and context tiers, so the
   arc is corroborated in *shape*, not as a controlled series. Table in §4.

**⚠ One live contradiction inside the estate's own record** (§5): the 2026-08-16
forensic says contention was **REFUTED** with a receipt; the 2026-08-21 A/B
run-note says the remainder of the jump "is consistent with the contention
candidate." They cannot both be right, and the article must pick or disclose.

---

## 1. What the two memory files claim

### 1.1 `the-speed-legend-memory.md`

Read at `[agent-memory]/the-speed-legend-memory.md`
(header says `modified: 2026-08-17T01:03:47.887Z`, `originSessionId:
65ce6e48-6b1c-4f74-835b-6cb52bd99c37`). Verbatim, lines 3 and 14-23:

> description: "⭐ the-gpu-box's +50% speed jump SOLVED to ONE upstream commit (forensic
> 2026-08-16, mechanism printed not inferred): ollama v0.32.10 stopped defaulting
> repeat_penalty=1.1 — and the same change silently turned that sampling OFF in
> live products, quality unmeasured"

> **The mechanism:** ollama **v0.32.10 stopped defaulting `repeat_penalty=1.1`**
> for models whose Modelfile sets none. gemma4:26b (128-expert MoE, 8
> experts/token) had been paying a once-per-token **full-vocab penalties pass** =
> **2.570 ms of every 4.776 ms token**.
>
> **The receipt:** byte-identical sealed-wire replay on the same runtime —
> **209.4 tok/s** with the new default vs **136.1 tok/s** with `repeat_penalty:1.1`
> re-added (= 97.5% of the historical 139.6). Apportionment: the **entire 1.538×**
> is the v0.32.10 default change; residual ≤2.5%. **CONTENTION REFUTED** — it ran
> the wrong way (the slow era had gemma solo; the fast era has 3 residents plus
> live traffic). Cutover timestamp of record: **2026-08-15 01:36:27Z**.

### 1.2 `the-speed-identity-memory.md`

Read at the same directory (`modified: 2026-08-17T17:43:03.714Z`,
`originSessionId: f0b22046-8b86-44ce-98f9-ea9c9233c41c`). Lines 23-28, verbatim:

> **Measured 07-16 (painted-real-hq, albedobaseXL, same prompts/seeds both boxes):**
> the-gpu-box world-tier 1536×1024@64 = **6.95s** vs the dev laptop 21.6s (the-gpu-box 3.1×);
> scene 1024²@32 = 3.28s vs 7.05s; sketch 512²@32 = 3.00s vs **1.98s local**
> (local wins tiny tiers).

**⚠ The two claims are unrelated and must not be fused.** "3.1× world tier" is an
**image-generation** (ComfyUI / SDXL) comparison between **two different boxes**
on 2026-07-16. It has nothing to do with tok/s, ollama, or the +50% legend. The
memory file merely happens to house both.

---

## 2. Job 1 — the legend, traced to its original sessions

### 2.1 The original observation (the operator's own words)

Raw transcript
`[agent-archive]/65ce6e48-6b1c-4f74-835b-6cb52bd99c37.jsonl`
(session "Benchmark Glimmer model against golden test sets", span 2026-08-12 →
2026-08-17). Extracted verbatim:

```
===== 2026-08-16T13:10:12.209Z  [user/text] =====
one more thing: after we updated ollama last night, i started see over
200tok/sec for gemma4:26b from the-gpu-box, :D is that just from ollama upgrade?
how is that possible?
```

The first response did **not** assert the upgrade. Verbatim, same file,
`2026-08-16T13:11:02.255Z [assistant/text]`:

> **On the 205 tok/s** — genuinely possible from the upgrade alone, but let me be
> house-honest about the three candidate causes before we celebrate a free +40%:
> 1. **The upgrade is a real candidate** […]
> 2. **But the 145 had company**: last night when the watchdog looked, the-gpu-box was
>    running *multi-tenant* — llama3.3:70b resident at 48 GB plus litellm traffic. […]
> 3. **And the skeptic's check**: confirm the context length didn't change […]

Later the same day the observation was reported on both product paths.
`2026-08-16T21:23:02.445Z [assistant/thinking]`:

> Both rulesage-live and amble now hit ~200 tok/s through the-gpu-box's ollama 0.32.13,
> so the gemma4:26b speedup seems to generalize across serving paths rather than
> being isolated.

### 2.2 The forensic that answered it

Dispatched `2026-08-16T21:24:12.999Z` as an Opus lane
(`tool_use:Agent`, description "The GPU box speed-jump forensic"). Its brief named five
candidate causes A–E, with contention (B) flagged as the leading suspect.

It returned `2026-08-16T21:38:55.536Z`. The report it wrote is **on disk**:

- `[workspace]/rs-thirteen-2026-08/recon/speed-forensic-2026-08-16.md` (20,460 bytes,
  mtime 2026-08-16 15:38), sha256 reported in the return as
  `999fe912b73ddd55adc6cc0b5ab87d559145cc0d84f6b8a5b137443ee4b17a4d`, commit `1afc375`.
- Raw probe data: `[workspace]/rs-thirteen-2026-08/recon/speed-probe-2026-08-16.json`.

### 2.3 What exactly was measured, when, and on what versions

**The decisive A/B — same binary, one option toggled.** From
`speed-probe-2026-08-16.json` (read directly, not via the report):

```
host: http://[gpu-box-daemon]      version: 0.32.13
bundle: S1-S1-ask-A
wire_sha256_sealed: 66e1a1989bd745fc34242e62d948ecb32f7fe5d5a2c0e4a08db5a05425021dbe
started_at: 2026-08-16T21:31:56Z   finished_at: 2026-08-16T21:32:08Z
```

| arm | `options_sent` | scored rows (`tok_s_eval`) | median |
|---|---|---|---|
| DEFAULT | `{"num_ctx": 32768}` | 209.79 · 208.31 · 209.37 | **209.37** |
| RP-1.1 | `{"repeat_penalty": 1.1, "num_ctx": 32768}` | 136.13 · 137.25 · 136.06 | **136.13** |

`prompt_eval_count` was 1617 on every row of both arms. The two arms ran **12
seconds apart** on one daemon. 209.37 ÷ 136.13 = **1.538×**.

**The historical baseline — independently recomputed by me from the sealed
records, not taken from the report.** Command and output:

```
$ python3 -c "... json.load over rs-open-call-2026-08-14/results/samples/S*-the-gpu-box/local-gemma4-26b.json ..."
== S1-the-gpu-box started_at 2026-08-14T13:43:20.418447+00:00
   wire 66e1a1989bd745fc ec 137 tok_s 140.14 | num_ctx 32768 keep_alive None
   wire 66e1a1989bd745fc ec 143 tok_s 140.10 | num_ctx 32768 keep_alive None
   wire dc345669c251b66b ec 157 tok_s 137.79 | num_ctx 32768 keep_alive None
   wire dc345669c251b66b ec 168 tok_s 137.75 | num_ctx 32768 keep_alive None
== S2-the-gpu-box started_at 2026-08-14T13:45:49.856330+00:00
   wire 4bd0fd9b179dfebf ec 188 tok_s 139.00 | num_ctx 32768 keep_alive None
== S3-the-gpu-box started_at 2026-08-14T13:44:15.734400+00:00
   wire 12ad77df8563b267 ec 135 tok_s 150.49 | num_ctx 32768 keep_alive None
MEDIAN over 6 = 139.55 min 137.75 max 150.49
```

Two things this independently confirms: the median **139.55** the memory cites,
and that the replayed bundle's hash `66e1a198…` **is byte-identical to the
historical request** (it appears in both the sealed bundle and the 08-14 sample
record). The replay was a genuine replay.

**The runtime pin for the old era** was inferred, then upgraded to a receipt by
process identity. From the forensic report, §2 footnote †:

> the-gpu-box's unit log shows **pid 189678** serving the sealed leg's calls at
> `2026-08-14T13:43:20Z`; that same pid was still serving at
> `2026-08-16T01:00:01Z`; its `/api/version` was read as `0.32.9` at
> `2026-08-16T01:27Z`; and `systemctl` puts its replacement (pid 231004) at
> `2026-08-16T01:36:27Z`.

**The cutover timestamp of record:** `2026-08-16T01:36:27Z` = **2026-08-16
01:36:27Z**. (The memory file writes it as "2026-08-15 01:36:27Z" — an off-by-one
day in the memory's UTC rendering. The forensic's own text and the
`ActiveEnterTimestamp` receipt the `ActiveEnterTimestamp` receipt reads `2026-08-16T01:36:27Z`. **Flag for the article: the memory's cutover date is wrong
by one day in UTC.**)

### 2.4 Does the archive support "sampling silently off" as the mechanism?

**Yes — and it is the rare case where the daemon printed the mechanism rather
than the lane inferring it.** From `speed-forensic-2026-08-16.md` §3, quoting the-gpu-box's
own `journalctl -u ollama` across the restart boundary (a `?` prefix means the
stage was skipped):

```
Before — pid 189678, ollama 0.32.9, 2026-08-16T01:00:01Z:
sampler chain: logits -> penalties -> ?dry -> ?top-n-sigma -> top-k -> ?typical -> ...
        repeat_last_n = 64, repeat_penalty = 1.100, ...

After — pid 231004, ollama 0.32.13, 2026-08-16T01:42:22Z:
sampler chain: logits -> ?penalties -> ?dry -> ?top-n-sigma -> top-k -> ?typical -> ...
        repeat_last_n = 64, repeat_penalty = 1.000, ...

Today's two probe arms, twelve seconds apart:
15:31:57  ?penalties …  repeat_penalty = 1.000   init sampler, took 0.21 ms   → 209 tok/s
15:32:01   penalties …  repeat_penalty = 1.100   init sampler, took 8.88 ms   → 136 tok/s
```

### 2.5 Upstream or estate-side? — **UPSTREAM. This is the load-bearing fact.**

Four independent receipts, all pointing the same way.

**(a) The vendor's own release note.** `speed-forensic-2026-08-16.md` §4 quotes
v0.32.10 (published 2026-08-12T22:36:49Z,
<https://github.com/ollama/ollama/releases/tag/v0.32.10>):

> "Models that don't set a `repeat_penalty` now default to 1.0 (off) instead of
> 1.1, matching other engines and speeding up speculative decoding; set a
> per-model parameter if an older model repeats itself."

The compare view `v0.32.9...v0.32.13` (16 commits) carries the commit
`api: stop applying repeat_penalty 1.1 to models that don't set one`.

**(b) The model's Modelfile sets none.** `curl /api/show gemma4:26b` returned only:

```
PARAMETER temperature 1
PARAMETER top_k 64
PARAMETER top_p 0.95
```

So the 1.1 the model was paying was **never configured by anyone at the estate** —
it was the daemon's default reaching a model that declared nothing.

**(c) No estate code sends any penalty.** My own grep across the three bench
workspaces plus the estate ops repo:

```
$ /bin/grep -rn --include=*.py --include=*.java --include=*.js --include=*.yaml \
    --include=*.yml --include=*.json -l "repeat_penalty" \
    [workspace]/rs-thirteen-2026-08 [workspace]/rs-open-call-2026-08-14 [workspace]/estate
[workspace]/rs-thirteen-2026-08/recon/speed-probe-2026-08-16.json
```

The **only** hit is the forensic probe that deliberately re-added it. And across
the three product repos:

```
$ /bin/grep -rn "repeat_penalty" [workspace]/projects/rulesage [workspace]/projects/amble [workspace]/projects/kiln
the doc-engine planning file PRODUCTION-LINE.md:311:  … the `repeat_penalty`-off quality hole under three live gemma surfaces
```

One hit, and it is a *planning note about the hole*, not a config. The house-law
request builder (`[workspace]/rs-open-call-2026-08-14/harness/ollama_houselaw.py:244`,
`build_chat_body`) pins `num_ctx` and requires `think` explicitly; it has no
penalty parameter at all.

**(d) The lane said so explicitly.** Follow-up recon agent, archive file
`…/65ce6e48-…/subagents/agent-af4ae0c48533568fe.jsonl`, `2026-08-16T22:07:18.820Z`:

> **No product path is affected** — and no estate code sends any penalty anywhere;
> the only mention is a *refusal* in `openai_api.py:270`. Not deliberate by us;
> **upstream**, and the vendor reversed it — `qwen3.8:27b` ships `presence_penalty 0`.

**Verdict on job 1's load-bearing question: the estate changed nothing. It
upgraded a runtime, and a vendor default moved underneath three live products.**
That is the article's spine and it is fully receipted.

### 2.6 Three corrections the article must apply to the memory as written

These are places where the archive has moved on and the memory file has not.

**Correction 1 — the "MoE is why" story is half right, and was reframed by a
second measurement the same evening.** The memory (and the forensic) attribute
the outsized gain to gemma4:26b being a 128-expert MoE with a small per-token
matmul budget. The follow-up lane then A/B'd a **dense** model and found the
mechanism generalizes. Archive, `…/subagents/agent-af4ae0c48533568fe.jsonl`,
`2026-08-16T22:07:18.820Z`, verbatim:

> **A second penalty measurement.** I A/B'd qwen3.6 (`recon/perf-probe-qwen-penalty.json`):
> 65.57 → 77.86 tok/s, **1.187×**, stage cost **2.407 ms/token** vs gemma's 2.570.
> That reframes the mechanism: the penalties tax is a near-**constant** full-vocab
> cost, so its relative bite is inversely proportional to model speed (35% on
> gemma, 15.8% on qwen). Gemma's MoE architecture isn't special — it just has the
> smallest budget for the tax to eat.

I verified the underlying probe file
`[workspace]/rs-thirteen-2026-08/recon/perf-probe-qwen-penalty.json`
(`started_utc: 2026-08-16T21:58:40Z`, model `qwen3.6:27b`, n=3 per arm):

| arm | `options_sent` | `tok_s_eval` rows | ms/token |
|---|---|---|---|
| DEFAULT (Modelfile `presence_penalty 1.5`) | `{}` | 65.58 · 65.55 · 65.57 | 15.248–15.255 |
| PP0 (stage dropped) | `{"presence_penalty": 0.0}` | 77.96 · 77.79 · … | 12.828–12.855 |

This is a **better** article beat than the MoE story: the penalty stage is a flat
per-token toll, so the *fastest* models lose the largest fraction of their budget
to it. It also makes the finding portable to any local stack, not a gemma quirk.

**Correction 2 — the "the-dev-laptop caller sends presence_penalty=1.500" bullet in the
memory is WRONG and was closed the same night.** Memory lines 32-34 say:

> A the-dev-laptop caller sends `presence_penalty=1.500` — 88 requests were still paying
> that stage tax on 08-16 (14:50Z signature; pre-dates the upgrade). Find
> whether 1.5 is load-bearing or inherited config.

The archive closed it 28 minutes after the memory's forensic landed
(`2026-08-16T22:07:18.820Z`, same message as above):

> **The `presence_penalty=1.5` caller — closed, and it isn't a caller.** It's
> `qwen3.6:27b`'s own stock Modelfile (`PARAMETER presence_penalty 1.5`), applied
> server-side. The GPU box's log ties it to the model one line before the sampler chain.
> The 88 taxed requests were **our own thirteen bench** (`local-qwen3.6-27b`, 84
> cells, 08:50–09:10 from this box).

**Do not repeat the "mystery caller" framing.** The true story is better: a
*second* vendor-shipped penalty default, on a different model, silently taxing
the estate's own bench arm — and quietly making that arm's published latency rows
non-comparable to every other local arm.

**Correction 3 — the cutover date in the memory is off by one day in UTC.** See
§2.3. Memory says `2026-08-15 01:36:27Z`; the receipt is
`2026-08-16T01:36:27Z` = **2026-08-16T01:36:27Z**.

### 2.7 The "3.1× world tier" claim — verified, and it is a different subject

I verified this from the render-test harness's own timing files rather than from
prose. Command and output:

```
$ python3 -c "... json.load timing-summary.json for both tests ..."
== the-gpu-box-speed_paintedreal-tiers      cold 8.74  albedobaseXL_v21.safetensors dpmpp_2m/karras cfg 6.0
   1536×1024@64 n= 6 median= 6.95  min= 6.87  max= 8.52
   1024×1024@32 n= 6 median= 3.28  min= 3.19  max= 3.56
   512×512@32   n= 6 median= 3.0   min= 2.93  max= 3.02
== the-dev-laptop-speed_paintedreal-tiers cold 35.05 albedobaseXL_v21.safetensors dpmpp_2m/karras cfg 6.0
   1536×1024@64 n= 6 median= 21.63 min= 21.54 max= 26.09
   1024×1024@32 n= 6 median= 7.05  min= 6.97  max= 7.1
   512×512@32   n= 6 median= 1.98  min= 1.94  max= 2.04
```

Paths: `[workspace]/render-tests/the-gpu-box-speed_paintedreal-tiers/timing-summary.json`
and `[workspace]/render-tests/the-dev-laptop-speed_paintedreal-tiers/timing-summary.json`
(both mtime 2026-07-16).

- 21.63 ÷ 6.95 = **3.113×** → the "3.1×" claim is **CONFIRMED** by arithmetic on
  the harness's own medians.
- **And it inverts at the small tier**: 1.98 ÷ 3.00 = **0.66×** — the laptop card
  *wins* 512²@32. That is a genuine NOT-claim in its own right (a bigger card is
  not uniformly faster; dispatch floor dominates tiny work).
- Cold start: the-gpu-box 8.74 s vs local 35.05 s; the-gpu-box's `8.74 − 3.00 = 5.74 s`
  checkpoint-load overhead matches the memory's "≈ +5.7s".

**UNVERIFIED:** the literal string "3.1×" does not appear in the archive — `scribe
search "3.1x world tier the-gpu-box"` returns only a 2026-08-21 message *about the memory
file* and one unrelated 2026-08-09 hit. The ratio is derived, not quoted. The
underlying medians are receipted; the label is the memory's own arithmetic (which
checks out).

---

## 3. Job 2 — the NOT-claims (upgrades are not always free)

### 3.1 (a) Muse Glimmer — **the mission's premise is wrong; here is what actually blocked it**

**There is no 412 in the Glimmer story.** `scribe search glimmer 412` returns four
hits; two are this very mission's own prompt text echoing into the archive, one is
a 2026-08-16 message about qwen, and one is a 2026-08-11 note about naming
discipline. The `412` receipt belongs to `qwen3.8:27b-nvfp4` — see §3.2.

Glimmer carries **two** distinct NOT-claims, a day apart:

**NOT-claim α — day zero: the model could not run on the estate's hardware at all.**
Session `ee1c1cb7-2933-48e1-b914-67b9641a2b94` ("Verify Ubuntu updates and service
health"), `2026-08-10T22:41:30.826Z [user/tool_result]`, quoting ollama's own
library page and release list verbatim:

```
=== releases ===
v0.32.7 (2026-08-10T10:49:15Z)
v0.32.6 (2026-08-04T18:49:20Z)
v0.32.5 (2026-07-27T01:25:40Z)

## Muse Glimmer
> Note: Muse Glimmer is currently available via initial support via Ollama's MLX
> engine on Apple Silicon. Additional support and optimizations for Apple Silicon,
> NVIDIA, AMD, and other platforms will be available in the coming days.
```

and a second fetch of the library page in the same session,
`2026-08-10T22:42:04.136Z [user/tool_result]`:

```
page bytes: 85429
Apple Silicon count: 2
promise phrasing count: 2
===
Note: Muse Glimmer is currently available via initial support via Ollama's MLX
engine on Apple Silicon. Support for NVIDIA, AMD, and other platforms will be
available in the coming days.
```

The lane's read, `2026-08-10T22:42:00.351Z [assistant/thinking]`:

> So the answer is clear — Muse Glimmer isn't available for CUDA yet. As of today,
> it's MLX/Apple-Silicon-only with NVIDIA support promised in the coming days.

The estate's response was to **build a daily watcher** rather than retry blindly:
`~/.config/systemd/user/glimmer-check.timer`, `OnCalendar=*-*-* 08:43:00`,
`Persistent=true` (cited in the archive at 2026-08-12T13:29:57Z). It fired on
2026-08-11 at 08:43 when the MLX-only disclaimer came off — archive,
`2026-08-11T15:48:19.682Z [claude-fable-5]`:

> **Glimmer is GO for CUDA** — the daily watcher fired this morning at 08:43
> (MLX-only disclaimer gone, NVIDIA in the release [notes])

with the version gate noted at `2026-08-11T15:48:15.901Z`:

> The GPU box is currently on version 0.32.5, which needs to be upgraded to at least
> 0.32.7 to support the glimmer architecture.

**NOT-claim β — the seat block itself: `think:false` silently drops the schema.**
The bench ran 2026-08-10/11. Its results file is on disk at
`[workspace]/rs-glimmer-bench-2026-08-10/results/RESULTS.md`, and it is explicitly
scoped (line 3):

> **This is EXPLORATORY — a directional read, not a seat bake-off.**

Verbatim rows from that file:

```
line 37:  | Format probes parsed | 2/6 | 6/6 |
line 39:  | **Format polarity** | **STANDARD TRAP** | **SCHEMA HELD BOTH WAYS** |
line 40:  | Median eval tok/s | 97.235 | 143.535 |
line 41:  | Median wall per fixture | 2.5805s | 0.9785s |
```

and the Glimmer probe table (lines 53-61):

```
| probe                     | think | format      | repeats parsed | schema enforced | stable |
| p1-schema-think-false     | false | json-schema | 0/2            | False           | True   |
| p2-schema-think-true      | true  | json-schema | 2/2            | True            | True   |
| p3-bare-json-think-false  | false | "json"      | 0/2            | n/a (no schema) | True   |

- `p1-schema-think-false` parse error: `JSONDecodeError: Extra data: line 5 column 2 (char 54)`
- `p1-schema-think-false` parse error: `JSONDecodeError: Extra data: line 5 column 2 (char 384)`
```

The "Extra data" is a template leak — the same file shows the returned body ending
in `}<|eot|>` (lines 73 and 91). Meanwhile `gemma4:26b` on the *same* ollama
version parsed 2/2 on all three probes (lines 100-104).

The ruling, written into the board the same day. Archive
`[agent-archive]/86119a9d-8ff4-4cc6-9651-5841aa4c630b.jsonl`,
`2026-08-11T16:14:51.678Z [assistant/tool_use:Edit]` on
`rulesage/planning/drafts/ORCHESTRATOR-SCREEN.md`:

> **FORMAT: STANDARD TRAP (think:false silently drops schema, 0/2 parsed — "Extra
> data" = the `<|eot|>` template leak) vs gemma4 SCHEMA-BOTH-WAYS on 0.32.9.**
> Seat runs think:false ⇒ **glimmer CANNOT wear the seat as packaged** — day-one
> template bugs, re-probe on ollama patches. VERDICT: no seat bake-off yet;
> gemma4 stays.

Note the qualifier that makes this a *good* NOT-claim: **"Day-one packaging, not
model quality"** (memory `muse-glimmer-seat-verdict.md`, line 21). On grounding it
was at parity (4/4 grounded, 2/2 refusals) and it *beat* gemma4 on list structure
(2/2 vs 1/2). It was disqualified on a serving-stack contract, not on ability.

**Versions and dates for the Glimmer NOT-claim:**

| when (UTC) | what | version |
|---|---|---|
| 2026-08-10 22:41 | library page + release notes read: MLX/Apple-Silicon only | ollama 0.32.7 latest; the-gpu-box on **0.32.5** |
| 2026-08-11 (watcher fire) | daily watcher fires: disclaimer gone, NVIDIA listed | needs **≥0.32.7** |
| 2026-08-11 (bench) | format probes: 0/2 parsed under `think:false` | the-gpu-box upgraded to **0.32.9** |
| 2026-08-11 16:14Z | seat ruled BLOCKED; gemma4:26b stays | 0.32.9 |

### 3.2 (a′) The actual `412` — a different model, a different week

`ollama pull qwen3.8:27b-nvfp4` answered `412: this model requires macOS`.
Discovered 2026-08-17 evening, during the qwen3.8 precision round.

Archive `[agent-archive]/fd51c578-80b7-439c-972b-6e2886534ae1.jsonl`,
`2026-08-17T21:57:24.257Z [assistant/text]`:

> **nvfp4 is unpullable on this box entirely** (the registry answers `412: this
> model requires macOS` — no quiet window fixes that, and our live stub currently
> promises its rows)

and the amendment it triggered, written the same minute
(`2026-08-17T21:57:37.013Z [tool_use:Write]` →
`[workspace]/rs-qwen38-quants-2026-08-17/PREREG-AMENDMENT-3.md`):

> `ollama pull qwen3.8:27b-nvfp4` answers `412: this model requires macOS`: the
> tag ships as 1,209 per-tensor layers rather than GGUF. This is not a box
> condition a maintenance window changes; the arm is WITHDRAWN, with this
> receipt, and the exhibit page says so rather than keeping a promise the
> platform cannot honor.

The *published* page was corrected the same evening rather than left promising —
`2026-08-17T21:57:41.195Z [tool_use:Edit]` on
`[workspace]/s2s-research-hub/site/the-new-kid/index.html`, new text:

> a third tag, `nvfp4`, answered its pull with `412: this model requires macOS` —
> it ships as per-tensor layers rather than GGUF, so that row is withdrawn with
> its receipt rather than promised

The irony worth naming in the article: v0.32.10's release notes — **the same
release that gave the estate its +50%** — also advertised *"Faster prefill on
NVFP4 MLX models with a global scale, about 7–8% on Qwen3.6 and Muse Glimmer"*
(`speed-forensic-2026-08-16.md` §4). That advertised win was **unreachable** on this
hardware. One release, one free gift and one closed door.

### 3.3 (b) The `think:false` kills `format` trap

Memory file `[agent-memory]/ollama-think-false-kills-format.md`
(`modified: 2026-08-17T17:42:20.833Z`, `originSessionId:
d982fb45-6ded-44f6-bdbf-9e5e52ce4044`). Verbatim, lines 11-21:

> **Measured on the-gpu-box (2026-07-28), same model, same prompt, same schema, one
> field different:**
>
> ```
> format alone                → '{"verdict": "PASS", "why": "2 + 2 equals 4…"}'   ✓
> format + "think": false     → '\nPASS'                                          ✗
> ```
>
> Reproduced on `nemotron3:33b` and consistent with `qwen3.5:27b` returning a
> bare `"FAIL"` (2 tokens, `done_reason:"stop"` — a complete, correct answer in
> the wrong shape). No error is returned; `format` is simply not applied.

**What it broke** (lines 27-32):

> the kiln-packs judge bake-off recorded BOTH reasoning candidates as UNMEASURABLE
> (129/129 format failures each) and would have eliminated them from the judge
> seat on a config artifact. The measurement law held — bare words were recorded
> as measurement failures, never laundered into verdicts — so no wrong number was
> published, only a wrong *exclusion*.

**Why it hides** (lines 24-26): `gemma4:12b` honored the schema perfectly (0/129
failures) in the same harness on the same day — so the bug presents as "some
models don't support structured output" when it is actually "our request disables
structured output for models that reason."

**The cure was itself corrected** (lines 34-43). The file's *original* advice —
omit `think` — was **retired** on 2026-08-17:

> do **NOT** omit `think` (this file's original cure — retired). Omitting it is
> not neutral: omitted ≡ `true`, and on `/api/generate` the reasoning is
> DISCARDED, so the caller sees an empty string that grounded pipelines read as
> an honest abstain. Keep `think` **explicit on every wire call**, and verify the
> `think:false` + `format` pairing **empirically per model and per ollama
> release**.

**It is per-model and per-release, with dated readings both ways** (lines 45-49):

> on ollama 0.32.9 gemma4 holds the schema under BOTH think polarities — its old
> inversion is GONE — while Glimmer 30B still fails. the-gpu-box now runs 0.32.13,
> un-re-probed against this trap; re-measure before assuming either result still holds.

Corroborated in the archive at `2026-08-16T01:54:37.706Z`:

> **`gemma4-26b` judge seat** parsed 2/2 on every config, stable. Notably p1
> (schema + `think:false`) showed `schema_enforced=True` — the pairing that once
> recorded two candidates 129/129 unmeasurable held here. Measured on two models
> on one release; not a claim [beyond that]

**Timeline for the article:**

| date | reading | version |
|---|---|---|
| 2026-07-28 | trap PROVEN: `format` silently unenforced under `think:false` | the-gpu-box, version **UNVERIFIED** in this file |
| (same era) | `gemma4:12b` unaffected, 0/129 failures | same harness/day |
| (earlier) | gemma4 had an **inverted** polarity | ollama **0.32.5** |
| 2026-08-11 | gemma4 inversion GONE, schema holds both ways; Glimmer still fails | ollama **0.32.9** |
| 2026-08-16 01:54Z | gemma4-26b p1 `schema_enforced=True`, 2/2 stable | ollama **0.32.13** |
| as of memory write 2026-08-17 | Glimmer **un-re-probed** on 0.32.13 | — |

### 3.4 (c) The bonus NOT-claim: the dividend is not durable

On 2026-08-17 the "+50% free" evaporated to **24.6 tok/s** — roughly an 8× loss
against the new baseline — for reasons that had nothing to do with any upgrade.
Archive `…/fd51c578-….jsonl`, `2026-08-17T21:48:58.602Z [assistant/thinking]`:

> The operator is reporting rulesage-live running at 24.6 tok/s, roughly 8x slower than
> its normal 200+ baseline.

Root cause, same session `2026-08-17T21:50:28.153Z`:

> Found it: gemma4 (RuleSage's answerer) is split 49% GPU/51% CPU at 24.6 tok/s
> because ComfyUI reclaimed 33GB of VRAM and reset gemma4's keep-alive to the
> 5-minute default.

This is banked as the "half-offload signature" law in
`[agent-memory]/the-vram-pressure-memory.md`
(lines 19-22): *"a 26-27B answering at ~25 tok/s with GPU util ~0% = layers on
CPU… `keep_alive -1` is a timer, NOT a pin."*

**The article beat:** a vendor gave the estate +50% for free on 08-16; a
neighbouring process took 8× of it back on 08-17. Free speed is a property of a
*runtime*, not of a *box*.

---

## 4. Job 3 — every dated `gemma4:26b` decode reading I could find, July → now

All values are decode-only (`eval_count ÷ eval_duration` from the daemon's own
counters) unless the row says otherwise. **The prompt-length / context-tier column
is the reason this is not a controlled series** — read it before reading the numbers.

| # | date (UTC) | reading | prompt / tier | posture + box | ollama | receipt |
|---|---|---|---|---|---|---|
| 1 | 2026-07-24 15:27:22Z | **146.0** warm (cold 140.5) | 5434 prompt tokens, 88 gen | amble answerer bake-off, prod payload, the-gpu-box | 0.32.x pre-.10 (**UNVERIFIED exact**) | transcript `6e8d9987-…jsonl` `2026-07-24T15:27:22.878Z`: `{"challenger_warm": {"model": "gemma4:26b", "wall_ms": 1007, "load_ms": 375, "prompt_tokens": 5434, "gen_tokens": 88, "tok_s": 146.0}}` |
| 2 | 2026-07-24 15:31:04Z | **142.4** | 5353 prompt tokens, 62 gen | full amble path (the-vps → tunnel → the-gpu-box) | same | same file `2026-07-24T15:31:04.096Z`: `{"ok": true, "model": "gemma4:26b", "total_ms": 1401, …, "tok_per_s": 142.4}` |
| 3 | (asserted, undated) | **146** | — | `amble/STATE.md:33` — *"Guide: gemma4:26b LIVE (MoE ~4b active; 146 tok/s, ~1.4s answers,"* | — | file read; traces to row 1 |
| 4 | 2026-08-12 18:53:24–36Z | **124.975** (n=10, 103.18–125.4) | **@32k tier** | chair-trials C5 stopwatch, 96G workstation | pre-cutover | `[workspace]/s2s-research-hub/site/chair-trials/data/c5-stopwatch.json`, row `gemma4:26b`, block `warm-32000` |
| 4b | same | **no @1k figure exists** | @1k tier | same | same | same file: `warm-1000` usable `0/10`, `"measurable": false` |
| 5 | 2026-08-14 13:43–13:45Z | **139.55** (n=6, 137.75–150.49) | 1617 prompt tokens | open-call sealed the-gpu-box leg; `resident models ['gemma4:26b']` | **0.32.9** (pid-identity receipt) | recomputed by me from `rs-open-call-2026-08-14/results/samples/S*-the-gpu-box/local-gemma4-26b.json` — see §2.3 |
| 6 | 2026-08-15 | **138.600** | production p50 over **n=1,982 live rulings** | rulesage-live, active path | pre-cutover | `[workspace]/s2s-research-hub/site/three-at-the-table/data/active-weights-derivation.md:358` — *"138.600 — a production observation, not a bench cell: derived p50 over n=1,982 live rulings, 2026-08-15, same eval_count/eval_duration definition"* |
| 7 | 2026-08-15 06:00:47Z | **~143** | — | session note | 0.32.9 | `scribe search "gemma4:26b runs around 143 tok/s"` → hit `[0.7856] 2026-08-15T06:00:47.210000+00:00 claude-opus-5` |
| — | **2026-08-16 01:36:27Z** | — | — | **CUTOVER**: ollama restarted 0.32.9 → 0.32.13 | — | `systemctl show ollama -p ActiveEnterTimestamp` → `2026-08-16T01:36:27Z` |
| 8 | 2026-08-16 13:10:12Z | **"over 200"** | live product traffic | operator's own observation on rulesage-live + amble | 0.32.13 | transcript, quoted in §2.1 |
| 9 | 2026-08-16 21:31:56–58Z | **209.37** (n=3, 208.31–209.79) | 1617 prompt tokens, `num_ctx` 32768 | forensic replay, DEFAULT arm | 0.32.13 | `speed-probe-2026-08-16.json` |
| 9b | 2026-08-16 21:32:01–08Z | **136.13** (n=3, 136.06–137.25) | same, `repeat_penalty:1.1` re-added | forensic replay, RP-1.1 arm | 0.32.13 | same |
| 10 | 2026-08-17 ~21:48Z | **24.6** | live rulesage ask | **half-offloaded, 49% GPU / 51% CPU** | 0.32.13 | §3.4 |
| 11 | 2026-08-21 07:22:51–07:29Z | **208.045** @1k (n=10, 206.63–209.31); **201.715** @8k (n=10, 200.35–203.86) | tiers 1000 / 8000, `num_ctx` 32768, temp 0, seed 0, `think:false`, npred 256 | three-at-the-table stopwatch window | 0.32.13 | computed by me from `[workspace]/s2s-research-hub/site/three-at-the-table/data/stopwatch-window-runs.json` (`run_id moe-dense-stopwatch-2026-08-19`, `written_utc 2026-08-21T07:29:17.511Z`) |
| 12 | 2026-08-21 07:44:45–07:55Z | **154.09** on 0.32.9 (n=10, 153.82–154.22) vs **203.00** on 0.32.13 (n=10, 201.37–203.63) | robber Q1, 68 prompt tokens, npred 220, ctx 32768, temp 0/seed 0 | **side-by-side runtime A/B, empty card** | both | `[workspace]/rs-runtime-dividend-2026-08-19/ab-runs.json` + `results-window/AB-RUN-2026-08-21.md` |
| 13 | 2026-08-21 09:39:39Z | **207.4** (200 tokens in 0.96 s) vs gemma4:12b **125.3** | paired probe, both warm | the-gpu-box | 0.32.13 | transcript `…/fd51c578-….jsonl` `2026-08-21T09:39:39.638Z [tool_result]`: `gemma4:26b eval: 200 tokens in 0.96 s = 207.4 tok/s \| prompt_eval: 0.035 s \| load: 0.464 s` |

**Reading of the table.**

- The pre-cutover cluster (rows 1, 2, 5, 6, 7) is tight: **138.6 – 146.0**, across
  three weeks, four prompt shapes, two serving paths, and one 1,982-call
  production p50. The memory's "~145" is the top of that band; **~140 is the
  better centre**, and row 6 (n=1,982 live) is the strongest single pre-upgrade
  receipt in the estate.
- The post-cutover cluster (rows 9, 11, 12-NEW, 13) is equally tight: **203.0 –
  209.4**. The memory's "~205" is right down the middle.
- **Row 4 (124.975) is the one that looks off, and it is explainable, not
  anomalous**: it is the only reading taken at the **32k** context tier. Row 11
  shows the within-run tier gradient on the *same* runtime (208.045 @1k →
  201.715 @8k), so a further drop at 32k is the expected shape.
- **A dense control exists and behaves as the mechanism predicts.** `gemma4:12b`,
  same two datasets: 105.665 @1k / 101.91 @8k on 2026-08-12 (pre) → 126.910 @1k /
  122.700 @8k on 2026-08-21 (post) = **+20%**, against the 26b's ~+50%. That is
  exactly the "flat per-token toll, bigger bite on faster models" reframe of
  §2.6/Correction 1. **⚠ This is NOT a controlled comparison** — different
  exhibits, different prompt builders, seven weeks apart — and must be presented
  as directional corroboration only.

---

## 5. ⚠ The contradiction the article has to resolve

Two estate artifacts disagree about whether **contention** contributed to the
~145 → ~205 arc.

**Artifact A — the forensic, 2026-08-16.** `speed-forensic-2026-08-16.md` §5, candidate B:

> **REFUTED as the cause of the baseline.** The 08-14 leg that produced 139.55
> recorded `resident models ['gemma4:26b']` — gemma was the *only* resident.
> Today's 209.37 was measured with **three** residents … and **live traffic in
> flight**… Today is the *busier* box and it is the *faster* one.

I verified the underlying field myself:

```
$ python3 -c "print(json.load(open('.../samples/S1-the-gpu-box/local-gemma4-26b.json'))['contention'])"
… OBSERVED AT RUN TIME (2026-08-14T13:43:08.697480+00:00) against the daemon this
leg's calls go to, over HTTP and nothing else: resident models ['gemma4:26b'],
shelf listing 19 tags. …
```

and the forensic's B′ row places the `llama3.3:70b` residency at
**2026-08-16T01:27–01:33Z** — *after* the cutover and after every ~140 reading.

**Artifact B — the A/B run note, 2026-08-21.**
`[workspace]/rs-runtime-dividend-2026-08-19/results-window/AB-RUN-2026-08-21.md`:

> **P1 CONFIRMED: +31.7% from the runtime alone** (>25% registered threshold).
> P2 (archive's ~145 for the OLD era): measured OLD = 154.09 on an EMPTY card —
> the archive's ~145 reading carried a 70B co-resident, so the remaining ~6% gap
> is consistent with the contention candidate, now bounded. The ~145→~205
> archive jump decomposes as: ~32 points runtime, the remainder contention

**The A/B note's premise is contradicted by a receipt.** No ~145-lineage reading
was taken with a 70B co-resident: row 5 recorded gemma solo, and rows 1, 2, 6 and
7 predate the 70B's residency window entirely.

**A simpler explanation for OLD = 154.09 vs the archive's ~140, offered as an
inference and labelled as such (NOT measured):** the A/B leg used a **68-token
prompt**; every ~140 reading used **1617–5434 tokens** of prompt. Row 11 shows
decode falling monotonically with context on this very model (208.0 @1k → 201.7
@8k on the same day). A 68-token prompt should sit *above* a 1617-token one, and
154.09 > 139.55 by 10.5% — the right sign and a plausible size. **This is an
inference from the tier gradient, not a measurement; a same-prompt run at both
prompt lengths on 0.32.9 would settle it and has not been done.**

**A second, quantitative tension worth disclosing.** The two experiments disagree
on the size of the sampler toll:

| experiment | design | ratio | implied stage cost |
|---|---|---|---|
| forensic, 2026-08-16 | **same binary** (0.32.13), option toggled, 1617-token prompt | 209.37 / 136.13 = **1.538×** | 7.346 − 4.776 = **2.570 ms/token** |
| A/B leg, 2026-08-21 | **different binaries** (0.32.9 vs 0.32.13), default options, 68-token prompt | 203.00 / 154.09 = **1.318×** | 6.490 − 4.926 = **1.564 ms/token** |

Both agree the runtime/default change is the dominant cause. They differ by ~1.0
ms/token on the toll's size. Candidate explanations (**all UNVERIFIED**): prompt
length changing the per-token baseline; the `llama.cpp` b10353→b10380 bump landing
differently under the two designs; or ordinary between-session variation. **The
honest article line is that two independent designs bracket the dividend at
1.32×–1.54×, both far outside noise, with the mechanism printed by the daemon in
one of them.**

---

## 6. Open holes in the estate record (as of the last archive entry I can see)

1. **The quality question is still unanswered.** The forensic's own §6 said
   *"Nothing here measured output quality… a 50% speed win that reintroduces
   looping is not a win."* The last archive mention is `2026-08-17T03:03:14.049Z`
   still listing the *"quality looping-check"* as owed. `scribe search "looping
   check repeat_penalty quality"` returns nothing later. **UNVERIFIED whether it
   has since run** — if the article claims the dividend is free, that claim is
   currently unbacked on the quality axis.
2. **`repeat_penalty` is still not pinned per seat.** No product config sets it
   (§2.5c). The doctrine the forensic earned — *"sampling parameters get pinned
   per seat, deliberately"* (`2026-08-16T21:39:48.693Z`) — is stated but, on the
   evidence of the grep, **not yet implemented**.
3. **[Item withheld at publication: an internal serving-posture question still open at press time, tracked in the estate's own records. Withholding disclosed here rather than silently cut.]**

---

## 7. Source index

**Memory files** (all under `[agent-memory]/`):
`the-speed-legend-memory.md` · `the-speed-identity-memory.md` ·
`muse-glimmer-seat-verdict.md` · `ollama-think-false-kills-format.md` ·
`the-vram-pressure-memory.md`

**Forensic and probe artifacts:**
`[workspace]/rs-thirteen-2026-08/recon/speed-forensic-2026-08-16.md` (commit `1afc375`) ·
`[workspace]/rs-thirteen-2026-08/recon/speed-probe-2026-08-16.json` ·
`[workspace]/rs-thirteen-2026-08/recon/perf-probe-qwen-penalty.json`

**Bench data read directly:**
`[workspace]/rs-open-call-2026-08-14/results/samples/S{1,2,3}-the-gpu-box/local-gemma4-26b.json` ·
`[workspace]/rs-open-call-2026-08-14/harness/ollama_houselaw.py` ·
`[workspace]/rs-glimmer-bench-2026-08-10/results/RESULTS.md` ·
`[workspace]/s2s-research-hub/site/chair-trials/data/c5-stopwatch.json` ·
`[workspace]/s2s-research-hub/site/three-at-the-table/data/stopwatch-window-runs.json` ·
`[workspace]/s2s-research-hub/site/three-at-the-table/data/active-weights-derivation.md` ·
`[workspace]/render-tests/{the-gpu-box-speed,the-dev-laptop-speed}_paintedreal-tiers/timing-summary.json` ·
`[workspace]/rs-runtime-dividend-2026-08-19/{AB-LEG-PREREG.md,ab-runs.json,results-window/AB-RUN-2026-08-21.md}` ·
`[workspace]/rs-qwen38-quants-2026-08-17/PREREG-AMENDMENT-3.md` (quoted from the archive write)

**Archive transcripts** (raw `.jsonl` under `[agent-archive]/`):
`65ce6e48-6b1c-4f74-835b-6cb52bd99c37` (the forensic session) and its subagent
`subagents/agent-af4ae0c48533568fe.jsonl` (the corrections) ·
`fd51c578-80b7-439c-972b-6e2886534ae1` (research-hub launch; 412, the 24.6 incident, the 08-21 probe) ·
`86119a9d-8ff4-4cc6-9651-5841aa4c630b` (Glimmer seat ruling) ·
`ee1c1cb7-2933-48e1-b914-67b9641a2b94` (Glimmer MLX-only day zero) ·
`6e8d9987-9e6e-4db1-8066-f3b1cb7bf44f` (07-24 amble bake-off)

**Method note.** `scribe search` / `scribe sessions` / `scribe status` located the
material; verbatim text was then read from the raw transcripts on disk with two
small read-only helpers in this session's scratchpad
(`reconC_extract.py`, `reconC_find.py`). A direct SQL path to the archive was
attempted and **denied by the harness classifier**; per house law that denial is
reported rather than worked around, and every finding here was obtained through
the sanctioned tools instead.
