# Window run — draft-head matrix, 2026-08-21

Summary for the window owner. Full analysis: `../matrix/RESULTS.md`.

- **Run id:** `drafthead-matrix-2×3-2026-08-19`
- **Started (UTC):** 2026-08-21T06:35:47Z
- **Ended (UTC):** 2026-08-21T07:09:30Z
- **Elapsed:** 33 m 43 s
- **Calls:** 270 scored (n=30 × 3 tiers × 3 blocks), plus 6 warm-ups, 2
  calibration probes, 1 end-state probe = 279 rows
- **Refusals:** **0** — no call errored, no REFUSAL row was written, no block
  was skipped or aborted

## Protocol sha verification

| check | value |
|---|---|
| pinned in the window brief's PREP RECEIPTS | `d4a44d904d925d5886fdbb825a8a02c0cdf549ae0622d7e63a60f383f0adda75` |
| measured before the first scored call | `d4a44d904d925d5886fdbb825a8a02c0cdf549ae0622d7e63a60f383f0adda75` |
| measured after the last scored call | `d4a44d904d925d5886fdbb825a8a02c0cdf549ae0622d7e63a60f383f0adda75` |

**Match, both times. Not amended. No REFUSED-if-amended condition arose.** The
pre-registered prediction was read and scored exactly as written; PROTOCOL.md
was not edited by this lane.

## Per-cell medians with ranges — decode tok/s

| block | `draft_num_predict` | tier | median | [min – max] | IQR | σ |
|---|---|---|---|---|---|---|
| A | 4 | ~1k | 109.13 | [101.32 – 114.14] | 106.22–111.09 | 3.04 |
| A | 4 | ~8k | 109.60 | [96.70 – 121.15] | 103.64–116.47 | 7.00 |
| A | 4 | ~32k | 94.41 | [86.05 – 103.99] | 89.93–95.86 | 3.86 |
| Z | 0 | ~1k | 76.11 | [75.89 – 76.36] | 76.00–76.28 | 0.14 |
| Z | 0 | ~8k | 74.46 | [74.27 – 74.70] | 74.37–74.63 | 0.13 |
| Z | 0 | ~32k | 68.94 | [68.78 – 69.18] | 68.89–69.04 | 0.10 |
| B | 4 | ~1k | 109.08 | [100.85 – 113.11] | 107.63–110.04 | 2.21 |
| B | 4 | ~8k | 103.50 | [95.82 – 122.57] | 100.19–110.98 | 6.65 |
| B | 4 | ~32k | 91.01 | [86.18 – 103.81] | 89.84–94.72 | 3.77 |

Prompt sizes were exact and identical across all 270 calls: 1041 / 8219 / 31015
tokens. Nonce at the prompt head on every call, so none was a cache hit.

## The registered prediction — verdict: **MIXED**

The prediction, as written and unedited:

> With the draft head disabled (`draft_num_predict 0`), the tier ordering will
> become monotonic and gently declining: 1k > 8k > 32k. With the draft head
> enabled (`draft_num_predict 4`), the 8k > 1k inversion will reappear.
> Disabling the draft head is expected to cost roughly 20–50% of decode
> throughput.

| clause | verdict | evidence |
|---|---|---|
| draft-0 becomes monotonic declining | **CONFIRMED** | 76.11 > 74.46 > 68.94; σ ≤ 0.14 tok/s per cell; bootstrap CI on 8k−1k is [−1.83, −1.45], excludes zero |
| the 8k > 1k inversion reappears at draft-4 | **NOT CONFIRMED** | pooled draft-4 is 109.08 > 105.14 > 92.44 — also monotone declining. Pooled 8k−1k = −3.61%, CI [−5.54, +1.13] spans zero. Block A alone shows +0.44%, which is a rounding error against that cell's σ of 7.00 and reverses in block B |
| disabling costs 20–50% of throughput | **CONFIRMED** | measured 30.2% / 29.2% / 25.4% by tier |

**Overall: MIXED.** Two clauses hold, the central one does not. No inversion
appeared anywhere — not in block A, block B, the pool, at n=10 or at n=30 — so
there was no anomaly in this run to attribute to the draft head. Register #8's
hypothesis is therefore **untested by this data** rather than confirmed or
refuted, and should be recorded that way.

## Bracket validity check — **DISAGREE**

| tier | block A | block B | gap | within ~3%? |
|---|---|---|---|---|
| ~1k | 109.13 | 109.08 | 0.04% | yes |
| ~8k | 109.60 | 103.50 | 5.72% | **NO** |
| ~32k | 94.41 | 91.01 | 3.67% | **NO** |

Per the brief's law, the draft-head contribution figures are therefore published
as a disagreement, not as a conclusion. `../matrix/RESULTS.md` §4 carries the
full treatment, including the bootstrap (the ~32k gap is statistically real; the
~8k gap is not), the controls showing the box itself did not drift (the draft-0
block between A and B is stable to ±0.25 tok/s, prompt-eval matches to 0.6%),
and the design finding that a ±3% tolerance cannot police an arm whose cv is
6.4%. At the pre-registered n=10 the bracket passes on all three tiers; that is
reported but explicitly **not** used to overturn the n=30 failure.

## Contribution figures (provisional — see bracket above)

| tier | draft-4 pooled | draft-0 | ratio | cost of disabling | share of the 2.95× ceiling |
|---|---|---|---|---|---|
| ~1k | 109.08 | 76.11 | 1.43× | 30.2% | 49% |
| ~8k | 105.14 | 74.46 | 1.41× | 29.2% | 48% |
| ~32k | 92.44 | 68.94 | 1.34× | 25.4% | 45% |

## The result that came out strongest

Not a pre-registered question, and the most solid thing in the run: **the draft
head multiplies run-to-run decode dispersion by 11× to 38×.** Draft-0 cells have
cv 0.15–0.19% (thirty consecutive ~8k calls spanned 0.43 tok/s over eleven
minutes); draft-4 cells have cv 2.0–6.4%. This supports the *mechanism* the
hypothesis proposed — acceptance is a property of generated text — while showing
that it expresses itself as scatter, not as a stable tier ordering.

Practical consequence: an n=10 median on a spec-decoding arm carries roughly
±2.6% standard error at ~8k. Several existing readings were taken at n=10.

## Unreproduced: C5's ~8k figure

C5 reported 133.54 tok/s at ~8k. The fastest single ~8k call in this run, across
60 draft-4 calls, was 122.57. C5's median sits **above the maximum of every 8k
measurement here** — the distributions do not overlap at that point, so this is
not sampling variation. At ~1k the two agree within 5.71% (109.08 vs 115.69).

Leading candidate, stated as a proposition and not a finding: prompt corpus.
This run's prompts are repetition-built; C5's were not, and draft acceptance
depends on the content being generated. The direct test is to replay C5's actual
corpus through the same bracketed design.

## Run-conduct receipts

- **Runner restarts: exactly 2**, as designed — into block Z (`load_duration`
  4590.3 ms) and into block B (5142.0 ms). Within blocks, every `load_duration`
  sat at 282–300 ms, the established no-restart signature.
- **End state:** `draft_num_predict` restored to the shipped default of 4 and
  verified — the 07:09:28Z end-state probe re-asserted 4 and returned
  `load_duration` **287.7 ms**, i.e. no relaunch, i.e. the runner was already at
  4. Had it been at any other value the probe would have reported 4000–5000 ms.
- **Residency:** `/api/ps` sampled at 8 points; at every one the loaded set was
  exactly the subject model, expiry 2318, and nothing else. No other model was
  ever loaded.
- **Health:** 8 front-page GETs on the live product at block boundaries — all
  HTTP 200, 0.134–0.148 s. No pause triggered.
- **Not touched:** no systemd unit, no warm timer, no daemon restart, no
  Modelfile edit, no model copy, no unnamed model. Serial calls only; runtime
  counters only (`eval_count` / `eval_duration` from the API, never wall-clock).
- **Sanitization:** every artifact in this directory and in `../matrix/` is
  box-neutral from the first byte, code included — no host, port, address, unit
  name, or hardware part. The endpoint is read from the environment at run time.

## Handoff

The matrix's end-state contract is met: draft default restored and verified,
residents untouched, results written. Per the brief, the tail (the-operator-49's
judge-panel smoke) and then the window's act-4 closing contract belong to the
window owner; this lane has not pinged the tail and has not started the close.
