# PREREG — the contention mini-bench: one card, both stacks, live (piece 4, "Inference vs diffusion")

*Registered 2026-08-26 ~13:4xZ by the orchestrator lane, BEFORE any scored call, per the week's prereg law.
For vision week piece 4 (releases tonight, dated 2026-08-27). The subject IS contention: the live serving
box (the 96 GB card) runs the estate's chat seat and its painter at once — this bench measures what each
stack costs the other, using only calls both stacks already serve in production. DESCRIPTIVE ROWS ONLY:
no gate, no seat verdict, no adoption decision rides on this bench — so no threshold table exists to move.*

## The window (declared)
2026-08-26, one contiguous window ~13:5xZ–15:0xZ, on the serving box, BESIDE live traffic with the live
seats UP — the quiet-hours law's beside-traffic arm, contention disclosed (it is the piece's subject).
Player impact honestly stated: chat replies may slow by the measured deltas for ~30 min. Serial legs.

## The two stacks (production shapes only — nothing pulled, nothing loaded that doesn't already serve)
- TOKENS: the resident chat model on the serving instance (the estate's 26B chat seat), /api/generate,
  stream=false, a fixed ~120-token prompt, num_predict=128, keep_alive unchanged. The runtime's own
  counters only: eval_count/eval_duration (tok/s), prompt_eval_*, load_duration.
- PIXELS: the game's own sketch tier (the shipped flux-dev sketch graph at its shipped size/steps) via
  ComfyUI's /prompt, seconds from ComfyUI's OWN execution timestamps (the render-tests rule).

## The rows (5 samples each, medians published with min–max)
R1 tokens ALONE ×5 · R2 renders ALONE ×5 · R3 tokens DURING an active render ×5 (each generate issued
while a render is mid-execution, verified by /queue) · R4 renders DURING a token stream ×5 (each render
submitted while a generate is in flight; the generate loop keeps the card busy) · R5 the RE-GRAB: the
chat model's load_duration on the first call after each render burst + free-VRAM read before/after
(the known residency alternation, measured at this posture rather than argued).

## Self-refutation checks (named before the run)
SR1: R1's five tok/s within ±10% of each other, else the box wasn't quiet — the run re-declares or the
variance publishes with the rows. SR2: R2's five seconds within ±15%, same rule. SR3: any failed render
or refused generate is EXCLUDED AND NAMED with its error verbatim (never silently resampled). SR4: if
R3/R4 show NO contention (deltas within SR1/SR2 noise), that IS the finding and publishes as such —
the piece does not get to assume drama.

## Publication
Rows land in the piece with this registration on the shelf beside it (sanitised per the house rules:
box named by VRAM class, no addresses). The piece re-quotes published token rows from the field guide
for the teaching half; THIS bench contributes only the contention and re-grab rows.

## ADDENDUM 1 — 2026-08-26 ~19:0xZ: the window as RUN differed from the window as DECLARED

*What this concerns, for the cold scroll: the registration above declared a ~13:5xZ–15:0xZ window with
"~30 min" of possible player impact.* The run actually landed 18:18:02Z–18:18:37Z — later than declared
(the first launch failed on a graph-templating defect in the harness, was re-pointed at the live graph
shape, and re-ran) — and the measured window was **35 seconds end to end**, not thirty minutes: the
per-call costs (0.3–0.6 s tokens, 1.5–1.9 s renders) were far below the padding the declaration assumed.
Player impact was therefore under a minute of possible contention, not ~30. No row, metric, or check changed;
the rows' own timestamps are the receipt. Compared to the declaration: shorter and later, nothing else.
*(Correction 2026-08-27 04:1xZ: this addendum first read 18:16:46Z and 111 seconds — the FAILED first launch's
timestamp conflated with the scored run's; the rows' own header stamps the scored run at 18:18:02Z, and its
footer at 18:18:37Z: 35 seconds. The failed launch made no scored call.)*
