# PRE-REGISTRATION — the robber quant ladder (written before any scored call)

Registered 2026-08-21 ~13:2X UTC, operator-blessed this morning ("we need to
keep the robber example … if there's a better way to show different
precisions, w/ receipts, let's do it — and roll it into the article").
Extends the four-arm robber mini-bench (ROBBER-PREREG.md, sha pinned in
PREREG-INDEX.txt) down and up the quantization ladder of the SAME model.
Runs beside live traffic on a quiet Friday morning (no reddit links posted
in days; the products stay up — contention disclosed, not engineered away).

## Arms — the gemma4 12B ladder (one model, every precision we can pull)
Candidates, by vendor tag: gemma4:12b-it-q2_K · gemma4:12b-it-q3_K_M ·
gemma4:12b (the resident q4_K_M seat build) · gemma4:12b-it-q5_K_M ·
gemma4:12b-it-q6_K · gemma4:12b-it-q8_0 (already local) ·
gemma4:12b-it-fp16. Availability rule: a tag the vendor library does not
carry is recorded as UNAVAILABLE (with the pull error verbatim) and is not
an arm; no substitute tags are hunted after this pin. Non-resident arms load
as visitors (keep_alive 5m), one at a time, released (keep_alive 0) before
the next load; the standing set is never unloaded; free-VRAM receipt
(including non-ollama holders) before each visitor load; /api/ps receipt at
end shows the standing four and nothing else.

## Questions (2) — verbatim the same bytes as ROBBER-PREREG.md's runner
Q1 (simple) and Q2 (complex), unchanged from robber_bench.py. No new
phrasings; the ladder varies ONLY the arm.

## Rule
n=5 scored per arm per question after 1 unscored warm-up per arm; serial,
2s pause; temperature 0, seed 0, think false, num_ctx 32768, num_predict
220 (Q1) / 320 (Q2); runtime counters only; medians reported with ranges.
The q4 and q8 arms re-run fresh this morning rather than re-quoting the
08-19 rows (same rule, new window — the two windows publish side by side,
labeled). Rows are an illustration under this rule, never a ranking.

## Predictions (falsifiable, written before any pull completes)
P1: at temperature 0 / seed 0, each arm's five replies are byte-identical
per question (determinism holds at every precision).
P2: every arm at q3_K_M and above gives the same VERDICT per question as
the 08-19 q4/q8 arms (No on Q1's main ask; Yes on Q2's move-back part) —
wording may drift, verdicts don't.
P3: the q2_K arm is where degradation becomes visible if it becomes
visible anywhere: its reply differs from the q4 reply in verdict,
coherence, or invented rules beyond the shared settlement-rule error.
P4: decode tok/s orders inversely with blob size, monotonically, within
each question.
Self-refutation honored: if P3 fails (q2_K holds the verdict), the page
says quantization survived to the bottom of this ladder on this ask — the
absence IS the finding; no threshold table is manufactured either way.

## What publishes
ladder-runs.json (all scored calls, counters, reply shas, verbatim
replies), the availability record, the pin of this file's sha alongside
ROBBER-PREREG.md's in PREREG-INDEX.txt and the ledger of ledgers, and a
page section in the-compressed-photograph rolling the replies out down the
ladder with the receipts inline.

---
APPENDED NOTE (before any pull completed or scored call ran, 2026-08-20
~12:58 UTC): the registration header above misdates itself "2026-08-21
~13:2X UTC" — a weekday-drift slip of exactly the kind the estate has been
bitten by before. The true registration and pin time is 2026-08-20 12:56:18
UTC, as the pin index records. Nothing else above this line changed; no
case, arm, question, rule, or prediction moved.
