# PRE-REGISTRATION — the robber quant ladder, registration #2
# (written before any scored call; supersedes the unrunnable #1)

Registered 2026-08-20 ~13:05 UTC. Registration #1's five candidate tags do
not exist in the vendor library (LADDER-AVAILABILITY.md holds the verbatim
404s); this registration names the four rungs the vendor actually ships for
this 12B, read from the library tags page BEFORE any rung produced a scored
call. Runs beside live traffic on a quiet Friday morning; products stay up;
contention disclosed, not engineered away.

## Arms — the vendor's real gemma4 12B ladder (4 rungs)
1. gemma4:12b — the resident q4_K_M seat build (post-hoc K-quant).
2. gemma4:12b-it-qat — the vendor's quantization-aware-trained 4-bit build:
   trained to be 4-bit, not compressed into it. Same headline bits as rung
   1, different recipe class — the two "4-bit"s are the point.
3. gemma4:12b-it-q8_0 — the 8-bit build (already local, re-run fresh).
4. gemma4:12b-it-bf16 — the uncompressed reference, ~24 GB.
Non-resident arms load as visitors (keep_alive 5m), ONE at a time, released
(keep_alive 0) before the next load; the standing set never unloads;
free-VRAM receipt (including non-ollama holders) before each visitor load;
/api/ps receipt at close shows the standing four and nothing else. bf16 at
~24 GB is the largest visitor and loads alone against ≥50 GiB free.

## Questions (2) — verbatim the same bytes as ROBBER-PREREG.md's runner
Q1 (simple) and Q2 (complex), unchanged from robber_bench.py. The ladder
varies ONLY the arm.

## Rule
n=5 scored per arm per question after 1 unscored warm-up per arm; serial,
2s pause; temperature 0, seed 0, think false, num_ctx 32768, num_predict
220 (Q1) / 320 (Q2); runtime counters only; medians reported with ranges.
The q4 and q8 rungs re-run fresh this morning rather than re-quoting the
08-19 rows; the two windows publish side by side, labeled. Rows are an
illustration under this rule, never a ranking.

## Predictions (falsifiable, written before any scored call)
P1: at temperature 0 / seed 0, each arm's five replies are byte-identical
per question (determinism holds at every precision, bf16 included).
P2: this morning's q4 and q8 replies are byte-identical to their own
2026-08-19 replies (same bytes, same question, same decoding — across days).
P3: bf16, the uncompressed reference, gives the same VERDICT as q4 on Q1's
main ask (the confident No with the invented settlement rule) — i.e. the
wrongness is the model's, not the compression's. If bf16 instead answers
Yes, compression changed a verdict and the page says so at full volume.
P4: the qat rung — same headline bits as q4_K_M, different recipe —
matches the q4 verdict on both questions; its exact wording is free to
differ (two different 4-bit brains need not share bytes).
P5: decode tok/s orders inversely with blob size across the four rungs,
within each question.
Self-refutation honored: whichever way P3/P4 land, the result publishes
with its receipt; no threshold table is manufactured for rungs that do not
exist (the vendor ships no q2/q3 — the ladder's bottom stays honestly
unmeasured and the where-it-stops-being-free bench stays OWED).

## What publishes
ladder-runs.json (all scored calls, counters, reply shas, verbatim
replies, the fresh-vs-08-19 byte comparison), LADDER-AVAILABILITY.md, both
registrations' shas in PREREG-INDEX.txt and the ledger of ledgers, and the
page section in the-compressed-photograph rolling the replies down the
ladder with receipts inline.
