# PRE-REGISTRATION — the robber mini-bench (written before any scored call)

Registered 2026-08-19 ~22:2X UTC, operator-blessed same hour ("if you need to
run any new benches for the robber catan stuff... let's do it"). Runs beside
live traffic (serial, paced, warm-resident-first); contention disclosed.

## Arms (4)
qwen3.8:27b (dense 27.3B, draft head as served) · gemma4:26b (MoE 25.8B/3.8B
active) — seventeen's pair, both resident · gemma4:12b (dense 11.9B Q4_K_M,
resident) · gemma4:12b-it-q8_0 (same 12B at Q8_0, 12.84 GB visitor in the
free fifth slot; fits-beside-the-set law satisfied by the pre-run free-VRAM
receipt incl. non-ollama holders) — the quant pair.

## Questions (2, verbatim in the runner)
Q1 (simple): the can-the-robber-go-back-to-the-desert question, family-night
phrasing, under 120 words.
Q2 (complex): robber on the desert on a rolled 7 — must it move; returning
later; stealing choice between two players' settlements; stealing from a
player with zero cards. Under 200 words.

## Rule
n=5 scored per arm per question after 1 unscored warm-up per arm; serial,
2s pause; temperature 0, seed 0, think false, num_ctx 32768, num_predict
220 (Q1) / 320 (Q2); runtime counters only. Residents called with
keep_alive -1; the q8 visitor 5m and released (keep_alive 0) at end with a
four-seat /api/ps receipt.

## Predictions (falsifiable)
P1: at temperature 0 / seed 0, each arm's five replies are byte-identical
per question (determinism receipt).
P2: the q8 12B decodes SLOWER than the q4 12B (more bytes per trip), with
byte-similar reply content.
P3: the MoE 26B decodes fastest of the four on both questions.
NO quality floors — replies are printed, not scored; the product's cited
answer (ruling 4615) referees the rule, not us.

## Artifacts
robber-runs.json (every call, counters, replies verbatim, shas) here;
seventeen's kit file data/robber-asks.json upgraded from its n=1
illustration to these registered rows for its two arms.
