# PREREG-C1 — judge trials (rulebook grounding) — FROZEN 2026-08-12 ~09:10, before any scored call

N = 21
(12 kill: contradiction 4, non-entailment 3, scope-shift 2, hedge-dropped 1, wrong-number 2 · 9 preservation)
Golden: `golden/judge-c1.json` sha256 `fede4e32154f92ac81ae68709fbde2cf1f91f3d1c76e7c83941023d986919702` (frozen from the verified draft; verify receipt: canary catch 6/6, 0 defects in the frozen 21 — verification-report.json). Scorer refuses to emit a verdict if the golden count ≠ 21.

## Floors (BY RULE, arithmetic shown — exhibit two's ratios rounded up on this population)
- kill-recall ≥ ceil(12 × 0.852) = **11/12**
- preservation ≥ ceil(9 × 0.8125) = **8/9**
- BOTH bind independently. Verdict = PASS/FAIL against floors ONLY — no ordering by margin; Wilson 95% intervals beside every rate; counts-of-N everywhere; differences under 2 cases read TIED.

## Roster (9, pre-registered order; quant published per row; cross-quant = named confound)
muse-glimmer:30b-q8_0-dflash (strength medium AND high — two rows, dial delta published) · gemma4:26b · gemma4:12b · qwen3.5:27b · qwen3.6:27b · nemotron-3.5-lightning:30b-a3b · granite4.1:30b-q8_0 · nemotron3:33b · olmo-3.1:32b-think-q4_K_M

## Protocol
temp 0, top_p 1, 3 repeats, strict-majority per case; think:false explicit for ALL (RATIONALE, published: the judge chair's production contract is think:false — a model that needs its visible trace to judge cannot sit the chair as contracted; glimmer's reasoning-strength dial rides the system prompt — REGISTERED strings, appended as the final system-prompt line for the two glimmer rows: "Reasoning strength: medium" and "Reasoning strength: high"; the full level vocabulary is low|medium|high|xhigh and no other level may be swept anywhere in this arc). Prompted schema mode; num_ctx 32768; num_predict 1024; timeout 180s. Polarity probes (6 calls) per model first, receipts published. Dual parse recorded; **strict parse binds**; leak-stripped is diagnostic only. Response-failure >10% ⇒ UNMEASURABLE (transport failures outside the ceiling, cure-by-retry). done_reason=length = own run-quality column. Kill counting: catch iff verdict != PASS (FAIL and UNCERTAIN both catch). Preservation: only PASS preserves. Self-consistency-across-repeats = named column. Outcome states: RANKED / EXPLORATORY / UNMEASURABLE / NOT-CARRIED / NOT-RUN.

## The prompt (frozen verbatim; pinned by this file's own sha in PREREG-INDEX)
SYSTEM:
You are a grounding judge for a rulebook. You will be given a CLAIM and a QUOTE from the rulebook. Decide whether the quote alone supports the claim.
Answer with ONLY a JSON object: {"verdict":"PASS"|"FAIL"|"UNCERTAIN","why":"<one sentence>"}
Rules: PASS only if the quote alone entails the claim exactly as stated. FAIL if the quote contradicts it. UNCERTAIN if the quote neither supports nor contradicts it. World knowledge must never rescue a claim the quote does not support. Numbers must match exactly.

USER:
CLAIM: {claim}
QUOTE:
{quote}

## Falsification honored
If every model clears both floors, the instrument is too easy and the article says so; if the fan's residual-risk cases (c1-k-con-s1, c1-k-bez-h1) drive systematic disagreement, they are named in the findings, not silently excused.
