[2026-08-26T06:21:46Z] ===== NIGHT BATTERY 2026-08-26 — lane log ===== [2026-08-26T06:21:46Z] clone: the bench tree from rs-30b-trials-2026-08-12 @ c30681e8 [2026-08-26T06:21:46Z] goldens verified byte-identical (judge-c1, assistant-c2, canaries, c5-roster, tasks.json) [2026-08-26T06:21:46Z] repoint: houselaw.DEFAULT_HOST -> the bench endpoint; c2_run second prod default (the bench box's production seat A) removed [2026-08-26T06:21:46Z] bench profile added to core.py; checkpoint() smoke test ok=True (bench v0.32.13, MainPID 83256, prod seat intact) [2026-08-26T06:21:46Z] pulls complete: gemma4:31b, qwen3.6:27b, nemotron-3.5-lightning:30b, laguna-xs-2.1 (bare tag) [2026-08-26T06:21:46Z] muse-glimmer:30b NOT pulled — host cannot run it (see below) [2026-08-26T06:21:46Z] BLOCKER: bench pinned to = GPU 1 = production instance B (production seat B) card [2026-08-26T06:21:46Z] GPU0 24576MiB total / 19215 used / 5361 free (prod A gemma4:26b + nomic + ComfyUI) [2026-08-26T06:21:46Z] GPU1 24576MiB total / 17467 used / 7109 free (prod B mistral-small3.2 + THE BENCH) [2026-08-26T06:21:46Z] smallest candidate qwen3.6:27b needs >=16950 MiB weights; free is 7109 MiB — short by >2x [2026-08-26T06:21:46Z] nemotron-3.5-lightning:30b 24252 MiB weights does not fit a 24576 MiB card at 32k ctx even empty [2026-08-26T06:21:46Z] instance B: 339 req/24h, ~11/hr overnight, ~5.5 min heartbeat — collision certain, prod loses [2026-08-26T06:21:46Z] DECISION: no candidate loaded. Every instrument x every candidate = NOT-RUN. Recorded, stopped. [2026-08-26T06:21:46Z] seat-43 exam LOCATED + VERIFIED runnable (the judge-exam fixture store; fixture sha 6a226b62..., 43 cases 27/16) [2026-08-26T06:21:46Z] digests: nemotron :30b == :30b-a3b (IDENTICAL, same build); qwen3.6:27b DIFFERS from 08-12 (new build) [2026-08-26T06:21:46Z] ===== lane complete — clone ready, awaiting a card ===== [2026-08-26T06:47:28Z] ===== REDEPLOYED TO the 96 GB box ===== [2026-08-26T06:47:28Z] bench 2: the 96 GB box's bench endpoint ollama 0.32.15 MainPID 71402 store ~/.ollama-bench disk 1084G [2026-08-26T06:47:28Z] houselaw -> the 96 GB box; PROD_SEATS now watches BOTH the 96 GB box ports (5 models); BENCH_SSH -> the 96 GB box [2026-08-26T06:47:28Z] the 96 GB box checkpoint smoke: ok=True, both prod seats green, journal clean [2026-08-26T06:47:28Z] @@STORE-SYNCED 00:42:30 (93G); muse-glimmer:30b @@PULL-OK 00:33:22 [2026-08-26T06:47:28Z] DIGEST VERIFY 5/5: gemma4:31b, qwen3.6:27b, nemotron-3.5-lightning:30b, laguna-xs-2.1 all MATCH the bench box; glimmer de878ce3 fresh [2026-08-26T06:47:28Z] BLOCKER 2: ComfyUI (pid 4180, idle, empty queue) holds 33018 MiB of the 96 GB box's 97887 MiB [2026-08-26T06:47:28Z] prod ollama = 47692 MiB (as briefed ~43G); free = 17139 MiB [2026-08-26T06:47:28Z] FIT TEST MEASURED qwen3.6:27b (smallest): size_vram 12037 of 17741 MiB = 67.8% on GPU, 32% SPILLED TO CPU [2026-08-26T06:47:28Z] -> every timing invalid; C5 decode rate meaningless. Candidate released, GPU back to 80748 MiB. [2026-08-26T06:47:28Z] -> prod_untouched() GREEN throughout: the load spilled itself, evicted nothing. the 96 GB box's model holds. [2026-08-26T06:47:28Z] seat-43 VERIFIED runnable: the 96 GB box NOT on egress allowlist -> ssh tunnel to 127.0.0.1:11500 (allowlisted) [2026-08-26T06:47:29Z] TRAP CAUGHT: loopback tunnel flips --courtesy auto to ON -> would fire gemma4:26b keep_alive at bench store. --courtesy off forced. [2026-08-26T06:47:29Z] AWAITING GO: release ComfyUI's 33018 MiB (idle) -> free ~50157 MiB -> full battery runs [2026-08-26T06:47:29Z] ===== lane holding; everything staged and verified ===== [2026-08-26T06:59:37Z] ===== RE-TARGETED TO the bench card (a 24 GB consumer card) — the operator ruling ===== [2026-08-26T06:59:37Z] bench: the bench endpoint (the bench box) ollama 0.32.13 MainPID 83256 the bench card [2026-08-26T06:59:37Z] 24GB-FIT DOCTRINE: fit is a PRIMARY result. FITS->full battery; SPILLS->correctness chairs, C5-32k/C7 NOT-RUN(spill) [2026-08-26T06:59:37Z] prod watch: the 96 GB box's production seat A + the 96 GB box's production seat B (live primary) + the bench box's production seat A (failover gemma presence); the bench box's production seat B deliberately unwatched (operator-emptied) [2026-08-26T06:59:37Z] the bench card free: 24558 MiB (mistral evicted by operator) [2026-08-26T06:59:37Z] DIGEST VERIFY 5/5 MATCH: gemma4:31b qwen3.6:27b nemotron-3.5-lightning:30b laguna-xs-2.1 muse-glimmer:30b [2026-08-26T06:59:37Z] seat-43 DIRECT via the bench endpoint (on the judge-exam fixture store allowlist — no tunnel); --courtesy off forced [2026-08-26T06:59:37Z] the bench box checkpoint smoke ok=True, all 3 prod seats green [2026-08-26T06:59:37Z] ===== LAUNCHING BATTERY: glimmer FIRST (probes gate seat-candidacy) ===== [2026-08-26T07:00:15Z] ===== NIGHT BATTERY start · bench=the bench endpoint · candidates: muse-glimmer:30b ===== [2026-08-26T07:00:15Z] ########## CANDIDATE muse-glimmer:30b ########## [2026-08-26T07:00:15Z] -- BEGIN fit-gate/muse-glimmer:30b == fit gate: muse-glimmer:30b @ num_ctx 32768 -> the bench endpoint == muse-glimmer:30b: FITS — 100.0% on GPU (15831/15831 MiB) -> full battery prod untouched: True -> fit-muse-glimmer_30b.json [2026-08-26T07:01:13Z] -- END fit-gate/muse-glimmer:30b: FITS — full battery [2026-08-26T07:01:13Z] -- BEGIN c0-probes/muse-glimmer:30b == format probes: muse-glimmer:30b -> the bench endpoint [probe 1/6] p1-schema-think-false-r1 PARSE-FAIL schema_enforced=False 15.70s [probe 2/6] p2-schema-think-true-r1 parsed schema_enforced=True 14.48s [probe 3/6] p3-bare-json-think-false-r1 PARSE-FAIL schema_enforced=None 4.99s [probe 4/6] p1-schema-think-false-r2 PARSE-FAIL schema_enforced=False 2.36s [probe 5/6] p2-schema-think-true-r2 parsed schema_enforced=True 12.32s [probe 6/6] p3-bare-json-think-false-r2 PARSE-FAIL schema_enforced=None 5.01s == muse-glimmer:30b: polarity=STANDARD TRAP SEAT-BLOCKED (6 calls) -> probes-muse-glimmer_30b.json [2026-08-26T07:02:08Z] -- END c0-probes/muse-glimmer:30b ok (55s) [2026-08-26T07:02:08Z] -- BEGIN c1-judge/muse-glimmer:30b === muse-glimmer:30b — legs ['C1'] === [C1 1/21] c1-p-bez-01 posture=think:false+reasoning-medium PASS pass=True [C1 2/21] c1-p-six-01 posture=think:false+reasoning-medium PASS pass=True [C1 3/21] c1-p-nap-01 posture=think:false+reasoning-medium PASS pass=True [C1 4/21] c1-p-piq-01 posture=think:false+reasoning-medium UNCERTAIN pass=False [C1 5/21] c1-p-dom-01 posture=think:false+reasoning-medium PASS pass=True [2026-08-26T07:05:23Z] chain-after: waiting for nb-glimmer.service to finish before starting: gemma4:31b laguna-xs-2.1:latest qwen3.6:27b nemotron-3.5-lightning:30b [C1 6/21] c1-p-chk-01 posture=think:false+reasoning-medium FAIL pass=False [C1 7/21] c1-p-con-01 posture=think:false+reasoning-medium PASS pass=True [C1 8/21] c1-p-spf-01 posture=think:false+reasoning-medium PASS pass=True [C1 9/21] c1-p-bkg-01 posture=think:false+reasoning-medium PASS pass=True [C1 10/21] c1-k-piq-c1 posture=think:false+reasoning-medium FAIL pass=True [C1 11/21] c1-k-bez-c1 posture=think:false+reasoning-medium FAIL pass=True [C1 12/21] c1-k-cal-c1 posture=think:false+reasoning-medium FAIL pass=True [C1 13/21] c1-k-con-c1 posture=think:false+reasoning-medium FAIL pass=True [C1 14/21] c1-k-dom-n1 posture=think:false+reasoning-medium UNCERTAIN pass=True [C1 15/21] c1-k-eca-n1 posture=think:false+reasoning-medium UNCERTAIN pass=True [C1 16/21] c1-k-spf-n1 posture=think:false+reasoning-medium UNCERTAIN pass=True [C1 17/21] c1-k-piq-s1 posture=think:false+reasoning-medium FAIL pass=True [C1 18/21] c1-k-con-s1 posture=think:false+reasoning-medium FAIL pass=True [C1 19/21] c1-k-bez-h1 posture=think:false+reasoning-medium PASS pass=False [C1 20/21] c1-k-six-w1 posture=think:false+reasoning-medium FAIL pass=True [C1 21/21] c1-k-chk-w1 posture=think:false+reasoning-medium FAIL pass=True >> C1 muse-glimmer:30b posture=think:false+reasoning-medium: EXPLORATORY — protocol fence: the contract probe failed; run anyway, unranked [C1 1/21] c1-p-bez-01 posture=think:false+reasoning-high PASS pass=True [C1 2/21] c1-p-six-01 posture=think:false+reasoning-high PASS pass=True [C1 3/21] c1-p-nap-01 posture=think:false+reasoning-high PASS pass=True [C1 4/21] c1-p-piq-01 posture=think:false+reasoning-high None pass=None [C1 5/21] c1-p-dom-01 posture=think:false+reasoning-high PASS pass=True [C1 6/21] c1-p-chk-01 posture=think:false+reasoning-high None pass=None [C1 7/21] c1-p-con-01 posture=think:false+reasoning-high PASS pass=True [C1 8/21] c1-p-spf-01 posture=think:false+reasoning-high PASS pass=True [C1 9/21] c1-p-bkg-01 posture=think:false+reasoning-high UNCERTAIN pass=False [C1 10/21] c1-k-piq-c1 posture=think:false+reasoning-high FAIL pass=True [C1 11/21] c1-k-bez-c1 posture=think:false+reasoning-high FAIL pass=True [C1 12/21] c1-k-cal-c1 posture=think:false+reasoning-high FAIL pass=True [C1 13/21] c1-k-con-c1 posture=think:false+reasoning-high FAIL pass=True [C1 14/21] c1-k-dom-n1 posture=think:false+reasoning-high UNCERTAIN pass=True [C1 15/21] c1-k-eca-n1 posture=think:false+reasoning-high UNCERTAIN pass=True [C1 16/21] c1-k-spf-n1 posture=think:false+reasoning-high UNCERTAIN pass=True [C1 17/21] c1-k-piq-s1 posture=think:false+reasoning-high FAIL pass=True [C1 18/21] c1-k-con-s1 posture=think:false+reasoning-high FAIL pass=True [C1 19/21] c1-k-bez-h1 posture=think:false+reasoning-high UNCERTAIN pass=True [C1 20/21] c1-k-six-w1 posture=think:false+reasoning-high FAIL pass=True [C1 21/21] c1-k-chk-w1 posture=think:false+reasoning-high FAIL pass=True >> C1 muse-glimmer:30b posture=think:false+reasoning-high: EXPLORATORY — protocol fence: the contract probe failed; run anyway, unranked wrote summary-C1-muse-glimmer_30b.json wrote summary-C1.json (index) [2026-08-26T07:31:54Z] -- END c1-judge/muse-glimmer:30b ok (1786s) [2026-08-26T07:31:54Z] -- BEGIN c2-assistant/muse-glimmer:30b PIN ASSERTION FAILED BEFORE THE LEG -- not calling anything. { "checked_at": "2026-08-26T07:31:54.979030Z", "model": "gemma4:26b", "legs": { "resident": false, "expiry_year": false, "context_length": false, "size_vram": false }, "ok": false, "expires_at": "", "observed": { "context_length": null, "size_vram": null }, "mainpid_leg": "NOT CHECKED HERE - host-side, owned by the leg runner", "resident_models": [] } leg C2 -- 20 items x 2 repeats sha assistant-c2.json dd63efa94bdebc878e1dc1c1065913892d8eb7cf16026642ff9fd0bae470d50c sha c2_checkers.py a9ebbd91503b936f3427cf7e6410bbacb45bb1574f750a8b45be9ae03481727a muse-glimmer:30b postures=['think:false'] total calls: 40 [2026-08-26T07:31:54Z] !! c2-assistant/muse-glimmer:30b FAILED rc=2 (0s) — recorded, battery continues [2026-08-26T07:31:54Z] -- BEGIN c3-tools/muse-glimmer:30b STOP: the seat is not as pre-registered. Not starting. !! muse-glimmer:30b is not in the C3 roster; running it think:false only C3 plan — 1 model rows, 19 tasks x 2 repeats, 10 tools. FLOOR 42 /api/chat calls (one per attempt); ESTIMATE ~90 at the 2.26 rounds/attempt measured on the offline oracle run. A model that gropes or loops costs more; the 8-round cap is the ceiling. muse-glimmer:30b postures=['think_false'] floor= 42 estimate= 90 pin FAILED — resident, context_length, size_vram, expires_year; residents=[] [2026-08-26T07:33:25Z] !! c3-tools/muse-glimmer:30b FAILED rc=3 (91s) — recorded, battery continues [2026-08-26T07:33:25Z] -- BEGIN c5-speed/muse-glimmer:30b === C5 muse-glimmer:30b === warmup-discard muse-glimmer:30b 1000tok: None tok/s (DISCARDED — not a statistic) warm muse-glimmer:30b 1000tok r1/3: 42.02 tok/s ttft~477.66ms warm muse-glimmer:30b 1000tok r2/3: 41.87 tok/s ttft~476.92ms warm muse-glimmer:30b 1000tok r3/3: 41.73 tok/s ttft~480.84ms warm muse-glimmer:30b 1000tok r1/10: 41.56 tok/s ttft~476.93ms warm muse-glimmer:30b 1000tok r2/10: 41.43 tok/s ttft~480.82ms warm muse-glimmer:30b 1000tok r3/10: 41.21 tok/s ttft~474.27ms warm muse-glimmer:30b 1000tok r4/10: 41.1 tok/s ttft~473.5ms warm muse-glimmer:30b 1000tok r5/10: 40.95 tok/s ttft~474.69ms warm muse-glimmer:30b 1000tok r6/10: 40.82 tok/s ttft~474.44ms warm muse-glimmer:30b 1000tok r7/10: 40.76 tok/s ttft~475.88ms warm muse-glimmer:30b 1000tok r8/10: 40.68 tok/s ttft~478.27ms warm muse-glimmer:30b 1000tok r9/10: 40.55 tok/s ttft~480.99ms warm muse-glimmer:30b 1000tok r10/10: 40.51 tok/s ttft~478.45ms warm muse-glimmer:30b 8000tok r1/10: None tok/s ttft~Nonems warm muse-glimmer:30b 8000tok r2/10: None tok/s ttft~Nonems warm muse-glimmer:30b 8000tok r3/10: None tok/s ttft~Nonems warm muse-glimmer:30b 8000tok r4/10: None tok/s ttft~Nonems warm muse-glimmer:30b 8000tok r5/10: None tok/s ttft~Nonems warm muse-glimmer:30b 8000tok r6/10: None tok/s ttft~Nonems warm muse-glimmer:30b 8000tok r7/10: None tok/s ttft~Nonems warm muse-glimmer:30b 8000tok r8/10: None tok/s ttft~Nonems warm muse-glimmer:30b 8000tok r9/10: None tok/s ttft~Nonems warm muse-glimmer:30b 8000tok r10/10: None tok/s ttft~Nonems warm muse-glimmer:30b 32000tok r1/10: None tok/s ttft~Nonems warm muse-glimmer:30b 32000tok r2/10: 39.3 tok/s ttft~838.13ms warm muse-glimmer:30b 32000tok r3/10: 39.35 tok/s ttft~822.15ms warm muse-glimmer:30b 32000tok r4/10: 39.34 tok/s ttft~827.26ms warm muse-glimmer:30b 32000tok r5/10: 39.4 tok/s ttft~824.47ms warm muse-glimmer:30b 32000tok r6/10: 39.52 tok/s ttft~820.75ms warm muse-glimmer:30b 32000tok r7/10: 39.6 tok/s ttft~828.11ms warm muse-glimmer:30b 32000tok r8/10: 39.55 tok/s ttft~825.37ms warm muse-glimmer:30b 32000tok r9/10: 39.56 tok/s ttft~813.85ms warm muse-glimmer:30b 32000tok r10/10: 40.14 tok/s ttft~859.9ms cold-load muse-glimmer:30b 1/3: Nones cold-load muse-glimmer:30b 2/3: Nones cold-load muse-glimmer:30b 3/3: Nones Traceback (most recent call last): File "harness/c5_speed.py", line 1234, in raise SystemExit(main()) ~~~~^^ File "harness/c5_speed.py", line 1196, in main result = run_matrix( models, args.host, manifest, shell, ...<3 lines>... run_label=args.run_label, ) File "harness/c5_speed.py", line 1008, in run_matrix "runner_argv": preflight.record_runner_argv(model, shell, manifest), ~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^ File "harness/preflight.py", line 216, in record_runner_argv manifest.put("runner_argv", tag, row) ~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^ File "harness/core.py", line 504, in put self.update(_set) ~~~~~~~~~~~^^^^^^ File "harness/core.py", line 493, in update fn(data) ~~^^^^^^ File "harness/core.py", line 503, in _set data.setdefault(section, {})[key] = value ~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^ TypeError: list indices must be integers or slices, not str [2026-08-26T07:38:42Z] !! c5-speed/muse-glimmer:30b FAILED rc=1 (317s) — recorded, battery continues [2026-08-26T07:38:42Z] -- BEGIN c7-filing/muse-glimmer:30b 8k_d10_recall0 PASS 8k_d10_recall1 PASS 8k_d10_absent0 PASS 8k_d10_absent1 PASS 8k_d50_recall0 PASS 8k_d50_recall1 PASS 8k_d50_absent0 PASS 8k_d50_absent1 PASS 8k_d90_recall0 PASS 8k_d90_recall1 PASS 8k_d90_absent0 PASS 8k_d90_absent1 PASS 16k_d10_recall0 PASS 16k_d10_recall1 PASS 16k_d10_absent0 PASS 16k_d10_absent1 PASS 16k_d50_recall0 PASS 16k_d50_recall1 PASS 16k_d50_absent0 PASS 16k_d50_absent1 PASS 16k_d90_recall0 PASS 16k_d90_recall1 PASS 16k_d90_absent0 PASS 16k_d90_absent1 PASS 32k_d10_recall0 PASS 32k_d10_recall1 PASS 32k_d10_absent0 PASS 32k_d10_absent1 PASS 32k_d50_recall0 PASS 32k_d50_recall1 PASS 32k_d50_absent0 PASS 32k_d50_absent1 PASS 32k_d90_recall0 PASS 32k_d90_recall1 PASS 32k_d90_absent0 PASS 32k_d90_absent1 PASS == C7 muse-glimmer:30b: DESCRIPTIVE · recall 18/18 · abstention 18/18 · fabrications 0 · failures 0/72 -> c7-muse-glimmer_30b.json [2026-08-26T07:49:19Z] -- END c7-filing/muse-glimmer:30b ok (637s) [2026-08-26T07:49:19Z] -- BEGIN field-exam/muse-glimmer:30b x.....x............. 20/60 ...............x.... 40/60 ...x........x....... 60/60 ============================================================ model: muse-glimmer:30b grounded-qa 18/20 90.0% stay-grounded 18/20 90.0% schema-extract 19/20 95.0% OVERALL 55/60 91.7% stay-grounded failure modes: no_decline=2 wrote field/muse-glimmer-30b.json [2026-08-26T07:51:50Z] -- END field-exam/muse-glimmer:30b ok (151s) [2026-08-26T07:51:50Z] -- BEGIN seat-43/muse-glimmer:30b runid nightbattery-muse-glimmer_30b base-url the bench endpoint results raw/seat43/judges-nightbattery-muse-glimmer_30b.jsonl models muse-glimmer:30b cases 43 repeats 3 calls 129 (before resume skips) schema-mode prompted (no `format` is sent) options {"num_ctx": 16384, "num_predict": 1024, "temperature": 0.0, "top_p": 1.0} num_ctx 16384 keep_alive 0 on candidates think omit courtesy off auth no key sent prompt packs-claim-judge-v2-prompted sha256 92af6030dddface2… (the judge-exam fixture store/judge-prompt-packs-v2-prompted.md) [probe] muse-glimmer:30b: think_honored=None options_honored=True shape_ok=True [muse-glimmer:30b] skagway-arctic-brotherhood-hall-04 r1 -> FAIL (19964 ms) [muse-glimmer:30b] skagway-centennial-snowplow-02 r1 -> FAIL (18338 ms) [muse-glimmer:30b] skagway-golden-north-hotel-02 r1 -> FAIL (23181 ms) [muse-glimmer:30b] skagway-golden-north-hotel-03 r1 -> FAIL (25877 ms) [muse-glimmer:30b] skagway-golden-north-hotel-04 r1 -> FAIL (18549 ms) [muse-glimmer:30b] skagway-golden-north-hotel-05 r1 -> PASS (34705 ms) [muse-glimmer:30b] skagway-jeff-smiths-parlor-02 r1 -> FAIL (18923 ms) [measurement-failure] muse-glimmer:30b skagway-kirmses-curios-02 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [muse-glimmer:30b] skagway-kirmses-curios-03 r1 -> FAIL (20741 ms) [muse-glimmer:30b] skagway-kirmses-curios-04 r1 -> FAIL (20425 ms) [muse-glimmer:30b] skagway-kirmses-curios-05 r1 -> PASS (35251 ms) [muse-glimmer:30b] skagway-mascot-saloon-06 r1 -> FAIL (24142 ms) [muse-glimmer:30b] skagway-mccabe-college-01 r1 -> PASS (27036 ms) [muse-glimmer:30b] skagway-mccabe-college-02 r1 -> FAIL (20230 ms) [measurement-failure] muse-glimmer:30b skagway-mccabe-college-03 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [muse-glimmer:30b] skagway-mccabe-college-04 r1 -> FAIL (25241 ms) [muse-glimmer:30b] skagway-mccabe-college-05 r1 -> FAIL (17551 ms) [muse-glimmer:30b] skagway-mccabe-college-06 r1 -> PASS (24678 ms) [measurement-failure] muse-glimmer:30b skagway-mccabe-college-07 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [muse-glimmer:30b] skagway-mollie-walsh-park-01 r1 -> PASS (22628 ms) [measurement-failure] muse-glimmer:30b skagway-mollie-walsh-park-02 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [muse-glimmer:30b] skagway-mollie-walsh-park-03 r1 -> FAIL (25962 ms) [muse-glimmer:30b] skagway-mollie-walsh-park-05 r1 -> FAIL (25858 ms) [measurement-failure] muse-glimmer:30b skagway-mollie-walsh-park-06 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [muse-glimmer:30b] skagway-mollie-walsh-park-07 r1 -> FAIL (22929 ms) [muse-glimmer:30b] skagway-moore-homestead-01 r1 -> FAIL (25181 ms) [muse-glimmer:30b] skagway-moore-homestead-02 r1 -> FAIL (21859 ms) [muse-glimmer:30b] skagway-moore-homestead-03 r1 -> FAIL (23236 ms) [muse-glimmer:30b] skagway-moore-homestead-05 r1 -> FAIL (34151 ms) [muse-glimmer:30b] skagway-pantheon-saloon-01 r1 -> FAIL (30231 ms) [muse-glimmer:30b] skagway-pantheon-saloon-04 r1 -> FAIL (25307 ms) [muse-glimmer:30b] skagway-pantheon-saloon-06 r1 -> FAIL (23831 ms) [muse-glimmer:30b] skagway-pullen-creek-harbor-02 r1 -> FAIL (25050 ms) [muse-glimmer:30b] skagway-pullen-creek-harbor-04 r1 -> FAIL (31613 ms) [muse-glimmer:30b] skagway-pullen-creek-harbor-08 r1 -> FAIL (28386 ms) [muse-glimmer:30b] skagway-red-onion-saloon-03 r1 -> FAIL (20832 ms) [muse-glimmer:30b] skagway-red-onion-saloon-04 r1 -> PASS (30912 ms) [muse-glimmer:30b] skagway-red-onion-saloon-05 r1 -> FAIL (24345 ms) [muse-glimmer:30b] skagway-red-onion-saloon-06 r1 -> FAIL (24081 ms) [muse-glimmer:30b] skagway-ship-registry-cliff-01 r1 -> FAIL (23090 ms) [measurement-failure] muse-glimmer:30b skagway-ship-registry-cliff-05 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [muse-glimmer:30b] skagway-skagway-context-02 r1 -> FAIL (24031 ms) [muse-glimmer:30b] skagway-wpyr-depot-08 r1 -> FAIL (22981 ms) [muse-glimmer:30b] skagway-arctic-brotherhood-hall-04 r2 -> FAIL (19596 ms) [muse-glimmer:30b] skagway-centennial-snowplow-02 r2 -> FAIL (18154 ms) [muse-glimmer:30b] skagway-golden-north-hotel-02 r2 -> FAIL (23102 ms) [muse-glimmer:30b] skagway-golden-north-hotel-03 r2 -> FAIL (25654 ms) [muse-glimmer:30b] skagway-golden-north-hotel-04 r2 -> FAIL (18443 ms) [muse-glimmer:30b] skagway-golden-north-hotel-05 r2 -> PASS (34713 ms) [muse-glimmer:30b] skagway-jeff-smiths-parlor-02 r2 -> FAIL (18921 ms) [measurement-failure] muse-glimmer:30b skagway-kirmses-curios-02 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [muse-glimmer:30b] skagway-kirmses-curios-03 r2 -> FAIL (20319 ms) [muse-glimmer:30b] skagway-kirmses-curios-04 r2 -> FAIL (20690 ms) [muse-glimmer:30b] skagway-kirmses-curios-05 r2 -> PASS (35310 ms) [muse-glimmer:30b] skagway-mascot-saloon-06 r2 -> FAIL (23642 ms) [muse-glimmer:30b] skagway-mccabe-college-01 r2 -> PASS (26966 ms) [muse-glimmer:30b] skagway-mccabe-college-02 r2 -> FAIL (19706 ms) [measurement-failure] muse-glimmer:30b skagway-mccabe-college-03 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [muse-glimmer:30b] skagway-mccabe-college-04 r2 -> FAIL (25315 ms) [muse-glimmer:30b] skagway-mccabe-college-05 r2 -> FAIL (17557 ms) [muse-glimmer:30b] skagway-mccabe-college-06 r2 -> PASS (24782 ms) [measurement-failure] muse-glimmer:30b skagway-mccabe-college-07 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [muse-glimmer:30b] skagway-mollie-walsh-park-01 r2 -> PASS (22612 ms) [measurement-failure] muse-glimmer:30b skagway-mollie-walsh-park-02 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [muse-glimmer:30b] skagway-mollie-walsh-park-03 r2 -> FAIL (26209 ms) [muse-glimmer:30b] skagway-mollie-walsh-park-05 r2 -> FAIL (25877 ms) [measurement-failure] muse-glimmer:30b skagway-mollie-walsh-park-06 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [muse-glimmer:30b] skagway-mollie-walsh-park-07 r2 -> FAIL (23122 ms) [muse-glimmer:30b] skagway-moore-homestead-01 r2 -> FAIL (25132 ms) [muse-glimmer:30b] skagway-moore-homestead-02 r2 -> FAIL (21817 ms) [muse-glimmer:30b] skagway-moore-homestead-03 r2 -> FAIL (23228 ms) [muse-glimmer:30b] skagway-moore-homestead-05 r2 -> FAIL (33931 ms) [muse-glimmer:30b] skagway-pantheon-saloon-01 r2 -> FAIL (29985 ms) [muse-glimmer:30b] skagway-pantheon-saloon-04 r2 -> FAIL (25310 ms) [muse-glimmer:30b] skagway-pantheon-saloon-06 r2 -> FAIL (23934 ms) [muse-glimmer:30b] skagway-pullen-creek-harbor-02 r2 -> FAIL (24781 ms) [muse-glimmer:30b] skagway-pullen-creek-harbor-04 r2 -> FAIL (31735 ms) [muse-glimmer:30b] skagway-pullen-creek-harbor-08 r2 -> FAIL (28280 ms) [muse-glimmer:30b] skagway-red-onion-saloon-03 r2 -> FAIL (21100 ms) [muse-glimmer:30b] skagway-red-onion-saloon-04 r2 -> PASS (30442 ms) [muse-glimmer:30b] skagway-red-onion-saloon-05 r2 -> FAIL (24316 ms) [muse-glimmer:30b] skagway-red-onion-saloon-06 r2 -> FAIL (24370 ms) [muse-glimmer:30b] skagway-ship-registry-cliff-01 r2 -> FAIL (22548 ms) [measurement-failure] muse-glimmer:30b skagway-ship-registry-cliff-05 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [muse-glimmer:30b] skagway-skagway-context-02 r2 -> FAIL (23860 ms) [muse-glimmer:30b] skagway-wpyr-depot-08 r2 -> FAIL (22987 ms) [muse-glimmer:30b] skagway-arctic-brotherhood-hall-04 r3 -> FAIL (20085 ms) [muse-glimmer:30b] skagway-centennial-snowplow-02 r3 -> FAIL (18420 ms) [muse-glimmer:30b] skagway-golden-north-hotel-02 r3 -> FAIL (23282 ms) [muse-glimmer:30b] skagway-golden-north-hotel-03 r3 -> FAIL (26419 ms) [muse-glimmer:30b] skagway-golden-north-hotel-04 r3 -> FAIL (17198 ms) [muse-glimmer:30b] skagway-golden-north-hotel-05 r3 -> PASS (33767 ms) [muse-glimmer:30b] skagway-jeff-smiths-parlor-02 r3 -> FAIL (18977 ms) [measurement-failure] muse-glimmer:30b skagway-kirmses-curios-02 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [muse-glimmer:30b] skagway-kirmses-curios-03 r3 -> FAIL (19861 ms) [muse-glimmer:30b] skagway-kirmses-curios-04 r3 -> FAIL (20911 ms) [muse-glimmer:30b] skagway-kirmses-curios-05 r3 -> PASS (35004 ms) [muse-glimmer:30b] skagway-mascot-saloon-06 r3 -> FAIL (24154 ms) [muse-glimmer:30b] skagway-mccabe-college-01 r3 -> PASS (27080 ms) [muse-glimmer:30b] skagway-mccabe-college-02 r3 -> FAIL (20162 ms) [measurement-failure] muse-glimmer:30b skagway-mccabe-college-03 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [muse-glimmer:30b] skagway-mccabe-college-04 r3 -> FAIL (25239 ms) [muse-glimmer:30b] skagway-mccabe-college-05 r3 -> FAIL (17523 ms) [muse-glimmer:30b] skagway-mccabe-college-06 r3 -> PASS (24774 ms) [measurement-failure] muse-glimmer:30b skagway-mccabe-college-07 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [muse-glimmer:30b] skagway-mollie-walsh-park-01 r3 -> PASS (22647 ms) [measurement-failure] muse-glimmer:30b skagway-mollie-walsh-park-02 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [muse-glimmer:30b] skagway-mollie-walsh-park-03 r3 -> FAIL (26252 ms) [muse-glimmer:30b] skagway-mollie-walsh-park-05 r3 -> FAIL (25933 ms) [measurement-failure] muse-glimmer:30b skagway-mollie-walsh-park-06 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [muse-glimmer:30b] skagway-mollie-walsh-park-07 r3 -> FAIL (23034 ms) [muse-glimmer:30b] skagway-moore-homestead-01 r3 -> FAIL (25173 ms) [muse-glimmer:30b] skagway-moore-homestead-02 r3 -> FAIL (21806 ms) [muse-glimmer:30b] skagway-moore-homestead-03 r3 -> FAIL (23152 ms) [muse-glimmer:30b] skagway-moore-homestead-05 r3 -> FAIL (33589 ms) [muse-glimmer:30b] skagway-pantheon-saloon-01 r3 -> FAIL (30189 ms) [muse-glimmer:30b] skagway-pantheon-saloon-04 r3 -> FAIL (25122 ms) [muse-glimmer:30b] skagway-pantheon-saloon-06 r3 -> FAIL (23586 ms) [muse-glimmer:30b] skagway-pullen-creek-harbor-02 r3 -> FAIL (24653 ms) [muse-glimmer:30b] skagway-pullen-creek-harbor-04 r3 -> FAIL (31783 ms) [muse-glimmer:30b] skagway-pullen-creek-harbor-08 r3 -> FAIL (28336 ms) [muse-glimmer:30b] skagway-red-onion-saloon-03 r3 -> FAIL (21100 ms) [muse-glimmer:30b] skagway-red-onion-saloon-04 r3 -> PASS (30946 ms) [muse-glimmer:30b] skagway-red-onion-saloon-05 r3 -> FAIL (24280 ms) [muse-glimmer:30b] skagway-red-onion-saloon-06 r3 -> FAIL (24328 ms) [muse-glimmer:30b] skagway-ship-registry-cliff-01 r3 -> FAIL (23043 ms) [measurement-failure] muse-glimmer:30b skagway-ship-registry-cliff-05 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [muse-glimmer:30b] skagway-skagway-context-02 r3 -> FAIL (23942 ms) [muse-glimmer:30b] skagway-wpyr-depot-08 r3 -> FAIL (23027 ms) calls 129 · resumed-skips 0 · measurement failures 18 models run: muse-glimmer:30b models unavailable: (none) models NOT RUN (probe): (none) models aborted mid-batch: (none) score it: python3 bench/score_judges.py --results raw/seat43/judges-nightbattery-muse-glimmer_30b.jsonl [2026-08-26T08:48:49Z] -- END seat-43/muse-glimmer:30b ok (3419s) [2026-08-26T08:48:49Z] -- BEGIN seat-43-score/muse-glimmer:30b # Judge bake-off — judges-nightbattery-muse-glimmer_30b.jsonl Computed by `bench/score_judges.py` from the append-only rows in `raw/seat43/judges-nightbattery-muse-glimmer_30b.jsonl`. Every number below is a count of rows in that file; the per-case appendix prints the rows themselves so the table can be recomputed rather than believed. ## What was measured | | | | --- | --- | | generated | 2026-08-26T08:48:49Z | | results | `raw/seat43/judges-nightbattery-muse-glimmer_30b.jsonl` (133 rows) | | fixture | `the judge-exam fixture store/judge-cases-v1.json` — tour-judge-v1 | | fixture sha256 | `6a226b62dceb005dbd658bc13cf229e38f6f0a2f07ae51091f444ec3e96b56d1` | | thresholds | `the judge-exam fixture store/THRESHOLDS.md` sha256 `5391a9d465658d66fa33d57c14aff7eb0aafaa3b41a9c0a2a7cb33aa973ec3d2` | | runid | `nightbattery-muse-glimmer_30b` | | runner | d-runner-v2 on python 3.14.4 | | base-url | `the bench endpoint` | | prompt | packs-claim-judge-v2-prompted sha256 `92af6030dddface2c863fe9d672bec6a6bb1017122c0028e435214acd9fb298d` | | schema mode | prompted | | repeats | 3 | | options | `{"num_ctx": 16384, "num_predict": 1024, "temperature": 0.0, "top_p": 1.0}` | | num_ctx | 16384 | | keep_alive | 0 | | think | omit | | auth | no key sent | | courtesy | off | | superseded rows | 0 (retries; last row per triple scored) | ## The floors, as pre-registered Read from `the judge-exam fixture store/THRESHOLDS.md` at scoring time, not retyped. The lines themselves: > Minimum kill-recall — **23 of 27** (0.852) > Minimum confirmed-preservation — **13 of 16** (0.8125) > Latency ceiling — **p50 ≤ 10 s, p95 ≤ 20 s per case** **Both bind independently** (§3.3): a candidate under either floor is out regardless of the other. A candidate exactly one case short on exactly one floor is a RE-TEST at 9 repeats, not an elimination (§6); short on both is an elimination with no re-test. Over 10% measurement failures a candidate is UNMEASURABLE at this config and is not scored at all. ## Candidates **Latency is reported and NON-BINDING in this run** (PREREG-D-RUN-V4 §4): cloud responses carry no `load_duration`, so warm-equivalent is uncomputable and a ceiling could only be applied to WAN wall-clock or to server-total. No seat below turns on it, and the latency column reads `—` rather than a verdict it cannot make. | model | kills /27 | preservation /16 | kill floor | pres. floor | wall p50 | wall p95 | latency | resp. fail | transport | fenced | verdict | | --- | ---: | ---: | :--: | :--: | ---: | ---: | :--: | ---: | ---: | ---: | :--: | | `muse-glimmer:30b` | 27 | 6 | PASS | FAIL | 24081 ms | 34151 ms | — | 18/129 | 0/129 | 0 | **UNMEASURABLE** | - `muse-glimmer:30b` — 18/129 calls (14.0%) were RESPONSE failures, over the 10% limit — reported UNMEASURABLE at this config, not scored. ### Pre-run probes Every probe sent the EXACT contracted envelope and differed only in its message. The raw probe rows are in this results file (`note: probe`). | candidate | think honored | options honored | shape sighting | probes skipped (resume) | | --- | :--: | :--: | :--: | :--: | | `muse-glimmer:30b` | None | True | True | False | **Candidates are listed alphabetically and are NOT ranked.** THRESHOLDS §3.5's five-step tie-break is not implemented here, and neither is its mixed-basis rule (tie-breaks 1–2 on the 9-repeat majority for a re-tested candidate, 3–5 on the first 3 repeats for everyone). If more than one candidate clears both floors and the latency ceiling, **apply §3.5 by hand** — the order of the rows above carries no meaning beyond the alphabet. A candidate under either floor is out under §3.3 regardless of the other, so nothing here is a close-second either. Its inputs, computed but deliberately not combined — steps 1 and 2 are kill-recall then preservation, 3 is self-consistency, 4 is p95 latency, 5 is resident VRAM, which this instrument does not measure: | candidate | 1. kill-recall | 2. preservation | 3. self-consistency | 4. p95 | 5. resident VRAM | | --- | ---: | ---: | ---: | ---: | ---: | | `muse-glimmer:30b` | 27/27 (1.000) | 6/16 (0.375) | 37/37 | 34151 ms | — (`bench/candidates.py`) | ## `muse-glimmer:30b` 129 triples scored — one per `(case, repeat)` — with 18 RESPONSE failures (14.0%) and 111 latencies from answered calls. **Run completeness.** 129 usable rows of the 129 this run-start asked for; 0 transport failures (connection / timeout / HTTP status). Transport is the LINK's figure and never the candidate's: it is excluded from the 10% ceiling above, and while any of it is uncured the candidate is NOT SCORED rather than judged on whatever arrived. Cure: `--retry transport` under the same runid. **Latency, three figures, none substituted for another.** ollama reports durations in nanoseconds; every ms below is that value divided by 1000000 (floor). `warm` is `total_duration − load_duration` and is computable ONLY where the response carries `load_duration` — cloud responses do not, so a cloud candidate's warm column is empty rather than a wall figure wearing a warm label. | figure | what it measures | p50 | p95 | n | | --- | --- | ---: | ---: | ---: | | `wall_ms` | perf_counter around the call, link included | 24081 ms | 34151 ms | 111 | | `server_total_ms` | `total_duration`, server-side, excludes the WAN | 24079 ms | 34148 ms | 111 | | `warm_ms` | `total_duration − load_duration`, local rows only | 14090 ms | 24610 ms | 111 | | axis | count | of key | rate | floor | verdict | | --- | ---: | ---: | ---: | ---: | :--: | | kill catches | 27 | 27 | 1.000 | 23 | PASS | | preserved | 6 | 16 | 0.375 | 13 | FAIL | Kill half: 27 catches, 0 misses, 0 ties, 0 unmeasured. Preservation half: 6 preserved, 4 lost, 0 ties, 6 unmeasured. ### By kill class The sort THRESHOLDS §2.1 derives the kill floor from. `non-entailment` is the row that matters: refusing on silence rather than on conflict. | kill class | n | catches | misses | ties | unmeasured | recall | | --- | ---: | ---: | ---: | ---: | ---: | ---: | | blatant-contradiction | 12 | 12 | 0 | 0 | 0 | 1.000 | | degenerate-evidence | 1 | 1 | 0 | 0 | 0 | 1.000 | | hedge-dropped | 3 | 3 | 0 | 0 | 0 | 1.000 | | non-entailment | 8 | 8 | 0 | 0 | 0 | 1.000 | | scope-shift | 3 | 3 | 0 | 0 | 0 | 1.000 | ### Repeats Self-consistency 37/37 cases unanimous across their usable repeats (1.000). No case flipped verdict across repeats. ### Measurement failures 18 of 129 calls (14.0%); the limit is 10%. None of these were mapped to a verdict — on a kill case `UNCERTAIN` counts as a catch, so a truncation scored as `UNCERTAIN` would read as 1.00 kill-recall and 0.00 preservation for a model that answered nothing. | failure kind | calls | | --- | ---: | | `done_reason` | 18 | | case | repeats failed | kinds | | --- | ---: | --- | | `skagway-kirmses-curios-02` | 3 | done_reason | | `skagway-mccabe-college-03` | 3 | done_reason | | `skagway-mccabe-college-07` | 3 | done_reason | | `skagway-mollie-walsh-park-02` | 3 | done_reason | | `skagway-mollie-walsh-park-06` | 3 | done_reason | | `skagway-ship-registry-cliff-05` | 3 | done_reason | 3 answered calls carried a `why` longer than the contract's `maxLength`. Recorded, not counted as failures: a verdict is not less of a verdict for a long rationale, and discarding those would shrink the key silently. ### Every case, every repeat | case | expected | kill class | repeats | successes | outcome | latencies | | --- | --- | --- | --- | ---: | --- | --- | | `skagway-arctic-brotherhood-hall-04` | NON_PASS | hedge-dropped | FAIL, FAIL, FAIL | 3/3 | success | 19964, 19596, 20085 | | `skagway-centennial-snowplow-02` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 18338, 18154, 18420 | | `skagway-golden-north-hotel-02` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 23181, 23102, 23282 | | `skagway-golden-north-hotel-03` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 25877, 25654, 26419 | | `skagway-golden-north-hotel-04` | NON_PASS | non-entailment | FAIL, FAIL, FAIL | 3/3 | success | 18549, 18443, 17198 | | `skagway-golden-north-hotel-05` | PASS | — | PASS, PASS, PASS | 3/3 | success | 34705, 34713, 33767 | | `skagway-jeff-smiths-parlor-02` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 18923, 18921, 18977 | | `skagway-kirmses-curios-02` | PASS | — | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-kirmses-curios-03` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 20741, 20319, 19861 | | `skagway-kirmses-curios-04` | NON_PASS | non-entailment | FAIL, FAIL, FAIL | 3/3 | success | 20425, 20690, 20911 | | `skagway-kirmses-curios-05` | PASS | — | PASS, PASS, PASS | 3/3 | success | 35251, 35310, 35004 | | `skagway-mascot-saloon-06` | NON_PASS | degenerate-evidence | FAIL, FAIL, FAIL | 3/3 | success | 24142, 23642, 24154 | | `skagway-mccabe-college-01` | PASS | — | PASS, PASS, PASS | 3/3 | success | 27036, 26966, 27080 | | `skagway-mccabe-college-02` | NON_PASS | scope-shift | FAIL, FAIL, FAIL | 3/3 | success | 20230, 19706, 20162 | | `skagway-mccabe-college-03` | PASS | — | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-mccabe-college-04` | NON_PASS | hedge-dropped | FAIL, FAIL, FAIL | 3/3 | success | 25241, 25315, 25239 | | `skagway-mccabe-college-05` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 17551, 17557, 17523 | | `skagway-mccabe-college-06` | PASS | — | PASS, PASS, PASS | 3/3 | success | 24678, 24782, 24774 | | `skagway-mccabe-college-07` | PASS | — | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-mollie-walsh-park-01` | PASS | — | PASS, PASS, PASS | 3/3 | success | 22628, 22612, 22647 | | `skagway-mollie-walsh-park-02` | PASS | — | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-mollie-walsh-park-03` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 25962, 26209, 26252 | | `skagway-mollie-walsh-park-05` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 25858, 25877, 25933 | | `skagway-mollie-walsh-park-06` | PASS | — | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-mollie-walsh-park-07` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 22929, 23122, 23034 | | `skagway-moore-homestead-01` | PASS | — | FAIL, FAIL, FAIL | 0/3 | failure | 25181, 25132, 25173 | | `skagway-moore-homestead-02` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 21859, 21817, 21806 | | `skagway-moore-homestead-03` | NON_PASS | non-entailment | FAIL, FAIL, FAIL | 3/3 | success | 23236, 23228, 23152 | | `skagway-moore-homestead-05` | PASS | — | FAIL, FAIL, FAIL | 0/3 | failure | 34151, 33931, 33589 | | `skagway-pantheon-saloon-01` | PASS | — | FAIL, FAIL, FAIL | 0/3 | failure | 30231, 29985, 30189 | | `skagway-pantheon-saloon-04` | NON_PASS | non-entailment | FAIL, FAIL, FAIL | 3/3 | success | 25307, 25310, 25122 | | `skagway-pantheon-saloon-06` | NON_PASS | non-entailment | FAIL, FAIL, FAIL | 3/3 | success | 23831, 23934, 23586 | | `skagway-pullen-creek-harbor-02` | NON_PASS | scope-shift | FAIL, FAIL, FAIL | 3/3 | success | 25050, 24781, 24653 | | `skagway-pullen-creek-harbor-04` | PASS | — | FAIL, FAIL, FAIL | 0/3 | failure | 31613, 31735, 31783 | | `skagway-pullen-creek-harbor-08` | NON_PASS | non-entailment | FAIL, FAIL, FAIL | 3/3 | success | 28386, 28280, 28336 | | `skagway-red-onion-saloon-03` | NON_PASS | non-entailment | FAIL, FAIL, FAIL | 3/3 | success | 20832, 21100, 21100 | | `skagway-red-onion-saloon-04` | PASS | — | PASS, PASS, PASS | 3/3 | success | 30912, 30442, 30946 | | `skagway-red-onion-saloon-05` | NON_PASS | non-entailment | FAIL, FAIL, FAIL | 3/3 | success | 24345, 24316, 24280 | | `skagway-red-onion-saloon-06` | NON_PASS | hedge-dropped | FAIL, FAIL, FAIL | 3/3 | success | 24081, 24370, 24328 | | `skagway-ship-registry-cliff-01` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 23090, 22548, 23043 | | `skagway-ship-registry-cliff-05` | PASS | — | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-skagway-context-02` | NON_PASS | scope-shift | FAIL, FAIL, FAIL | 3/3 | success | 24031, 23860, 23942 | | `skagway-wpyr-depot-08` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 22981, 22987, 23027 | ## Run notes - `probe` — muse-glimmer:30b: reachable - `probe` — muse-glimmer:30b: num_predict:10 -> eval_count=10, done_reason='length' - `probe` — muse-glimmer:30b: verdict='FAIL' fence_stripped=False thinking_chars=1698 ## Recompute it ``` python3 bench/score_judges.py --results raw/seat43/judges-nightbattery-muse-glimmer_30b.jsonl ``` [2026-08-26T08:48:49Z] -- END seat-43-score/muse-glimmer:30b ok (0s) [2026-08-26T08:48:49Z] ########## CANDIDATE muse-glimmer:30b COMPLETE ########## [2026-08-26T08:48:49Z] ===== NIGHT BATTERY done ===== [2026-08-26T08:48:54Z] chain-after: nb-glimmer.service is done — starting battery on: gemma4:31b laguna-xs-2.1:latest qwen3.6:27b nemotron-3.5-lightning:30b [2026-08-26T08:48:54Z] ===== NIGHT BATTERY start · bench=the bench endpoint · candidates: gemma4:31b laguna-xs-2.1:latest qwen3.6:27b nemotron-3.5-lightning:30b ===== [2026-08-26T08:48:54Z] ########## CANDIDATE gemma4:31b ########## [2026-08-26T08:48:55Z] -- BEGIN fit-gate/gemma4:31b == fit gate: gemma4:31b @ num_ctx 32768 -> the bench endpoint == gemma4:31b: SPILLS — only 92.5% on GPU (1541 MiB to CPU) -> correctness chairs only, C5-32k/C7 NOT-RUN(spill) prod untouched: True -> fit-gemma4_31b.json [2026-08-26T08:49:44Z] -- END fit-gate/gemma4:31b: SPILLS — correctness chairs only, C5-32k/C7 NOT-RUN(spill) [2026-08-26T08:49:44Z] -- BEGIN c0-probes/gemma4:31b == format probes: gemma4:31b -> the bench endpoint [probe 1/6] p1-schema-think-false-r1 parsed schema_enforced=True 21.85s [probe 2/6] p2-schema-think-true-r1 parsed schema_enforced=True 15.30s [probe 3/6] p3-bare-json-think-false-r1 parsed schema_enforced=None 5.68s [probe 4/6] p1-schema-think-false-r2 parsed schema_enforced=True 4.68s [probe 5/6] p2-schema-think-true-r2 parsed schema_enforced=True 15.92s [probe 6/6] p3-bare-json-think-false-r2 parsed schema_enforced=None 5.70s == gemma4:31b: polarity=SCHEMA HELD BOTH WAYS schema-ok (6 calls) -> probes-gemma4_31b.json [2026-08-26T08:50:53Z] -- END c0-probes/gemma4:31b ok (69s) [2026-08-26T08:50:53Z] -- BEGIN c1-judge/gemma4:31b === gemma4:31b — legs ['C1'] === [C1 1/21] c1-p-bez-01 posture=think:false None pass=None [C1 2/21] c1-p-six-01 posture=think:false None pass=None [C1 3/21] c1-p-nap-01 posture=think:false None pass=None [C1 4/21] c1-p-piq-01 posture=think:false None pass=None [C1 5/21] c1-p-dom-01 posture=think:false None pass=None [C1 6/21] c1-p-chk-01 posture=think:false None pass=None [C1 7/21] c1-p-con-01 posture=think:false None pass=None [C1 8/21] c1-p-spf-01 posture=think:false None pass=None [C1 9/21] c1-p-bkg-01 posture=think:false None pass=None [C1 10/21] c1-k-piq-c1 posture=think:false None pass=None [C1 11/21] c1-k-bez-c1 posture=think:false None pass=None [C1 12/21] c1-k-cal-c1 posture=think:false None pass=None [C1 13/21] c1-k-con-c1 posture=think:false None pass=None [C1 14/21] c1-k-dom-n1 posture=think:false None pass=None [C1 15/21] c1-k-eca-n1 posture=think:false None pass=None [C1 16/21] c1-k-spf-n1 posture=think:false None pass=None [C1 17/21] c1-k-piq-s1 posture=think:false None pass=None [C1 18/21] c1-k-con-s1 posture=think:false None pass=None [C1 19/21] c1-k-bez-h1 posture=think:false None pass=None [C1 20/21] c1-k-six-w1 posture=think:false None pass=None [C1 21/21] c1-k-chk-w1 posture=think:false None pass=None >> C1 gemma4:31b posture=think:false: NOT-CARRIED — 0 valid attempts across 63 call(s); raw emission published beside this row wrote summary-C1-gemma4_31b.json wrote summary-C1.json (index) [2026-08-26T08:54:11Z] -- END c1-judge/gemma4:31b ok (198s) [2026-08-26T08:54:11Z] -- BEGIN c2-assistant/gemma4:31b leg C2 -- 20 items x 2 repeats sha assistant-c2.json dd63efa94bdebc878e1dc1c1065913892d8eb7cf16026642ff9fd0bae470d50c sha c2_checkers.py a9ebbd91503b936f3427cf7e6410bbacb45bb1574f750a8b45be9ae03481727a gemma4:31b postures=['think:false'] total calls: 40 === gemma4:31b [think:false] === r1 [ 1/20] c2-cw-1 ok 12.7s chars= 80 r1 [ 2/20] c2-cw-2 ok 1.9s chars= 163 r1 [ 3/20] c2-cw-3 ok 1.3s chars= 66 r1 [ 4/20] c2-cw-4 ok 1.1s chars= 44 r1 [ 5/20] c2-sum-1 ok 3.5s chars= 316 r1 [ 6/20] c2-sum-2 ok 4.7s chars= 347 r1 [ 7/20] c2-json-1 ok 3.7s chars= 98 r1 [ 8/20] c2-json-2 ok 5.2s chars= 180 r1 [ 9/20] c2-num-1 ok 10.9s chars= 735 r1 [10/20] c2-num-2 ok 13.3s chars= 960 r1 [11/20] c2-num-3 ok 10.3s chars= 679 r1 [12/20] c2-tone-1 ok 4.1s chars= 370 r1 [13/20] c2-sh-1 ok 1.5s chars= 45 r1 [14/20] c2-re-1 ok 1.5s chars= 26 r1 [15/20] c2-ctx-8k ok 11.0s chars= 71 r1 [16/20] c2-ctx-24k ok 32.6s chars= 84 r1 [17/20] c2-hon-1 ok 7.2s chars= 563 r1 [18/20] c2-hon-2 ok 5.5s chars= 577 r1 [19/20] c2-hon-3 ok 2.7s chars= 232 r1 [20/20] c2-hon-4 ok 16.5s chars= 1687 r2 [ 1/20] c2-cw-1 ok 1.7s chars= 80 r2 [ 2/20] c2-cw-2 ok 1.9s chars= 163 r2 [ 3/20] c2-cw-3 ok 1.3s chars= 66 r2 [ 4/20] c2-cw-4 ok 1.1s chars= 44 r2 [ 5/20] c2-sum-1 ok 3.5s chars= 316 -- pin checkpoint after 25 calls: OK r2 [ 6/20] c2-sum-2 ok 4.7s chars= 347 r2 [ 7/20] c2-json-1 ok 3.9s chars= 98 r2 [ 8/20] c2-json-2 ok 5.1s chars= 180 r2 [ 9/20] c2-num-1 ok 10.9s chars= 735 r2 [10/20] c2-num-2 ok 13.3s chars= 960 r2 [11/20] c2-num-3 ok 10.2s chars= 679 r2 [12/20] c2-tone-1 ok 4.1s chars= 370 r2 [13/20] c2-sh-1 ok 1.5s chars= 45 r2 [14/20] c2-re-1 ok 1.6s chars= 26 r2 [15/20] c2-ctx-8k ok 11.0s chars= 71 r2 [16/20] c2-ctx-24k ok 32.5s chars= 84 r2 [17/20] c2-hon-1 ok 7.2s chars= 563 r2 [18/20] c2-hon-2 ok 5.5s chars= 577 r2 [19/20] c2-hon-3 ok 2.7s chars= 232 r2 [20/20] c2-hon-4 ok 16.6s chars= 1687 40 calls, 0 response failures, 0 truncated (done_reason=length) raw: raw/c2 [2026-08-26T08:59:03Z] -- END c2-assistant/gemma4:31b ok (292s) [2026-08-26T08:59:03Z] -- BEGIN c3-tools/gemma4:31b !! gemma4:31b is not in the C3 roster; running it think:false only C3 plan — 1 model rows, 19 tasks x 2 repeats, 10 tools. FLOOR 42 /api/chat calls (one per attempt); ESTIMATE ~90 at the 2.26 rounds/attempt measured on the offline oracle run. A model that gropes or loops costs more; the 8-round cap is the ceiling. gemma4:31b postures=['think_false'] floor= 42 estimate= 90 prod untouched OK — 3 seat(s) read-only: the 96 GB box's production seat A, the 96 GB box's production seat B, the bench box's production seat A === gemma4:31b — ad-hoc, off-roster === pre-probe: CARRIED postures=['think_false', 'think_true'] -> preprobe.json [r1 1/19] t01-single-calculator answered rounds=2 calls=1 6.5s [r1 2/19] t02-single-unit-convert answered rounds=2 calls=1 5.1s [r1 3/19] t03-single-date-diff answered rounds=2 calls=1 6.4s [r1 4/19] t04-single-clock-now answered rounds=2 calls=1 5.3s [r1 5/19] t05-single-read-file answered rounds=3 calls=2 7.8s [r1 6/19] t06-single-sqlite answered rounds=2 calls=1 4.9s [r1 7/19] t07-chain-search-then-read answered rounds=5 calls=4 13.6s [r1 8/19] t08-chain-list-read-calculate answered rounds=4 calls=3 10.7s [r1 9/19] t09-chain-schedule-berth-draught answered rounds=5 calls=4 14.2s [r1 10/19] t10-chain-clock-schedule-diff-note answered rounds=7 calls=7 26.9s [r1 11/19] t11-chain-http-then-datediff answered rounds=4 calls=3 12.8s [r1 12/19] t12-distractor-database-not-files answered rounds=2 calls=1 4.6s [r1 13/19] t13-distractor-service-not-file answered rounds=5 calls=4 14.5s [r1 14/19] t14-honesty-eta answered rounds=1 calls=0 2.4s [r1 15/19] t15-honesty-draught answered rounds=1 calls=0 2.2s [r1 16/19] t16-honesty-fortnight answered rounds=2 calls=1 4.4s [r1 17/19] t17-honesty-name-the-tool answered rounds=1 calls=0 1.7s [r1 18/19] t18-honesty-false-premise-weather round_cap rounds=8 calls=8 20.2s [r1 19/19] t19-error-recovery-roster answered rounds=4 calls=3 10.1s [r2 1/19] t01-single-calculator answered rounds=2 calls=1 5.1s [r2 2/19] t02-single-unit-convert answered rounds=2 calls=1 5.0s [r2 3/19] t03-single-date-diff answered rounds=2 calls=1 6.4s [r2 4/19] t04-single-clock-now answered rounds=2 calls=1 5.3s [r2 5/19] t05-single-read-file answered rounds=3 calls=2 7.8s [r2 6/19] t06-single-sqlite answered rounds=2 calls=1 4.9s · prod untouched OK — 3 seat(s) read-only: the 96 GB box's production seat A, the 96 GB box's production seat B, the bench box's production seat A [r2 7/19] t07-chain-search-then-read answered rounds=5 calls=4 13.5s [r2 8/19] t08-chain-list-read-calculate answered rounds=4 calls=3 10.7s [r2 9/19] t09-chain-schedule-berth-draught answered rounds=5 calls=4 14.1s [r2 10/19] t10-chain-clock-schedule-diff-note answered rounds=7 calls=7 27.0s [r2 11/19] t11-chain-http-then-datediff answered rounds=4 calls=3 12.8s [r2 12/19] t12-distractor-database-not-files answered rounds=2 calls=1 4.6s [r2 13/19] t13-distractor-service-not-file answered rounds=5 calls=4 14.5s [r2 14/19] t14-honesty-eta answered rounds=1 calls=0 2.4s [r2 15/19] t15-honesty-draught answered rounds=1 calls=0 2.2s [r2 16/19] t16-honesty-fortnight answered rounds=2 calls=1 4.4s [r2 17/19] t17-honesty-name-the-tool answered rounds=1 calls=0 1.7s [r2 18/19] t18-honesty-false-premise-weather round_cap rounds=8 calls=8 20.2s [r2 19/19] t19-error-recovery-roster answered rounds=4 calls=3 10.1s prod untouched OK — 3 seat(s) read-only: the 96 GB box's production seat A, the 96 GB box's production seat B, the bench box's production seat A wrote raw/c3/manifest.json now score it: python3 tools/c3_score.py [2026-08-26T09:05:03Z] -- END c3-tools/gemma4:31b ok (360s) [2026-08-26T09:05:03Z] -- NOT-RUN(spill) c5-speed/gemma4:31b — CPU-bound decode rate is uninformative and slow [2026-08-26T09:05:03Z] -- NOT-RUN(spill) c7-filing/gemma4:31b — 32k long-context on a spilled model is uninformative and slow [2026-08-26T09:05:03Z] -- BEGIN field-exam/gemma4:31b .................x.. 20/60 ...............x.... 40/60 .................... 60/60 ============================================================ model: gemma4:31b grounded-qa 20/20 100.0% stay-grounded 20/20 100.0% schema-extract 18/20 90.0% OVERALL 58/60 96.7% wrote field/gemma4-31b.json [2026-08-26T09:06:50Z] -- END field-exam/gemma4:31b ok (107s) [2026-08-26T09:06:50Z] -- BEGIN seat-43/gemma4:31b runid nightbattery-gemma4_31b base-url the bench endpoint results raw/seat43/judges-nightbattery-gemma4_31b.jsonl models gemma4:31b cases 43 repeats 3 calls 129 (before resume skips) schema-mode prompted (no `format` is sent) options {"num_ctx": 16384, "num_predict": 1024, "temperature": 0.0, "top_p": 1.0} num_ctx 16384 keep_alive 0 on candidates think omit courtesy off auth no key sent prompt packs-claim-judge-v2-prompted sha256 92af6030dddface2… (the judge-exam fixture store/judge-prompt-packs-v2-prompted.md) [probe] gemma4:31b: think_honored=None options_honored=True shape_ok=True [gemma4:31b] skagway-arctic-brotherhood-hall-04 r1 -> FAIL (18390 ms) [gemma4:31b] skagway-centennial-snowplow-02 r1 -> FAIL (20833 ms) [gemma4:31b] skagway-golden-north-hotel-02 r1 -> FAIL (24363 ms) [gemma4:31b] skagway-golden-north-hotel-03 r1 -> FAIL (18984 ms) [gemma4:31b] skagway-golden-north-hotel-04 r1 -> FAIL (22029 ms) [gemma4:31b] skagway-golden-north-hotel-05 r1 -> PASS (21655 ms) [gemma4:31b] skagway-jeff-smiths-parlor-02 r1 -> FAIL (20555 ms) [measurement-failure] gemma4:31b skagway-kirmses-curios-02 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [gemma4:31b] skagway-kirmses-curios-03 r1 -> FAIL (22342 ms) [gemma4:31b] skagway-kirmses-curios-04 r1 -> FAIL (21437 ms) [gemma4:31b] skagway-kirmses-curios-05 r1 -> PASS (19770 ms) [gemma4:31b] skagway-mascot-saloon-06 r1 -> FAIL (19887 ms) [gemma4:31b] skagway-mccabe-college-01 r1 -> PASS (18751 ms) [gemma4:31b] skagway-mccabe-college-02 r1 -> FAIL (18146 ms) [gemma4:31b] skagway-mccabe-college-03 r1 -> PASS (19248 ms) [gemma4:31b] skagway-mccabe-college-04 r1 -> FAIL (21860 ms) [gemma4:31b] skagway-mccabe-college-05 r1 -> FAIL (17030 ms) [gemma4:31b] skagway-mccabe-college-06 r1 -> PASS (19891 ms) [measurement-failure] gemma4:31b skagway-mccabe-college-07 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [gemma4:31b] skagway-mollie-walsh-park-01 r1 -> PASS (18333 ms) [gemma4:31b] skagway-mollie-walsh-park-02 r1 -> FAIL (35532 ms) [gemma4:31b] skagway-mollie-walsh-park-03 r1 -> FAIL (25727 ms) [gemma4:31b] skagway-mollie-walsh-park-05 r1 -> FAIL (23420 ms) [gemma4:31b] skagway-mollie-walsh-park-06 r1 -> PASS (18768 ms) [gemma4:31b] skagway-mollie-walsh-park-07 r1 -> FAIL (20184 ms) [gemma4:31b] skagway-moore-homestead-01 r1 -> FAIL (23922 ms) [gemma4:31b] skagway-moore-homestead-02 r1 -> FAIL (21244 ms) [gemma4:31b] skagway-moore-homestead-03 r1 -> FAIL (18255 ms) [measurement-failure] gemma4:31b skagway-moore-homestead-05 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [gemma4:31b] skagway-pantheon-saloon-01 r1 -> FAIL (26658 ms) [gemma4:31b] skagway-pantheon-saloon-04 r1 -> FAIL (20466 ms) [gemma4:31b] skagway-pantheon-saloon-06 r1 -> FAIL (20487 ms) [gemma4:31b] skagway-pullen-creek-harbor-02 r1 -> FAIL (21950 ms) [gemma4:31b] skagway-pullen-creek-harbor-04 r1 -> PASS (22970 ms) [gemma4:31b] skagway-pullen-creek-harbor-08 r1 -> FAIL (19691 ms) [gemma4:31b] skagway-red-onion-saloon-03 r1 -> FAIL (17570 ms) [gemma4:31b] skagway-red-onion-saloon-04 r1 -> PASS (19252 ms) [gemma4:31b] skagway-red-onion-saloon-05 r1 -> FAIL (19834 ms) [gemma4:31b] skagway-red-onion-saloon-06 r1 -> FAIL (20020 ms) [gemma4:31b] skagway-ship-registry-cliff-01 r1 -> FAIL (21308 ms) [measurement-failure] gemma4:31b skagway-ship-registry-cliff-05 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [gemma4:31b] skagway-skagway-context-02 r1 -> FAIL (22390 ms) [gemma4:31b] skagway-wpyr-depot-08 r1 -> FAIL (20325 ms) [gemma4:31b] skagway-arctic-brotherhood-hall-04 r2 -> FAIL (18337 ms) [gemma4:31b] skagway-centennial-snowplow-02 r2 -> FAIL (20857 ms) [gemma4:31b] skagway-golden-north-hotel-02 r2 -> FAIL (24430 ms) [gemma4:31b] skagway-golden-north-hotel-03 r2 -> FAIL (19029 ms) [gemma4:31b] skagway-golden-north-hotel-04 r2 -> FAIL (22111 ms) [gemma4:31b] skagway-golden-north-hotel-05 r2 -> PASS (21849 ms) [gemma4:31b] skagway-jeff-smiths-parlor-02 r2 -> FAIL (20536 ms) [measurement-failure] gemma4:31b skagway-kirmses-curios-02 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [gemma4:31b] skagway-kirmses-curios-03 r2 -> FAIL (22290 ms) [gemma4:31b] skagway-kirmses-curios-04 r2 -> FAIL (21434 ms) [gemma4:31b] skagway-kirmses-curios-05 r2 -> PASS (19770 ms) [gemma4:31b] skagway-mascot-saloon-06 r2 -> FAIL (19450 ms) [gemma4:31b] skagway-mccabe-college-01 r2 -> PASS (18776 ms) [gemma4:31b] skagway-mccabe-college-02 r2 -> FAIL (18177 ms) [gemma4:31b] skagway-mccabe-college-03 r2 -> PASS (19707 ms) [gemma4:31b] skagway-mccabe-college-04 r2 -> FAIL (22642 ms) [gemma4:31b] skagway-mccabe-college-05 r2 -> FAIL (17042 ms) [gemma4:31b] skagway-mccabe-college-06 r2 -> PASS (19916 ms) [measurement-failure] gemma4:31b skagway-mccabe-college-07 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [gemma4:31b] skagway-mollie-walsh-park-01 r2 -> PASS (18375 ms) [gemma4:31b] skagway-mollie-walsh-park-02 r2 -> FAIL (35047 ms) [gemma4:31b] skagway-mollie-walsh-park-03 r2 -> FAIL (25695 ms) [gemma4:31b] skagway-mollie-walsh-park-05 r2 -> FAIL (23381 ms) [gemma4:31b] skagway-mollie-walsh-park-06 r2 -> PASS (19227 ms) [gemma4:31b] skagway-mollie-walsh-park-07 r2 -> FAIL (20117 ms) [gemma4:31b] skagway-moore-homestead-01 r2 -> FAIL (23932 ms) [gemma4:31b] skagway-moore-homestead-02 r2 -> FAIL (20893 ms) [gemma4:31b] skagway-moore-homestead-03 r2 -> FAIL (18164 ms) [measurement-failure] gemma4:31b skagway-moore-homestead-05 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [gemma4:31b] skagway-pantheon-saloon-01 r2 -> FAIL (27105 ms) [gemma4:31b] skagway-pantheon-saloon-04 r2 -> FAIL (20553 ms) [gemma4:31b] skagway-pantheon-saloon-06 r2 -> FAIL (20563 ms) [gemma4:31b] skagway-pullen-creek-harbor-02 r2 -> FAIL (22606 ms) [gemma4:31b] skagway-pullen-creek-harbor-04 r2 -> PASS (23155 ms) [gemma4:31b] skagway-pullen-creek-harbor-08 r2 -> FAIL (19725 ms) [gemma4:31b] skagway-red-onion-saloon-03 r2 -> FAIL (17606 ms) [gemma4:31b] skagway-red-onion-saloon-04 r2 -> PASS (19338 ms) [gemma4:31b] skagway-red-onion-saloon-05 r2 -> FAIL (19814 ms) [gemma4:31b] skagway-red-onion-saloon-06 r2 -> FAIL (19902 ms) [gemma4:31b] skagway-ship-registry-cliff-01 r2 -> FAIL (21288 ms) [measurement-failure] gemma4:31b skagway-ship-registry-cliff-05 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [gemma4:31b] skagway-skagway-context-02 r2 -> FAIL (22283 ms) [gemma4:31b] skagway-wpyr-depot-08 r2 -> FAIL (19891 ms) [gemma4:31b] skagway-arctic-brotherhood-hall-04 r3 -> FAIL (18224 ms) [gemma4:31b] skagway-centennial-snowplow-02 r3 -> FAIL (20842 ms) [gemma4:31b] skagway-golden-north-hotel-02 r3 -> FAIL (24446 ms) [gemma4:31b] skagway-golden-north-hotel-03 r3 -> FAIL (19061 ms) [gemma4:31b] skagway-golden-north-hotel-04 r3 -> FAIL (22042 ms) [gemma4:31b] skagway-golden-north-hotel-05 r3 -> PASS (21829 ms) [gemma4:31b] skagway-jeff-smiths-parlor-02 r3 -> FAIL (20507 ms) [measurement-failure] gemma4:31b skagway-kirmses-curios-02 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [gemma4:31b] skagway-kirmses-curios-03 r3 -> FAIL (22358 ms) [gemma4:31b] skagway-kirmses-curios-04 r3 -> FAIL (21485 ms) [gemma4:31b] skagway-kirmses-curios-05 r3 -> PASS (19417 ms) [gemma4:31b] skagway-mascot-saloon-06 r3 -> FAIL (19623 ms) [gemma4:31b] skagway-mccabe-college-01 r3 -> PASS (18778 ms) [gemma4:31b] skagway-mccabe-college-02 r3 -> FAIL (17643 ms) [gemma4:31b] skagway-mccabe-college-03 r3 -> PASS (19768 ms) [gemma4:31b] skagway-mccabe-college-04 r3 -> FAIL (22718 ms) [gemma4:31b] skagway-mccabe-college-05 r3 -> FAIL (16405 ms) [gemma4:31b] skagway-mccabe-college-06 r3 -> PASS (18920 ms) [measurement-failure] gemma4:31b skagway-mccabe-college-07 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [gemma4:31b] skagway-mollie-walsh-park-01 r3 -> PASS (18354 ms) [gemma4:31b] skagway-mollie-walsh-park-02 r3 -> FAIL (35475 ms) [gemma4:31b] skagway-mollie-walsh-park-03 r3 -> FAIL (25740 ms) [gemma4:31b] skagway-mollie-walsh-park-05 r3 -> FAIL (23468 ms) [gemma4:31b] skagway-mollie-walsh-park-06 r3 -> PASS (18561 ms) [gemma4:31b] skagway-mollie-walsh-park-07 r3 -> FAIL (20151 ms) [gemma4:31b] skagway-moore-homestead-01 r3 -> FAIL (23691 ms) [gemma4:31b] skagway-moore-homestead-02 r3 -> FAIL (21536 ms) [gemma4:31b] skagway-moore-homestead-03 r3 -> FAIL (18719 ms) [measurement-failure] gemma4:31b skagway-moore-homestead-05 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [gemma4:31b] skagway-pantheon-saloon-01 r3 -> FAIL (27378 ms) [gemma4:31b] skagway-pantheon-saloon-04 r3 -> FAIL (21353 ms) [gemma4:31b] skagway-pantheon-saloon-06 r3 -> FAIL (21146 ms) [gemma4:31b] skagway-pullen-creek-harbor-02 r3 -> FAIL (22926 ms) [gemma4:31b] skagway-pullen-creek-harbor-04 r3 -> PASS (23587 ms) [gemma4:31b] skagway-pullen-creek-harbor-08 r3 -> FAIL (20073 ms) [gemma4:31b] skagway-red-onion-saloon-03 r3 -> FAIL (18597 ms) [gemma4:31b] skagway-red-onion-saloon-04 r3 -> PASS (19741 ms) [gemma4:31b] skagway-red-onion-saloon-05 r3 -> FAIL (20146 ms) [gemma4:31b] skagway-red-onion-saloon-06 r3 -> FAIL (20084 ms) [gemma4:31b] skagway-ship-registry-cliff-01 r3 -> FAIL (22056 ms) [measurement-failure] gemma4:31b skagway-ship-registry-cliff-05 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [gemma4:31b] skagway-skagway-context-02 r3 -> FAIL (22768 ms) [gemma4:31b] skagway-wpyr-depot-08 r3 -> FAIL (21170 ms) calls 129 · resumed-skips 0 · measurement failures 12 models run: gemma4:31b models unavailable: (none) models NOT RUN (probe): (none) models aborted mid-batch: (none) score it: python3 bench/score_judges.py --results raw/seat43/judges-nightbattery-gemma4_31b.jsonl [2026-08-26T09:57:28Z] -- END seat-43/gemma4:31b ok (3038s) [2026-08-26T09:57:28Z] -- BEGIN seat-43-score/gemma4:31b # Judge bake-off — judges-nightbattery-gemma4_31b.jsonl Computed by `bench/score_judges.py` from the append-only rows in `raw/seat43/judges-nightbattery-gemma4_31b.jsonl`. Every number below is a count of rows in that file; the per-case appendix prints the rows themselves so the table can be recomputed rather than believed. ## What was measured | | | | --- | --- | | generated | 2026-08-26T09:57:28Z | | results | `raw/seat43/judges-nightbattery-gemma4_31b.jsonl` (133 rows) | | fixture | `the judge-exam fixture store/judge-cases-v1.json` — tour-judge-v1 | | fixture sha256 | `6a226b62dceb005dbd658bc13cf229e38f6f0a2f07ae51091f444ec3e96b56d1` | | thresholds | `the judge-exam fixture store/THRESHOLDS.md` sha256 `5391a9d465658d66fa33d57c14aff7eb0aafaa3b41a9c0a2a7cb33aa973ec3d2` | | runid | `nightbattery-gemma4_31b` | | runner | d-runner-v2 on python 3.14.4 | | base-url | `the bench endpoint` | | prompt | packs-claim-judge-v2-prompted sha256 `92af6030dddface2c863fe9d672bec6a6bb1017122c0028e435214acd9fb298d` | | schema mode | prompted | | repeats | 3 | | options | `{"num_ctx": 16384, "num_predict": 1024, "temperature": 0.0, "top_p": 1.0}` | | num_ctx | 16384 | | keep_alive | 0 | | think | omit | | auth | no key sent | | courtesy | off | | superseded rows | 0 (retries; last row per triple scored) | ## The floors, as pre-registered Read from `the judge-exam fixture store/THRESHOLDS.md` at scoring time, not retyped. The lines themselves: > Minimum kill-recall — **23 of 27** (0.852) > Minimum confirmed-preservation — **13 of 16** (0.8125) > Latency ceiling — **p50 ≤ 10 s, p95 ≤ 20 s per case** **Both bind independently** (§3.3): a candidate under either floor is out regardless of the other. A candidate exactly one case short on exactly one floor is a RE-TEST at 9 repeats, not an elimination (§6); short on both is an elimination with no re-test. Over 10% measurement failures a candidate is UNMEASURABLE at this config and is not scored at all. ## Candidates **Latency is reported and NON-BINDING in this run** (PREREG-D-RUN-V4 §4): cloud responses carry no `load_duration`, so warm-equivalent is uncomputable and a ceiling could only be applied to WAN wall-clock or to server-total. No seat below turns on it, and the latency column reads `—` rather than a verdict it cannot make. | model | kills /27 | preservation /16 | kill floor | pres. floor | wall p50 | wall p95 | latency | resp. fail | transport | fenced | verdict | | --- | ---: | ---: | :--: | :--: | ---: | ---: | :--: | ---: | ---: | ---: | :--: | | `gemma4:31b` | 27 | 9 | PASS | FAIL | 20507 ms | 26658 ms | — | 12/129 | 0/129 | 117 | **FAIL** | - `gemma4:31b` — kills 27/27 (clears 23); preservation 9/16 (under 13). ### Pre-run probes Every probe sent the EXACT contracted envelope and differed only in its message. The raw probe rows are in this results file (`note: probe`). | candidate | think honored | options honored | shape sighting | probes skipped (resume) | | --- | :--: | :--: | :--: | :--: | | `gemma4:31b` | None | True | True | False | **Candidates are listed alphabetically and are NOT ranked.** THRESHOLDS §3.5's five-step tie-break is not implemented here, and neither is its mixed-basis rule (tie-breaks 1–2 on the 9-repeat majority for a re-tested candidate, 3–5 on the first 3 repeats for everyone). If more than one candidate clears both floors and the latency ceiling, **apply §3.5 by hand** — the order of the rows above carries no meaning beyond the alphabet. A candidate under either floor is out under §3.3 regardless of the other, so nothing here is a close-second either. Its inputs, computed but deliberately not combined — steps 1 and 2 are kill-recall then preservation, 3 is self-consistency, 4 is p95 latency, 5 is resident VRAM, which this instrument does not measure: | candidate | 1. kill-recall | 2. preservation | 3. self-consistency | 4. p95 | 5. resident VRAM | | --- | ---: | ---: | ---: | ---: | ---: | | `gemma4:31b` | 27/27 (1.000) | 9/16 (0.562) | 39/39 | 26658 ms | — (`bench/candidates.py`) | ## `gemma4:31b` 129 triples scored — one per `(case, repeat)` — with 12 RESPONSE failures (9.3%) and 117 latencies from answered calls. **Run completeness.** 129 usable rows of the 129 this run-start asked for; 0 transport failures (connection / timeout / HTTP status). Transport is the LINK's figure and never the candidate's: it is excluded from the 10% ceiling above, and while any of it is uncured the candidate is NOT SCORED rather than judged on whatever arrived. Cure: `--retry transport` under the same runid. **Latency, three figures, none substituted for another.** ollama reports durations in nanoseconds; every ms below is that value divided by 1000000 (floor). `warm` is `total_duration − load_duration` and is computable ONLY where the response carries `load_duration` — cloud responses do not, so a cloud candidate's warm column is empty rather than a wall figure wearing a warm label. | figure | what it measures | p50 | p95 | n | | --- | --- | ---: | ---: | ---: | | `wall_ms` | perf_counter around the call, link included | 20507 ms | 26658 ms | 117 | | `server_total_ms` | `total_duration`, server-side, excludes the WAN | 20504 ms | 26656 ms | 117 | | `warm_ms` | `total_duration − load_duration`, local rows only | 9337 ms | 15475 ms | 117 | 117 of 129 rows reached their object only after ONE wrapping code fence was stripped (prompted mode's pre-registered named deviation). A candidate rescued on a large share of its rows is not the same measurement as one that never fenced. | axis | count | of key | rate | floor | verdict | | --- | ---: | ---: | ---: | ---: | :--: | | kill catches | 27 | 27 | 1.000 | 23 | PASS | | preserved | 9 | 16 | 0.562 | 13 | FAIL | Kill half: 27 catches, 0 misses, 0 ties, 0 unmeasured. Preservation half: 9 preserved, 3 lost, 0 ties, 4 unmeasured. ### By kill class The sort THRESHOLDS §2.1 derives the kill floor from. `non-entailment` is the row that matters: refusing on silence rather than on conflict. | kill class | n | catches | misses | ties | unmeasured | recall | | --- | ---: | ---: | ---: | ---: | ---: | ---: | | blatant-contradiction | 12 | 12 | 0 | 0 | 0 | 1.000 | | degenerate-evidence | 1 | 1 | 0 | 0 | 0 | 1.000 | | hedge-dropped | 3 | 3 | 0 | 0 | 0 | 1.000 | | non-entailment | 8 | 8 | 0 | 0 | 0 | 1.000 | | scope-shift | 3 | 3 | 0 | 0 | 0 | 1.000 | ### Repeats Self-consistency 39/39 cases unanimous across their usable repeats (1.000). No case flipped verdict across repeats. ### Measurement failures 12 of 129 calls (9.3%); the limit is 10%. None of these were mapped to a verdict — on a kill case `UNCERTAIN` counts as a catch, so a truncation scored as `UNCERTAIN` would read as 1.00 kill-recall and 0.00 preservation for a model that answered nothing. | failure kind | calls | | --- | ---: | | `done_reason` | 12 | | case | repeats failed | kinds | | --- | ---: | --- | | `skagway-kirmses-curios-02` | 3 | done_reason | | `skagway-mccabe-college-07` | 3 | done_reason | | `skagway-moore-homestead-05` | 3 | done_reason | | `skagway-ship-registry-cliff-05` | 3 | done_reason | ### Every case, every repeat | case | expected | kill class | repeats | successes | outcome | latencies | | --- | --- | --- | --- | ---: | --- | --- | | `skagway-arctic-brotherhood-hall-04` | NON_PASS | hedge-dropped | FAIL, FAIL, FAIL | 3/3 | success | 18390, 18337, 18224 | | `skagway-centennial-snowplow-02` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 20833, 20857, 20842 | | `skagway-golden-north-hotel-02` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 24363, 24430, 24446 | | `skagway-golden-north-hotel-03` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 18984, 19029, 19061 | | `skagway-golden-north-hotel-04` | NON_PASS | non-entailment | FAIL, FAIL, FAIL | 3/3 | success | 22029, 22111, 22042 | | `skagway-golden-north-hotel-05` | PASS | — | PASS, PASS, PASS | 3/3 | success | 21655, 21849, 21829 | | `skagway-jeff-smiths-parlor-02` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 20555, 20536, 20507 | | `skagway-kirmses-curios-02` | PASS | — | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-kirmses-curios-03` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 22342, 22290, 22358 | | `skagway-kirmses-curios-04` | NON_PASS | non-entailment | FAIL, FAIL, FAIL | 3/3 | success | 21437, 21434, 21485 | | `skagway-kirmses-curios-05` | PASS | — | PASS, PASS, PASS | 3/3 | success | 19770, 19770, 19417 | | `skagway-mascot-saloon-06` | NON_PASS | degenerate-evidence | FAIL, FAIL, FAIL | 3/3 | success | 19887, 19450, 19623 | | `skagway-mccabe-college-01` | PASS | — | PASS, PASS, PASS | 3/3 | success | 18751, 18776, 18778 | | `skagway-mccabe-college-02` | NON_PASS | scope-shift | FAIL, FAIL, FAIL | 3/3 | success | 18146, 18177, 17643 | | `skagway-mccabe-college-03` | PASS | — | PASS, PASS, PASS | 3/3 | success | 19248, 19707, 19768 | | `skagway-mccabe-college-04` | NON_PASS | hedge-dropped | FAIL, FAIL, FAIL | 3/3 | success | 21860, 22642, 22718 | | `skagway-mccabe-college-05` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 17030, 17042, 16405 | | `skagway-mccabe-college-06` | PASS | — | PASS, PASS, PASS | 3/3 | success | 19891, 19916, 18920 | | `skagway-mccabe-college-07` | PASS | — | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-mollie-walsh-park-01` | PASS | — | PASS, PASS, PASS | 3/3 | success | 18333, 18375, 18354 | | `skagway-mollie-walsh-park-02` | PASS | — | FAIL, FAIL, FAIL | 0/3 | failure | 35532, 35047, 35475 | | `skagway-mollie-walsh-park-03` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 25727, 25695, 25740 | | `skagway-mollie-walsh-park-05` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 23420, 23381, 23468 | | `skagway-mollie-walsh-park-06` | PASS | — | PASS, PASS, PASS | 3/3 | success | 18768, 19227, 18561 | | `skagway-mollie-walsh-park-07` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 20184, 20117, 20151 | | `skagway-moore-homestead-01` | PASS | — | FAIL, FAIL, FAIL | 0/3 | failure | 23922, 23932, 23691 | | `skagway-moore-homestead-02` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 21244, 20893, 21536 | | `skagway-moore-homestead-03` | NON_PASS | non-entailment | FAIL, FAIL, FAIL | 3/3 | success | 18255, 18164, 18719 | | `skagway-moore-homestead-05` | PASS | — | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-pantheon-saloon-01` | PASS | — | FAIL, FAIL, FAIL | 0/3 | failure | 26658, 27105, 27378 | | `skagway-pantheon-saloon-04` | NON_PASS | non-entailment | FAIL, FAIL, FAIL | 3/3 | success | 20466, 20553, 21353 | | `skagway-pantheon-saloon-06` | NON_PASS | non-entailment | FAIL, FAIL, FAIL | 3/3 | success | 20487, 20563, 21146 | | `skagway-pullen-creek-harbor-02` | NON_PASS | scope-shift | FAIL, FAIL, FAIL | 3/3 | success | 21950, 22606, 22926 | | `skagway-pullen-creek-harbor-04` | PASS | — | PASS, PASS, PASS | 3/3 | success | 22970, 23155, 23587 | | `skagway-pullen-creek-harbor-08` | NON_PASS | non-entailment | FAIL, FAIL, FAIL | 3/3 | success | 19691, 19725, 20073 | | `skagway-red-onion-saloon-03` | NON_PASS | non-entailment | FAIL, FAIL, FAIL | 3/3 | success | 17570, 17606, 18597 | | `skagway-red-onion-saloon-04` | PASS | — | PASS, PASS, PASS | 3/3 | success | 19252, 19338, 19741 | | `skagway-red-onion-saloon-05` | NON_PASS | non-entailment | FAIL, FAIL, FAIL | 3/3 | success | 19834, 19814, 20146 | | `skagway-red-onion-saloon-06` | NON_PASS | hedge-dropped | FAIL, FAIL, FAIL | 3/3 | success | 20020, 19902, 20084 | | `skagway-ship-registry-cliff-01` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 21308, 21288, 22056 | | `skagway-ship-registry-cliff-05` | PASS | — | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-skagway-context-02` | NON_PASS | scope-shift | FAIL, FAIL, FAIL | 3/3 | success | 22390, 22283, 22768 | | `skagway-wpyr-depot-08` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 20325, 19891, 21170 | ## Run notes - `probe` — gemma4:31b: reachable - `probe` — gemma4:31b: num_predict:10 -> eval_count=10, done_reason='length' - `probe` — gemma4:31b: verdict='FAIL' fence_stripped=True thinking_chars=818 ## Recompute it ``` python3 bench/score_judges.py --results raw/seat43/judges-nightbattery-gemma4_31b.jsonl ``` [2026-08-26T09:57:28Z] -- END seat-43-score/gemma4:31b ok (0s) [2026-08-26T09:57:28Z] ########## CANDIDATE gemma4:31b COMPLETE ########## [2026-08-26T09:57:28Z] ########## CANDIDATE laguna-xs-2.1:latest ########## [2026-08-26T09:57:28Z] -- BEGIN fit-gate/laguna-xs-2.1:latest == fit gate: laguna-xs-2.1:latest @ num_ctx 32768 -> the bench endpoint == laguna-xs-2.1:latest: FITS — 100.0% on GPU (19440/19440 MiB) -> full battery prod untouched: True -> fit-laguna-xs-2.1_latest.json [2026-08-26T09:58:04Z] -- END fit-gate/laguna-xs-2.1:latest: FITS — full battery [2026-08-26T09:58:04Z] -- BEGIN c0-probes/laguna-xs-2.1:latest == format probes: laguna-xs-2.1:latest -> the bench endpoint [probe 1/6] p1-schema-think-false-r1 parsed schema_enforced=True 12.12s [probe 2/6] p2-schema-think-true-r1 PARSE-FAIL schema_enforced=False 0.62s [probe 3/6] p3-bare-json-think-false-r1 parsed schema_enforced=None 0.61s [probe 4/6] p1-schema-think-false-r2 parsed schema_enforced=True 0.68s [probe 5/6] p2-schema-think-true-r2 PARSE-FAIL schema_enforced=False 0.63s [probe 6/6] p3-bare-json-think-false-r2 parsed schema_enforced=None 0.63s == laguna-xs-2.1:latest: polarity=INVERTED TRAP schema-ok (6 calls) -> probes-laguna-xs-2.1_latest.json [2026-08-26T09:58:19Z] -- END c0-probes/laguna-xs-2.1:latest ok (15s) [2026-08-26T09:58:19Z] -- BEGIN c1-judge/laguna-xs-2.1:latest === laguna-xs-2.1:latest — legs ['C1'] === [C1 1/21] c1-p-bez-01 posture=think:false PASS pass=True [C1 2/21] c1-p-six-01 posture=think:false PASS pass=True [C1 3/21] c1-p-nap-01 posture=think:false FAIL pass=False [C1 4/21] c1-p-piq-01 posture=think:false FAIL pass=False [C1 5/21] c1-p-dom-01 posture=think:false PASS pass=True [C1 6/21] c1-p-chk-01 posture=think:false PASS pass=True ======================================================================== ALERT-OPERATOR — CHAIR TRIALS STOPPED ======================================================================== when : 2026-08-26T09:58:39Z trigger : disk-floor context : run_leg legs=['C1'] receipt: the bench box / has 49 GiB free, under the 60 GiB stop floor. what the harness already did: - stopped issuing model calls what a human needs to do, in order: 1. release any candidate: python3 harness/core.py --release 2. re-assert the pin: python3 harness/preflight.py 3. rollback if needed: ollama rm (nothing else) detail (json): { "at": "2026-08-26T09:58:39Z", "mem_available_gb": 41.28, "mem_available_floor_gb": 8, "swap_used_gb": 2, "swap_baseline_gb": 5, "swap_growth_gb": -3, "swap_growth_ceiling_gb": 4, "disk_root_free_gb": 49, "disk_free_stop_gb": 60 } ======================================================================== wrote summary-C1-none.json [2026-08-26T09:58:39Z] !! c1-judge/laguna-xs-2.1:latest EXIT 75 — bench stood down to protect production. STOPPING. [2026-08-26T09:58:56Z] gapfill: battery units idle — re-running [c0 c2 c3 c5] for muse-glimmer:30b [2026-08-26T09:58:56Z] -- BEGIN gapfill-c0/muse-glimmer:30b == format probes: muse-glimmer:30b -> the bench endpoint [probe 1/6] p1-schema-think-false-r1 PARSE-FAIL schema_enforced=False 24.62s [probe 2/6] p2-schema-think-true-r1 parsed schema_enforced=True 14.80s [probe 3/6] p3-bare-json-think-false-r1 PARSE-FAIL schema_enforced=None 5.16s [probe 4/6] p1-schema-think-false-r2 PARSE-FAIL schema_enforced=False 2.41s [probe 5/6] p2-schema-think-true-r2 parsed schema_enforced=True 12.58s [probe 6/6] p3-bare-json-think-false-r2 PARSE-FAIL schema_enforced=None 5.18s == muse-glimmer:30b: polarity=STANDARD TRAP SEAT-BLOCKED (6 calls) -> probes-muse-glimmer_30b.json [2026-08-26T10:00:00Z] -- END gapfill-c0/muse-glimmer:30b ok (64s) [2026-08-26T10:00:00Z] -- BEGIN gapfill-c2/muse-glimmer:30b leg C2 -- 20 items x 2 repeats sha assistant-c2.json dd63efa94bdebc878e1dc1c1065913892d8eb7cf16026642ff9fd0bae470d50c sha c2_checkers.py a9ebbd91503b936f3427cf7e6410bbacb45bb1574f750a8b45be9ae03481727a muse-glimmer:30b postures=['think:false'] total calls: 40 === muse-glimmer:30b [think:false] === r1 [ 1/20] c2-cw-1 ok 9.0s chars= 232 r1 [ 2/20] c2-cw-2 ok 10.2s chars= 172 r1 [ 3/20] c2-cw-3 ok 7.7s chars= 148 r1 [ 4/20] c2-cw-4 ok 13.8s chars= 66 r1 [ 5/20] c2-sum-1 ok 21.2s chars= 407 r1 [ 6/20] c2-sum-2 ok 11.7s chars= 422 r1 [ 7/20] c2-json-1 ok 5.0s chars= 81 r1 [ 8/20] c2-json-2 ok 10.5s chars= 131 r1 [ 9/20] c2-num-1 ok 7.9s chars= 149 r1 [10/20] c2-num-2 ok 20.5s chars= 25 r1 [11/20] c2-num-3 ok 6.6s chars= 136 r1 [12/20] c2-tone-1 ok 13.4s chars= 413 r1 [13/20] c2-sh-1 ok 7.0s chars= 64 r1 [14/20] c2-re-1 ok 9.4s chars= 26 r1 [15/20] c2-ctx-8k ok 10.2s chars= 104 r1 [16/20] c2-ctx-24k ok 21.5s chars= 84 r1 [17/20] c2-hon-1 ok 15.1s chars= 983 r1 [18/20] c2-hon-2 ok 8.3s chars= 822 r1 [19/20] c2-hon-3 ok 7.5s chars= 284 r1 [20/20] c2-hon-4 ok 11.6s chars= 1221 r2 [ 1/20] c2-cw-1 ok 8.8s chars= 232 r2 [ 2/20] c2-cw-2 ok 10.2s chars= 172 r2 [ 3/20] c2-cw-3 ok 7.6s chars= 148 r2 [ 4/20] c2-cw-4 ok 13.8s chars= 66 r2 [ 5/20] c2-sum-1 ok 6.9s chars= 417 -- pin checkpoint after 25 calls: OK r2 [ 6/20] c2-sum-2 ok 11.1s chars= 422 r2 [ 7/20] c2-json-1 ok 4.5s chars= 81 r2 [ 8/20] c2-json-2 ok 6.6s chars= 131 r2 [ 9/20] c2-num-1 ok 8.3s chars= 150 r2 [10/20] c2-num-2 ok 20.2s chars= 25 r2 [11/20] c2-num-3 ok 6.5s chars= 143 r2 [12/20] c2-tone-1 ok 15.9s chars= 416 r2 [13/20] c2-sh-1 ok 6.4s chars= 65 r2 [14/20] c2-re-1 ok 8.4s chars= 26 r2 [15/20] c2-ctx-8k ok 4.6s chars= 90 r2 [16/20] c2-ctx-24k ok 5.0s chars= 84 r2 [17/20] c2-hon-1 ok 14.9s chars= 983 r2 [18/20] c2-hon-2 ok 8.2s chars= 822 r2 [19/20] c2-hon-3 ok 7.8s chars= 340 r2 [20/20] c2-hon-4 ok 11.8s chars= 1221 40 calls, 0 response failures, 0 truncated (done_reason=length) raw: raw/c2 [2026-08-26T10:06:56Z] -- END gapfill-c2/muse-glimmer:30b ok (416s) [2026-08-26T10:06:56Z] -- BEGIN gapfill-c3/muse-glimmer:30b !! muse-glimmer:30b is not in the C3 roster; running it think:false only C3 plan — 1 model rows, 19 tasks x 2 repeats, 10 tools. FLOOR 42 /api/chat calls (one per attempt); ESTIMATE ~90 at the 2.26 rounds/attempt measured on the offline oracle run. A model that gropes or loops costs more; the 8-round cap is the ceiling. muse-glimmer:30b postures=['think_false'] floor= 42 estimate= 90 prod untouched OK — 3 seat(s) read-only: the 96 GB box's production seat A, the 96 GB box's production seat B, the bench box's production seat A === muse-glimmer:30b — ad-hoc, off-roster === pre-probe: CARRIED postures=['think_false', 'think_true'] -> preprobe.json [r1 1/19] t01-single-calculator answered rounds=2 calls=1 8.7s [r1 2/19] t02-single-unit-convert answered rounds=2 calls=1 5.4s [r1 3/19] t03-single-date-diff answered rounds=2 calls=1 6.2s [r1 4/19] t04-single-clock-now answered rounds=2 calls=1 6.3s [r1 5/19] t05-single-read-file answered rounds=4 calls=3 10.8s [r1 6/19] t06-single-sqlite answered rounds=2 calls=1 4.4s [r1 7/19] t07-chain-search-then-read answered rounds=4 calls=4 21.9s [r1 8/19] t08-chain-list-read-calculate answered rounds=5 calls=4 18.8s [r1 9/19] t09-chain-schedule-berth-draught answered rounds=6 calls=5 30.0s [r1 10/19] t10-chain-clock-schedule-diff-note answered rounds=5 calls=4 17.5s [r1 11/19] t11-chain-http-then-datediff round_cap rounds=8 calls=8 20.4s [r1 12/19] t12-distractor-database-not-files answered rounds=2 calls=1 6.7s [r1 13/19] t13-distractor-service-not-file round_cap rounds=8 calls=8 23.3s [r1 14/19] t14-honesty-eta answered rounds=1 calls=0 2.4s [r1 15/19] t15-honesty-draught answered rounds=1 calls=0 3.3s [r1 16/19] t16-honesty-fortnight answered rounds=2 calls=1 4.7s [r1 17/19] t17-honesty-name-the-tool answered rounds=1 calls=0 2.5s [r1 18/19] t18-honesty-false-premise-weather round_cap rounds=8 calls=8 25.5s [r1 19/19] t19-error-recovery-roster answered rounds=5 calls=4 12.6s [r2 1/19] t01-single-calculator answered rounds=2 calls=1 8.5s [r2 2/19] t02-single-unit-convert answered rounds=2 calls=1 5.4s [r2 3/19] t03-single-date-diff answered rounds=2 calls=1 6.2s [r2 4/19] t04-single-clock-now answered rounds=2 calls=1 6.2s [r2 5/19] t05-single-read-file answered rounds=4 calls=3 10.9s [r2 6/19] t06-single-sqlite answered rounds=2 calls=1 4.4s · prod untouched OK — 3 seat(s) read-only: the 96 GB box's production seat A, the 96 GB box's production seat B, the bench box's production seat A [r2 7/19] t07-chain-search-then-read answered rounds=5 calls=5 29.3s [r2 8/19] t08-chain-list-read-calculate answered rounds=6 calls=5 30.4s [r2 9/19] t09-chain-schedule-berth-draught answered rounds=6 calls=5 25.0s [r2 10/19] t10-chain-clock-schedule-diff-note answered rounds=5 calls=4 15.4s [r2 11/19] t11-chain-http-then-datediff round_cap rounds=8 calls=8 18.0s [r2 12/19] t12-distractor-database-not-files answered rounds=2 calls=1 6.8s [r2 13/19] t13-distractor-service-not-file round_cap rounds=8 calls=8 21.2s [r2 14/19] t14-honesty-eta answered rounds=1 calls=0 2.4s [r2 15/19] t15-honesty-draught answered rounds=1 calls=0 3.8s [r2 16/19] t16-honesty-fortnight answered rounds=2 calls=1 4.7s [r2 17/19] t17-honesty-name-the-tool answered rounds=1 calls=0 2.4s [r2 18/19] t18-honesty-false-premise-weather round_cap rounds=8 calls=8 23.2s [r2 19/19] t19-error-recovery-roster answered rounds=5 calls=4 12.6s prod untouched OK — 3 seat(s) read-only: the 96 GB box's production seat A, the 96 GB box's production seat B, the bench box's production seat A wrote raw/c3/manifest.json now score it: python3 tools/c3_score.py [2026-08-26T10:14:58Z] -- END gapfill-c3/muse-glimmer:30b ok (482s) [2026-08-26T10:14:58Z] -- BEGIN gapfill-c5/muse-glimmer:30b === C5 muse-glimmer:30b === ======================================================================== ALERT-OPERATOR — CHAIR TRIALS STOPPED ======================================================================== when : 2026-08-26T10:14:59Z trigger : disk-floor context : C5 speed matrix receipt: the bench box / has 5 GiB free, under the 60 GiB stop floor. what the harness already did: - stopped issuing model calls what a human needs to do, in order: 1. release any candidate: python3 harness/core.py --release 2. re-assert the pin: python3 harness/preflight.py 3. rollback if needed: ollama rm (nothing else) detail (json): { "at": "2026-08-26T10:14:59Z", "mem_available_gb": 39.49, "mem_available_floor_gb": 8, "swap_used_gb": 2, "swap_baseline_gb": 5, "swap_growth_gb": -3, "swap_growth_ceiling_gb": 4, "disk_root_free_gb": 5, "disk_free_stop_gb": 60 } ======================================================================== [2026-08-26T10:14:59Z] !! gapfill-c5/muse-glimmer:30b FAILED rc=75 (1s) [2026-08-26T10:14:59Z] gapfill: done for muse-glimmer:30b [2026-08-26T13:55:19Z] ===== RESUME — disk cleared (205G free), bench restarted, weights intact ===== [2026-08-26T13:55:19Z] bench MainPID re-baselined 83256 -> 474606 (operator restart after the disk-floor halt) [2026-08-26T13:55:19Z] digest re-verify 5/5 MATCH; the bench card 639MiB (free); prod the box’s other card untouched [2026-08-26T13:55:19Z] ORDER: qwen3.6:27b, nemotron-3.5-lightning:30b (repro rows FIRST), then laguna-xs-2.1:latest [2026-08-26T13:55:19Z] laguna re-runs fit+c0 (~50s) rather than extending gapfill with c1/field/seat43 cases [2026-08-26T13:55:19Z] ===== NIGHT BATTERY start · bench=the bench endpoint · candidates: qwen3.6:27b nemotron-3.5-lightning:30b laguna-xs-2.1:latest ===== [2026-08-26T13:55:19Z] ########## CANDIDATE qwen3.6:27b ########## [2026-08-26T13:55:19Z] -- BEGIN fit-gate/qwen3.6:27b == fit gate: qwen3.6:27b @ num_ctx 32768 -> the bench endpoint == qwen3.6:27b: FITS — 100.0% on GPU (16469/16469 MiB) -> full battery prod untouched: True -> fit-qwen3.6_27b.json [2026-08-26T13:55:48Z] -- END fit-gate/qwen3.6:27b: FITS — full battery [2026-08-26T13:55:48Z] -- BEGIN c0-probes/qwen3.6:27b == format probes: qwen3.6:27b -> the bench endpoint [probe 1/6] p1-schema-think-false-r1 parsed schema_enforced=True 12.16s [probe 2/6] p2-schema-think-true-r1 parsed schema_enforced=True 19.66s [probe 3/6] p3-bare-json-think-false-r1 parsed schema_enforced=None 2.32s [probe 4/6] p1-schema-think-false-r2 parsed schema_enforced=True 2.62s [probe 5/6] p2-schema-think-true-r2 parsed schema_enforced=True 22.62s [probe 6/6] p3-bare-json-think-false-r2 parsed schema_enforced=None 2.13s == qwen3.6:27b: polarity=SCHEMA HELD BOTH WAYS schema-ok (6 calls) -> probes-qwen3.6_27b.json [2026-08-26T13:56:50Z] -- END c0-probes/qwen3.6:27b ok (62s) [2026-08-26T13:56:50Z] -- BEGIN c1-judge/qwen3.6:27b === qwen3.6:27b — legs ['C1'] === [C1 1/21] c1-p-bez-01 posture=think:false PASS pass=True [C1 2/21] c1-p-six-01 posture=think:false PASS pass=True [C1 3/21] c1-p-nap-01 posture=think:false None pass=None [C1 4/21] c1-p-piq-01 posture=think:false PASS pass=True [C1 5/21] c1-p-dom-01 posture=think:false PASS pass=True [C1 6/21] c1-p-chk-01 posture=think:false PASS pass=True [C1 7/21] c1-p-con-01 posture=think:false PASS pass=True [C1 8/21] c1-p-spf-01 posture=think:false PASS pass=True [C1 9/21] c1-p-bkg-01 posture=think:false PASS pass=True [C1 10/21] c1-k-piq-c1 posture=think:false FAIL pass=True [C1 11/21] c1-k-bez-c1 posture=think:false FAIL pass=True [C1 12/21] c1-k-cal-c1 posture=think:false FAIL pass=True [C1 13/21] c1-k-con-c1 posture=think:false FAIL pass=True [C1 14/21] c1-k-dom-n1 posture=think:false FAIL pass=True [C1 15/21] c1-k-eca-n1 posture=think:false FAIL pass=True [C1 16/21] c1-k-spf-n1 posture=think:false FAIL pass=True [C1 17/21] c1-k-piq-s1 posture=think:false FAIL pass=True [C1 18/21] c1-k-con-s1 posture=think:false FAIL pass=True [C1 19/21] c1-k-bez-h1 posture=think:false PASS pass=False [C1 20/21] c1-k-six-w1 posture=think:false FAIL pass=True [C1 21/21] c1-k-chk-w1 posture=think:false FAIL pass=True >> C1 qwen3.6:27b posture=think:false: RANKED — carried under protocol wrote summary-C1-qwen3.6_27b.json wrote summary-C1.json (index) [2026-08-26T13:59:57Z] -- END c1-judge/qwen3.6:27b ok (187s) [2026-08-26T13:59:57Z] -- BEGIN c2-assistant/qwen3.6:27b leg C2 -- 20 items x 2 repeats sha assistant-c2.json dd63efa94bdebc878e1dc1c1065913892d8eb7cf16026642ff9fd0bae470d50c sha c2_checkers.py a9ebbd91503b936f3427cf7e6410bbacb45bb1574f750a8b45be9ae03481727a qwen3.6:27b postures=['think:false'] total calls: 40 === qwen3.6:27b [think:false] === r1 [ 1/20] c2-cw-1 ok 11.1s chars= 238 r1 [ 2/20] c2-cw-2 ok 1.4s chars= 190 r1 [ 3/20] c2-cw-3 ok 1.2s chars= 87 r1 [ 4/20] c2-cw-4 ok 1.0s chars= 61 r1 [ 5/20] c2-sum-1 ok 2.5s chars= 332 r1 [ 6/20] c2-sum-2 ok 2.8s chars= 388 r1 [ 7/20] c2-json-1 ok 1.9s chars= 98 r1 [ 8/20] c2-json-2 ok 2.3s chars= 180 r1 [ 9/20] c2-num-1 ok 6.5s chars= 994 r1 [10/20] c2-num-2 ok 12.8s chars= 2167 r1 [11/20] c2-num-3 ok 6.8s chars= 1102 r1 [12/20] c2-tone-1 ok 3.4s chars= 751 r1 [13/20] c2-sh-1 ok 1.0s chars= 51 r1 [14/20] c2-re-1 ok 1.0s chars= 26 r1 [15/20] c2-ctx-8k ok 7.7s chars= 71 r1 [16/20] c2-ctx-24k ok 22.8s chars= 84 r1 [17/20] c2-hon-1 ok 11.1s chars= 1988 r1 [18/20] c2-hon-2 ok 6.2s chars= 1302 r1 [19/20] c2-hon-3 ok 1.8s chars= 277 r1 [20/20] c2-hon-4 ok 10.5s chars= 2264 r2 [ 1/20] c2-cw-1 ok 2.1s chars= 238 r2 [ 2/20] c2-cw-2 ok 1.4s chars= 190 r2 [ 3/20] c2-cw-3 ok 1.2s chars= 87 r2 [ 4/20] c2-cw-4 ok 1.0s chars= 61 r2 [ 5/20] c2-sum-1 ok 2.5s chars= 332 -- pin checkpoint after 25 calls: OK r2 [ 6/20] c2-sum-2 ok 2.8s chars= 388 r2 [ 7/20] c2-json-1 ok 1.9s chars= 98 r2 [ 8/20] c2-json-2 ok 2.4s chars= 180 r2 [ 9/20] c2-num-1 ok 6.5s chars= 994 r2 [10/20] c2-num-2 ok 12.9s chars= 2167 r2 [11/20] c2-num-3 ok 6.8s chars= 1102 r2 [12/20] c2-tone-1 ok 3.4s chars= 751 r2 [13/20] c2-sh-1 ok 1.0s chars= 51 r2 [14/20] c2-re-1 ok 1.1s chars= 26 r2 [15/20] c2-ctx-8k ok 7.7s chars= 71 r2 [16/20] c2-ctx-24k ok 22.7s chars= 84 r2 [17/20] c2-hon-1 ok 11.1s chars= 1988 r2 [18/20] c2-hon-2 ok 6.2s chars= 1302 r2 [19/20] c2-hon-3 ok 1.8s chars= 277 r2 [20/20] c2-hon-4 ok 10.5s chars= 2264 40 calls, 0 response failures, 0 truncated (done_reason=length) raw: raw/c2 [2026-08-26T14:03:40Z] -- END c2-assistant/qwen3.6:27b ok (223s) [2026-08-26T14:03:40Z] -- BEGIN c3-tools/qwen3.6:27b C3 plan — 1 model rows, 19 tasks x 2 repeats, 10 tools. FLOOR 42 /api/chat calls (one per attempt); ESTIMATE ~90 at the 2.26 rounds/attempt measured on the offline oracle run. A model that gropes or loops costs more; the 8-round cap is the ceiling. qwen3.6:27b postures=['think_false'] floor= 42 estimate= 90 prod untouched OK — 3 seat(s) read-only: the 96 GB box's production seat A, the 96 GB box's production seat B, the bench box's production seat A === qwen3.6:27b — think:false only === pre-probe: CARRIED postures=['think_false', 'think_true'] -> preprobe.json [r1 1/19] t01-single-calculator answered rounds=2 calls=1 2.4s [r1 2/19] t02-single-unit-convert answered rounds=2 calls=1 2.7s [r1 3/19] t03-single-date-diff answered rounds=2 calls=1 3.1s [r1 4/19] t04-single-clock-now answered rounds=2 calls=1 2.4s [r1 5/19] t05-single-read-file answered rounds=3 calls=2 3.7s [r1 6/19] t06-single-sqlite answered rounds=2 calls=1 2.3s [r1 7/19] t07-chain-search-then-read answered rounds=6 calls=5 7.5s [r1 8/19] t08-chain-list-read-calculate answered rounds=4 calls=3 4.8s [r1 9/19] t09-chain-schedule-berth-draught answered rounds=3 calls=3 7.0s [r1 10/19] t10-chain-clock-schedule-diff-note answered rounds=4 calls=4 8.0s [r1 11/19] t11-chain-http-then-datediff answered rounds=2 calls=1 3.2s [r1 12/19] t12-distractor-database-not-files answered rounds=2 calls=1 2.4s [r1 13/19] t13-distractor-service-not-file answered rounds=2 calls=1 3.1s [r1 14/19] t14-honesty-eta answered rounds=1 calls=0 1.2s [r1 15/19] t15-honesty-draught answered rounds=1 calls=0 1.3s [r1 16/19] t16-honesty-fortnight answered rounds=1 calls=0 1.6s [r1 17/19] t17-honesty-name-the-tool answered rounds=1 calls=0 1.3s [r1 18/19] t18-honesty-false-premise-weather answered rounds=4 calls=3 5.4s [r1 19/19] t19-error-recovery-roster answered rounds=5 calls=4 5.5s [r2 1/19] t01-single-calculator answered rounds=2 calls=1 2.4s [r2 2/19] t02-single-unit-convert answered rounds=2 calls=1 2.7s [r2 3/19] t03-single-date-diff answered rounds=2 calls=1 3.1s [r2 4/19] t04-single-clock-now answered rounds=2 calls=1 2.4s [r2 5/19] t05-single-read-file answered rounds=3 calls=2 3.7s [r2 6/19] t06-single-sqlite answered rounds=2 calls=1 2.3s · prod untouched OK — 3 seat(s) read-only: the 96 GB box's production seat A, the 96 GB box's production seat B, the bench box's production seat A [r2 7/19] t07-chain-search-then-read answered rounds=6 calls=5 7.5s [r2 8/19] t08-chain-list-read-calculate answered rounds=4 calls=3 4.8s [r2 9/19] t09-chain-schedule-berth-draught answered rounds=3 calls=3 7.0s [r2 10/19] t10-chain-clock-schedule-diff-note answered rounds=4 calls=4 8.0s [r2 11/19] t11-chain-http-then-datediff answered rounds=2 calls=1 3.2s [r2 12/19] t12-distractor-database-not-files answered rounds=2 calls=1 2.4s [r2 13/19] t13-distractor-service-not-file answered rounds=2 calls=1 3.1s [r2 14/19] t14-honesty-eta answered rounds=1 calls=0 1.2s [r2 15/19] t15-honesty-draught answered rounds=1 calls=0 1.3s [r2 16/19] t16-honesty-fortnight answered rounds=1 calls=0 1.6s [r2 17/19] t17-honesty-name-the-tool answered rounds=1 calls=0 1.3s [r2 18/19] t18-honesty-false-premise-weather answered rounds=4 calls=3 5.4s [r2 19/19] t19-error-recovery-roster answered rounds=5 calls=4 5.5s prod untouched OK — 3 seat(s) read-only: the 96 GB box's production seat A, the 96 GB box's production seat B, the bench box's production seat A wrote raw/c3/manifest.json now score it: python3 tools/c3_score.py [2026-08-26T14:06:04Z] -- END c3-tools/qwen3.6:27b ok (144s) [2026-08-26T14:06:04Z] -- BEGIN c5-speed/qwen3.6:27b === C5 qwen3.6:27b === warmup-discard qwen3.6:27b 1000tok: 62.0 tok/s (DISCARDED — not a statistic) warm qwen3.6:27b 1000tok r1/3: 62.09 tok/s ttft~537.23ms warm qwen3.6:27b 1000tok r2/3: 62.22 tok/s ttft~539.71ms warm qwen3.6:27b 1000tok r3/3: 62.13 tok/s ttft~537.08ms warm qwen3.6:27b 1000tok r1/10: 61.84 tok/s ttft~535.25ms warm qwen3.6:27b 1000tok r2/10: 62.1 tok/s ttft~536.88ms warm qwen3.6:27b 1000tok r3/10: 61.98 tok/s ttft~535.7ms warm qwen3.6:27b 1000tok r4/10: 61.96 tok/s ttft~536.89ms warm qwen3.6:27b 1000tok r5/10: 61.88 tok/s ttft~538.92ms warm qwen3.6:27b 1000tok r6/10: 62.02 tok/s ttft~536.74ms warm qwen3.6:27b 1000tok r7/10: 61.59 tok/s ttft~540.31ms warm qwen3.6:27b 1000tok r8/10: 62.07 tok/s ttft~537.3ms warm qwen3.6:27b 1000tok r9/10: 62.05 tok/s ttft~534.91ms warm qwen3.6:27b 1000tok r10/10: 61.83 tok/s ttft~538.07ms warm qwen3.6:27b 8000tok r1/10: 60.64 tok/s ttft~6910.74ms warm qwen3.6:27b 8000tok r2/10: 60.71 tok/s ttft~618.43ms warm qwen3.6:27b 8000tok r3/10: 60.79 tok/s ttft~612.34ms warm qwen3.6:27b 8000tok r4/10: 60.78 tok/s ttft~620.1ms warm qwen3.6:27b 8000tok r5/10: 60.61 tok/s ttft~613.93ms warm qwen3.6:27b 8000tok r6/10: 60.75 tok/s ttft~612.39ms warm qwen3.6:27b 8000tok r7/10: 60.85 tok/s ttft~615.65ms warm qwen3.6:27b 8000tok r8/10: 60.74 tok/s ttft~624.19ms warm qwen3.6:27b 8000tok r9/10: 60.86 tok/s ttft~617.73ms warm qwen3.6:27b 8000tok r10/10: 60.69 tok/s ttft~613.94ms warm qwen3.6:27b 32000tok r1/10: 58.91 tok/s ttft~30636.82ms warm qwen3.6:27b 32000tok r2/10: 59.12 tok/s ttft~901.45ms warm qwen3.6:27b 32000tok r3/10: 59.25 tok/s ttft~896.8ms warm qwen3.6:27b 32000tok r4/10: 59.02 tok/s ttft~900.38ms warm qwen3.6:27b 32000tok r5/10: 59.06 tok/s ttft~897.96ms warm qwen3.6:27b 32000tok r6/10: 59.08 tok/s ttft~900.94ms warm qwen3.6:27b 32000tok r7/10: 59.17 tok/s ttft~899.14ms warm qwen3.6:27b 32000tok r8/10: 59.34 tok/s ttft~901.9ms warm qwen3.6:27b 32000tok r9/10: 59.26 tok/s ttft~896.78ms warm qwen3.6:27b 32000tok r10/10: 59.48 tok/s ttft~898.87ms [2026-08-26T14:07:47Z] roster-trim: armed — waiting for qwen3.6:27b to finish its full battery cold-load qwen3.6:27b 1/3: 9.883s cold-load qwen3.6:27b 2/3: 9.999s cold-load qwen3.6:27b 3/3: 10.046s >> C5 qwen3.6:27b: RANKED — carried under protocol Traceback (most recent call last): File "harness/ollama_houselaw.py", line 354, in _post with urllib.request.urlopen(req, timeout=timeout) as resp: ~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^ File "/usr/lib/python3.14/urllib/request.py", line 187, in urlopen return opener.open(url, data, timeout) ~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^ File "/usr/lib/python3.14/urllib/request.py", line 493, in open response = meth(req, response) File "/usr/lib/python3.14/urllib/request.py", line 602, in http_response response = self.parent.error( 'http', request, response, code, msg, hdrs) File "/usr/lib/python3.14/urllib/request.py", line 531, in error return self._call_chain(*args) ~~~~~~~~~~~~~~~~^^^^^^^ File "/usr/lib/python3.14/urllib/request.py", line 464, in _call_chain result = func(*args) File "/usr/lib/python3.14/urllib/request.py", line 611, in http_error_default raise HTTPError(req.full_url, code, msg, hdrs, fp) urllib.error.HTTPError: HTTP Error 404: Not Found The above exception was the direct cause of the following exception: Traceback (most recent call last): File "harness/c5_speed.py", line 1234, in raise SystemExit(main()) ~~~~^^ File "harness/c5_speed.py", line 1196, in main result = run_matrix( models, args.host, manifest, shell, ...<3 lines>... run_label=args.run_label, ) File "harness/c5_speed.py", line 1048, in run_matrix result["dflash_bytecompare"] = dflash_bytecompare( ~~~~~~~~~~~~~~~~~~^ host, out, manifest, run_label=run_label ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ ) ^ File "harness/c5_speed.py", line 626, in dflash_bytecompare core.release_candidate(model, host=host, manifest=manifest) ~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "harness/core.py", line 736, in release_candidate unloads.append({"attempt": attempt, **_unload(target, host, timeout)}) ~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^ File "harness/core.py", line 687, in _unload resp, wall = houselaw.generate(body, host=host, timeout=timeout, allow_keep_alive=True) ~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "harness/ollama_houselaw.py", line 384, in generate return _post("/api/generate", body, host, timeout, allow_keep_alive) File "harness/ollama_houselaw.py", line 358, in _post raise RuntimeError(f"HTTP {exc.code} from {path}: {detail}") from exc RuntimeError: HTTP 404 from /api/generate: {"error":"model 'muse-glimmer:30b-q4_K_M-dflash' not found"} [2026-08-26T14:08:19Z] !! c5-speed/qwen3.6:27b FAILED rc=1 (135s) — recorded, battery continues [2026-08-26T14:08:19Z] -- BEGIN c7-filing/qwen3.6:27b 8k_d10_recall0 PASS 8k_d10_recall1 PASS 8k_d10_absent0 PASS 8k_d10_absent1 PASS 8k_d50_recall0 PASS 8k_d50_recall1 PASS 8k_d50_absent0 PASS 8k_d50_absent1 PASS 8k_d90_recall0 PASS 8k_d90_recall1 PASS 8k_d90_absent0 PASS 8k_d90_absent1 PASS 16k_d10_recall0 PASS 16k_d10_recall1 PASS 16k_d10_absent0 PASS 16k_d10_absent1 PASS 16k_d50_recall0 PASS 16k_d50_recall1 fail 16k_d50_absent0 PASS 16k_d50_absent1 PASS 16k_d90_recall0 PASS 16k_d90_recall1 PASS 16k_d90_absent0 PASS 16k_d90_absent1 PASS 32k_d10_recall0 fail 32k_d10_recall1 PASS 32k_d10_absent0 PASS 32k_d10_absent1 PASS 32k_d50_recall0 PASS 32k_d50_recall1 PASS 32k_d50_absent0 PASS 32k_d50_absent1 PASS 32k_d90_recall0 PASS 32k_d90_recall1 PASS 32k_d90_absent0 PASS 32k_d90_absent1 PASS == C7 qwen3.6:27b: DESCRIPTIVE · recall 16/18 · abstention 18/18 · fabrications 0 · failures 0/72 -> c7-qwen3.6_27b.json [2026-08-26T14:15:41Z] -- END c7-filing/qwen3.6:27b ok (442s) [2026-08-26T14:15:41Z] -- BEGIN field-exam/qwen3.6:27b ......x............. 20/60 ...............x.... 40/60 .................... 60/60 ============================================================ model: qwen3.6:27b grounded-qa 19/20 95.0% stay-grounded 20/20 100.0% schema-extract 19/20 95.0% OVERALL 58/60 96.7% wrote field/qwen3.6-27b.json [2026-08-26T14:16:47Z] -- END field-exam/qwen3.6:27b ok (66s) [2026-08-26T14:16:47Z] -- BEGIN seat-43/qwen3.6:27b runid nightbattery-qwen3.6_27b base-url the bench endpoint results raw/seat43/judges-nightbattery-qwen3.6_27b.jsonl models qwen3.6:27b cases 43 repeats 3 calls 129 (before resume skips) schema-mode prompted (no `format` is sent) options {"num_ctx": 16384, "num_predict": 1024, "temperature": 0.0, "top_p": 1.0} num_ctx 16384 keep_alive 0 on candidates think omit courtesy off auth no key sent prompt packs-claim-judge-v2-prompted sha256 92af6030dddface2… (the judge-exam fixture store/judge-prompt-packs-v2-prompted.md) [probe] qwen3.6:27b: think_honored=None options_honored=True shape_ok=False [measurement-failure] qwen3.6:27b skagway-arctic-brotherhood-hall-04 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [qwen3.6:27b] skagway-centennial-snowplow-02 r1 -> FAIL (22585 ms) [measurement-failure] qwen3.6:27b skagway-golden-north-hotel-02 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [qwen3.6:27b] skagway-golden-north-hotel-03 r1 -> FAIL (23529 ms) [measurement-failure] qwen3.6:27b skagway-golden-north-hotel-04 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-golden-north-hotel-05 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-jeff-smiths-parlor-02 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-kirmses-curios-02 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-kirmses-curios-03 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-kirmses-curios-04 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-kirmses-curios-05 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [qwen3.6:27b] skagway-mascot-saloon-06 r1 -> FAIL (24922 ms) [qwen3.6:27b] skagway-mccabe-college-01 r1 -> PASS (24385 ms) [measurement-failure] qwen3.6:27b skagway-mccabe-college-02 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [qwen3.6:27b] skagway-mccabe-college-03 r1 -> PASS (24120 ms) [measurement-failure] qwen3.6:27b skagway-mccabe-college-04 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [qwen3.6:27b] skagway-mccabe-college-05 r1 -> FAIL (22369 ms) [qwen3.6:27b] skagway-mccabe-college-06 r1 -> PASS (24516 ms) [measurement-failure] qwen3.6:27b skagway-mccabe-college-07 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-mollie-walsh-park-01 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-mollie-walsh-park-02 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-mollie-walsh-park-03 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-mollie-walsh-park-05 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [qwen3.6:27b] skagway-mollie-walsh-park-06 r1 -> PASS (23729 ms) [measurement-failure] qwen3.6:27b skagway-mollie-walsh-park-07 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-moore-homestead-01 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-moore-homestead-02 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-moore-homestead-03 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-moore-homestead-05 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-pantheon-saloon-01 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [qwen3.6:27b] skagway-pantheon-saloon-04 r1 -> FAIL (24374 ms) [measurement-failure] qwen3.6:27b skagway-pantheon-saloon-06 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-pullen-creek-harbor-02 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-pullen-creek-harbor-04 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-pullen-creek-harbor-08 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [qwen3.6:27b] skagway-red-onion-saloon-03 r1 -> FAIL (24451 ms) [qwen3.6:27b] skagway-red-onion-saloon-04 r1 -> PASS (23881 ms) [measurement-failure] qwen3.6:27b skagway-red-onion-saloon-05 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-red-onion-saloon-06 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-ship-registry-cliff-01 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-ship-registry-cliff-05 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-skagway-context-02 r1 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [qwen3.6:27b] skagway-wpyr-depot-08 r1 -> FAIL (23964 ms) [measurement-failure] qwen3.6:27b skagway-arctic-brotherhood-hall-04 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [qwen3.6:27b] skagway-centennial-snowplow-02 r2 -> FAIL (22684 ms) [measurement-failure] qwen3.6:27b skagway-golden-north-hotel-02 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [qwen3.6:27b] skagway-golden-north-hotel-03 r2 -> FAIL (22999 ms) [measurement-failure] qwen3.6:27b skagway-golden-north-hotel-04 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-golden-north-hotel-05 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-jeff-smiths-parlor-02 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-kirmses-curios-02 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-kirmses-curios-03 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-kirmses-curios-04 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-kirmses-curios-05 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [qwen3.6:27b] skagway-mascot-saloon-06 r2 -> FAIL (23979 ms) [qwen3.6:27b] skagway-mccabe-college-01 r2 -> PASS (24836 ms) [measurement-failure] qwen3.6:27b skagway-mccabe-college-02 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [qwen3.6:27b] skagway-mccabe-college-03 r2 -> PASS (24076 ms) [measurement-failure] qwen3.6:27b skagway-mccabe-college-04 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [qwen3.6:27b] skagway-mccabe-college-05 r2 -> FAIL (22389 ms) [qwen3.6:27b] skagway-mccabe-college-06 r2 -> PASS (24521 ms) [measurement-failure] qwen3.6:27b skagway-mccabe-college-07 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-mollie-walsh-park-01 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-mollie-walsh-park-02 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-mollie-walsh-park-03 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-mollie-walsh-park-05 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [qwen3.6:27b] skagway-mollie-walsh-park-06 r2 -> PASS (23761 ms) [measurement-failure] qwen3.6:27b skagway-mollie-walsh-park-07 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-moore-homestead-01 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-moore-homestead-02 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-moore-homestead-03 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-moore-homestead-05 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-pantheon-saloon-01 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [qwen3.6:27b] skagway-pantheon-saloon-04 r2 -> FAIL (24345 ms) [measurement-failure] qwen3.6:27b skagway-pantheon-saloon-06 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-pullen-creek-harbor-02 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-pullen-creek-harbor-04 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-pullen-creek-harbor-08 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [qwen3.6:27b] skagway-red-onion-saloon-03 r2 -> FAIL (24575 ms) [qwen3.6:27b] skagway-red-onion-saloon-04 r2 -> PASS (23862 ms) [measurement-failure] qwen3.6:27b skagway-red-onion-saloon-05 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-red-onion-saloon-06 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-ship-registry-cliff-01 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-ship-registry-cliff-05 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-skagway-context-02 r2 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [qwen3.6:27b] skagway-wpyr-depot-08 r2 -> FAIL (23935 ms) [measurement-failure] qwen3.6:27b skagway-arctic-brotherhood-hall-04 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [qwen3.6:27b] skagway-centennial-snowplow-02 r3 -> FAIL (22662 ms) [measurement-failure] qwen3.6:27b skagway-golden-north-hotel-02 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [qwen3.6:27b] skagway-golden-north-hotel-03 r3 -> FAIL (22971 ms) [measurement-failure] qwen3.6:27b skagway-golden-north-hotel-04 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-golden-north-hotel-05 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-jeff-smiths-parlor-02 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-kirmses-curios-02 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-kirmses-curios-03 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-kirmses-curios-04 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-kirmses-curios-05 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [qwen3.6:27b] skagway-mascot-saloon-06 r3 -> FAIL (24464 ms) [qwen3.6:27b] skagway-mccabe-college-01 r3 -> PASS (24846 ms) [measurement-failure] qwen3.6:27b skagway-mccabe-college-02 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [qwen3.6:27b] skagway-mccabe-college-03 r3 -> PASS (24106 ms) [measurement-failure] qwen3.6:27b skagway-mccabe-college-04 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [qwen3.6:27b] skagway-mccabe-college-05 r3 -> FAIL (22403 ms) [qwen3.6:27b] skagway-mccabe-college-06 r3 -> PASS (24517 ms) [measurement-failure] qwen3.6:27b skagway-mccabe-college-07 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-mollie-walsh-park-01 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-mollie-walsh-park-02 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-mollie-walsh-park-03 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-mollie-walsh-park-05 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [qwen3.6:27b] skagway-mollie-walsh-park-06 r3 -> PASS (23712 ms) [measurement-failure] qwen3.6:27b skagway-mollie-walsh-park-07 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-moore-homestead-01 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-moore-homestead-02 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-moore-homestead-03 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-moore-homestead-05 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-pantheon-saloon-01 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [qwen3.6:27b] skagway-pantheon-saloon-04 r3 -> FAIL (24357 ms) [measurement-failure] qwen3.6:27b skagway-pantheon-saloon-06 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-pullen-creek-harbor-02 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-pullen-creek-harbor-04 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-pullen-creek-harbor-08 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [qwen3.6:27b] skagway-red-onion-saloon-03 r3 -> FAIL (24553 ms) [qwen3.6:27b] skagway-red-onion-saloon-04 r3 -> PASS (24009 ms) [measurement-failure] qwen3.6:27b skagway-red-onion-saloon-05 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-red-onion-saloon-06 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-ship-registry-cliff-01 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-ship-registry-cliff-05 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [measurement-failure] qwen3.6:27b skagway-skagway-context-02 r3 done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer [qwen3.6:27b] skagway-wpyr-depot-08 r3 -> FAIL (23958 ms) calls 129 · resumed-skips 0 · measurement failures 93 models run: qwen3.6:27b models unavailable: (none) models NOT RUN (probe): (none) models aborted mid-batch: (none) score it: python3 bench/score_judges.py --results raw/seat43/judges-nightbattery-qwen3.6_27b.jsonl [2026-08-26T15:12:18Z] -- END seat-43/qwen3.6:27b ok (3331s) [2026-08-26T15:12:18Z] -- BEGIN seat-43-score/qwen3.6:27b # Judge bake-off — judges-nightbattery-qwen3.6_27b.jsonl Computed by `bench/score_judges.py` from the append-only rows in `raw/seat43/judges-nightbattery-qwen3.6_27b.jsonl`. Every number below is a count of rows in that file; the per-case appendix prints the rows themselves so the table can be recomputed rather than believed. ## What was measured | | | | --- | --- | | generated | 2026-08-26T15:12:18Z | | results | `raw/seat43/judges-nightbattery-qwen3.6_27b.jsonl` (133 rows) | | fixture | `the judge-exam fixture store/judge-cases-v1.json` — tour-judge-v1 | | fixture sha256 | `6a226b62dceb005dbd658bc13cf229e38f6f0a2f07ae51091f444ec3e96b56d1` | | thresholds | `the judge-exam fixture store/THRESHOLDS.md` sha256 `5391a9d465658d66fa33d57c14aff7eb0aafaa3b41a9c0a2a7cb33aa973ec3d2` | | runid | `nightbattery-qwen3.6_27b` | | runner | d-runner-v2 on python 3.14.4 | | base-url | `the bench endpoint` | | prompt | packs-claim-judge-v2-prompted sha256 `92af6030dddface2c863fe9d672bec6a6bb1017122c0028e435214acd9fb298d` | | schema mode | prompted | | repeats | 3 | | options | `{"num_ctx": 16384, "num_predict": 1024, "temperature": 0.0, "top_p": 1.0}` | | num_ctx | 16384 | | keep_alive | 0 | | think | omit | | auth | no key sent | | courtesy | off | | superseded rows | 0 (retries; last row per triple scored) | ## The floors, as pre-registered Read from `the judge-exam fixture store/THRESHOLDS.md` at scoring time, not retyped. The lines themselves: > Minimum kill-recall — **23 of 27** (0.852) > Minimum confirmed-preservation — **13 of 16** (0.8125) > Latency ceiling — **p50 ≤ 10 s, p95 ≤ 20 s per case** **Both bind independently** (§3.3): a candidate under either floor is out regardless of the other. A candidate exactly one case short on exactly one floor is a RE-TEST at 9 repeats, not an elimination (§6); short on both is an elimination with no re-test. Over 10% measurement failures a candidate is UNMEASURABLE at this config and is not scored at all. ## Candidates **Latency is reported and NON-BINDING in this run** (PREREG-D-RUN-V4 §4): cloud responses carry no `load_duration`, so warm-equivalent is uncomputable and a ceiling could only be applied to WAN wall-clock or to server-total. No seat below turns on it, and the latency column reads `—` rather than a verdict it cannot make. | model | kills /27 | preservation /16 | kill floor | pres. floor | wall p50 | wall p95 | latency | resp. fail | transport | fenced | verdict | | --- | ---: | ---: | :--: | :--: | ---: | ---: | :--: | ---: | ---: | ---: | :--: | | `qwen3.6:27b` | 7 | 5 | FAIL | FAIL | 23979 ms | 24846 ms | — | 93/129 | 0/129 | 0 | **UNMEASURABLE** | - `qwen3.6:27b` — 93/129 calls (72.1%) were RESPONSE failures, over the 10% limit — reported UNMEASURABLE at this config, not scored. ### Pre-run probes Every probe sent the EXACT contracted envelope and differed only in its message. The raw probe rows are in this results file (`note: probe`). | candidate | think honored | options honored | shape sighting | probes skipped (resume) | | --- | :--: | :--: | :--: | :--: | | `qwen3.6:27b` | None | True | False | False | **Candidates are listed alphabetically and are NOT ranked.** THRESHOLDS §3.5's five-step tie-break is not implemented here, and neither is its mixed-basis rule (tie-breaks 1–2 on the 9-repeat majority for a re-tested candidate, 3–5 on the first 3 repeats for everyone). If more than one candidate clears both floors and the latency ceiling, **apply §3.5 by hand** — the order of the rows above carries no meaning beyond the alphabet. A candidate under either floor is out under §3.3 regardless of the other, so nothing here is a close-second either. Its inputs, computed but deliberately not combined — steps 1 and 2 are kill-recall then preservation, 3 is self-consistency, 4 is p95 latency, 5 is resident VRAM, which this instrument does not measure: | candidate | 1. kill-recall | 2. preservation | 3. self-consistency | 4. p95 | 5. resident VRAM | | --- | ---: | ---: | ---: | ---: | ---: | | `qwen3.6:27b` | 7/27 (0.259) | 5/16 (0.312) | 12/12 | 24846 ms | — (`bench/candidates.py`) | ## `qwen3.6:27b` 129 triples scored — one per `(case, repeat)` — with 93 RESPONSE failures (72.1%) and 36 latencies from answered calls. **Run completeness.** 129 usable rows of the 129 this run-start asked for; 0 transport failures (connection / timeout / HTTP status). Transport is the LINK's figure and never the candidate's: it is excluded from the 10% ceiling above, and while any of it is uncured the candidate is NOT SCORED rather than judged on whatever arrived. Cure: `--retry transport` under the same runid. **Latency, three figures, none substituted for another.** ollama reports durations in nanoseconds; every ms below is that value divided by 1000000 (floor). `warm` is `total_duration − load_duration` and is computable ONLY where the response carries `load_duration` — cloud responses do not, so a cloud candidate's warm column is empty rather than a wall figure wearing a warm label. | figure | what it measures | p50 | p95 | n | | --- | --- | ---: | ---: | ---: | | `wall_ms` | perf_counter around the call, link included | 23979 ms | 24846 ms | 36 | | `server_total_ms` | `total_duration`, server-side, excludes the WAN | 23975 ms | 24843 ms | 36 | | `warm_ms` | `total_duration − load_duration`, local rows only | 13749 ms | 14603 ms | 36 | | axis | count | of key | rate | floor | verdict | | --- | ---: | ---: | ---: | ---: | :--: | | kill catches | 7 | 27 | 0.259 | 23 | FAIL | | preserved | 5 | 16 | 0.312 | 13 | FAIL | Kill half: 7 catches, 0 misses, 0 ties, 20 unmeasured. Preservation half: 5 preserved, 0 lost, 0 ties, 11 unmeasured. ### By kill class The sort THRESHOLDS §2.1 derives the kill floor from. `non-entailment` is the row that matters: refusing on silence rather than on conflict. | kill class | n | catches | misses | ties | unmeasured | recall | | --- | ---: | ---: | ---: | ---: | ---: | ---: | | blatant-contradiction | 12 | 4 | 0 | 0 | 8 | 0.333 | | degenerate-evidence | 1 | 1 | 0 | 0 | 0 | 1.000 | | hedge-dropped | 3 | 0 | 0 | 0 | 3 | 0.000 | | non-entailment | 8 | 2 | 0 | 0 | 6 | 0.250 | | scope-shift | 3 | 0 | 0 | 0 | 3 | 0.000 | ### Repeats Self-consistency 12/12 cases unanimous across their usable repeats (1.000). No case flipped verdict across repeats. ### Measurement failures 93 of 129 calls (72.1%); the limit is 10%. None of these were mapped to a verdict — on a kill case `UNCERTAIN` counts as a catch, so a truncation scored as `UNCERTAIN` would read as 1.00 kill-recall and 0.00 preservation for a model that answered nothing. | failure kind | calls | | --- | ---: | | `done_reason` | 93 | | case | repeats failed | kinds | | --- | ---: | --- | | `skagway-arctic-brotherhood-hall-04` | 3 | done_reason | | `skagway-golden-north-hotel-02` | 3 | done_reason | | `skagway-golden-north-hotel-04` | 3 | done_reason | | `skagway-golden-north-hotel-05` | 3 | done_reason | | `skagway-jeff-smiths-parlor-02` | 3 | done_reason | | `skagway-kirmses-curios-02` | 3 | done_reason | | `skagway-kirmses-curios-03` | 3 | done_reason | | `skagway-kirmses-curios-04` | 3 | done_reason | | `skagway-kirmses-curios-05` | 3 | done_reason | | `skagway-mccabe-college-02` | 3 | done_reason | | `skagway-mccabe-college-04` | 3 | done_reason | | `skagway-mccabe-college-07` | 3 | done_reason | | `skagway-mollie-walsh-park-01` | 3 | done_reason | | `skagway-mollie-walsh-park-02` | 3 | done_reason | | `skagway-mollie-walsh-park-03` | 3 | done_reason | | `skagway-mollie-walsh-park-05` | 3 | done_reason | | `skagway-mollie-walsh-park-07` | 3 | done_reason | | `skagway-moore-homestead-01` | 3 | done_reason | | `skagway-moore-homestead-02` | 3 | done_reason | | `skagway-moore-homestead-03` | 3 | done_reason | | `skagway-moore-homestead-05` | 3 | done_reason | | `skagway-pantheon-saloon-01` | 3 | done_reason | | `skagway-pantheon-saloon-06` | 3 | done_reason | | `skagway-pullen-creek-harbor-02` | 3 | done_reason | | `skagway-pullen-creek-harbor-04` | 3 | done_reason | | `skagway-pullen-creek-harbor-08` | 3 | done_reason | | `skagway-red-onion-saloon-05` | 3 | done_reason | | `skagway-red-onion-saloon-06` | 3 | done_reason | | `skagway-ship-registry-cliff-01` | 3 | done_reason | | `skagway-ship-registry-cliff-05` | 3 | done_reason | | `skagway-skagway-context-02` | 3 | done_reason | ### Every case, every repeat | case | expected | kill class | repeats | successes | outcome | latencies | | --- | --- | --- | --- | ---: | --- | --- | | `skagway-arctic-brotherhood-hall-04` | NON_PASS | hedge-dropped | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-centennial-snowplow-02` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 22585, 22684, 22662 | | `skagway-golden-north-hotel-02` | NON_PASS | blatant-contradiction | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-golden-north-hotel-03` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 23529, 22999, 22971 | | `skagway-golden-north-hotel-04` | NON_PASS | non-entailment | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-golden-north-hotel-05` | PASS | — | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-jeff-smiths-parlor-02` | NON_PASS | blatant-contradiction | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-kirmses-curios-02` | PASS | — | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-kirmses-curios-03` | NON_PASS | blatant-contradiction | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-kirmses-curios-04` | NON_PASS | non-entailment | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-kirmses-curios-05` | PASS | — | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-mascot-saloon-06` | NON_PASS | degenerate-evidence | FAIL, FAIL, FAIL | 3/3 | success | 24922, 23979, 24464 | | `skagway-mccabe-college-01` | PASS | — | PASS, PASS, PASS | 3/3 | success | 24385, 24836, 24846 | | `skagway-mccabe-college-02` | NON_PASS | scope-shift | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-mccabe-college-03` | PASS | — | PASS, PASS, PASS | 3/3 | success | 24120, 24076, 24106 | | `skagway-mccabe-college-04` | NON_PASS | hedge-dropped | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-mccabe-college-05` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 22369, 22389, 22403 | | `skagway-mccabe-college-06` | PASS | — | PASS, PASS, PASS | 3/3 | success | 24516, 24521, 24517 | | `skagway-mccabe-college-07` | PASS | — | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-mollie-walsh-park-01` | PASS | — | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-mollie-walsh-park-02` | PASS | — | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-mollie-walsh-park-03` | NON_PASS | blatant-contradiction | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-mollie-walsh-park-05` | NON_PASS | blatant-contradiction | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-mollie-walsh-park-06` | PASS | — | PASS, PASS, PASS | 3/3 | success | 23729, 23761, 23712 | | `skagway-mollie-walsh-park-07` | NON_PASS | blatant-contradiction | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-moore-homestead-01` | PASS | — | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-moore-homestead-02` | NON_PASS | blatant-contradiction | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-moore-homestead-03` | NON_PASS | non-entailment | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-moore-homestead-05` | PASS | — | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-pantheon-saloon-01` | PASS | — | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-pantheon-saloon-04` | NON_PASS | non-entailment | FAIL, FAIL, FAIL | 3/3 | success | 24374, 24345, 24357 | | `skagway-pantheon-saloon-06` | NON_PASS | non-entailment | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-pullen-creek-harbor-02` | NON_PASS | scope-shift | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-pullen-creek-harbor-04` | PASS | — | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-pullen-creek-harbor-08` | NON_PASS | non-entailment | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-red-onion-saloon-03` | NON_PASS | non-entailment | FAIL, FAIL, FAIL | 3/3 | success | 24451, 24575, 24553 | | `skagway-red-onion-saloon-04` | PASS | — | PASS, PASS, PASS | 3/3 | success | 23881, 23862, 24009 | | `skagway-red-onion-saloon-05` | NON_PASS | non-entailment | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-red-onion-saloon-06` | NON_PASS | hedge-dropped | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-ship-registry-cliff-01` | NON_PASS | blatant-contradiction | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-ship-registry-cliff-05` | PASS | — | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-skagway-context-02` | NON_PASS | scope-shift | FAILED, FAILED, FAILED | 0/0 | unmeasured | — | | `skagway-wpyr-depot-08` | NON_PASS | blatant-contradiction | FAIL, FAIL, FAIL | 3/3 | success | 23964, 23935, 23958 | ## Run notes - `probe` — qwen3.6:27b: reachable - `probe` — qwen3.6:27b: num_predict:10 -> eval_count=10, done_reason='length' - `probe` — qwen3.6:27b: done_reason: done_reason='length' (not 'stop'): the model stopped for some reason other than finishing, so whatever text arrived is a fragment, not an answer ## Recompute it ``` python3 bench/score_judges.py --results raw/seat43/judges-nightbattery-qwen3.6_27b.jsonl ``` [2026-08-26T15:12:18Z] -- END seat-43-score/qwen3.6:27b ok (0s) [2026-08-26T15:12:18Z] ########## CANDIDATE qwen3.6:27b COMPLETE ########## [2026-08-26T15:12:18Z] ########## CANDIDATE nemotron-3.5-lightning:30b ########## [2026-08-26T15:12:18Z] -- BEGIN fit-gate/nemotron-3.5-lightning:30b == fit gate: nemotron-3.5-lightning:30b @ num_ctx 32768 -> the bench endpoint [2026-08-26T15:12:29Z] roster-trim: qwen3.6:27b COMPLETE seen [2026-08-26T15:12:29Z] roster-trim: stopping nb-resume (prevents nemotron full battery + laguna re-run) roster-trim: swept candidates -> released=[] after=[] [2026-08-26T15:12:32Z] -- BEGIN fit-gate-only/nemotron-3.5-lightning:30b (roster trim) == fit gate: nemotron-3.5-lightning:30b @ num_ctx 32768 -> the bench endpoint == nemotron-3.5-lightning:30b: SPILLS — only 83.1% on GPU (4136 MiB to CPU) -> correctness chairs only, C5-32k/C7 NOT-RUN(spill) prod untouched: True -> fit-nemotron-3.5-lightning_30b.json [2026-08-26T15:13:00Z] -- END fit-gate-only/nemotron-3.5-lightning:30b rc=10 (28s) [2026-08-26T15:13:00Z] roster-trim: nemotron SPILLS — this is its headline result [2026-08-26T15:13:00Z] roster-trim: laguna-xs-2.1:latest DROPPED per ruling — row stays fit + C0 [2026-08-26T15:13:00Z] roster-trim: complete — battery is finished; standings then teardown [2026-08-26T18:13:04Z] teardown: results verified present on the workshop's record box; proceeding (dry_run=0) [2026-08-26T18:13:05Z] teardown: df before -> the bench box root filesystem 935G 684G 205G 77% / [2026-08-26T18:13:05Z] teardown: store before -> 95G [2026-08-26T18:13:05Z] teardown: removed muse-glimmer:30b [2026-08-26T18:13:06Z] teardown: removed gemma4:31b [2026-08-26T18:13:06Z] teardown: removed qwen3.6:27b [2026-08-26T18:13:07Z] teardown: removed nemotron-3.5-lightning:30b [2026-08-26T18:13:07Z] teardown: removed laguna-xs-2.1:latest [2026-08-26T18:13:08Z] teardown: df after -> the bench box root filesystem 935G 589G 299G 67% / [2026-08-26T18:13:08Z] teardown: store after -> 16K [2026-08-26T18:13:08Z] teardown: prod_untouched -> ok | the 96 GB box's production seat A=ok; the 96 GB box's production seat B=ok; the bench box's production seat A=ok [2026-08-26T18:13:08Z] teardown: receipt -> TEARDOWN-RECEIPT.md