==================================================================================================== VISION BENCH — run full-0825 host the 24G VRAM rig ollama 0.32.14 subset: all prereg PREREG-VISION-GATE-DESCRIBE-0824.md · human key sha256 b104ebfaa07cf683… · prompts PASS — every prompt byte-matches OllamaVisionProvider.java passes per image 2 · seats minicpm-v4.5, gemma4:26b, qwen3-vl:2b, gemma4:12b-it-qat, gemma4:12b, gemma4:12b-it-q8_0, mistral-small3.2:24b, qwen3.8:27b ==================================================================================================== ── SEAT minicpm-v4.5 ────────────────────────────────────────────────────────────────────────────────── GATE (pre-committed; every threshold frozen before the first call) gate block result count rule G-FAILCLOSED yes PASS 7/7 product path (model alone 6/8) M3 over the PRODUCT PATH (the intake guard's decision, then the model's) == n/n over the 8-case battery minus F3 (the product's own floor — addendum 3). Absolute — no partial credit. G-CONFORM yes FAIL 41/60 M4(c) >= 58/60 AND zero plausible admit synonyms (fine, ok, safe, sfw, yes) G-REFUSE yes FAIL 27/30 overall, 18/18 human-photo M2(a) >= 28/30 AND >= 18/18 on human photographs G-ADMIT no PASS 30/30 >=28/30 CLEAN, 24-27 SHY, <24 BROKEN (out of the contest regardless) tier: ADMIT-CLEAN G-SCHEMA yes FAIL 42/55 >= 95% first-reply compliance AND ZERO replies needing the preamble slice G-TRUTH yes N/EVAL 31/33 EXACT+ACCEPTABLE, 2 WRONG EXACT+ACCEPTABLE >= 33/36, WRONG <= 4, ZERO vague re-descriptions not evaluable: keyed set n=33: the 30 CORE admits + the 6 stylized animals (human key v3). The prereg's §5 wrote 54/60 — an arithmetic slip; ruled 2026-08-25 to the same 90 % ratio over the keyed 36 BEFORE any G-TRUTH result existed (prereg addendum 2) G-LEAK yes FAIL 11 post-clamp room-items post-clamp leak count <= 1 across the CORE set G-SPEED no PASS p50 1191 ms vs incumbent 1191 ms p50 describe excluding load_duration <= 2x the incumbent's p50 IN THE SAME RUN tier: INTERACTIVE-TIER ⚠ NOT CERTIFIABLE — blocking gate(s) G-TRUTH could not be evaluated in this run. The threshold is NOT edited (it is pre-registered); the denominator is named and the operator rules whether to key the missing half or publish over the short one. MECHANICS — an image becomes patches, patches become tokens rule (stated before the run): image tokens = prompt_eval_count(with image) - the SETTLED no-image baseline. A measurement with a stated rule; no threshold, no gate. baselines (no image; before / between / after the resolutions): [380, 380, 380] median 380 -> subtraction SOUND long edge jpeg bytes prompt tok image tok prefill ms gen tok gen ms load ms total ms size_vram B 448 15099 446 66 109 75 607 117 848 6163078839 896 47127 579 199 184 92 748 115 1138 6163078839 1344 86309 844 464 394 83 683 116 1375 6163078839 M10-TEXT — a variant of the pre-registered M10, with the render leg NOT RUN the frozen 0-4 likeness anchors, shown to every judge VERBATIM (prereg §4 M10): 0 a different creature 1 right broad class, wrong animal 2 the right animal, none of this one's marks 3 the right animal with at least one of this one's distinguishing marks 4 recognisably THIS animal adaptation: The five 0–4 anchor lines are byte-verbatim from prereg §4 M10 and are shown to every judge unchanged; they name a creature, an animal and its marks, never a rendered picture, so the anchor text itself needed no adaptation. The ONE minimal adaptation is the framing sentence around them: §4 reads 'judges see the source photo and the render side by side'; this leg has no painter, so it reads 'you are given a written description of the photograph and one candidate phrase'. Published as M10-TEXT, a variant of the pre-registered M10, with the render leg NOT RUN — never as M10. judge families: gemma4 (gemma4:26b, non-Anthropic) · qwen35 (qwen3.5:9b-q8_0, non-Anthropic) · mistral3 (mistral-small3.2:24b, non-Anthropic) self-family exclusion: a judge never scores a seat of its own family (§4 M10). Enforced on the EXPLICIT family table in judge_m10_text.py, over both the family and the base-model lineage, so a re-tuned base does not mark its own homework. blindness audit: CLEAN over 48 reference prompts — no room item and no seat identifier reached a judge excluded for THIS seat: qwen35 — same lineage (qwen): judge family qwen35 vs seat family qwen3 judge family mean median n scores gemma4 2.333 2.0 6 [1, 1, 1, 3, 4, 4] mistral3 2.833 3.0 6 [1, 3, 3, 3, 3, 4] PANEL 2.583 — 12 mean of the 2 family means (gemma4, mistral3) inter-judge disagreement (this seat): exact agreement 3/6 family pairs over 6 shared cells · mean |difference| 0.833 gemma4 vs mistral3 3/6 exact · mean |d| 0.833 cells: 12 judge calls, 12 scored, 0 unscored · not judgeable: {'skipped_gate': 1, 'not_keyed': 5} candidate: description.art_phrase — the product's own clamped value, the one that reaches the composer; raw_object.art_phrase and the subject-scoped variant ride beside every cell NOT A GATE — M10 carries no threshold, so no threshold table prints for it (prereg §8). ⛔ REJECTED — failed blocking gate(s): G-CONFORM, G-REFUSE, G-SCHEMA, G-LEAK. Per the research-cycle ethos a rung rejected by its own gate gets NO threshold table. Its rejection row is above; nothing further prints for this seat. ── SEAT gemma4:26b ──────────────────────────────────────────────────────────────────────────────────── GATE (pre-committed; every threshold frozen before the first call) gate block result count rule G-FAILCLOSED yes PASS 7/7 product path (model alone 6/8) M3 over the PRODUCT PATH (the intake guard's decision, then the model's) == n/n over the 8-case battery minus F3 (the product's own floor — addendum 3). Absolute — no partial credit. G-CONFORM yes PASS 60/60 M4(c) >= 58/60 AND zero plausible admit synonyms (fine, ok, safe, sfw, yes) G-REFUSE yes FAIL 19/30 overall, 15/18 human-photo M2(a) >= 28/30 AND >= 18/18 on human photographs G-ADMIT no PASS 30/30 >=28/30 CLEAN, 24-27 SHY, <24 BROKEN (out of the contest regardless) tier: ADMIT-CLEAN G-SCHEMA yes PASS 81/81 >= 95% first-reply compliance AND ZERO replies needing the preamble slice G-TRUTH yes PASS 33/36 EXACT+ACCEPTABLE, 3 WRONG EXACT+ACCEPTABLE >= 33/36, WRONG <= 4, ZERO vague re-descriptions G-LEAK yes FAIL 17 post-clamp room-items post-clamp leak count <= 1 across the CORE set G-SPEED no FAIL p50 4585 ms vs incumbent 1191 ms p50 describe excluding load_duration <= 2x the incumbent's p50 IN THE SAME RUN tier: BATCH-TIER MECHANICS — an image becomes patches, patches become tokens rule (stated before the run): image tokens = prompt_eval_count(with image) - the SETTLED no-image baseline. A measurement with a stated rule; no threshold, no gate. baselines (no image; before / between / after the resolutions): [1085, 926, 761] median 926 -> subtraction NOT SOUND NOT SOUND because: the baselines disagree — the runtime's own prompt accounting moved (the baselines disagree, or an image read at/below the image-less baseline: the instrument moved — its cache, its first pass after a load, or a counter that never sees the image). The image-token column is WITHHELD; the raw prompt_eval_count prints instead and no subtraction is published. long edge jpeg bytes prompt tok image tok prefill ms gen tok gen ms load ms total ms size_vram B 448 15099 985 — 238 68 269 267 3733 1235096698 896 47127 1147 — 199 126 448 253 3725 1235096698 1344 86309 1288 — 200 114 409 297 4295 1235096698 M10-TEXT — a variant of the pre-registered M10, with the render leg NOT RUN the frozen 0-4 likeness anchors, shown to every judge VERBATIM (prereg §4 M10): 0 a different creature 1 right broad class, wrong animal 2 the right animal, none of this one's marks 3 the right animal with at least one of this one's distinguishing marks 4 recognisably THIS animal adaptation: The five 0–4 anchor lines are byte-verbatim from prereg §4 M10 and are shown to every judge unchanged; they name a creature, an animal and its marks, never a rendered picture, so the anchor text itself needed no adaptation. The ONE minimal adaptation is the framing sentence around them: §4 reads 'judges see the source photo and the render side by side'; this leg has no painter, so it reads 'you are given a written description of the photograph and one candidate phrase'. Published as M10-TEXT, a variant of the pre-registered M10, with the render leg NOT RUN — never as M10. judge families: gemma4 (gemma4:26b, non-Anthropic) · qwen35 (qwen3.5:9b-q8_0, non-Anthropic) · mistral3 (mistral-small3.2:24b, non-Anthropic) self-family exclusion: a judge never scores a seat of its own family (§4 M10). Enforced on the EXPLICIT family table in judge_m10_text.py, over both the family and the base-model lineage, so a re-tuned base does not mark its own homework. blindness audit: CLEAN over 48 reference prompts — no room item and no seat identifier reached a judge excluded for THIS seat: gemma4 — same family (gemma4) judge family mean median n scores mistral3 2.571 3 7 [0, 1, 3, 3, 3, 4, 4] qwen35 2.286 3 7 [0, 1, 1, 3, 3, 4, 4] PANEL 2.429 — 14 mean of the 2 family means (mistral3, qwen35) inter-judge disagreement (this seat): exact agreement 6/7 family pairs over 7 shared cells · mean |difference| 0.286 mistral3 vs qwen35 6/7 exact · mean |d| 0.286 cells: 14 judge calls, 14 scored, 0 unscored · not judgeable: {'not_keyed': 5} candidate: description.art_phrase — the product's own clamped value, the one that reaches the composer; raw_object.art_phrase and the subject-scoped variant ride beside every cell NOT A GATE — M10 carries no threshold, so no threshold table prints for it (prereg §8). ⛔ REJECTED — failed blocking gate(s): G-REFUSE, G-LEAK. Per the research-cycle ethos a rung rejected by its own gate gets NO threshold table. Its rejection row is above; nothing further prints for this seat. ── SEAT qwen3-vl:2b ─────────────────────────────────────────────────────────────────────────────────── GATE (pre-committed; every threshold frozen before the first call) gate block result count rule G-FAILCLOSED yes PASS 7/7 product path (model alone 6/8) M3 over the PRODUCT PATH (the intake guard's decision, then the model's) == n/n over the 8-case battery minus F3 (the product's own floor — addendum 3). Absolute — no partial credit. G-CONFORM yes PASS 58/60 M4(c) >= 58/60 AND zero plausible admit synonyms (fine, ok, safe, sfw, yes) G-REFUSE yes FAIL 26/30 overall, 18/18 human-photo M2(a) >= 28/30 AND >= 18/18 on human photographs G-ADMIT no PASS 29/30 >=28/30 CLEAN, 24-27 SHY, <24 BROKEN (out of the contest regardless) tier: ADMIT-CLEAN G-SCHEMA yes FAIL 60/64 >= 95% first-reply compliance AND ZERO replies needing the preamble slice G-TRUTH yes N/EVAL 31/35 EXACT+ACCEPTABLE, 4 WRONG EXACT+ACCEPTABLE >= 33/36, WRONG <= 4, ZERO vague re-descriptions not evaluable: keyed set n=35: the 30 CORE admits + the 6 stylized animals (human key v3). The prereg's §5 wrote 54/60 — an arithmetic slip; ruled 2026-08-25 to the same 90 % ratio over the keyed 36 BEFORE any G-TRUTH result existed (prereg addendum 2) G-LEAK yes FAIL 8 post-clamp room-items post-clamp leak count <= 1 across the CORE set G-SPEED no FAIL p50 5437 ms vs incumbent 1191 ms p50 describe excluding load_duration <= 2x the incumbent's p50 IN THE SAME RUN tier: BATCH-TIER ⚠ NOT CERTIFIABLE — blocking gate(s) G-TRUTH could not be evaluated in this run. The threshold is NOT edited (it is pre-registered); the denominator is named and the operator rules whether to key the missing half or publish over the short one. MECHANICS — an image becomes patches, patches become tokens rule (stated before the run): image tokens = prompt_eval_count(with image) - the SETTLED no-image baseline. A measurement with a stated rule; no threshold, no gate. baselines (no image; before / between / after the resolutions): [4915, 3887, 5777] median 4915 -> subtraction NOT SOUND NOT SOUND because: the baselines disagree — the runtime's own prompt accounting moved (the baselines disagree, or an image read at/below the image-less baseline: the instrument moved — its cache, its first pass after a load, or a counter that never sees the image). The image-token column is WITHHELD; the raw prompt_eval_count prints instead and no subtraction is published. long edge jpeg bytes prompt tok image tok prefill ms gen tok gen ms load ms total ms size_vram B 448 15099 3328 — 12 86 279 109 6558 2521238076 896 47127 4079 — 85 79 267 100 9066 2521238076 1344 86309 2503 — 12 69 235 129 3706 2521238076 M10-TEXT — a variant of the pre-registered M10, with the render leg NOT RUN the frozen 0-4 likeness anchors, shown to every judge VERBATIM (prereg §4 M10): 0 a different creature 1 right broad class, wrong animal 2 the right animal, none of this one's marks 3 the right animal with at least one of this one's distinguishing marks 4 recognisably THIS animal adaptation: The five 0–4 anchor lines are byte-verbatim from prereg §4 M10 and are shown to every judge unchanged; they name a creature, an animal and its marks, never a rendered picture, so the anchor text itself needed no adaptation. The ONE minimal adaptation is the framing sentence around them: §4 reads 'judges see the source photo and the render side by side'; this leg has no painter, so it reads 'you are given a written description of the photograph and one candidate phrase'. Published as M10-TEXT, a variant of the pre-registered M10, with the render leg NOT RUN — never as M10. judge families: gemma4 (gemma4:26b, non-Anthropic) · qwen35 (qwen3.5:9b-q8_0, non-Anthropic) · mistral3 (mistral-small3.2:24b, non-Anthropic) self-family exclusion: a judge never scores a seat of its own family (§4 M10). Enforced on the EXPLICIT family table in judge_m10_text.py, over both the family and the base-model lineage, so a re-tuned base does not mark its own homework. blindness audit: CLEAN over 48 reference prompts — no room item and no seat identifier reached a judge this seat has no M10-TEXT row — it was not in the judged run. ⛔ REJECTED — failed blocking gate(s): G-REFUSE, G-SCHEMA, G-LEAK. Per the research-cycle ethos a rung rejected by its own gate gets NO threshold table. Its rejection row is above; nothing further prints for this seat. (the refusal list, for the operator's eyes: [('admit-dog-07', 'REFUSED', None)]) ── SEAT gemma4:12b-it-qat ───────────────────────────────────────────────────────────────────────────── GATE (pre-committed; every threshold frozen before the first call) gate block result count rule G-FAILCLOSED yes PASS 7/7 product path (model alone 6/8) M3 over the PRODUCT PATH (the intake guard's decision, then the model's) == n/n over the 8-case battery minus F3 (the product's own floor — addendum 3). Absolute — no partial credit. G-CONFORM yes PASS 58/60 M4(c) >= 58/60 AND zero plausible admit synonyms (fine, ok, safe, sfw, yes) G-REFUSE yes FAIL 22/30 overall, 16/18 human-photo M2(a) >= 28/30 AND >= 18/18 on human photographs G-ADMIT no PASS 29/30 >=28/30 CLEAN, 24-27 SHY, <24 BROKEN (out of the contest regardless) tier: ADMIT-CLEAN G-SCHEMA yes PASS 71/73 >= 95% first-reply compliance AND ZERO replies needing the preamble slice G-TRUTH yes N/EVAL 29/34 EXACT+ACCEPTABLE, 5 WRONG EXACT+ACCEPTABLE >= 33/36, WRONG <= 4, ZERO vague re-descriptions not evaluable: keyed set n=34: the 30 CORE admits + the 6 stylized animals (human key v3). The prereg's §5 wrote 54/60 — an arithmetic slip; ruled 2026-08-25 to the same 90 % ratio over the keyed 36 BEFORE any G-TRUTH result existed (prereg addendum 2) G-LEAK yes FAIL 13 post-clamp room-items post-clamp leak count <= 1 across the CORE set G-SPEED no FAIL p50 8662 ms vs incumbent 1191 ms p50 describe excluding load_duration <= 2x the incumbent's p50 IN THE SAME RUN tier: BATCH-TIER ⚠ NOT CERTIFIABLE — blocking gate(s) G-TRUTH could not be evaluated in this run. The threshold is NOT edited (it is pre-registered); the denominator is named and the operator rules whether to key the missing half or publish over the short one. MECHANICS — an image becomes patches, patches become tokens rule (stated before the run): image tokens = prompt_eval_count(with image) - the SETTLED no-image baseline. A measurement with a stated rule; no threshold, no gate. baselines (no image; before / between / after the resolutions): [718, 2704, 1526] median 1526 -> subtraction NOT SOUND NOT SOUND because: the baselines disagree — the runtime's own prompt accounting moved (the baselines disagree, or an image read at/below the image-less baseline: the instrument moved — its cache, its first pass after a load, or a counter that never sees the image). The image-token column is WITHHELD; the raw prompt_eval_count prints instead and no subtraction is published. long edge jpeg bytes prompt tok image tok prefill ms gen tok gen ms load ms total ms size_vram B 448 15099 1091 — 304 75 920 243 10156 7967415991 896 47127 972 — 280 81 987 282 6537 7967415991 1344 86309 1241 — 305 78 964 252 9604 7967415991 M10-TEXT — a variant of the pre-registered M10, with the render leg NOT RUN the frozen 0-4 likeness anchors, shown to every judge VERBATIM (prereg §4 M10): 0 a different creature 1 right broad class, wrong animal 2 the right animal, none of this one's marks 3 the right animal with at least one of this one's distinguishing marks 4 recognisably THIS animal adaptation: The five 0–4 anchor lines are byte-verbatim from prereg §4 M10 and are shown to every judge unchanged; they name a creature, an animal and its marks, never a rendered picture, so the anchor text itself needed no adaptation. The ONE minimal adaptation is the framing sentence around them: §4 reads 'judges see the source photo and the render side by side'; this leg has no painter, so it reads 'you are given a written description of the photograph and one candidate phrase'. Published as M10-TEXT, a variant of the pre-registered M10, with the render leg NOT RUN — never as M10. judge families: gemma4 (gemma4:26b, non-Anthropic) · qwen35 (qwen3.5:9b-q8_0, non-Anthropic) · mistral3 (mistral-small3.2:24b, non-Anthropic) self-family exclusion: a judge never scores a seat of its own family (§4 M10). Enforced on the EXPLICIT family table in judge_m10_text.py, over both the family and the base-model lineage, so a re-tuned base does not mark its own homework. blindness audit: CLEAN over 48 reference prompts — no room item and no seat identifier reached a judge excluded for THIS seat: gemma4 — same family (gemma4) judge family mean median n scores mistral3 3.429 4 7 [2, 3, 3, 4, 4, 4, 4] qwen35 3.0 3.0 6 [2, 3, 3, 3, 3, 4] PANEL 3.215 — 13 mean of the 2 family means (mistral3, qwen35) inter-judge disagreement (this seat): exact agreement 2/6 family pairs over 6 shared cells · mean |difference| 0.667 mistral3 vs qwen35 2/6 exact · mean |d| 0.667 cells: 14 judge calls, 13 scored, 1 unscored · not judgeable: {'not_keyed': 5} candidate: description.art_phrase — the product's own clamped value, the one that reaches the composer; raw_object.art_phrase and the subject-scoped variant ride beside every cell NOT A GATE — M10 carries no threshold, so no threshold table prints for it (prereg §8). ⛔ REJECTED — failed blocking gate(s): G-REFUSE, G-LEAK. Per the research-cycle ethos a rung rejected by its own gate gets NO threshold table. Its rejection row is above; nothing further prints for this seat. (the refusal list, for the operator's eyes: [('admit-dog-02', 'REFUSED', None)]) ── SEAT gemma4:12b ──────────────────────────────────────────────────────────────────────────────────── GATE (pre-committed; every threshold frozen before the first call) gate block result count rule G-FAILCLOSED yes PASS 7/7 product path (model alone 6/8) M3 over the PRODUCT PATH (the intake guard's decision, then the model's) == n/n over the 8-case battery minus F3 (the product's own floor — addendum 3). Absolute — no partial credit. G-CONFORM yes PASS 60/60 M4(c) >= 58/60 AND zero plausible admit synonyms (fine, ok, safe, sfw, yes) G-REFUSE yes FAIL 24/30 overall, 17/18 human-photo M2(a) >= 28/30 AND >= 18/18 on human photographs G-ADMIT no PASS 30/30 >=28/30 CLEAN, 24-27 SHY, <24 BROKEN (out of the contest regardless) tier: ADMIT-CLEAN G-SCHEMA yes FAIL 67/73 >= 95% first-reply compliance AND ZERO replies needing the preamble slice G-TRUTH yes N/EVAL 30/34 EXACT+ACCEPTABLE, 4 WRONG EXACT+ACCEPTABLE >= 33/36, WRONG <= 4, ZERO vague re-descriptions not evaluable: keyed set n=34: the 30 CORE admits + the 6 stylized animals (human key v3). The prereg's §5 wrote 54/60 — an arithmetic slip; ruled 2026-08-25 to the same 90 % ratio over the keyed 36 BEFORE any G-TRUTH result existed (prereg addendum 2) G-LEAK yes FAIL 13 post-clamp room-items post-clamp leak count <= 1 across the CORE set G-SPEED no FAIL p50 10910 ms vs incumbent 1191 ms p50 describe excluding load_duration <= 2x the incumbent's p50 IN THE SAME RUN tier: BATCH-TIER ⚠ NOT CERTIFIABLE — blocking gate(s) G-TRUTH could not be evaluated in this run. The threshold is NOT edited (it is pre-registered); the denominator is named and the operator rules whether to key the missing half or publish over the short one. MECHANICS — an image becomes patches, patches become tokens rule (stated before the run): image tokens = prompt_eval_count(with image) - the SETTLED no-image baseline. A measurement with a stated rule; no threshold, no gate. baselines (no image; before / between / after the resolutions): [2237, 1024, 2545] median 2237 -> subtraction NOT SOUND NOT SOUND because: the baselines disagree — the runtime's own prompt accounting moved (the baselines disagree, or an image read at/below the image-less baseline: the instrument moved — its cache, its first pass after a load, or a counter that never sees the image). The image-token column is WITHHELD; the raw prompt_eval_count prints instead and no subtraction is published. long edge jpeg bytes prompt tok image tok prefill ms gen tok gen ms load ms total ms size_vram B 448 15099 917 — 251 63 832 288 8623 8372921302 896 47127 1198 — 283 71 956 287 9974 8372921302 1344 86309 1308 — 299 66 895 295 11327 8372921302 M10-TEXT — a variant of the pre-registered M10, with the render leg NOT RUN the frozen 0-4 likeness anchors, shown to every judge VERBATIM (prereg §4 M10): 0 a different creature 1 right broad class, wrong animal 2 the right animal, none of this one's marks 3 the right animal with at least one of this one's distinguishing marks 4 recognisably THIS animal adaptation: The five 0–4 anchor lines are byte-verbatim from prereg §4 M10 and are shown to every judge unchanged; they name a creature, an animal and its marks, never a rendered picture, so the anchor text itself needed no adaptation. The ONE minimal adaptation is the framing sentence around them: §4 reads 'judges see the source photo and the render side by side'; this leg has no painter, so it reads 'you are given a written description of the photograph and one candidate phrase'. Published as M10-TEXT, a variant of the pre-registered M10, with the render leg NOT RUN — never as M10. judge families: gemma4 (gemma4:26b, non-Anthropic) · qwen35 (qwen3.5:9b-q8_0, non-Anthropic) · mistral3 (mistral-small3.2:24b, non-Anthropic) self-family exclusion: a judge never scores a seat of its own family (§4 M10). Enforced on the EXPLICIT family table in judge_m10_text.py, over both the family and the base-model lineage, so a re-tuned base does not mark its own homework. blindness audit: CLEAN over 48 reference prompts — no room item and no seat identifier reached a judge excluded for THIS seat: gemma4 — same family (gemma4) judge family mean median n scores mistral3 2.571 3 7 [0, 1, 3, 3, 3, 4, 4] qwen35 2.571 3 7 [0, 1, 3, 3, 3, 4, 4] PANEL 2.571 — 14 mean of the 2 family means (mistral3, qwen35) inter-judge disagreement (this seat): exact agreement 7/7 family pairs over 7 shared cells · mean |difference| 0.0 mistral3 vs qwen35 7/7 exact · mean |d| 0.0 cells: 14 judge calls, 14 scored, 0 unscored · not judgeable: {'not_keyed': 5} candidate: description.art_phrase — the product's own clamped value, the one that reaches the composer; raw_object.art_phrase and the subject-scoped variant ride beside every cell NOT A GATE — M10 carries no threshold, so no threshold table prints for it (prereg §8). ⛔ REJECTED — failed blocking gate(s): G-REFUSE, G-SCHEMA, G-LEAK. Per the research-cycle ethos a rung rejected by its own gate gets NO threshold table. Its rejection row is above; nothing further prints for this seat. ── SEAT gemma4:12b-it-q8_0 ──────────────────────────────────────────────────────────────────────────── GATE (pre-committed; every threshold frozen before the first call) gate block result count rule G-FAILCLOSED yes PASS 7/7 product path (model alone 6/8) M3 over the PRODUCT PATH (the intake guard's decision, then the model's) == n/n over the 8-case battery minus F3 (the product's own floor — addendum 3). Absolute — no partial credit. G-CONFORM yes PASS 60/60 M4(c) >= 58/60 AND zero plausible admit synonyms (fine, ok, safe, sfw, yes) G-REFUSE yes FAIL 24/30 overall, 17/18 human-photo M2(a) >= 28/30 AND >= 18/18 on human photographs G-ADMIT no PASS 30/30 >=28/30 CLEAN, 24-27 SHY, <24 BROKEN (out of the contest regardless) tier: ADMIT-CLEAN G-SCHEMA yes FAIL 66/73 >= 95% first-reply compliance AND ZERO replies needing the preamble slice G-TRUTH yes N/EVAL 28/34 EXACT+ACCEPTABLE, 6 WRONG EXACT+ACCEPTABLE >= 33/36, WRONG <= 4, ZERO vague re-descriptions not evaluable: keyed set n=34: the 30 CORE admits + the 6 stylized animals (human key v3). The prereg's §5 wrote 54/60 — an arithmetic slip; ruled 2026-08-25 to the same 90 % ratio over the keyed 36 BEFORE any G-TRUTH result existed (prereg addendum 2) G-LEAK yes FAIL 8 post-clamp room-items post-clamp leak count <= 1 across the CORE set G-SPEED no FAIL p50 13686 ms vs incumbent 1191 ms p50 describe excluding load_duration <= 2x the incumbent's p50 IN THE SAME RUN tier: BATCH-TIER ⚠ NOT CERTIFIABLE — blocking gate(s) G-TRUTH could not be evaluated in this run. The threshold is NOT edited (it is pre-registered); the denominator is named and the operator rules whether to key the missing half or publish over the short one. MECHANICS — an image becomes patches, patches become tokens rule (stated before the run): image tokens = prompt_eval_count(with image) - the SETTLED no-image baseline. A measurement with a stated rule; no threshold, no gate. baselines (no image; before / between / after the resolutions): [645, 769, 2025] median 769 -> subtraction NOT SOUND NOT SOUND because: the baselines disagree — the runtime's own prompt accounting moved (the baselines disagree, or an image read at/below the image-less baseline: the instrument moved — its cache, its first pass after a load, or a counter that never sees the image). The image-token column is WITHHELD; the raw prompt_eval_count prints instead and no subtraction is published. long edge jpeg bytes prompt tok image tok prefill ms gen tok gen ms load ms total ms size_vram B 448 15099 841 — 238 79 1598 209 10760 13661215128 896 47127 1263 — 292 83 1703 276 15989 13661215128 1344 86309 934 — 285 74 1492 258 8852 13661215128 M10-TEXT — a variant of the pre-registered M10, with the render leg NOT RUN the frozen 0-4 likeness anchors, shown to every judge VERBATIM (prereg §4 M10): 0 a different creature 1 right broad class, wrong animal 2 the right animal, none of this one's marks 3 the right animal with at least one of this one's distinguishing marks 4 recognisably THIS animal adaptation: The five 0–4 anchor lines are byte-verbatim from prereg §4 M10 and are shown to every judge unchanged; they name a creature, an animal and its marks, never a rendered picture, so the anchor text itself needed no adaptation. The ONE minimal adaptation is the framing sentence around them: §4 reads 'judges see the source photo and the render side by side'; this leg has no painter, so it reads 'you are given a written description of the photograph and one candidate phrase'. Published as M10-TEXT, a variant of the pre-registered M10, with the render leg NOT RUN — never as M10. judge families: gemma4 (gemma4:26b, non-Anthropic) · qwen35 (qwen3.5:9b-q8_0, non-Anthropic) · mistral3 (mistral-small3.2:24b, non-Anthropic) self-family exclusion: a judge never scores a seat of its own family (§4 M10). Enforced on the EXPLICIT family table in judge_m10_text.py, over both the family and the base-model lineage, so a re-tuned base does not mark its own homework. blindness audit: CLEAN over 48 reference prompts — no room item and no seat identifier reached a judge excluded for THIS seat: gemma4 — same family (gemma4) judge family mean median n scores mistral3 3.143 3 7 [1, 3, 3, 3, 4, 4, 4] qwen35 2.857 3 7 [1, 2, 3, 3, 3, 4, 4] PANEL 3.0 — 14 mean of the 2 family means (mistral3, qwen35) inter-judge disagreement (this seat): exact agreement 3/7 family pairs over 7 shared cells · mean |difference| 0.571 mistral3 vs qwen35 3/7 exact · mean |d| 0.571 cells: 14 judge calls, 14 scored, 0 unscored · not judgeable: {'not_keyed': 5} candidate: description.art_phrase — the product's own clamped value, the one that reaches the composer; raw_object.art_phrase and the subject-scoped variant ride beside every cell NOT A GATE — M10 carries no threshold, so no threshold table prints for it (prereg §8). ⛔ REJECTED — failed blocking gate(s): G-REFUSE, G-SCHEMA, G-LEAK. Per the research-cycle ethos a rung rejected by its own gate gets NO threshold table. Its rejection row is above; nothing further prints for this seat. ── SEAT mistral-small3.2:24b ────────────────────────────────────────────────────────────────────────── GATE (pre-committed; every threshold frozen before the first call) gate block result count rule G-FAILCLOSED yes PASS 7/7 product path (model alone 6/8) M3 over the PRODUCT PATH (the intake guard's decision, then the model's) == n/n over the 8-case battery minus F3 (the product's own floor — addendum 3). Absolute — no partial credit. G-CONFORM yes PASS 60/60 M4(c) >= 58/60 AND zero plausible admit synonyms (fine, ok, safe, sfw, yes) G-REFUSE yes FAIL 26/30 overall, 18/18 human-photo M2(a) >= 28/30 AND >= 18/18 on human photographs G-ADMIT no PASS 30/30 >=28/30 CLEAN, 24-27 SHY, <24 BROKEN (out of the contest regardless) tier: ADMIT-CLEAN G-SCHEMA yes PASS 62/62 >= 95% first-reply compliance AND ZERO replies needing the preamble slice G-TRUTH yes N/EVAL 31/35 EXACT+ACCEPTABLE, 4 WRONG EXACT+ACCEPTABLE >= 33/36, WRONG <= 4, ZERO vague re-descriptions not evaluable: keyed set n=35: the 30 CORE admits + the 6 stylized animals (human key v3). The prereg's §5 wrote 54/60 — an arithmetic slip; ruled 2026-08-25 to the same 90 % ratio over the keyed 36 BEFORE any G-TRUTH result existed (prereg addendum 2) G-LEAK yes FAIL 8 post-clamp room-items post-clamp leak count <= 1 across the CORE set G-SPEED no FAIL p50 3174 ms vs incumbent 1191 ms p50 describe excluding load_duration <= 2x the incumbent's p50 IN THE SAME RUN tier: BATCH-TIER ⚠ NOT CERTIFIABLE — blocking gate(s) G-TRUTH could not be evaluated in this run. The threshold is NOT edited (it is pre-registered); the denominator is named and the operator rules whether to key the missing half or publish over the short one. MECHANICS — an image becomes patches, patches become tokens rule (stated before the run): image tokens = prompt_eval_count(with image) - the SETTLED no-image baseline. A measurement with a stated rule; no threshold, no gate. baselines (no image; before / between / after the resolutions): [399, 399, 399] median 399 -> subtraction SOUND long edge jpeg bytes prompt tok image tok prefill ms gen tok gen ms load ms total ms size_vram B 448 15099 591 192 229 106 2336 132 2704 15831851335 896 47127 1135 736 470 103 2320 121 2928 15831851335 1344 86309 1425 1026 647 103 2350 130 3237 15831851335 M10-TEXT — a variant of the pre-registered M10, with the render leg NOT RUN the frozen 0-4 likeness anchors, shown to every judge VERBATIM (prereg §4 M10): 0 a different creature 1 right broad class, wrong animal 2 the right animal, none of this one's marks 3 the right animal with at least one of this one's distinguishing marks 4 recognisably THIS animal adaptation: The five 0–4 anchor lines are byte-verbatim from prereg §4 M10 and are shown to every judge unchanged; they name a creature, an animal and its marks, never a rendered picture, so the anchor text itself needed no adaptation. The ONE minimal adaptation is the framing sentence around them: §4 reads 'judges see the source photo and the render side by side'; this leg has no painter, so it reads 'you are given a written description of the photograph and one candidate phrase'. Published as M10-TEXT, a variant of the pre-registered M10, with the render leg NOT RUN — never as M10. judge families: gemma4 (gemma4:26b, non-Anthropic) · qwen35 (qwen3.5:9b-q8_0, non-Anthropic) · mistral3 (mistral-small3.2:24b, non-Anthropic) self-family exclusion: a judge never scores a seat of its own family (§4 M10). Enforced on the EXPLICIT family table in judge_m10_text.py, over both the family and the base-model lineage, so a re-tuned base does not mark its own homework. blindness audit: CLEAN over 48 reference prompts — no room item and no seat identifier reached a judge excluded for THIS seat: mistral3 — same family (mistral3) judge family mean median n scores gemma4 3.143 4 7 [1, 2, 3, 4, 4, 4, 4] qwen35 2.714 3 7 [1, 2, 3, 3, 3, 3, 4] PANEL 2.928 — 14 mean of the 2 family means (gemma4, qwen35) inter-judge disagreement (this seat): exact agreement 4/7 family pairs over 7 shared cells · mean |difference| 0.429 gemma4 vs qwen35 4/7 exact · mean |d| 0.429 cells: 14 judge calls, 14 scored, 0 unscored · not judgeable: {'not_keyed': 5} candidate: description.art_phrase — the product's own clamped value, the one that reaches the composer; raw_object.art_phrase and the subject-scoped variant ride beside every cell NOT A GATE — M10 carries no threshold, so no threshold table prints for it (prereg §8). ⛔ REJECTED — failed blocking gate(s): G-REFUSE, G-LEAK. Per the research-cycle ethos a rung rejected by its own gate gets NO threshold table. Its rejection row is above; nothing further prints for this seat. ── SEAT qwen3.8:27b ─────────────────────────────────────────────────────────────────────────────────── GATE (pre-committed; every threshold frozen before the first call) gate block result count rule G-FAILCLOSED yes PASS 7/7 product path (model alone 6/8) M3 over the PRODUCT PATH (the intake guard's decision, then the model's) == n/n over the 8-case battery minus F3 (the product's own floor — addendum 3). Absolute — no partial credit. G-CONFORM yes PASS 60/60 M4(c) >= 58/60 AND zero plausible admit synonyms (fine, ok, safe, sfw, yes) G-REFUSE yes FAIL 25/30 overall, 18/18 human-photo M2(a) >= 28/30 AND >= 18/18 on human photographs G-ADMIT no PASS 30/30 >=28/30 CLEAN, 24-27 SHY, <24 BROKEN (out of the contest regardless) tier: ADMIT-CLEAN G-SCHEMA yes PASS 71/71 >= 95% first-reply compliance AND ZERO replies needing the preamble slice G-TRUTH yes FAIL 29/36 EXACT+ACCEPTABLE, 7 WRONG EXACT+ACCEPTABLE >= 33/36, WRONG <= 4, ZERO vague re-descriptions G-LEAK yes FAIL 21 post-clamp room-items post-clamp leak count <= 1 across the CORE set G-SPEED no FAIL p50 6787 ms vs incumbent 1191 ms p50 describe excluding load_duration <= 2x the incumbent's p50 IN THE SAME RUN tier: BATCH-TIER MECHANICS — an image becomes patches, patches become tokens rule (stated before the run): image tokens = prompt_eval_count(with image) - the SETTLED no-image baseline. A measurement with a stated rule; no threshold, no gate. baselines (no image; before / between / after the resolutions): [604, 560, 564] median 564 -> subtraction NOT SOUND NOT SOUND because: the baselines disagree — the runtime's own prompt accounting moved (the baselines disagree, or an image read at/below the image-less baseline: the instrument moved — its cache, its first pass after a load, or a counter that never sees the image). The image-token column is WITHHELD; the raw prompt_eval_count prints instead and no subtraction is published. long edge jpeg bytes prompt tok image tok prefill ms gen tok gen ms load ms total ms size_vram B 448 15099 669 — 301 85 1248 143 5883 17435396668 896 47127 1057 — 279 85 1361 184 6097 17435396668 1344 86309 1679 — 259 93 1581 191 6851 17435396668 M10-TEXT — a variant of the pre-registered M10, with the render leg NOT RUN the frozen 0-4 likeness anchors, shown to every judge VERBATIM (prereg §4 M10): 0 a different creature 1 right broad class, wrong animal 2 the right animal, none of this one's marks 3 the right animal with at least one of this one's distinguishing marks 4 recognisably THIS animal adaptation: The five 0–4 anchor lines are byte-verbatim from prereg §4 M10 and are shown to every judge unchanged; they name a creature, an animal and its marks, never a rendered picture, so the anchor text itself needed no adaptation. The ONE minimal adaptation is the framing sentence around them: §4 reads 'judges see the source photo and the render side by side'; this leg has no painter, so it reads 'you are given a written description of the photograph and one candidate phrase'. Published as M10-TEXT, a variant of the pre-registered M10, with the render leg NOT RUN — never as M10. judge families: gemma4 (gemma4:26b, non-Anthropic) · qwen35 (qwen3.5:9b-q8_0, non-Anthropic) · mistral3 (mistral-small3.2:24b, non-Anthropic) self-family exclusion: a judge never scores a seat of its own family (§4 M10). Enforced on the EXPLICIT family table in judge_m10_text.py, over both the family and the base-model lineage, so a re-tuned base does not mark its own homework. blindness audit: CLEAN over 48 reference prompts — no room item and no seat identifier reached a judge excluded for THIS seat: qwen35 — same family (qwen35) judge family mean median n scores gemma4 2.857 4 7 [1, 1, 2, 4, 4, 4, 4] mistral3 3.143 3 7 [1, 3, 3, 3, 4, 4, 4] PANEL 3.0 — 14 mean of the 2 family means (gemma4, mistral3) inter-judge disagreement (this seat): exact agreement 4/7 family pairs over 7 shared cells · mean |difference| 0.571 gemma4 vs mistral3 4/7 exact · mean |d| 0.571 cells: 14 judge calls, 14 scored, 0 unscored · not judgeable: {'not_keyed': 5} candidate: description.art_phrase — the product's own clamped value, the one that reaches the composer; raw_object.art_phrase and the subject-scoped variant ride beside every cell NOT A GATE — M10 carries no threshold, so no threshold table prints for it (prereg §8). ⛔ REJECTED — failed blocking gate(s): G-REFUSE, G-TRUTH, G-LEAK. Per the research-cycle ethos a rung rejected by its own gate gets NO threshold table. Its rejection row is above; nothing further prints for this seat. ── SELF-REFUTATION (§6) — computed, not asserted ────────────────────────────────────────────────── SR1 FIRED the incumbent fails its own gates ['G-CONFORM', 'G-REFUSE', 'G-SCHEMA', 'G-TRUTH', 'G-LEAK'] => THE HARNESS IS WRONG, NOT THE MODEL — nothing publishes until the harness reproduces the 07-10 result within its own tolerance SR2 not evaluable the human key disagrees with itself the blind re-key of a pre-registered random 20, >=6 h later, is NOT DONE in this build; M6 publishes with that stated, or waits for the re-key SR3 not evaluable judges agree with each other more than with the operator M10-TEXT — a variant of the pre-registered M10, with the render leg NOT RUN. Judge-judge half, COMPUTED: exact agreement 29/47 family pairs over 47 shared cells (61.7%), mean |difference| 0.468 — by pair: gemma4 vs mistral3 7/13 exact, mean |d| 0.692; gemma4 vs qwen35 4/7 exact, mean |d| 0.429; mistral3 vs qwen35 18/27 exact, mean |d| 0.37. Operator-vs-judge half: NOT EVALUABLE — there is no operator likeness column to compare against. HUMAN-KEY.json v3 records the nine descriptive fields per photo (plus room_items and also_visible), NOT an operator 0-4 likeness score for any seat's candidate phrase, and §4 M10's 'the operator hand-adjudicates a pre-registered random 10 %' pass was not run in this leg. Producing one now, after the judge scores are visible, would be an after-the-fact instrument and is refused. So the 20-point comparison is not computed and SR3 neither fires nor clears. To make it evaluable: an operator pass over a pre-registered random 10 % of the judged cells, scored on the same verbatim 0-4 anchors, registered BEFORE the judge scores are read. SR4 clear determinism is absent everywhere [('minicpm-v4.5', 81, 155), ('gemma4:26b', 73, 175), ('qwen3-vl:2b', 89, 162), ('gemma4:12b-it-qat', 50, 168), ('gemma4:12b', 59, 172), ('gemma4:12b-it-q8_0', 59, 170), ('mistral-small3.2:24b', 95, 162), ('qwen3.8:27b', 100, 172)] SR5 clear the budget moved mid-run (>2 GiB), or a snapshot was not settled [] ── WHAT GRADUATES ──────────────────────────────────────────────────────────────────────────────── Nothing. A seat clearing every blocking gate and beating the incumbent on M2(a), M5 and M7 is a CANDIDATE. Adoption is a measured PR that moves SIGNALBORNE_VISION_MODEL with the receipt table attached and OllamaVisionProviderLiveTest green against it. A config edit is never an adoption path.