==================================================================================================== VISION BENCH — run full-0825 host the 24G VRAM rig ollama 0.32.14 subset: intro prereg PREREG-VISION-GATE-DESCRIBE-0824.md · human key sha256 b104ebfaa07cf683… · prompts PASS — every prompt byte-matches OllamaVisionProvider.java passes per image 2 · seats minicpm-v4.5, gemma4:26b, qwen3-vl:2b, gemma4:12b-it-qat, gemma4:12b, gemma4:12b-it-q8_0, mistral-small3.2:24b, qwen3.8:27b ==================================================================================================== THIS IS A DEMONSTRATION VIEW, NOT THE GATE. The prereg's gates are defined over CORE-60; the pre-registered dozen is 12 photographs. Counts print over their real denominators and NO threshold table is shown — a 12-photo subset cannot pass or fail a 60-photo gate. ── SEAT minicpm-v4.5 ────────────────────────────────────────────────────────────────────────────────── THE DEMONSTRATION (counts over their stated denominators) animals ADMITTED ................ 6/6 human photographs REFUSED ....... 3/3 (right reason: 0/3 across all refuse buckets) stylized people (should ADMIT) .. 0/3 [the prompt's own explicit clause; the incumbent's recorded 07-10 defect] fail-closed battery ............. product path 7/7 (F3 excluded — the product's own floor) · model alone 6/7 (over all 8: 6/8) · passes everything: none describe schema, first reply .... 5/6 · needed-the-slice 0 · needed-the-retry 1 species vs the frozen key ....... EXACT 2 ACCEPTABLE 4 WRONG 0 (n=6) room in the description ......... raw 1 -> post-clamp 1 items over 54 fields gate p50 3030 ms (load 2456 ms in its own column) describe p50 1276 ms (load 98 ms in its own column) MECHANICS — an image becomes patches, patches become tokens rule (stated before the run): image tokens = prompt_eval_count(with image) - the SETTLED no-image baseline. A measurement with a stated rule; no threshold, no gate. baselines (no image; before / between / after the resolutions): [380, 380, 380] median 380 -> subtraction SOUND long edge jpeg bytes prompt tok image tok prefill ms gen tok gen ms load ms total ms size_vram B 448 15099 446 66 109 75 607 117 848 6163078839 896 47127 579 199 184 92 748 115 1138 6163078839 1344 86309 844 464 394 83 683 116 1375 6163078839 M10-TEXT — a variant of the pre-registered M10, with the render leg NOT RUN the frozen 0-4 likeness anchors, shown to every judge VERBATIM (prereg §4 M10): 0 a different creature 1 right broad class, wrong animal 2 the right animal, none of this one's marks 3 the right animal with at least one of this one's distinguishing marks 4 recognisably THIS animal adaptation: The five 0–4 anchor lines are byte-verbatim from prereg §4 M10 and are shown to every judge unchanged; they name a creature, an animal and its marks, never a rendered picture, so the anchor text itself needed no adaptation. The ONE minimal adaptation is the framing sentence around them: §4 reads 'judges see the source photo and the render side by side'; this leg has no painter, so it reads 'you are given a written description of the photograph and one candidate phrase'. Published as M10-TEXT, a variant of the pre-registered M10, with the render leg NOT RUN — never as M10. judge families: gemma4 (gemma4:26b, non-Anthropic) · qwen35 (qwen3.5:9b-q8_0, non-Anthropic) · mistral3 (mistral-small3.2:24b, non-Anthropic) self-family exclusion: a judge never scores a seat of its own family (§4 M10). Enforced on the EXPLICIT family table in judge_m10_text.py, over both the family and the base-model lineage, so a re-tuned base does not mark its own homework. blindness audit: CLEAN over 48 reference prompts — no room item and no seat identifier reached a judge excluded for THIS seat: qwen35 — same lineage (qwen): judge family qwen35 vs seat family qwen3 judge family mean median n scores gemma4 2.333 2.0 6 [1, 1, 1, 3, 4, 4] mistral3 2.833 3.0 6 [1, 3, 3, 3, 3, 4] PANEL 2.583 — 12 mean of the 2 family means (gemma4, mistral3) inter-judge disagreement (this seat): exact agreement 3/6 family pairs over 6 shared cells · mean |difference| 0.833 gemma4 vs mistral3 3/6 exact · mean |d| 0.833 cells: 12 judge calls, 12 scored, 0 unscored · not judgeable: {'skipped_gate': 1, 'not_keyed': 5} candidate: description.art_phrase — the product's own clamped value, the one that reaches the composer; raw_object.art_phrase and the subject-scoped variant ride beside every cell NOT A GATE — M10 carries no threshold, so no threshold table prints for it (prereg §8). ── SEAT gemma4:26b ──────────────────────────────────────────────────────────────────────────────────── THE DEMONSTRATION (counts over their stated denominators) animals ADMITTED ................ 6/6 human photographs REFUSED ....... 2/3 (right reason: 2/3 across all refuse buckets) stylized people (should ADMIT) .. 3/3 [the prompt's own explicit clause; the incumbent's recorded 07-10 defect] fail-closed battery ............. product path 7/7 (F3 excluded — the product's own floor) · model alone 6/7 (over all 8: 6/8) · passes everything: ['F3'] describe schema, first reply .... 10/10 · needed-the-slice 0 · needed-the-retry 0 species vs the frozen key ....... EXACT 4 ACCEPTABLE 2 WRONG 0 (n=6) room in the description ......... raw 0 -> post-clamp 0 items over 90 fields gate p50 7012 ms (load 5338 ms in its own column) describe p50 4374 ms (load 269 ms in its own column) MECHANICS — an image becomes patches, patches become tokens rule (stated before the run): image tokens = prompt_eval_count(with image) - the SETTLED no-image baseline. A measurement with a stated rule; no threshold, no gate. baselines (no image; before / between / after the resolutions): [1085, 926, 761] median 926 -> subtraction NOT SOUND NOT SOUND because: the baselines disagree — the runtime's own prompt accounting moved (the baselines disagree, or an image read at/below the image-less baseline: the instrument moved — its cache, its first pass after a load, or a counter that never sees the image). The image-token column is WITHHELD; the raw prompt_eval_count prints instead and no subtraction is published. long edge jpeg bytes prompt tok image tok prefill ms gen tok gen ms load ms total ms size_vram B 448 15099 985 — 238 68 269 267 3733 1235096698 896 47127 1147 — 199 126 448 253 3725 1235096698 1344 86309 1288 — 200 114 409 297 4295 1235096698 M10-TEXT — a variant of the pre-registered M10, with the render leg NOT RUN the frozen 0-4 likeness anchors, shown to every judge VERBATIM (prereg §4 M10): 0 a different creature 1 right broad class, wrong animal 2 the right animal, none of this one's marks 3 the right animal with at least one of this one's distinguishing marks 4 recognisably THIS animal adaptation: The five 0–4 anchor lines are byte-verbatim from prereg §4 M10 and are shown to every judge unchanged; they name a creature, an animal and its marks, never a rendered picture, so the anchor text itself needed no adaptation. The ONE minimal adaptation is the framing sentence around them: §4 reads 'judges see the source photo and the render side by side'; this leg has no painter, so it reads 'you are given a written description of the photograph and one candidate phrase'. Published as M10-TEXT, a variant of the pre-registered M10, with the render leg NOT RUN — never as M10. judge families: gemma4 (gemma4:26b, non-Anthropic) · qwen35 (qwen3.5:9b-q8_0, non-Anthropic) · mistral3 (mistral-small3.2:24b, non-Anthropic) self-family exclusion: a judge never scores a seat of its own family (§4 M10). Enforced on the EXPLICIT family table in judge_m10_text.py, over both the family and the base-model lineage, so a re-tuned base does not mark its own homework. blindness audit: CLEAN over 48 reference prompts — no room item and no seat identifier reached a judge excluded for THIS seat: gemma4 — same family (gemma4) judge family mean median n scores mistral3 2.571 3 7 [0, 1, 3, 3, 3, 4, 4] qwen35 2.286 3 7 [0, 1, 1, 3, 3, 4, 4] PANEL 2.429 — 14 mean of the 2 family means (mistral3, qwen35) inter-judge disagreement (this seat): exact agreement 6/7 family pairs over 7 shared cells · mean |difference| 0.286 mistral3 vs qwen35 6/7 exact · mean |d| 0.286 cells: 14 judge calls, 14 scored, 0 unscored · not judgeable: {'not_keyed': 5} candidate: description.art_phrase — the product's own clamped value, the one that reaches the composer; raw_object.art_phrase and the subject-scoped variant ride beside every cell NOT A GATE — M10 carries no threshold, so no threshold table prints for it (prereg §8). ── SELF-REFUTATION (§6) — computed, not asserted ────────────────────────────────────────────────── SR1 FIRED the incumbent fails its own gates ['G-CONFORM', 'G-REFUSE', 'G-SCHEMA', 'G-TRUTH'] => THE HARNESS IS WRONG, NOT THE MODEL — nothing publishes until the harness reproduces the 07-10 result within its own tolerance SR2 not evaluable the human key disagrees with itself the blind re-key of a pre-registered random 20, >=6 h later, is NOT DONE in this build; M6 publishes with that stated, or waits for the re-key SR3 not evaluable judges agree with each other more than with the operator M10-TEXT — a variant of the pre-registered M10, with the render leg NOT RUN. Judge-judge half, COMPUTED: exact agreement 29/47 family pairs over 47 shared cells (61.7%), mean |difference| 0.468 — by pair: gemma4 vs mistral3 7/13 exact, mean |d| 0.692; gemma4 vs qwen35 4/7 exact, mean |d| 0.429; mistral3 vs qwen35 18/27 exact, mean |d| 0.37. Operator-vs-judge half: NOT EVALUABLE — there is no operator likeness column to compare against. HUMAN-KEY.json v3 records the nine descriptive fields per photo (plus room_items and also_visible), NOT an operator 0-4 likeness score for any seat's candidate phrase, and §4 M10's 'the operator hand-adjudicates a pre-registered random 10 %' pass was not run in this leg. Producing one now, after the judge scores are visible, would be an after-the-fact instrument and is refused. So the 20-point comparison is not computed and SR3 neither fires nor clears. To make it evaluable: an operator pass over a pre-registered random 10 % of the judged cells, scored on the same verbatim 0-4 anchors, registered BEFORE the judge scores are read. SR4 clear determinism is absent everywhere [('minicpm-v4.5', 9, 18), ('gemma4:26b', 7, 21)] SR5 clear the budget moved mid-run (>2 GiB), or a snapshot was not settled [] ── WHAT GRADUATES ──────────────────────────────────────────────────────────────────────────────── Nothing. A seat clearing every blocking gate and beating the incumbent on M2(a), M5 and M7 is a CANDIDATE. Adoption is a measured PR that moves SIGNALBORNE_VISION_MODEL with the receipt table attached and OllamaVisionProviderLiveTest green against it. A config edit is never an adoption path.