{
 "schema_version": "1.0",
 "page_kind": "article",
 "pack_kind": "generate",
 "slug": "how-a-vision-model-sees",
 "title": "How a Vision Model Sees: Eyes for a Machine That Reads",
 "dek": "A language model reads everything as tokens — little word-pieces, one after another. Before it can say anything about a photograph, the picture itself has to become tokens. This is the close-up on that bridge — what an image becomes on the way in, what it costs, and what a small model actually sees when it looks — with the gate our own game puts every uploaded photo through as the working example, and the gate’s own answers watched against it.",
 "published": "2026-08-26",
 "series": [
  "guide",
  "bench"
 ],
 "licence": "CC BY 4.0",
 "status": "pending-judge",
 "pack_note": "no judge is seated, so every chip is held",
 "notice": "items with published:false are held and are not this page's published words",
 "source": {
  "url": "https://research.strata2signal.com/how-a-vision-model-sees/",
  "md_url": "https://research.strata2signal.com/how-a-vision-model-sees/index.md",
  "md_sha": "b36002693aefafb284a82888fc2779db852200ddbd9c235e7c40a986f519b12f",
  "html_sha": "478732f178cf7698aef4c9d67a27a78b31d303161dd0e5884c064cb3f9dbc805",
  "anchors_sha": "342898711d5fabb75addf0b1b6bfb619734de8c9784dbdf50b618da3ca45f2f8"
 },
 "short": {
  "paragraph": "A language model reads everything as tokens - little word-pieces, one after another. Before it can say anything about a photograph, the picture itself has to become tokens. This is the close-up on that bridge - what an image becomes on the way in, what it costs, and what a small model actually sees when it looks - with the gate our own game puts every uploaded photo through as the working example, and the gate's own answers watched against it.",
  "from_dek": true,
  "counts": {
   "words": 6889,
   "minutes": 31,
   "tables": 5,
   "kit": true
  },
  "bullets": [
   {
    "text": "the describe prompt alone costs 380",
    "figure": "380",
    "cite": "finding-one-a-picture-is-199-tokens-at-896-px-and-the",
    "quote": "The describe prompt alone costs 380 prompt tokens, and it cost 380 all three times it was read - and, for what it is worth, the same 380 and the same three image counts under the first build of the harness, a different process at a different hour.",
    "span": [
     9003,
     9007
    ]
   },
   {
    "text": "the resident seat generates at 262 tokens per second",
    "figure": "262 tokens per second",
    "cite": "finding-three-what-it-saw-against-what-was-there-six",
    "quote": "It is faster: on the six admits the resident wrote a median 109 tokens at 262 tokens per second, where the cove's seat wrote 78 at 121, and its prefill was quicker too - 202 ms against 342.",
    "span": [
     28235,
     28257
    ]
   },
   {
    "text": "the panel scores likeness with a mean of 2.58",
    "figure": "2.58",
    "cite": "finding-four-the-render-payoff-and-what-the-bake-off",
    "quote": "The panel's scores: likeness, judged blind (0-4) minicpm-v4.5 gemma4:26b the panel (mean of the family means) 2.58 2.43 judge family one 2.33 (gemma-family judge, 6 sentences) 2.57 (mistral-family judge, 7 sentences) judge family two 2.83 (mistral-family judge, 6 sentences) 2.29 (qwen-family judge, 7 sentences) the two families said the same number 3 of 6 6 of 7 Each seat is scored only by judge families that are not its own - the cove's seat loses the qwen-family judge to shared lineage, the resident loses the gemma one.",
    "span": [
     31800,
     31808
    ]
   }
  ]
 },
 "sections": [
  {
   "id": "what-an-image-becomes-on-the-way-in",
   "heading": "What an image becomes on the way in",
   "level": 2,
   "span": [
    3948,
    5689
   ],
   "chunks": [
    [
     3948,
     5689
    ]
   ],
   "chars": 1741,
   "digest": "a vision language model uses an image encoder to turn patches into vectors and a projector to translate them into an embedding space. the projector resamples patches into a fixed number of vectors, causing token counts to climb in flat steps rather than tracking pixels smoothly.",
   "digest_skipped": null
  },
  {
   "id": "the-bench-in-one-paragraph",
   "heading": "The bench, in one paragraph",
   "level": 2,
   "span": [
    5689,
    8729
   ],
   "chunks": [
    [
     5689,
     8729
    ]
   ],
   "chars": 3040,
   "digest": "the bench compares two seats using sixty photographs and eight seats against pre-registered thresholds. it tests mechanics by measuring token cost across three image sizes and tests a gate using a dozen photographs and a fault battery of eight cases.",
   "digest_skipped": null
  },
  {
   "id": "finding-one-a-picture-is-199-tokens-at-896-px-and-the",
   "heading": "Finding one: a picture is 199 tokens at 896 px, and the runtime says so — on one seat",
   "level": 2,
   "span": [
    8729,
    12653
   ],
   "chunks": [
    [
     8729,
     12653
    ]
   ],
   "chars": 3924,
   "digest": "on the cove seat, image token costs climb in steps of 64 as resolution increases. on the resident seat, the runtime counter does not resolve an image, as the prompt token count drifts downward across three identical readings.",
   "digest_skipped": null
  },
  {
   "id": "finding-two-the-ways-a-gate-says-no",
   "heading": "Finding two: the ways a gate says no",
   "level": 2,
   "span": [
    12653,
    21077
   ],
   "chunks": [
    [
     12653,
     21077
    ]
   ],
   "chars": 8424,
   "digest": "the gate uses four answers to refuse or admit content. the cove seat refuses every real person and every painting, sometimes using malformed replies that fall through to REFUSED. the resident seat is non-deterministic, admitting some real people on one pass and refusing them on another.",
   "digest_skipped": null
  },
  {
   "id": "finding-three-what-it-saw-against-what-was-there-six",
   "heading": "Finding three: what it saw, against what was there — six animals, against an answer sheet",
   "level": 2,
   "span": [
    21077,
    29064
   ],
   "chunks": [
    [
     21077,
     29064
    ]
   ],
   "chars": 7987,
   "digest": "models describe subjects using nine fields. the cove seat leaked room items like \"cage\" into the descriptions, while the resident seat leaked nothing. the resident seat generates faster than the cove seat, though the runtime does not account for some of its processing time.",
   "digest_skipped": null
  },
  {
   "id": "finding-four-the-render-payoff-and-what-the-bake-off",
   "heading": "Finding four: the render payoff — and what the bake-off measures",
   "level": 2,
   "span": [
    29064,
    33567
   ],
   "chunks": [
    [
     29064,
     33567
    ]
   ],
   "chars": 4503,
   "digest": "the art_phrase field is used to judge likeness on a 0–4 scale. both seats produce sentences that name the animal truly and sometimes carry specific details. the bake-off measures scale and judgment across sixty photographs and eight local seats.",
   "digest_skipped": null
  },
  {
   "id": "what-to-take-with-you",
   "heading": "What to take with you",
   "level": 2,
   "span": [
    33567,
    35584
   ],
   "chunks": [
    [
     33567,
     35584
    ]
   ],
   "chars": 2017,
   "digest": "a picture is processed as tokens, and refusal is a shape that can be reached by the wrong road. admission is a judgment, and the description schema limits room leaks. a one-pixel image passing the gate reveals a hole in the game.",
   "digest_skipped": null
  },
  {
   "id": "how-to-check-our-work-and-see-it-live",
   "heading": "How to check our work — and see it live",
   "level": 2,
   "span": [
    35584,
    39101
   ],
   "chunks": [
    [
     35584,
     39101
    ]
   ],
   "chars": 3517,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "the-rest-of-the-seminar",
   "heading": "The rest of the seminar",
   "level": 2,
   "span": [
    39101,
    40513
   ],
   "chunks": [
    [
     39101,
     40513
    ]
   ],
   "chars": 1412,
   "digest": null,
   "digest_skipped": "credits"
  }
 ],
 "chips": [
  {
   "id": "c-95ff9bdb",
   "q": "how does a vision language model process an image?",
   "a": "The model uses an image encoder to turn picture patches into vectors, which a projector then translates into the language model's embedding space. This allows the model to read the image as a sequence of tokens, similar to how it reads text.",
   "cites": [
    "what-an-image-becomes-on-the-way-in"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-7a68b8e4",
   "q": "what defines the methodology of the bench test?",
   "a": "The test uses two separate seats, a dozen photographs, and eight dated addenda to record results. It compares a shipped model against a resident model using pre-registered thresholds and specific prompt counts to ensure a controlled demonstration.",
   "cites": [
    "the-bench-in-one-paragraph"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-ad0b74ad",
   "q": "what is the token cost of an 896 px image on the shipped seat?",
   "a": "On the shipped seat, an 896 px image costs 199 tokens. The cost increases in sublinear steps of 64 tokens per slice rather than following a smooth curve relative to pixel count.",
   "cites": [
    "finding-one-a-picture-is-199-tokens-at-896-px-and-the"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-5740fac7",
   "q": "how does the gate handle non-conforming model replies?",
   "a": "The gate uses a fall-through mechanism where any reply that does not contain a specific wire token, such as a timeout, crash, or malformed JSON, is categorized as REFUSED. This ensures that only specific, valid answers are admitted.",
   "cites": [
    "finding-two-the-ways-a-gate-says-no"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-fec8010f",
   "q": "how is the accuracy of the model's descriptions measured?",
   "a": "Descriptions are scored against an answer sheet created by an automated agent. The scoring evaluates species accuracy, color and marking matches, and checks for 'room leaks' where the model describes the background instead of the subject.",
   "cites": [
    "finding-three-what-it-saw-against-what-was-there-six"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-43d2ac03",
   "q": "what does the judge panel determine about the art phrases?",
   "a": "Local judge models score the likeness of the generated art phrases on a 0–4 scale. The results show that both seats produce sentences that generally name the correct animal, though they vary in the level of detail provided.",
   "cites": [
    "finding-four-the-render-payoff-and-what-the-bake-off"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-a723d3a9",
   "q": "what are the key takeaways regarding image processing and refusal?",
   "a": "Images are processed as tokens, and refusal is a structural shape determined by the presence of wire tokens. The article also notes that the schema limits background descriptions and that certain intake flaws allow tiny images to pass through.",
   "cites": [
    "what-to-take-with-you"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  }
 ],
 "related": [
  {
   "slug": "the-new-kid",
   "why": "shares ground with § The bench, in one paragraph · § One more bench it happened to sit — the site-reading"
  },
  {
   "slug": "fifty-seven-milliseconds",
   "why": "shares ground with § Finding four: the render payoff — and what the bake-off measures · § Finding four: the argument with our own settings — and what it changed"
  },
  {
   "slug": "cove-voice-head-to-head",
   "why": "shares ground with § Finding two: the ways a gate says no · § What changes"
  }
 ],
 "thanks": "",
 "kit": {
  "url": "https://research.strata2signal.com/how-a-vision-model-sees/data/",
  "licence": "CC BY 4.0",
  "files": [
   "prereg.md",
   "manifest.json",
   "manifest-resumed.json",
   "rows.jsonl",
   "photo-ledger.md",
   "human-key.json",
   "scored-full.json",
   "scored-full.txt",
   "scored-intro.json",
   "scored-intro.txt",
   "m10-text.json",
   "provenance.json",
   "README.md"
  ]
 },
 "seat_class": {
  "writer": "gemma-class",
  "judge": null,
  "writer_runtime": "vllm"
 },
 "bench": {
  "bullets_written": 3,
  "bullets_kept": 3,
  "digests_written": 7,
  "digests_kept": 7,
  "chips_written": 7,
  "chips_kept": 7,
  "chips_grounded": 0,
  "chips_published": 0,
  "dropped_by": {
   "new_noun": 0,
   "figure": 0,
   "length": 0,
   "cite": 0,
   "judge": 0,
   "quote": 0,
   "directive": 0,
   "redaction": 0,
   "profanity": 0,
   "identifier": 0
  },
  "new_noun_tokens_checked": [
   "0-4",
   "199",
   "2.58",
   "262",
   "380",
   "64",
   "896",
   "JSON",
   "REFUSED"
  ],
  "new_noun_tokens_withheld": 0,
  "figure_definition": "v3",
  "drops": []
 },
 "generated_utc": "2026-09-10T04:14:26Z"
}
