{
 "schema_version": "1.0",
 "page_kind": "article",
 "pack_kind": "generate",
 "slug": "the-same-sixteen",
 "title": "The Same Sixteen",
 "dek": "We put glm-5.3-flash — ollama’s new 320B-total cloud model — on two frozen house instruments; it tied our local 12B and 27B at 16/19, and the finding is about the ruler.",
 "published": "2026-08-31",
 "series": [
  "bench"
 ],
 "licence": "CC BY 4.0",
 "status": "pending-judge",
 "pack_note": "no judge is seated, so every chip is held",
 "notice": "items with published:false are held and are not this page's published words",
 "source": {
  "url": "https://research.strata2signal.com/the-same-sixteen/",
  "md_url": "https://research.strata2signal.com/the-same-sixteen/index.md",
  "md_sha": "a0fe32240543414d1e06a9c2486fad38eeb3930e55f6661d513e3ca254b481c4",
  "html_sha": "a20c2f7430792b1705bc7148f62d180b8fa0ef2f6a847259dd203efdc42eb11d",
  "anchors_sha": "ce716aa11dac6a20535156bbbce5dd87a98ac72cad0d6b56521788e277c63e6e"
 },
 "short": {
  "paragraph": "We put glm-5.3-flash - ollama's new 320B-total cloud model - on two frozen house instruments; it tied our local 12B and 27B at 16/19, and the finding is about the ruler.",
  "from_dek": true,
  "counts": {
   "words": 3937,
   "minutes": 18,
   "tables": 1,
   "kit": true
  },
  "bullets": [
   {
    "text": "the arm scored on the tool bench 16/19",
    "figure": "16/19",
    "cite": "sixteen-of-nineteen",
    "quote": "The arm scored 16/19 - 84.2%, Wilson 95% interval [62.4%, 94.5%] (a Wilson interval: the range of underlying rates consistent with this count.",
    "span": [
     7425,
     7433
    ]
   },
   {
    "text": "the arm scored on the field exam 39/40",
    "figure": "39/40",
    "cite": "the-exam-that-stopped-discriminating",
    "quote": "The arm scored 39/40 - 97.5%, Wilson 95% [87.1%, 99.6%].",
    "span": [
     16266,
     16274
    ]
   },
   {
    "text": "prompt tokens ran on the toolbench 24×",
    "figure": "24×",
    "cite": "what-a-toolbench-costs-for-anyone-sizing-one",
    "quote": "Across the 104 toolbench requests, prompt tokens ran 24× completion tokens - 195,532 against 8,115 - because a toolbench resends the tool schemas and the accumulated transcript on every round.",
    "span": [
     18807,
     18813
    ]
   }
  ]
 },
 "sections": [
  {
   "id": "a-different-kind-of-row",
   "heading": "A different kind of row",
   "level": 2,
   "span": [
    1791,
    3717
   ],
   "chunks": [
    [
     1791,
     3717
    ]
   ],
   "chars": 1926,
   "digest": "the authors use a reference arm, a dated reading of a hosted model, to measure how far their local fleet is from a model of 18 billion active parameters per token, 320 billion total. this arm is not a candidate for a production seat and is recorded beside candidates but never scored into their standings.",
   "digest_skipped": null
  },
  {
   "id": "three-checks-bracket-the-run",
   "heading": "Three checks bracket the run",
   "level": 2,
   "span": [
    3717,
    5999
   ],
   "chunks": [
    [
     3717,
     5999
    ]
   ],
   "chars": 2282,
   "digest": "three checks bracket the run to ensure the instrument remains the same. these include verifying that think:false does not leak reasoning into the answer field, confirming the path emits tool calls, and re-proving that the format parameter is not used to enforce JSON-schema grammar on the cloud path.",
   "digest_skipped": null
  },
  {
   "id": "the-ledger-of-what-did-not-read",
   "heading": "The ledger of what did not read",
   "level": 2,
   "span": [
    5999,
    7085
   ],
   "chunks": [
    [
     5999,
     7085
    ]
   ],
   "chars": 1086,
   "digest": "the arm's row includes nine instruments, but only two read: the tool bench and the field exam. four could not honestly be read due to unenforced schema, speed issues, or memory constraints, while three more did not apply to a model that cannot hold a seat.",
   "digest_skipped": null
  },
  {
   "id": "sixteen-of-nineteen",
   "heading": "Sixteen of nineteen",
   "level": 2,
   "span": [
    7085,
    12087
   ],
   "chunks": [
    [
     7085,
     12087
    ]
   ],
   "chars": 5002,
   "digest": "the tool bench arm scored 16/19 — 84.2%, Wilson 95% interval [62.4%, 94.5%]. while it achieved five of five on call chains, it had 39 of 88 unnecessary tool calls and 49 of 66 argument slots right. two tasks are scored by literal-string checkers that fail correct answers.",
   "digest_skipped": null
  },
  {
   "id": "two-local-models-sixteen-days-earlier",
   "heading": "Two local models, sixteen days earlier",
   "level": 2,
   "span": [
    12087,
    15908
   ],
   "chunks": [
    [
     12087,
     15908
    ]
   ],
   "chars": 3821,
   "digest": "the arm's 16/19 count matches two local models, gemma4:12b and qwen3.6:27b, recorded sixteen days earlier. the bench cannot separate a comparator only at a gap of seven tasks or more, meaning the differences between these models may fall inside the instrument's own error bars.",
   "digest_skipped": null
  },
  {
   "id": "the-exam-that-stopped-discriminating",
   "heading": "The exam that stopped discriminating",
   "level": 2,
   "span": [
    15908,
    17384
   ],
   "chunks": [
    [
     15908,
     17384
    ]
   ],
   "chars": 1476,
   "digest": "the field exam arm scored 39/40 — 97.5%, Wilson 95% [87.1%, 99.6%]. because local models also score 39/40 or better, the exam has stopped discriminating at this level. one miss involved a model providing an illustration that contained a forbidden value.",
   "digest_skipped": null
  },
  {
   "id": "a-probe-not-a-score",
   "heading": "A probe, not a score",
   "level": 2,
   "span": [
    17384,
    18620
   ],
   "chunks": [
    [
     17384,
     18620
    ]
   ],
   "chars": 1236,
   "digest": "a probe was run on the field exam's JSON-schema task because cloud enforcement failed. all twenty replies parsed as JSON and nineteen of twenty carried the right values, but this was recorded as a probe and added to no total to avoid corrupting counts.",
   "digest_skipped": null
  },
  {
   "id": "what-a-toolbench-costs-for-anyone-sizing-one",
   "heading": "What a toolbench costs, for anyone sizing one",
   "level": 2,
   "span": [
    18620,
    19448
   ],
   "chunks": [
    [
     18620,
     19448
    ]
   ],
   "chars": 828,
   "digest": "across 104 toolbench requests, prompt tokens ran 24× completion tokens. across all 164 requests, the pooled ratio is 17×. a hosted tool-calling workload is a prompt-token cost long before it is a completion-token cost.",
   "digest_skipped": null
  },
  {
   "id": "what-a-reference-reading-is-worth",
   "heading": "What a reference reading is worth",
   "level": 2,
   "span": [
    19448,
    20354
   ],
   "chunks": [
    [
     19448,
     20354
    ]
   ],
   "chars": 906,
   "digest": "the reading is dated 2026-08-28 and provides three findings about the authors' instruments. the tool bench needs more items, the field exam needs harder ones, and two checkers fail correct answers on string matches, making those rows a floor.",
   "digest_skipped": null
  },
  {
   "id": "what-to-take-with-you",
   "heading": "What to take with you",
   "level": 2,
   "span": [
    20354,
    22476
   ],
   "chunks": [
    [
     20354,
     22476
    ]
   ],
   "chars": 2122,
   "digest": "the bench cannot separate 16/19 from anything above 10/19. exams may not survive a new serving path, and checkers should be grepped for string literals. toolbenches are a prompt-token bill, with a 24× ratio on the toolbench leg.",
   "digest_skipped": null
  },
  {
   "id": "how-to-check-our-work",
   "heading": "How to check our work",
   "level": 2,
   "span": [
    22476,
    23351
   ],
   "chunks": [
    [
     22476,
     23351
    ]
   ],
   "chars": 875,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "the-rest-of-the-seminar",
   "heading": "The rest of the seminar",
   "level": 2,
   "span": [
    23351,
    24392
   ],
   "chunks": [
    [
     23351,
     24392
    ]
   ],
   "chars": 1041,
   "digest": null,
   "digest_skipped": "credits"
  }
 ],
 "chips": [
  {
   "id": "c-a3710d4b",
   "q": "what purpose does the reference arm serve in this testing process?",
   "a": "The reference arm provides a dated reading of a large hosted model's performance on frozen instruments. It is used to measure how far the local fleet is from a model of that size on specific tool tasks and rulebook-reading items, without being scored as a candidate for a production seat.",
   "cites": [
    "a-different-kind-of-row"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-5f01e108",
   "q": "how does the cloud path affect the think:false setting?",
   "a": "On the cloud path, the think:false setting fails to strip reasoning from the answer field, causing reasoning to be emitted into the answer with stray tags or no delimiter at all. This results in a measurement failure rather than a low score, as the scorer takes the entire answer field.",
   "cites": [
    "three-checks-bracket-the-run"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-1d254c55",
   "q": "which instruments were excluded from the reference arm's reading?",
   "a": "Four instruments could not be honestly read due to unenforced schemas, speed measurement issues, or memory requirements. Three others did not apply because the model cannot hold a seat, and two were successfully read: the tool bench and the field exam.",
   "cites": [
    "the-ledger-of-what-did-not-read"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-36b4a6c5",
   "q": "what was the arm's score on the tool bench instrument?",
   "a": "The arm scored 16/19, which is 84.2%, with a Wilson 95% interval of [62.4%, 94.5%]. While it achieved the best chain results on five of five tasks, it also produced 39 unnecessary tool calls out of 88, a rate of 44.3%.",
   "cites": [
    "sixteen-of-nineteen"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-a64186cb",
   "q": "how do the local models compare to the arm on the tool bench?",
   "a": "Two local models, gemma4:12b and qwen3.6:27b, landed on the same 16/19 count as the arm. However, the bench lacks the resolution to separate these models, as a gap of seven tasks or more is required to distinguish them with statistical significance.",
   "cites": [
    "two-local-models-sixteen-days-earlier"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-35d059ea",
   "q": "why did the field exam stop providing useful discrimination?",
   "a": "The field exam reached a saturation point where most high-performing local models scored 39/40 or better. Because the arm also scored 39/40, the instrument failed to differentiate between the models and requires harder items to remain useful.",
   "cites": [
    "the-exam-that-stopped-discriminating"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-d0c7b0b3",
   "q": "what is the relationship between prompt and completion tokens in a toolbench?",
   "a": "In the toolbench, prompt tokens ran 24× completion tokens because the tool schemas and accumulated transcripts are resent on every round. Across all requests, the pooled ratio was 17×, indicating that hosted tool-calling workloads are primarily a prompt-token cost.",
   "cites": [
    "what-a-toolbench-costs-for-anyone-sizing-one"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  }
 ],
 "related": [
  {
   "slug": "how-a-vision-model-sees",
   "why": "shares ground with § What to take with you · § The bench, in one paragraph"
  },
  {
   "slug": "the-new-kid",
   "why": "shares ground with § The ledger of what did not read · § The four answers"
  },
  {
   "slug": "three-new-voices-at-the-narrators-chair",
   "why": "shares ground with § A different kind of row · § Who's at the chair tonight"
  }
 ],
 "thanks": "",
 "kit": {
  "url": "https://research.strata2signal.com/the-same-sixteen/data/",
  "licence": "CC BY 4.0",
  "files": [
   "C3-RESULTS.md",
   "README.md",
   "REFERENCE-glm-5.3-flash-2026-08-28.gate.json",
   "REFERENCE-glm-5.3-flash-2026-08-28.md",
   "RESULTS-TABLES.md",
   "c3-primary-rows-2026-08-12.json",
   "counting-rules.md",
   "field/glm-5.3-flash-cloud.json",
   "prereg-reference-arm.md",
   "provenance.json",
   "raw/c3/glm-5.3-flash_cloud/preprobe.json",
   "raw/c3/glm-5.3-flash_cloud/think_true/calls.jsonl",
   "raw/c3/glm-5.3-flash_cloud/think_true/trace.jsonl",
   "raw/c3/manifest-glm-5.3-flash-cloud-2026-08-28.json",
   "reference-arm-row.json",
   "tasks.json",
   "tools-manifest.json"
  ]
 },
 "seat_class": {
  "writer": "gemma-class",
  "judge": null,
  "writer_runtime": "vllm"
 },
 "bench": {
  "bullets_written": 3,
  "bullets_kept": 3,
  "digests_written": 10,
  "digests_kept": 10,
  "chips_written": 7,
  "chips_kept": 7,
  "chips_grounded": 0,
  "chips_published": 0,
  "dropped_by": {
   "new_noun": 0,
   "figure": 0,
   "length": 0,
   "cite": 0,
   "judge": 0,
   "quote": 0,
   "directive": 0,
   "redaction": 0,
   "profanity": 0
  },
  "new_noun_tokens_checked": [
   "10/19",
   "104",
   "12b",
   "16/19",
   "164",
   "17×",
   "18",
   "2026-08-28",
   "24×",
   "27b",
   "320",
   "39",
   "39/40",
   "44.3%",
   "49",
   "62.4%",
   "66",
   "84.2%",
   "87.1%",
   "88",
   "94.5%",
   "95%",
   "97.5%",
   "99.6%",
   "JSON",
   "JSON-schema",
   "Wilson",
   "gemma4",
   "qwen3.6"
  ],
  "new_noun_tokens_withheld": 0,
  "figure_definition": "v3",
  "drops": []
 },
 "generated_utc": "2026-09-09T01:26:01Z"
}
