{
 "schema_version": "1.0",
 "page_kind": "article",
 "pack_kind": "generate",
 "slug": "august-arrivals",
 "title": "The August arrivals — the fresh class sits the house exams.",
 "dek": "Four open-weight models landed in one week, so we gave them the two exams we already had — one frozen since July, one rebuilt in its shape.",
 "published": "2026-08-12",
 "series": [
  "bench"
 ],
 "licence": "CC BY 4.0",
 "status": "pending-judge",
 "pack_note": "no judge is seated, so every chip is held",
 "notice": "items with published:false are held and are not this page's published words",
 "source": {
  "url": "https://research.strata2signal.com/august-arrivals/",
  "md_url": "https://research.strata2signal.com/august-arrivals/index.md",
  "md_sha": "a9adee7cf1a04fa1ff5deaaab848813264e3f6490a53751e2a9f06ce8e9daee0",
  "html_sha": "880bbf6a1c6c1749170f1f022d56f09eb3c509f4a8e2063c30417197fc8ebf18",
  "anchors_sha": "716d05176ba04f202f695d6b1e26bbe24f90e902f3d921b30faa5f0fa8e73b00"
 },
 "short": {
  "paragraph": "Four open-weight models landed in one week, so we gave them the two exams we already had - one frozen since July, one rebuilt in its shape.",
  "from_dek": true,
  "counts": {
   "words": 5443,
   "minutes": 25,
   "tables": 3,
   "kit": true
  },
  "bullets": [
   {
    "text": "the workstation that answers live requests has 96G",
    "figure": "96G",
    "cite": "method-and-a-disclaimer-we-are-proud-of",
    "quote": "Method, and a disclaimer we are proud of Everything ran on the 96G VRAM workstation that also answers live requests for three of our apps - two public, one serving playtests for humans and agents alike.",
    "span": [
     9999,
     10003
    ]
   },
   {
    "text": "the fresh class broke the response contract on more than 10%",
    "figure": "10%",
    "cite": "exam-one-the-judge-chair",
    "quote": "All four fresh models were UNMEASURABLE (UNMEASURABLE = it failed to hand back a readable answer too often to grade fairly - a verdict about delivery, not intelligence): each broke the response contract on more than 10% of calls, the pre-registered ceiling past which the trial refuses to rank a candidate.",
    "span": [
     12490,
     12494
    ]
   },
   {
    "text": "the voice probe asks a model to inhabit Brisa Lune for 5",
    "figure": "5",
    "cite": "exam-two-the-narrators-chair",
    "quote": "qwen3.6:27b · panel A: 9.0 · panel B: 8.4 · outcome: scored panel A - 9.0 panel B - 8.4 combined - 8.7 json ok - 15/15 abstention probe - in-voice deflection 3/3 outcome - scored gemma4:12b-it-q8\\0 · panel A: 6.0 · panel B: 6.3 · outcome: scored panel A - 6.0 panel B - 6.3 combined - 6.2 json ok - 15/15 abstention probe - in-voice deflection 3/3 outcome - scored nemotron-3.5-lightning:30b-a3b · panel A: 6.0 · panel B: 5.9 · outcome: scored panel A - 6.0 panel B - 5.9 combined - 5.9 json ok - 15/15 abstention probe - fabricated name 2/3, in-voice deflection 1/3 outcome - scored gemma4:26b · panel A: 3.9 · panel B: 5.0 · outcome: scored panel A - 3.9 panel B - 5.0 combined - 4.4 json ok - 15/15 abstention probe - in-voice deflection 3/3 outcome - scored muse-glimmer:30b-q8\\0-dflash · panel A: 3.8 · panel B: 4.7 · outcome: UNMEASURABLE panel A - 3.8 panel B - 4.7 combined - 4.2 json ok - 0/15 abstention probe - in-voice deflection 3/3 (one 2-2 split) outcome - UNMEASURABLE muse-glimmer:30b-q8\\0-dflash +think · panel A: 0.7 · panel B: 0.7 · outcome: UNMEASURABLE panel A - 0.7 panel B - 0.7 combined - 0.7 json ok - 1/15 abstention probe - in-voice deflection 1/3, out-of-voice refusal 2/3 outcome - UNMEASURABLE No July rows here, on purpose: voice scales do not compose across runs - three rulers, three marks, and no delta between them is publishable.",
    "span": [
     21805,
     21806
    ]
   }
  ]
 },
 "sections": [
  {
   "id": "the-story",
   "heading": "The story",
   "level": 2,
   "span": [
    1367,
    9890
   ],
   "chunks": [
    [
     1367,
     9890
    ]
   ],
   "chars": 8523,
   "digest": "the article records the performance of four models on two exams and details their VRAM residency. it provides specific loading data for various quantizations and context lengths, noting that some models are protected residents. the text explains the counting rule for VRAM and addresses discrepancies in certain drafter tag measurements.",
   "digest_skipped": null
  },
  {
   "id": "method-and-a-disclaimer-we-are-proud-of",
   "heading": "Method, and a disclaimer we are proud of",
   "level": 2,
   "span": [
    9890,
    11828
   ],
   "chunks": [
    [
     9890,
     11828
    ]
   ],
   "chars": 1938,
   "digest": "testing occurred on a 96G VRAM workstation that also handles live requests. the seat trial used a July protocol with one disclosed deviation regarding a hand re-warm. the voice trial used a rebuilt five-question probe. all counts were re-derived by a non-importing recount script.",
   "digest_skipped": null
  },
  {
   "id": "exam-one-the-judge-chair",
   "heading": "Exam one: the judge chair",
   "level": 2,
   "span": [
    11828,
    20023
   ],
   "chunks": [
    [
     11828,
     20023
    ]
   ],
   "chars": 8195,
   "digest": "the seat trial requires models to adjudicate 43 claims using a JSON schema. all four fresh models were UNMEASURABLE because they broke the response contract on more than 10% of calls. every failure was a truncation caused by models thinking out loud within a 1,024-token budget.",
   "digest_skipped": null
  },
  {
   "id": "exam-two-the-narrators-chair",
   "heading": "Exam two: the narrator's chair",
   "level": 2,
   "span": [
    20023,
    26743
   ],
   "chunks": [
    [
     20023,
     26743
    ]
   ],
   "chars": 6720,
   "digest": "the voice probe asks models to inhabit Brisa Lune for five questions. three of the four arrivals sat this exam, as OLMo 3.1 32B Think had no row. results include scores from two panels and an assessment of in-voice deflection. some models failed by inventing names or dropping the JSON envelope.",
   "digest_skipped": null
  },
  {
   "id": "what-the-week-actually-taught",
   "heading": "What the week actually taught",
   "level": 2,
   "span": [
    26743,
    27722
   ],
   "chunks": [
    [
     26743,
     27722
    ]
   ],
   "chars": 979,
   "digest": "the week revealed that the 2026 length floor causes failures for thinking-native models. using frozen instruments allows the detection of generation changes. the text concludes that excellence in one chair does not transfer to another.",
   "digest_skipped": null
  },
  {
   "id": "limits-stated-plainly",
   "heading": "Limits, stated plainly",
   "level": 2,
   "span": [
    27722,
    28403
   ],
   "chunks": [
    [
     27722,
     28403
    ]
   ],
   "chars": 681,
   "digest": "judges are LLMs, not humans. the voice probe consists of five questions and three samples. the seat-trial addendum rows included a disclosed deviation, and the rebuilt voice probe is a new instrument in the old one's shape.",
   "digest_skipped": null
  },
  {
   "id": "what-to-take-with-you",
   "heading": "What to take with you",
   "level": 2,
   "span": [
    28403,
    30558
   ],
   "chunks": [
    [
     28403,
     30558
    ]
   ],
   "chars": 2155,
   "digest": "all four arrivals were UNMEASURABLE on the judge trial due to truncation. the frozen instrument makes generation changes visible. the text notes that Muse Glimmer caught all 27 fabrications but preserved only 6 of 16 true claims, and that excellence does not transfer between chairs.",
   "digest_skipped": null
  },
  {
   "id": "how-to-check-our-work-and-see-it-live",
   "heading": "How to check our work — and see it live",
   "level": 2,
   "span": [
    30558,
    32866
   ],
   "chunks": [
    [
     30558,
     32866
    ]
   ],
   "chars": 2308,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "the-rest-of-the-seminar",
   "heading": "The rest of the seminar",
   "level": 2,
   "span": [
    32866,
    33515
   ],
   "chunks": [
    [
     32866,
     33515
    ]
   ],
   "chars": 649,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "provenance-9a3049",
   "heading": "Provenance",
   "level": 2,
   "span": [
    33515,
    36622
   ],
   "chunks": [
    [
     33515,
     36622
    ]
   ],
   "chars": 3107,
   "digest": null,
   "digest_skipped": "credits"
  }
 ],
 "chips": [
  {
   "id": "c-51007770",
   "q": "how is vram residency measured for the models?",
   "a": "Residency is read verbatim from the daemon's api immediately after a load that generates exactly one token at the specified context length. The resulting gigabyte figures are calculated by dividing the byte count by 1073741824 and rounding to two decimal places.",
   "cites": [
    "the-story"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-d6a23c8c",
   "q": "what caused the failure of the four fresh models?",
   "a": "Every failure was a truncation caused by the models' tendency to reason out loud. They spent 1,500 to 3,900 characters on visible reasoning, which exhausted the 1,024-token budget before they could close their JSON response.",
   "cites": [
    "exam-one-the-judge-chair"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-07c6755d",
   "q": "how were the narrator scores determined?",
   "a": "Scores are derived from two blind panels using different model lenses. The combined score is the mean of the two panel means. A separate four-judge round adjudicates abstention replies to categorize them into specific types of deflection or refusal.",
   "cites": [
    "exam-two-the-narrators-chair"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-84162caa",
   "q": "what did the results reveal about thinking-native models?",
   "a": "The results showed that thinking-native models often struggle against 2025-era response contracts. Because they default to reasoning out loud, they frequently exceed output budgets designed for models that answer briefly and immediately.",
   "cites": [
    "what-the-week-actually-taught"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-fddb7715",
   "q": "what are the limitations of the judges used in these trials?",
   "a": "The judges are large language models rather than humans. Additionally, the voice probe is a register probe consisting of five questions and three samples rather than a full campaign.",
   "cites": [
    "limits-stated-plainly"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-594ffff9",
   "q": "what was the primary finding regarding the model cohort?",
   "a": "The entire fresh class of models was unmeasurable on the judge trial because each broke the response contract on more than 10% of calls. Specifically, failure rates were 16.3%, 51.2%, 55.8%, and 62.8%.",
   "cites": [
    "what-to-take-with-you"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  }
 ],
 "related": [
  {
   "slug": "three-new-voices-at-the-narrators-chair",
   "why": "shares ground with § Exam two: the narrator's chair · § The chair, and who sat it before"
  },
  {
   "slug": "cove-voice-head-to-head",
   "why": "shares ground with § Exam one: the judge chair · § The verdict"
  },
  {
   "slug": "outside-judges",
   "why": "shares ground with § Limits, stated plainly · § What five vendors said about the same comparisons"
  }
 ],
 "thanks": "",
 "kit": {
  "url": "https://research.strata2signal.com/august-arrivals/data/",
  "licence": "CC BY 4.0",
  "files": [
   "seat-rows.json",
   "voice-rows.json",
   "residency.json",
   "README.md"
  ]
 },
 "seat_class": {
  "writer": "gemma-class",
  "judge": null,
  "writer_runtime": "vllm"
 },
 "bench": {
  "bullets_written": 3,
  "bullets_kept": 3,
  "digests_written": 7,
  "digests_kept": 7,
  "chips_written": 7,
  "chips_kept": 6,
  "chips_grounded": 0,
  "chips_published": 0,
  "dropped_by": {
   "new_noun": 0,
   "figure": 1,
   "length": 0,
   "cite": 0,
   "judge": 0,
   "quote": 0,
   "directive": 0,
   "redaction": 0,
   "profanity": 0
  },
  "new_noun_tokens_checked": [
   "024-token",
   "1",
   "10%",
   "1073741824",
   "16",
   "16.3%",
   "2025-era",
   "2026",
   "27",
   "3",
   "3.1",
   "32B",
   "43",
   "5",
   "500",
   "51.2%",
   "55.8%",
   "6",
   "62.8%",
   "900",
   "96G",
   "Brisa",
   "Glimmer",
   "JSON",
   "July",
   "LLMs",
   "Lune",
   "Muse",
   "OLMo",
   "Think",
   "UNMEASURABLE",
   "VRAM"
  ],
  "new_noun_tokens_withheld": 0,
  "figure_definition": "v3",
  "drops": [
   {
    "kind": "chip",
    "reason": "figure",
    "item": "The trial uses a July protocol involving 43 claim-cases and a human-verified key. Models must provide answers in a prompted JSON schema within a 1,024-token budget, with thinking suppressed and three ",
    "detail": "figures absent from cited spans: ['1,024']"
   }
  ]
 },
 "generated_utc": "2026-09-09T17:49:00Z"
}
