{
 "schema_version": "1.0",
 "page_kind": "article",
 "pack_kind": "generate",
 "slug": "the-new-kid",
 "title": "qwen3.8:27b across four house benches: one seat filled, one floor missed",
 "dek": "qwen3.8:27b landed on our box the day after it shipped, and we sat it in the exams we already had rather than designing anything for it — it filled a screening seat that had never been filled, and it missed a floor in the judge seat.",
 "published": "2026-08-17",
 "series": [
  "bench"
 ],
 "licence": "CC BY 4.0",
 "status": "pending-judge",
 "pack_note": "no judge is seated, so every chip is held",
 "notice": "items with published:false are held and are not this page's published words",
 "source": {
  "url": "https://research.strata2signal.com/the-new-kid/",
  "md_url": "https://research.strata2signal.com/the-new-kid/index.md",
  "md_sha": "13fb660e7da6683f87e1be695ef283253a9e9eb5d14015c0dd8b7ac6ce0c476c",
  "html_sha": "16bba9bad700d53d839aca6d2a9f48f572d97d292aa217bbf77668fd469398b4",
  "anchors_sha": "d0fdbb5fee6ebae7fe5f165193b47a72469a6656382d55306bccbefa239959f7"
 },
 "short": {
  "paragraph": "qwen3.8:27b landed on our box the day after it shipped, and we sat it in the exams we already had rather than designing anything for it - it filled a screening seat that had never been filled, and it missed a floor in the judge seat.",
  "from_dek": true,
  "counts": {
   "words": 8853,
   "minutes": 40,
   "tables": 4,
   "kit": false
  },
  "bullets": [
   {
    "text": "the model weights were pulled on 2026-08-16",
    "figure": "2026-08-16",
    "cite": "what-it-ran-on",
    "quote": "Pulled 2026-08-16T01:43:41Z.",
    "span": [
     4538,
     4549
    ]
   },
   {
    "text": "the classifier seat achieved a median speed of 406 ms",
    "figure": "406 ms",
    "cite": "the-classifier-seat-five-gates-all-pass-faster",
    "quote": "(The same standard cuts the other way later on this page, at the judge seat's 11-of-16, and it is applied there too.) What is not tied is speed: a median of 406 ms per verdict against qwen3.6's 557 ms and gemma4's 541 ms, p95 530 ms against 679 ms and 552 ms - about 27% faster than the sibling at the median, over 1,122 responses each, all three at num\\ctx 32768 with think explicitly false.",
    "span": [
     13537,
     13544
    ]
   },
   {
    "text": "the narrator's open call scored 7.33",
    "figure": "7.33",
    "cite": "the-narrators-open-call-mid-table-uncertainty-printed",
    "quote": "The number: 7.33 out of 10, from thirty cells - five judges over six replies - averaged within each judging family first, then across the five families, so no one vendor's seat count can tilt it (7.333 before rounding).",
    "span": [
     18249,
     18256
    ]
   }
  ]
 },
 "sections": [
  {
   "id": "the-four-answers",
   "heading": "The four answers",
   "level": 2,
   "span": [
    1774,
    3160
   ],
   "chunks": [
    [
     1774,
     3160
    ]
   ],
   "chars": 1386,
   "digest": "the article presents results for four instruments: the classifier seat, the narrator's open call, the judge seat, and the site-reading bench. results include a PASS for the classifier seat, a TIED 7.33 of 10 for the open call, a FAIL for the judge seat, and a 0 of 20 for the site-reading bench.",
   "digest_skipped": null
  },
  {
   "id": "what-it-ran-on",
   "heading": "What it ran on",
   "level": 2,
   "span": [
    3160,
    6392
   ],
   "chunks": [
    [
     3160,
     6392
    ]
   ],
   "chars": 3232,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "the-arrival",
   "heading": "The arrival",
   "level": 2,
   "span": [
    6392,
    8709
   ],
   "chunks": [
    [
     6392,
     8709
    ]
   ],
   "chars": 2317,
   "digest": "qwen3.8:27b arrived on the inference box on 2026-08-16. the text distinguishes between three chairs—the classifier seat, the judge seat, and the narrator's chair—and the site-reading bench. it notes that two of the four kits are published and two are not.",
   "digest_skipped": null
  },
  {
   "id": "the-classifier-seat-five-gates-all-pass-faster",
   "heading": "The classifier seat — five gates, all pass, faster",
   "level": 2,
   "span": [
    8709,
    16778
   ],
   "chunks": [
    [
     8709,
     16778
    ]
   ],
   "chars": 8069,
   "digest": "the classifier seat for RuleSage saw three candidates pass all five gates, including G1, G2, G3, G4, and G5. qwen3.8:27b achieved a median of 406 ms, which is 27% faster than its sibling at the median. the gates read TIED because differences in G2 were within the interval.",
   "digest_skipped": null
  },
  {
   "id": "the-narrators-open-call-mid-table-uncertainty-printed",
   "heading": "The narrator's open call — mid-table, uncertainty printed",
   "level": 2,
   "span": [
    16778,
    23654
   ],
   "chunks": [
    [
     16778,
     23654
    ]
   ],
   "chars": 6876,
   "digest": "the narrator's open call scored 7.33 of 10, which is TIED inside a 0.5 band. the five family means behind this score were 7.0, 7.5, 7.5, 7.917 and 6.75. the text discusses judge dispersion, the recusal law, and the impact of using outside judges.",
   "digest_skipped": null
  },
  {
   "id": "the-judge-seat-and-the-floor-it-missed",
   "heading": "The judge seat — and the floor it missed",
   "level": 2,
   "span": [
    23654,
    31424
   ],
   "chunks": [
    [
     23654,
     31424
    ]
   ],
   "chars": 7770,
   "digest": "the judge seat resulted in a FAIL because preservation was 11 of 16 against a floor of 13. kill-recall passed at 25 of 27. the section also covers chair trials C1 through C5, noting results like RANKED for C1 and DESCRIPTIVE for C2.",
   "digest_skipped": null
  },
  {
   "id": "one-more-bench-it-happened-to-sit-the-site-reading",
   "heading": "One more bench it happened to sit — the site-reading",
   "level": 2,
   "span": [
    31424,
    33787
   ],
   "chunks": [
    [
     31424,
     33787
    ]
   ],
   "chars": 2363,
   "digest": "the site-reading bench results for qwen3.8 show 0 of 20 closed-book, 10 of 20 with the map, and 6 of 20 with the pages. the results match qwen3.6:27b on all four headline columns.",
   "digest_skipped": null
  },
  {
   "id": "could-the-exams-have-been-rigged-for-it",
   "heading": "Could the exams have been rigged for it?",
   "level": 2,
   "span": [
    33787,
    37005
   ],
   "chunks": [
    [
     33787,
     37005
    ]
   ],
   "chars": 3218,
   "digest": "the text addresses whether exams were rigged, noting the open call sealed on 2026-08-14, before the model was pulled. the site-reading bench sealed on 2026-08-16, after the model landed. the classifier gates were pre-registered on 2026-08-16.",
   "digest_skipped": null
  },
  {
   "id": "the-seat-question",
   "heading": "The seat question",
   "level": 2,
   "span": [
    37005,
    40996
   ],
   "chunks": [
    [
     37005,
     40996
    ]
   ],
   "chars": 3991,
   "digest": "the classifier seat is filled by qwen3.8, which is recommended for the seat due to speed and current generation. the judge seat is not filled because it missed the preservation floor. the narrator's chair remains unchanged as C4 was not run.",
   "digest_skipped": null
  },
  {
   "id": "the-other-precisions",
   "heading": "The other precisions",
   "level": 2,
   "span": [
    40996,
    44060
   ],
   "chunks": [
    [
     40996,
     44060
    ]
   ],
   "chars": 3064,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "what-this-page-does-not-say",
   "heading": "What this page does not say",
   "level": 2,
   "span": [
    44060,
    44592
   ],
   "chunks": [
    [
     44060,
     44592
    ]
   ],
   "chars": 532,
   "digest": "this page does not say if qwen3.8 is good or bad, but provides counts for the instruments on specific dates and quants.",
   "digest_skipped": null
  },
  {
   "id": "what-to-take-with-you",
   "heading": "What to take with you",
   "level": 2,
   "span": [
    44592,
    46544
   ],
   "chunks": [
    [
     44592,
     46544
    ]
   ],
   "chars": 1952,
   "digest": "key takeaways include that a model can be excellent in one chair and unfit for the next, and that extra precision bought no accuracy at all. it also notes that the gates read TIED and a benchmark median is a warm number.",
   "digest_skipped": null
  },
  {
   "id": "how-to-check-our-work-and-see-it-live",
   "heading": "How to check our work — and see it live",
   "level": 2,
   "span": [
    46544,
    48558
   ],
   "chunks": [
    [
     46544,
     48558
    ]
   ],
   "chars": 2014,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "the-rest-of-the-seminar",
   "heading": "The rest of the seminar",
   "level": 2,
   "span": [
    48558,
    49332
   ],
   "chunks": [
    [
     48558,
     49332
    ]
   ],
   "chars": 774,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "the-files-behind-the-figures",
   "heading": "The files behind the figures",
   "level": 2,
   "span": [
    49332,
    53610
   ],
   "chunks": [
    [
     49332,
     53610
    ]
   ],
   "chars": 4278,
   "digest": "the section details which kits are open, such as the open call's addendum kit and exhibit thirteen's kit, and which are closed, like the classifier bench and judge-seat rows. it provides sha256 digests for several files.",
   "digest_skipped": null
  }
 ],
 "chips": [
  {
   "id": "c-36118d2e",
   "q": "how did the four instruments perform on qwen3.8?",
   "a": "The classifier seat passed all five gates, the narrator's open call tied at 7.33 of 10, the judge seat failed preservation with 11 of 16, and the site-reading bench scored 0 of 20 closed-book, 10 of 20 with the map, and 6 of 20 with the pages.",
   "cites": [
    "the-four-answers"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-b34610b2",
   "q": "what distinction exists between the three chairs?",
   "a": "The classifier seat screens incoming reader questions for RuleSage, the judge seat decides whether to kill or leave a claim standing, and the narrator's chair voices townsfolk in a living-world game.",
   "cites": [
    "the-arrival"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-dda77624",
   "q": "how much faster was qwen3.8 in the classifier seat?",
   "a": "The model achieved a median of 406 ms per verdict, which was 27% faster than its sibling at the median.",
   "cites": [
    "the-classifier-seat-five-gates-all-pass-faster"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-201bde01",
   "q": "what was the score for the narrator's open call?",
   "a": "The score was 7.33 of 10, which was tied inside a 0.5 band. The five family means behind this score were 7.0, 7.5, 7.5, 7.917, and 6.75.",
   "cites": [
    "the-narrators-open-call-mid-table-uncertainty-printed"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-46294a3a",
   "q": "why did the judge seat result in a fail?",
   "a": "The model failed because preservation was 11 of 16, missing the pre-registered floor of 13. Kill-recall passed with 25 of 27.",
   "cites": [
    "the-judge-seat-and-the-floor-it-missed"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-96f51373",
   "q": "what was the result of the site-reading bench?",
   "a": "The model scored 0 of 20 closed-book, 10 of 20 with the llms.txt map, and 6 of 20 with the site's own pages. It scored 8 of 8 on map-only questions and 4 of 5 on prose-only questions.",
   "cites": [
    "one-more-bench-it-happened-to-sit-the-site-reading"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  }
 ],
 "related": [
  {
   "slug": "the-instrument-travels",
   "why": "shares ground with § The judge seat — and the floor it missed · § The judge's chair scored one candidate of three"
  },
  {
   "slug": "three-new-voices-at-the-narrators-chair",
   "why": "shares ground with § The four answers · § Seven seats, six carried, and where their zeros sit"
  },
  {
   "slug": "outside-judges",
   "why": "shares ground with § The judge seat — and the floor it missed · § The gauntlet: our judges, examined"
  }
 ],
 "thanks": "",
 "kit": null,
 "seat_class": {
  "writer": "gemma-class",
  "judge": null,
  "writer_runtime": "vllm"
 },
 "bench": {
  "bullets_written": 3,
  "bullets_kept": 3,
  "digests_written": 13,
  "digests_kept": 11,
  "chips_written": 8,
  "chips_kept": 6,
  "chips_grounded": 0,
  "chips_published": 0,
  "dropped_by": {
   "new_noun": 4,
   "figure": 0,
   "length": 0,
   "cite": 0,
   "judge": 0,
   "quote": 0,
   "directive": 0,
   "redaction": 0,
   "profanity": 0
  },
  "new_noun_tokens_checked": [
   "0",
   "0.5",
   "10",
   "11",
   "13",
   "16",
   "20",
   "2026-08-14",
   "2026-08-16",
   "25",
   "27",
   "27%",
   "27b",
   "4",
   "406",
   "5",
   "6",
   "6.75",
   "7.0",
   "7.33",
   "7.5",
   "7.917",
   "8",
   "96G",
   "Box",
   "C1",
   "C2",
   "C4",
   "C5",
   "DESCRIPTIVE",
   "FAIL",
   "G1",
   "G2",
   "G3",
   "G4",
   "G5",
   "PASS",
   "RANKED",
   "RealKeep",
   "RuleSage",
   "TIED",
   "VRAM",
   "qwen3.6",
   "qwen3.8",
   "sha256"
  ],
  "new_noun_tokens_withheld": 4,
  "figure_definition": "v3",
  "drops": [
   {
    "kind": "digest",
    "reason": "new_noun",
    "item": null,
    "item_chars": 4,
    "withheld": "a string naming something the article does not",
    "detail": ""
   },
   {
    "kind": "digest",
    "reason": "new_noun",
    "item": null,
    "item_chars": 3,
    "withheld": "a string naming something the article does not",
    "detail": ""
   },
   {
    "kind": "chip",
    "reason": "new_noun",
    "item": null,
    "item_chars": 4,
    "withheld": "a string naming something the article does not",
    "detail": ""
   },
   {
    "kind": "chip",
    "reason": "new_noun",
    "item": null,
    "item_chars": 3,
    "withheld": "a string naming something the article does not",
    "detail": ""
   }
  ]
 },
 "generated_utc": "2026-09-09T01:28:36Z"
}
