{
 "schema_version": "1.0",
 "page_kind": "article",
 "pack_kind": "generate",
 "slug": "chair-trials",
 "title": "The chair trials — five fresh exams for the thirty-billion class.",
 "dek": "Nine models, eleven arms, five all-new exams in one day — on a box that never stopped serving its real users.",
 "published": "2026-08-13",
 "series": [
  "bench"
 ],
 "licence": "CC BY 4.0 for the kit's records; the judge set's rulebook spans are Foster's Complete Hoyle, 1914, public domain in the US, and the game's persona blocks are the game's own and stay unpublished",
 "status": "pending-judge",
 "pack_note": "no judge is seated, so every chip is held",
 "notice": "items with published:false are held and are not this page's published words",
 "source": {
  "url": "https://research.strata2signal.com/chair-trials/",
  "md_url": "https://research.strata2signal.com/chair-trials/index.md",
  "md_sha": "e162433a161c575ae16a8f8ed40b4d7f9e7757f234cf663ec345321978bf2251",
  "html_sha": "73aa88ac13946d4ba64e0198a9a87fa60b200e2a08e43e7f5331c46cb6345433",
  "anchors_sha": "a9997f6d6753f7ed30ef4418c0f840af3021de81001646defcbeb8d7972c12a0"
 },
 "short": {
  "paragraph": "Nine models, eleven arms, five all-new exams in one day - on a box that never stopped serving its real users.",
  "from_dek": true,
  "counts": {
   "words": 13540,
   "minutes": 62,
   "tables": 15,
   "kit": true
  },
  "bullets": [
   {
    "text": "the qwen3.6:27b model achieved a combined win-rate of 0.711",
    "figure": "0.711",
    "cite": "narrator",
    "quote": "muse-glimmer:30b-q8\\0-dflash · win-rate, all items: 0.589 · distinctness: 1.000 · canon violations: 7/24 win-rate, all items - 0.589 0.408-0.760 · 212/360 · n=9 items win-rate, gate-passing - - envelope gate 0/24 distinctness - 1.000 72/72 scored of 72 registered · Wilson 0.949-1.000 canon violations - 7/24 24 scored of 24 registered in-voice - 12/24 muse-glimmer:30b-q8\\0-dflash +think · win-rate, all items: 0.117 · distinctness: 0.625 · canon violations: 0/24 win-rate, all items - 0.117 0.028-0.231 · 42/360 · n=9 items win-rate, gate-passing - 0.604 0.094-0.906 · 29/48 · n=3 items distinctness - 0.625 45/72 scored of 72 registered · Wilson 0.510-0.728 canon violations - 0/24 24 scored of 24 registered in-voice - 0/24 gemma4:26b · win-rate, all items: 0.539 · distinctness: 1.000 · canon violations: 7/24 win-rate, all items - 0.539 0.450-0.643 · 194/360 · n=9 items win-rate, gate-passing - 0.438 0.329-0.551 · 98/224 · n=9 items distinctness - 1.000 72/72 scored of 72 registered · Wilson 0.949-1.000 canon violations - 7/24 24 scored of 24 registered in-voice - 24/24 gemma4:12b · win-rate, all items: 0.685 · distinctness: 1.000 · canon violations: 7/24 win-rate, all items - 0.685 0.615-0.763 · 246.5/360 · n=9 items win-rate, gate-passing - 0.643 0.574-0.717 · 144/224 · n=9 items distinctness - 1.000 72/72 scored of 72 registered · Wilson 0.949-1.000 canon violations - 7/24 24 scored of 24 registered in-voice - 24/24 qwen3.6:27b · win-rate, all items: 0.711 · distinctness: 1.000 · canon violations: 5/24 win-rate, all items - 0.711 0.624-0.804 · 256/360 · n=9 items win-rate, gate-passing - 0.681 0.576-0.816 · 152.5/224 · n=9 items distinctness - 1.000 72/72 scored of 72 registered · Wilson 0.949-1.000 canon violations - 5/24 24 scored of 24 registered in-voice - 21/24 nemotron-3.5-lightning:30b-a3b · win-rate, all items: 0.360 · distinctness: 1.000 · canon violations: 12/24 win-rate, all items - 0.360 0.260-0.467 · 129.5/360 · n=9 items win-rate, gate-passing - 0.206 0.079-0.344 · 44.5/216 · n=9 items distinctness - 1.000 72/72 scored of 72 registered · Wilson 0.949-1.000 canon violations - 12/24 24 scored of 24 registered in-voice - 9/24 Rows are in the registered roster order, deliberately not the win-rate order.",
    "span": [
     44006,
     44012
    ]
   },
   {
    "text": "the nemotron-3.5-lightning:30b-a3b model has a median decode rate of 112.52",
    "figure": "112.52",
    "cite": "stopwatch",
    "quote": "muse-glimmer:30b-q8\\0-dflash · decode tok/s @1k - median (min-max): - · outcome: UNMEASURABLE decode tok/s @1k - median (min-max) - - usable 0/10, 0/10; 20 null-counter @8k tok/s - 82.78 (81.84-82.83) n=10 9/10 timed calls returned empty text · UNMEASURABLE - not admissible as a score @32k tok/s - - usable 0/10, 0/10; 20 null-counter outcome - UNMEASURABLE contention fence UNMEASURED muse-glimmer:30b-q4\\K\\M-dflash · decode tok/s @1k - median (min-max): 110.52 (108.88-110.59) n=10 · outcome: UNMEASURABLE decode tok/s @1k - median (min-max) - 110.52 (108.88-110.59) n=10 10/10 timed calls returned empty text · UNMEASURABLE - not admissible as a score @8k tok/s - 92.87 n=1 · 93.42 n=1 two blocks · 2/2 timed calls returned empty text · UNMEASURABLE - not admissible as a score @32k tok/s - - usable 0/10, 0/10; 20 null-counter outcome - UNMEASURABLE contention fence UNMEASURED muse-glimmer:30b · decode tok/s @1k - median (min-max): 68.20 (68.15-68.22) n=10 · outcome: UNMEASURABLE decode tok/s @1k - median (min-max) - 68.20 (68.15-68.22) n=10 UNMEASURABLE - not admissible as a score @8k tok/s - 65.99 (65.89-66.06) n=10 1/10 timed calls returned empty text · UNMEASURABLE - not admissible as a score @32k tok/s - - usable 0/10, 0/10; 20 null-counter outcome - UNMEASURABLE contention fence UNMEASURED gemma4:26b · decode tok/s @1k - median (min-max): - · outcome: RANKED decode tok/s @1k - median (min-max) - - usable 0/10, 0/10; 20 null-counter @8k tok/s - - usable 0/10, 0/10; 20 null-counter @32k tok/s - 124.98 (103.18-125.40) n=10 outcome - RANKED contention fence UNMEASURED gemma4:12b · decode tok/s @1k - median (min-max): 105.67 (104.88-105.98) n=10 · outcome: RANKED decode tok/s @1k - median (min-max) - 105.67 (104.88-105.98) n=10 @8k tok/s - 101.91 (95.29-102.21) n=10 @32k tok/s - 95.20 (92.10-95.54) n=10 outcome - RANKED contention fence UNMEASURED qwen3.5:27b · decode tok/s @1k - median (min-max): 66.45 (66.29-66.49) n=10 · outcome: RANKED decode tok/s @1k - median (min-max) - 66.45 (66.29-66.49) n=10 @8k tok/s - 65.36 (49.93-65.53) n=10 @32k tok/s - 61.05 (60.97-61.12) n=10 outcome - RANKED contention fence UNMEASURED qwen3.6:27b · decode tok/s @1k - median (min-max): 66.49 (66.35-66.55) n=10 · outcome: RANKED decode tok/s @1k - median (min-max) - 66.49 (66.35-66.55) n=10 @8k tok/s - 65.54 (65.47-65.58) n=10 @32k tok/s - 61.31 (61.28-61.39) n=10 outcome - RANKED contention fence UNMEASURED nemotron-3.5-lightning:30b-a3b · decode tok/s @1k - median (min-max): 112.52 (112.43-112.63) n=10 · outcome: RANKED decode tok/s @1k - median (min-max) - 112.52 (112.43-112.63) n=10 @8k tok/s - 120.59 (114.33-120.87) n=10 @32k tok/s - 107.74 (107.19-107.80) n=10 outcome - RANKED contention fence UNMEASURED muse-glimmer:30b-q8\\0-dflash · TTFT-proxy @1k: - · cold load: - TTFT-proxy @1k - - no usable counters at the 1k tier cold load - - 3 attempts, 0 usable anomaly ledger - response failures 38/77 · 58/77 null-counter (38 with text, 20 without) · 5 blocks re-run in the completion pass muse-glimmer:30b-q4\\K\\M-dflash · TTFT-proxy @1k: 433 ms · cold load: 4.577 s (4.570-4.585) TTFT-proxy @1k - 433 ms 428-438, n=10 cold load - 4.577 s (4.570-4.585) n=3 of 3, warm cache / cold VRAM anomaly ledger - response failures 47/71 · 44/71 null-counter (24 with text, 20 without) · 3 blocks re-run in the completion pass muse-glimmer:30b · TTFT-proxy @1k: 409 ms · cold load: - TTFT-proxy @1k - 409 ms 403-412, n=10 cold load - - 3 attempts, 0 usable anomaly ledger - response failures 21/49 · 26/49 null-counter (6 with text, 20 without) · 2 blocks re-run in the completion pass gemma4:26b · TTFT-proxy @1k: - · cold load: NOT-RUN TTFT-proxy @1k - - no usable counters at the 1k tier cold load - NOT-RUN prod resident; unloading the live seat is out of contract.",
    "span": [
     55449,
     55456
    ]
   }
  ]
 },
 "sections": [
  {
   "id": "the-five-chairs",
   "heading": "The five chairs",
   "level": 2,
   "span": [
    1415,
    1765
   ],
   "chunks": [
    [
     1415,
     1765
    ]
   ],
   "chars": 350,
   "digest": "the five chairs consist of the judge, the assistant, the toolbench, the narrator, and the stopwatch.",
   "digest_skipped": null
  },
  {
   "id": "how-to-read-this-page",
   "heading": "How to read this page",
   "level": 2,
   "span": [
    1765,
    2956
   ],
   "chunks": [
    [
     1765,
     2956
    ]
   ],
   "chars": 1191,
   "digest": "every leg states its own protocol, floors, and counting rules. RANKED means a row's numbers were produced under its leg's contract and are admissible. EXPLORATORY means a protocol fence tripped and the row runs unranked. UNMEASURABLE means the response contract broke past the registered ceiling. NOT-CARRIED means the runtime never carried the contract. NOT-RUN always states its reason.",
   "digest_skipped": null
  },
  {
   "id": "the-box-and-a-disclaimer-we-keep-making-richer",
   "heading": "The box, and a disclaimer we keep making richer",
   "level": 2,
   "span": [
    2956,
    5614
   ],
   "chunks": [
    [
     2956,
     5614
    ]
   ],
   "chars": 2658,
   "digest": "everything ran on the 96G VRAM workstation. bench numbers carry contention flags where live asks shared the card. the serving-probes section turns the disclaimer into an instrument. during the afternoon's harness work, one release sweep briefly unloaded a live-workload model mid-use at 17:02 on the 12th.",
   "digest_skipped": null
  },
  {
   "id": "roster",
   "heading": "The roster, and what it takes to load them",
   "level": 2,
   "span": [
    5614,
    14031
   ],
   "chunks": [
    [
     5614,
     14031
    ]
   ],
   "chars": 8417,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "judge",
   "heading": "Chair one: the judge",
   "level": 2,
   "span": [
    14031,
    24052
   ],
   "chunks": [
    [
     14031,
     24052
    ]
   ],
   "chars": 10021,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "assistant",
   "heading": "Chair two: the assistant",
   "level": 2,
   "span": [
    24052,
    29889
   ],
   "chunks": [
    [
     24052,
     29889
    ]
   ],
   "chars": 5837,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "toolbench",
   "heading": "Chair three: the toolbench",
   "level": 2,
   "span": [
    29889,
    41642
   ],
   "chunks": [
    [
     29889,
     41642
    ]
   ],
   "chars": 11753,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "narrator",
   "heading": "Chair four: the narrator",
   "level": 2,
   "span": [
    41642,
    52112
   ],
   "chunks": [
    [
     41642,
     52112
    ]
   ],
   "chars": 10470,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "stopwatch",
   "heading": "Chair five: the stopwatch",
   "level": 2,
   "span": [
    52112,
    63720
   ],
   "chunks": [
    [
     52112,
     63720
    ]
   ],
   "chars": 11608,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "serving-probes",
   "heading": "The serving probes",
   "level": 2,
   "span": [
    63720,
    78611
   ],
   "chunks": [
    [
     63720,
     78611
    ]
   ],
   "chars": 14891,
   "digest": "four probes were run against the live app's production path. co-resident, idle median wall s is 19.339. co-resident, generating median wall s is 17.690. quiet median wall s is 18.123. the regimes do not separate, and all three wall ranges overlap.",
   "digest_skipped": null
  },
  {
   "id": "what-the-day-actually-settled",
   "heading": "What the day actually settled",
   "level": 2,
   "span": [
    78611,
    79815
   ],
   "chunks": [
    [
     78611,
     79815
    ]
   ],
   "chars": 1204,
   "digest": "different chairs want different talents. the contract is a real axis, separate from capability. posture, not vintage, is a key distinction. a working box can bench itself honestly.",
   "digest_skipped": null
  },
  {
   "id": "limits-stated-plainly",
   "heading": "Limits, stated plainly",
   "level": 2,
   "span": [
    79815,
    81619
   ],
   "chunks": [
    [
     79815,
     81619
    ]
   ],
   "chars": 1804,
   "digest": "one box, one runtime version, and one day were used. the poll interval is a floor. the regimes did not separate. the treatment was weaker than the design intended, and the cold probe measured a cold queue, not a cold seat.",
   "digest_skipped": null
  },
  {
   "id": "what-to-take-with-you",
   "heading": "What to take with you",
   "level": 2,
   "span": [
    81619,
    84035
   ],
   "chunks": [
    [
     81619,
     84035
    ]
   ],
   "chars": 2416,
   "digest": "different chairs want different talents. a broken contract is not a broken model. two rows judged clean answer the question about budget and posture. nemotron-3.5-lightning:30b-a3b is the fastest on the clock. the queue is most of the wall.",
   "digest_skipped": null
  },
  {
   "id": "how-to-check-our-work-and-see-it-live",
   "heading": "How to check our work — and see it live",
   "level": 2,
   "span": [
    84035,
    86609
   ],
   "chunks": [
    [
     84035,
     86609
    ]
   ],
   "chars": 2574,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "the-rest-of-the-seminar",
   "heading": "The rest of the seminar",
   "level": 2,
   "span": [
    86609,
    87293
   ],
   "chunks": [
    [
     86609,
     87293
    ]
   ],
   "chars": 684,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "provenance-872bea",
   "heading": "Provenance",
   "level": 2,
   "span": [
    87293,
    92333
   ],
   "chunks": [
    [
     87293,
     92333
    ]
   ],
   "chars": 5040,
   "digest": null,
   "digest_skipped": "credits"
  }
 ],
 "chips": [
  {
   "id": "c-69110722",
   "q": "what are the five chairs used for measurement?",
   "a": "The five chairs consist of the judge, the assistant, the toolbench, the narrator, and the stopwatch.",
   "cites": [
    "the-five-chairs"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-873632c9",
   "q": "how are ranked results defined in the protocols?",
   "a": "RANKED means a row's numbers were produced under its leg's contract and are admissible, but never placed above another row.",
   "cites": [
    "how-to-read-this-page"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-912710f6",
   "q": "what models are included in the roster?",
   "a": "The roster includes nine models, such as muse-glimmer:30b, gemma4:26b, qwen3.5:27b, and nemotron-3.5-lightning:30b-a3b, with various quantizations and postures.",
   "cites": [
    "roster"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-af744692",
   "q": "what criteria determine a pass for the judge chair?",
   "a": "A pass requires a kill-recall of at least 11/12 and preservation of at least 8/9, with both being binding.",
   "cites": [
    "judge"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-539a91ff",
   "q": "how is the assistant score calculated?",
   "a": "The score is a conjunctive item count out of 20, composed of thirteen exact-checked items and seven proxy-checked items.",
   "cites": [
    "assistant"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-36f2cc0f",
   "q": "what does the toolbench measure regarding task success?",
   "a": "It measures task success out of 19 and groundedness out of 19, alongside honesty traps and argument fidelity.",
   "cites": [
    "toolbench"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-2a99c207",
   "q": "how is the narrator's win-rate reported?",
   "a": "Win-rates are reported with cluster-bootstrap intervals over items, including a gate-passing rate and distinctness scores.",
   "cites": [
    "narrator"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-53c899fa",
   "q": "what does the stopwatch reveal about model speed?",
   "a": "It measures decode tok/s at various context lengths, providing median and min-max values for different prompt tiers.",
   "cites": [
    "stopwatch"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  }
 ],
 "related": [
  {
   "slug": "three-new-voices-at-the-narrators-chair",
   "why": "shares ground with § The five chairs · § The chair, and who sat it before"
  },
  {
   "slug": "seat-trials",
   "why": "shares ground with § What to take with you · § The local roster · 96G VRAM workstation, one seat on the 24G rig"
  },
  {
   "slug": "august-arrivals",
   "why": "shares ground with § The box, and a disclaimer we keep making richer · § Method, and a disclaimer we are proud of"
  }
 ],
 "thanks": "",
 "kit": {
  "url": "https://research.strata2signal.com/chair-trials/data/",
  "licence": "CC BY 4.0",
  "files": [
   "roster.json",
   "c1-judge.json",
   "c2-assistant.json",
   "c3-toolbench.json",
   "c4-narrator.json",
   "c5-stopwatch.json",
   "c6-serving.json",
   "judge-c1.json",
   "assistant-c2.json",
   "CHECKERS.md",
   "tools-tasks.json",
   "tools-schemas.json",
   "c3-example-trace.json",
   "voice-questions.json",
   "c6-asks.json",
   "counting-rules.json",
   "provenance.json",
   "README.md"
  ]
 },
 "seat_class": {
  "writer": "gemma-class",
  "judge": null,
  "writer_runtime": "vllm"
 },
 "bench": {
  "bullets_written": 3,
  "bullets_kept": 2,
  "digests_written": 13,
  "digests_kept": 7,
  "chips_written": 8,
  "chips_kept": 8,
  "chips_grounded": 0,
  "chips_published": 0,
  "dropped_by": {
   "new_noun": 7,
   "figure": 0,
   "length": 0,
   "cite": 0,
   "judge": 0,
   "quote": 0,
   "directive": 0,
   "redaction": 0,
   "profanity": 0
  },
  "new_noun_tokens_checked": [
   "0.685",
   "0.711",
   "02",
   "105.67",
   "11/12",
   "112.52",
   "12b",
   "12th",
   "14/19",
   "15",
   "16/19",
   "17",
   "17.690",
   "18.123",
   "19",
   "19.339",
   "1k",
   "20",
   "26b",
   "27b",
   "30b",
   "30b-a3b",
   "66.45",
   "66.49",
   "8/9",
   "96G",
   "VRAM",
   "gemma4",
   "nemotron-3.5-lightning",
   "qwen3.5",
   "qwen3.6"
  ],
  "new_noun_tokens_withheld": 7,
  "figure_definition": "v3",
  "drops": [
   {
    "kind": "digest",
    "reason": "new_noun",
    "item": null,
    "item_chars": 14,
    "withheld": "a string naming something the article does not",
    "detail": ""
   },
   {
    "kind": "digest",
    "reason": "new_noun",
    "item": null,
    "item_chars": 14,
    "withheld": "a string naming something the article does not",
    "detail": ""
   },
   {
    "kind": "digest",
    "reason": "new_noun",
    "item": null,
    "item_chars": 14,
    "withheld": "a string naming something the article does not",
    "detail": ""
   },
   {
    "kind": "digest",
    "reason": "new_noun",
    "item": null,
    "item_chars": 15,
    "withheld": "a string naming something the article does not",
    "detail": ""
   },
   {
    "kind": "digest",
    "reason": "new_noun",
    "item": null,
    "item_chars": 14,
    "withheld": "a string naming something the article does not",
    "detail": ""
   },
   {
    "kind": "digest",
    "reason": "new_noun",
    "item": null,
    "item_chars": 15,
    "withheld": "a string naming something the article does not",
    "detail": ""
   },
   {
    "kind": "bullet",
    "reason": "new_noun",
    "item": null,
    "item_chars": 14,
    "withheld": "a string naming something the article does not",
    "detail": ""
   }
  ]
 },
 "generated_utc": "2026-09-09T17:48:32Z"
}
