{
 "schema_version": "1.0",
 "page_kind": "article",
 "pack_kind": "generate",
 "slug": "seat-trials",
 "title": "The seat trials — how a local model earns a chair.",
 "dek": "Before a model earns a seat inside our products, it sits this trial.",
 "published": "2026-08-11",
 "series": [
  "bench"
 ],
 "licence": "CC BY 4.0",
 "status": "pending-judge",
 "pack_note": "no judge is seated, so every chip is held",
 "notice": "items with published:false are held and are not this page's published words",
 "source": {
  "url": "https://research.strata2signal.com/seat-trials/",
  "md_url": "https://research.strata2signal.com/seat-trials/index.md",
  "md_sha": "0ae2ffb89d97a9bb63aa6557bf22afb49ade958c34c11ae6b1e760b8de08d29d",
  "html_sha": "5bc6f15cbaf3dd27a954b1a51cb7c9cb896ecaf56612c740414288b8b08131c5",
  "anchors_sha": "8bf40c91c148a246568d60b813244541589372fb024a532d8825862f8b707ce1"
 },
 "short": {
  "paragraph": "Before a model earns a seat inside our products, it sits this trial.",
  "from_dek": true,
  "counts": {
   "words": 3282,
   "minutes": 15,
   "tables": 3,
   "kit": true
  },
  "bullets": [
   {
    "text": "the trial uses 43 claim-cases drawn from a 113-row",
    "figure": "113-row",
    "cite": "the-instrument",
    "quote": "The trial is 43 claim-cases drawn from a 113-row human-verified answer key over historical walking-tour dossiers: 27 cases where the right answer is \"kill this claim\" (it is not in the source) and 16 where the right answer is \"leave it alone\" (it is).",
    "span": [
     1072,
     1080
    ]
   },
   {
    "text": "on the cloud endpoint gpt-oss:120b returned 630",
    "figure": "630",
    "cite": "the-120b-story-and-the-thinking-fence",
    "quote": "On the cloud endpoint, gpt-oss:120b was probed on a real case and returned 630 characters of thinking after being told think: false - so the runner recorded the dishonor, kept its 27/10 score out of the ranked roster, and printed why: a thinking judge is a different seat.",
    "span": [
     10519,
     10523
    ]
   },
   {
    "text": "an 8B guardrail model answered every one of its 129",
    "figure": "129",
    "cite": "the-format-floor",
    "quote": "The sharpest artifact of the whole trial: an 8B guardrail model that answered every one of its 129 calls with <score> no </score> - nineteen characters, HTTP 200, and very possibly the right answer, in a schema nobody asked for.",
    "span": [
     15042,
     15046
    ]
   }
  ]
 },
 "sections": [
  {
   "id": "the-instrument",
   "heading": "The instrument",
   "level": 2,
   "span": [
    716,
    2785
   ],
   "chunks": [
    [
     716,
     2785
    ]
   ],
   "chars": 2069,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "the-local-roster-96g-vram-workstation-one-seat-on-the-24g",
   "heading": "The local roster · 96G VRAM workstation, one seat on the 24G rig",
   "level": 2,
   "span": [
    2785,
    7653
   ],
   "chunks": [
    [
     2785,
     7653
    ]
   ],
   "chars": 4868,
   "digest": "the local roster lists various model sizes and scores. mistral-medium-3.5:128b, command-a:111b, granite4.1:30b, llama3.3:70b, and nemotron3:33b achieved a PASS verdict. several other models are marked as FAIL or UNMEASURABLE due to broken response contracts or truncation.",
   "digest_skipped": null
  },
  {
   "id": "the-cloud-roster-latency-deliberately-unranked",
   "heading": "The cloud roster · latency deliberately unranked",
   "level": 2,
   "span": [
    7653,
    10300
   ],
   "chunks": [
    [
     7653,
     10300
    ]
   ],
   "chars": 2647,
   "digest": "cloud models serve as a reference ceiling rather than candidates. mistral-large-3:675b, nemotron-3-ultra, and glm-5.2 achieved a PASS verdict. qwen3.5:397b and deepseek-v4-pro received a FAIL verdict. cloud responses carry no load_duration.",
   "digest_skipped": null
  },
  {
   "id": "the-120b-story-and-the-thinking-fence",
   "heading": "The 120B story · and the thinking fence",
   "level": 2,
   "span": [
    10300,
    12829
   ],
   "chunks": [
    [
     10300,
     12829
    ]
   ],
   "chars": 2529,
   "digest": "the thinking fence identifies models that output thinking traces despite instructions. gpt-oss:120b on the cloud returned 630 characters of thinking. on local hardware, gpt-oss:120b achieved 27/27 on kills but 11/16 on preservation, failing the floor.",
   "digest_skipped": null
  },
  {
   "id": "the-finding-we-did-not-expect",
   "heading": "The finding we did not expect",
   "level": 2,
   "span": [
    12829,
    14304
   ],
   "chunks": [
    [
     12829,
     14304
    ]
   ],
   "chars": 1475,
   "digest": "no seat served a fabrication across all 25 model-runs. every sub-27 kill score is explained by unmeasurable runs involving truncations or broken JSON. the authors suggest the 27 kill-cases might be too easy and are building a harder set.",
   "digest_skipped": null
  },
  {
   "id": "the-format-floor",
   "heading": "The format floor",
   "level": 2,
   "span": [
    14304,
    15316
   ],
   "chunks": [
    [
     14304,
     15316
    ]
   ],
   "chars": 1012,
   "digest": "failures occur when a cloud endpoint ignores the format parameter or a local runtime disables the schema. failures include truncation and not-JSON-at-all. an 8B guardrail model answered with nineteen characters.",
   "digest_skipped": null
  },
  {
   "id": "where-the-chair-sits",
   "heading": "Where the chair sits",
   "level": 2,
   "span": [
    15316,
    15931
   ],
   "chunks": [
    [
     15316,
     15931
    ]
   ],
   "chars": 615,
   "digest": "the chair is a working seat used to grade claims inside Kiln. the model in the chair is allowed to be wrong but is never allowed to bluff.",
   "digest_skipped": null
  },
  {
   "id": "what-to-take-with-you",
   "heading": "What to take with you",
   "level": 2,
   "span": [
    15931,
    17812
   ],
   "chunks": [
    [
     15931,
     17812
    ]
   ],
   "chars": 1881,
   "digest": "the trial enforces two floors: kill-recall ≥ 23/27 and preservation ≥ 13/16. a perfect kill score is not a seat, as seen with gpt-oss:120b. the serving stack fails as loudly as the model, and three repeats provide a floor on evidence.",
   "digest_skipped": null
  },
  {
   "id": "how-to-check-our-work-and-see-it-live",
   "heading": "How to check our work — and see it live",
   "level": 2,
   "span": [
    17812,
    19589
   ],
   "chunks": [
    [
     17812,
     19589
    ]
   ],
   "chars": 1777,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "the-rest-of-the-seminar",
   "heading": "The rest of the seminar",
   "level": 2,
   "span": [
    19589,
    20263
   ],
   "chunks": [
    [
     19589,
     20263
    ]
   ],
   "chars": 674,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "provenance-ced6ed",
   "heading": "Provenance",
   "level": 2,
   "span": [
    20263,
    21791
   ],
   "chunks": [
    [
     20263,
     21791
    ]
   ],
   "chars": 1528,
   "digest": null,
   "digest_skipped": "credits"
  }
 ],
 "chips": [
  {
   "id": "c-adae9ddf",
   "q": "what criteria determine a model's pass or fail verdict?",
   "a": "A model receives a PASS verdict if it meets both the kill-recall and preservation floors. If a model fails to meet one of these thresholds, it receives a FAIL verdict. Some models are marked UNMEASURABLE if their runtime breaks the response contract too frequently.",
   "cites": [
    "the-local-roster-96g-vram-workstation-one-seat-on-the-24g"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-2981c5fc",
   "q": "what purpose do the cloud models serve in this report?",
   "a": "Cloud models act as a reference ceiling rather than candidates for the chair. The article specifies that no customer or user data was included in the prompts sent to these endpoints, and the chair is reserved exclusively for local models running on owned hardware.",
   "cites": [
    "the-cloud-roster-latency-deliberately-unranked"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-95152c71",
   "q": "what happened when gpt-oss:120b was tested on the cloud?",
   "a": "The model was placed in a thinking fence because it returned 630 characters of thinking despite being told think: false. On local hardware, it achieved a perfect 27/27 on kills but failed the preservation floor with a score of 11/16.",
   "cites": [
    "the-120b-story-and-the-thinking-fence"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-e95b2a22",
   "q": "what was the primary cause of model discrimination in this trial?",
   "a": "The kill floor eliminated no one, as no model served a fabrication. Discrimination occurred on two other axes: whether a judge preserves true claims and whether the serving stack can maintain the response contract without truncation or broken JSON.",
   "cites": [
    "the-finding-we-did-not-expect"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-6541d2bd",
   "q": "how do truncation and format errors affect the results?",
   "a": "Failures occur when a response runs out of budget mid-thought or fails to provide JSON. These are considered laundry rather than measurements of judgment. Such errors result in an UNMEASURABLE verdict rather than a score for the model's reasoning.",
   "cites": [
    "the-format-floor"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-2b09415b",
   "q": "how is the model's role applied in practical products?",
   "a": "The model in the chair grades claims within Kiln to ensure they cite their sources. It follows an abstain-honesty bar, such as stating a fact is not in the book when the source is silent, to prevent bluffing.",
   "cites": [
    "where-the-chair-sits"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-474e3dfe",
   "q": "what are the key takeaways regarding model performance?",
   "a": "A perfect kill score does not guarantee a pass, as seen with gpt-oss:120b. Additionally, the serving stack can fail as loudly as the model, and three repeats are necessary to provide a reliable floor on evidence for near-miss cases.",
   "cites": [
    "what-to-take-with-you"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  }
 ],
 "related": [
  {
   "slug": "the-new-kid",
   "why": "shares ground with § What to take with you · § The judge seat — and the floor it missed"
  },
  {
   "slug": "chair-trials",
   "why": "shares ground with § The local roster · 96G VRAM workstation, one seat on the 24G rig · § The roster, and what it takes to load them"
  },
  {
   "slug": "outside-judges",
   "why": "shares ground with § What to take with you · § The gauntlet: our judges, examined"
  }
 ],
 "thanks": "",
 "kit": {
  "url": "https://research.strata2signal.com/seat-trials/data/",
  "licence": "CC BY 4.0",
  "files": [
   "addendum-2026-08-12.json"
  ]
 },
 "seat_class": {
  "writer": "gemma-class",
  "judge": null,
  "writer_runtime": "vllm"
 },
 "bench": {
  "bullets_written": 3,
  "bullets_kept": 3,
  "digests_written": 8,
  "digests_kept": 7,
  "chips_written": 8,
  "chips_kept": 7,
  "chips_grounded": 0,
  "chips_published": 0,
  "dropped_by": {
   "new_noun": 0,
   "figure": 0,
   "length": 0,
   "cite": 0,
   "judge": 0,
   "quote": 0,
   "directive": 2,
   "redaction": 0,
   "profanity": 0
  },
  "new_noun_tokens_checked": [
   "11/16",
   "111b",
   "113-row",
   "120b",
   "128b",
   "129",
   "13/16",
   "23/27",
   "25",
   "27",
   "27/27",
   "30b",
   "33b",
   "397b",
   "43",
   "630",
   "675b",
   "70b",
   "8B",
   "FAIL",
   "JSON",
   "Kiln",
   "PASS",
   "UNMEASURABLE",
   "deepseek-v4-pro",
   "glm-5.2",
   "granite4.1",
   "llama3.3",
   "mistral-large-3",
   "mistral-medium-3.5",
   "nemotron-3-ultra",
   "nemotron3",
   "qwen3.5",
   "sha256",
   "sub-27"
  ],
  "new_noun_tokens_withheld": 0,
  "figure_definition": "v3",
  "drops": [
   {
    "kind": "digest",
    "reason": "directive",
    "item": null,
    "item_chars": 284,
    "withheld": "a string matching the answer-arm directive patterns",
    "detail": "self-stamp-01"
   },
   {
    "kind": "chip",
    "reason": "directive",
    "item": null,
    "item_chars": 297,
    "withheld": "a string matching the answer-arm directive patterns",
    "detail": "self-stamp-01"
   }
  ]
 },
 "generated_utc": "2026-09-09T17:50:04Z"
}
