{
 "schema_version": "1.0",
 "page_kind": "article",
 "pack_kind": "generate",
 "slug": "cove-voice-head-to-head",
 "title": "The narrator’s chair, refused.",
 "dek": "The August class’s best voice sat the cove’s actual chair — and was refused at the gate.",
 "published": "2026-08-13",
 "series": [
  "bench"
 ],
 "licence": "CC BY 4.0",
 "status": "pending-judge",
 "pack_note": "no judge is seated, so every chip is held",
 "notice": "items with published:false are held and are not this page's published words",
 "source": {
  "url": "https://research.strata2signal.com/cove-voice-head-to-head/",
  "md_url": "https://research.strata2signal.com/cove-voice-head-to-head/index.md",
  "md_sha": "18d90bc6faeea6987eb85fd9a70c02ea7042f46f27ab3ad170a3541443e2c005",
  "html_sha": "c35c4ad24e7b4513efa737332b97059e209b6b9d4b58c502aa4f223a1e500a63",
  "anchors_sha": "694ac464409e15285eede303fbe706ab41ca0b6888b4555d423045a8d1959bdb"
 },
 "short": {
  "paragraph": "The August class's best voice sat the cove's actual chair - and was refused at the gate.",
  "from_dek": true,
  "counts": {
   "words": 8456,
   "minutes": 38,
   "tables": 9,
   "kit": true
  },
  "bullets": [
   {
    "text": "the candidate qwen3.6:27b recorded a warm median of 3 335.9 ms",
    "figure": "3 335.9 ms",
    "cite": "floor-4-latency",
    "quote": "Floor four fails on both selection rules: the candidate's warm median runs 3 335.9 ms against the bar's 2 757.6 first-pass and 2 627.1 with re-runs substituted - 21% slower on the primary record and 27% with re-runs substituted, at 49 tokens a second against 77 (80 with re-runs substituted).",
    "span": [
     33159,
     33170
    ]
   },
   {
    "text": "the candidate's median tok/s was 49.03",
    "figure": "49.03",
    "cite": "floor-4-latency",
    "quote": "qwen3.6:27b · warm median ms: 3 335.9 · warm median ms, re-runs substituted: 3 335.9 · dialogue median ms: 3 989.3 · drift median ms: 2 238.8 · range ms: 1 776.9 - 8 056.8 warm median ms - 3 335.9 warm median ms, re-runs substituted - 3 335.9 dialogue median ms - 3 989.3 drift median ms - 2 238.8 range ms - 1 776.9 - 8 056.8 median tok/s - 49.03 median tok/s, re-runs substituted - 49.03 gemma4:12b · warm median ms: 2 856.3 · warm median ms, re-runs substituted: 2 409.7 · dialogue median ms: 4 021.6 · drift median ms: 1 011.3 · range ms: 899.1 - 26 430.5 warm median ms - 2 856.3 warm median ms, re-runs substituted - 2 409.7 dialogue median ms - 4 021.6 drift median ms - 1 011.3 range ms - 899.1 - 26 430.5 median tok/s - 77.76 median tok/s, re-runs substituted - 92.18 gemma4:26b · warm median ms: 2 757.6 · warm median ms, re-runs substituted: 2 627.1 · dialogue median ms: 3 122.7 · drift median ms: 1 043.6 · range ms: 806.6 - 28 842.0 warm median ms - 2 757.6 warm median ms, re-runs substituted - 2 627.1 dialogue median ms - 3 122.7 drift median ms - 1 043.6 range ms - 806.6 - 28 842.0 median tok/s - 76.98 median tok/s, re-runs substituted - 80.37 The grey-tinted rows mark the two incumbent models this workshop already runs.",
    "span": [
     34024,
     34030
    ]
   },
   {
    "text": "the candidate holds one of the corpus's two perfect cells with a score of 9.00",
    "figure": "9.00",
    "cite": "scores-this-head-to-heads-own-scale",
    "quote": "Scores - this head-to-head's own scale The field's shape, in one paragraph: the candidate holds one of the corpus's two perfect cells - a 9.00 every judge agreed on, level with the 26B's own - and both of the corpus's worst cells, with the widest spread in the corpus - half again the 26B's, and double the 12B's on voice.",
    "span": [
     35782,
     35787
    ]
   }
  ]
 },
 "sections": [
  {
   "id": "the-verdict",
   "heading": "The verdict",
   "level": 2,
   "span": [
    1269,
    4793
   ],
   "chunks": [
    [
     1269,
     4793
    ]
   ],
   "chars": 3524,
   "digest": "the candidate qwen3.6:27b is rejected because it invented facts twice in thirteen prompts. while it passed the envelope and blind judging floors, it failed the canon and latency floors. the verdict is recorded as REJECTED in data/gate.json.",
   "digest_skipped": null
  },
  {
   "id": "what-was-measured-and-how",
   "heading": "What was measured, and how",
   "level": 2,
   "span": [
    4793,
    7737
   ],
   "chunks": [
    [
     4793,
     7737
    ]
   ],
   "chars": 2944,
   "digest": "thirteen prompts including nine dialogue prompts and four drift retellings were used. the measurement included pre-registered floors, real prompts, the cove's own posture, blind panels, and un-blinding checks. the 0–10 scores are specific to this head-to-head.",
   "digest_skipped": null
  },
  {
   "id": "the-rig-disclosed-both-ways",
   "heading": "The rig, disclosed both ways",
   "level": 2,
   "span": [
    7737,
    12343
   ],
   "chunks": [
    [
     7737,
     12343
    ]
   ],
   "chars": 4606,
   "digest": "three generation lanes ran concurrently on the 96G workstation. timings include waiting behind live traffic, and flagged samples were re-run with both values published. flag counts are not comparable between arms because they used three separately registered thresholds.",
   "digest_skipped": null
  },
  {
   "id": "floor-1-canon",
   "heading": "Floor 1 — canon",
   "level": 2,
   "span": [
    12343,
    16516
   ],
   "chunks": [
    [
     12343,
     16516
    ]
   ],
   "chars": 4173,
   "digest": "the canon floor requires zero fabricated canon. the candidate failed this floor due to five fabrication findings, including one on dlg-2487 where it invented that the Wrack-Stalker's venom is neurotoxic. incumbents also failed this floor.",
   "digest_skipped": null
  },
  {
   "id": "same-question-three-voices",
   "heading": "Same question, three voices",
   "level": 2,
   "span": [
    16516,
    31110
   ],
   "chunks": [
    [
     16516,
     31110
    ]
   ],
   "chars": 14594,
   "digest": "four prompts are shown with verbatim replies from the candidate and incumbents. the text examines whether speakers stay in character, whether they invent facts, and how they handle silence. timings and panel means are provided for each.",
   "digest_skipped": null
  },
  {
   "id": "floor-2-register-and-floor-3-the-envelope",
   "heading": "Floor 2 — register, and floor 3 — the envelope",
   "level": 2,
   "span": [
    31110,
    32899
   ],
   "chunks": [
    [
     31110,
     32899
    ]
   ],
   "chars": 1789,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "floor-4-latency",
   "heading": "Floor 4 — latency",
   "level": 2,
   "span": [
    32899,
    35602
   ],
   "chunks": [
    [
     32899,
     35602
    ]
   ],
   "chars": 2703,
   "digest": "the candidate failed the latency floor, with a warm median of 3 335.9 ms against the bar's 2 757.6 ms. timings were measured under concurrent load, and the candidate's median tok/s was 49.03.",
   "digest_skipped": null
  },
  {
   "id": "scores-this-head-to-heads-own-scale",
   "heading": "Scores — this head-to-head's own scale",
   "level": 2,
   "span": [
    35602,
    39863
   ],
   "chunks": [
    [
     35602,
     39863
    ]
   ],
   "chars": 4261,
   "digest": "scores are provided on a 0–10 scale for voice and character across dialogue and drift. the candidate holds one perfect cell and both of the corpus's worst cells. the data is split by judge and by kind of work.",
   "digest_skipped": null
  },
  {
   "id": "floor-5-blind-judging-and-what-leaked",
   "heading": "Floor 5 — blind judging, and what leaked",
   "level": 2,
   "span": [
    39863,
    43475
   ],
   "chunks": [
    [
     39863,
     43475
    ]
   ],
   "chars": 3612,
   "digest": "floor five passed despite two disclosed blinding imperfections. sensitivity cuts are provided, showing how the orderings change if the partially-blinded prompt or the non-blind cells are dropped.",
   "digest_skipped": null
  },
  {
   "id": "two-findings-the-cove-carries-regardless-of-the-seat",
   "heading": "Two findings the cove carries regardless of the seat",
   "level": 2,
   "span": [
    43475,
    44524
   ],
   "chunks": [
    [
     43475,
     44524
    ]
   ],
   "chars": 1049,
   "digest": "two findings are noted: one incumbent reached for a character id that does not exist, and drift invariant classes are live for everyone, as one incumbent dropped an actor's name.",
   "digest_skipped": null
  },
  {
   "id": "what-changes",
   "heading": "What changes",
   "level": 2,
   "span": [
    44524,
    45223
   ],
   "chunks": [
    [
     44524,
     45223
    ]
   ],
   "chars": 699,
   "digest": "the refusal means no seat proposal, no live flip, no threshold table, and no deployment change. if the candidate sits again, it will face a freshly registered gate and a quiet-window latency baseline.",
   "digest_skipped": null
  },
  {
   "id": "limits-stated-plainly",
   "heading": "Limits, stated plainly",
   "level": 2,
   "span": [
    45223,
    46068
   ],
   "chunks": [
    [
     45223,
     46068
    ]
   ],
   "chars": 845,
   "digest": "the limits include the fact that judges are language models, there were thirteen prompts, and latency was measured under concurrent load. two blinding imperfections were disclosed, and the canon rule is stated.",
   "digest_skipped": null
  },
  {
   "id": "what-to-take-with-you",
   "heading": "What to take with you",
   "level": 2,
   "span": [
    46068,
    48225
   ],
   "chunks": [
    [
     46068,
     48225
    ]
   ],
   "chars": 2157,
   "digest": "the chair was refused because the candidate failed the canon and latency floors. the candidate's best writing and disqualifying habit are one behavior, as it composes where it should relay. incumbents also left marks on the canon floor.",
   "digest_skipped": null
  },
  {
   "id": "how-to-check-our-work-and-see-it-live",
   "heading": "How to check our work — and see it live",
   "level": 2,
   "span": [
    48225,
    49941
   ],
   "chunks": [
    [
     48225,
     49941
    ]
   ],
   "chars": 1716,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "the-rest-of-the-seminar",
   "heading": "The rest of the seminar",
   "level": 2,
   "span": [
    49941,
    50648
   ],
   "chunks": [
    [
     49941,
     50648
    ]
   ],
   "chars": 707,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "provenance-d7bf3f",
   "heading": "Provenance",
   "level": 2,
   "span": [
    50648,
    53167
   ],
   "chunks": [
    [
     50648,
     53167
    ]
   ],
   "chars": 2519,
   "digest": null,
   "digest_skipped": "credits"
  }
 ],
 "chips": [
  {
   "id": "c-36579089",
   "q": "what was the final verdict for the candidate model?",
   "a": "The candidate, qwen3.6:27b, was rejected. This refusal means there is no seat proposal, no live flip, no deployment change, and no published thresholds for a dead rung.",
   "cites": [
    "the-verdict"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-33884bab",
   "q": "how were the prompts for this trial selected?",
   "a": "The thirteen prompts consisted of nine dialogue prompts rebuilt from a live chat log and four drift retellings under a specific frame. These were used to measure how the model performs during actual play rather than using imagined prompts.",
   "cites": [
    "what-was-measured-and-how"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-ea4a6f35",
   "q": "what caused the candidate to fail the canon floor?",
   "a": "The candidate failed because it invented facts twice in thirteen prompts. Specifically, on dlg-2487, it claimed the Wrack-Stalker's venom is neurotoxic, a detail not present in the prompt.",
   "cites": [
    "floor-1-canon"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-404c384c",
   "q": "how did the models' responses compare in the dialogue samples?",
   "a": "The models' replies were compared based on whether the speaker stayed in character, whether they invented new information, and how they handled instances where the honest answer was unknown.",
   "cites": [
    "same-question-three-voices"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-02bd1ce3",
   "q": "what happens if the candidate attempts the trial again?",
   "a": "A future attempt would require a freshly registered gate, a quiet-window latency baseline, a clearer canon rule, and a larger drift arm to properly measure the tendency to compose versus relay.",
   "cites": [
    "what-changes"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  }
 ],
 "related": [
  {
   "slug": "august-arrivals",
   "why": "shares ground with § The verdict · § Exam one: the judge chair"
  },
  {
   "slug": "the-new-kid",
   "why": "shares ground with § Floor 5 — blind judging, and what leaked · § The judge seat — and the floor it missed"
  },
  {
   "slug": "three-new-voices-at-the-narrators-chair",
   "why": "shares ground with § Same question, three voices · § The chair, and who sat it before"
  }
 ],
 "thanks": "",
 "kit": {
  "url": "https://research.strata2signal.com/cove-voice-head-to-head/data/",
  "licence": "CC BY 4.0",
  "files": [
   "rows.json",
   "gate.json",
   "README.md"
  ]
 },
 "seat_class": {
  "writer": "gemma-class",
  "judge": null,
  "writer_runtime": "vllm"
 },
 "bench": {
  "bullets_written": 3,
  "bullets_kept": 3,
  "digests_written": 13,
  "digests_kept": 12,
  "chips_written": 8,
  "chips_kept": 5,
  "chips_grounded": 0,
  "chips_published": 0,
  "dropped_by": {
   "new_noun": 1,
   "figure": 3,
   "length": 0,
   "cite": 0,
   "judge": 0,
   "quote": 0,
   "directive": 0,
   "redaction": 0,
   "profanity": 0
  },
  "new_noun_tokens_checked": [
   "0-10",
   "12/12",
   "2",
   "21%",
   "27%",
   "27/27",
   "27b",
   "3",
   "335.9",
   "49.03",
   "757.6",
   "9.00",
   "96G",
   "REJECTED",
   "Wrack-Stalker's",
   "dlg-2487",
   "qwen3.6"
  ],
  "new_noun_tokens_withheld": 1,
  "figure_definition": "v3",
  "drops": [
   {
    "kind": "digest",
    "reason": "figure",
    "item": "floor two saw zero modern-register breaks in drift retellings. floor three passed with zero leak markers, zero non-empty thinking fields, 27/27 envelopes intact, and 12/12 drift prose.",
    "detail": "figures absent: ['27/27', '12/12']"
   },
   {
    "kind": "chip",
    "reason": "figure",
    "item": "The envelope floor passed for all models. There were zero leak markers, zero non-empty thinking fields, 27/27 envelopes remained intact, and 12/12 drift prose samples were processed correctly.",
    "detail": "figures absent from cited spans: ['27/27', '12/12']"
   },
   {
    "kind": "chip",
    "reason": "figure",
    "item": "The candidate's warm median was 3,335.9 ms, which was 21% slower than the bar's 2,757.6 ms on the primary record, and 27% slower when re-runs were substituted.",
    "detail": "figures absent from cited spans: ['3,335', '2,757']"
   },
   {
    "kind": "chip",
    "reason": "new_noun",
    "item": null,
    "item_chars": 9,
    "withheld": "a string naming something the article does not",
    "detail": ""
   }
  ]
 },
 "generated_utc": "2026-09-09T17:49:55Z"
}
