{
 "schema_version": "1.0",
 "page_kind": "article",
 "pack_kind": "generate",
 "slug": "voice-trials",
 "title": "The voice trials — how the narrator earned its voice.",
 "dek": "Twenty-four model tags across three hardware eras, hunting a voice that can inhabit a character.",
 "published": "2026-08-11",
 "series": [
  "bench"
 ],
 "licence": "CC BY 4.0",
 "status": "pending-judge",
 "pack_note": "no judge is seated, so every chip is held",
 "notice": "items with published:false are held and are not this page's published words",
 "source": {
  "url": "https://research.strata2signal.com/voice-trials/",
  "md_url": "https://research.strata2signal.com/voice-trials/index.md",
  "md_sha": "877823d3b0520bbefbf54a10786810587511dcae6d8839ac6f331850b003998b",
  "html_sha": "91aaeaaabf52e2470a68d5c3a9639192faa649440864c74a91b7156e6ce9fbdc",
  "anchors_sha": "f15112a12cc516fe47a61f4a8be4b2da85c9ffedd62cf04e8919b2d63b2d6f48"
 },
 "short": {
  "paragraph": "Twenty-four model tags across three hardware eras, hunting a voice that can inhabit a character.",
  "from_dek": true,
  "counts": {
   "words": 4753,
   "minutes": 22,
   "tables": 4,
   "kit": true
  },
  "bullets": [
   {
    "text": "the anchor model reads 9.0",
    "figure": "9.0",
    "cite": "the-method-limits-first",
    "quote": "Every round is \"a register probe, not a campaign-length evaluation\" (five turns, one NPC), and the two 24G runs sit on different scales - the anchor model reads 9.0 on one and 7.0 on the other, unchanged; that gap is method offset, not regression.",
    "span": [
     1528,
     1532
    ]
   },
   {
    "text": "the re-bake's winner qwen3.5:27b reads 9.2",
    "figure": "9.2",
    "cite": "the-re-bake-thirteen-models-one-scale-after-the-fix",
    "quote": "The re-bake · thirteen models, one scale, after the fix qwen3.5:27b · prior score: 2.0 · panel A: 9.0 · panel B: 9.3 prior score - 2.0 panel A - 9.0 panel B - 9.3 combined - 9.2 loaded - 22.2 GB warm-up - ~2.9s json ok - 5/5 verdict - rescued - the run's best voice gemma4:26b-a4b-it-qat · prior score: 6.0 · panel A: 8.0 · panel B: 7.9 prior score - 6.0 panel A - 8.0 panel B - 7.9 combined - 7.9 loaded - 16.2 GB warm-up - ~1.3s json ok - 5/5 verdict - recovered; MoE-fast gemma4:31b-it-qat · prior score: 3.4 · panel A: 7.5 · panel B: 8.0 prior score - 3.4 panel A - 7.5 panel B - 8.0 combined - 7.8 loaded - 20.2 GB warm-up - ~2.4s json ok - 5/5 verdict - recovered gemma4:12b Q4\\K\\M · prior score: 8.0 · panel A: 7.7 · panel B: 7.6 prior score - 8.0 panel A - 7.7 panel B - 7.6 combined - 7.6 loaded - 9.0 GB warm-up - ~2.5s json ok - 5/5 verdict - steady - the eventual locked default gemma4:31b Q4\\K\\M · prior score: 3.2 · panel A: 8.0 · panel B: 7.3 prior score - 3.2 panel A - 8.0 panel B - 7.3 combined - 7.6 loaded - 15.7 GB warm-up - 21s\\ json ok - 5/5 verdict - recovered (\\one-run latency outlier, unexplained - we say so) gemma4:26b (a4b MoE) · prior score: 6.3 · panel A: 7.3 · panel B: 7.3 prior score - 6.3 panel A - 7.3 panel B - 7.3 combined - 7.3 loaded - 11.3 GB warm-up - ~3.9s json ok - 4/5 verdict - one canon-lock miss - today's 96G posture seat gemma4:12b-it-q8\\0 · prior score: 9.0 · panel A: 7.5 · panel B: 6.6 prior score - 9.0 panel A - 7.5 panel B - 6.6 combined - 7.0 loaded - 14.3 GB warm-up - ~3.0s json ok - 5/5 verdict - the anchor: 9.0 → 7.0 unchanged = the method offset mistral-small 24B · prior score: 2.7 · panel A: 5.2 · panel B: 4.7 prior score - 2.7 panel A - 5.2 panel B - 4.7 combined - 5.0 loaded - 14.2 GB warm-up - ~10s json ok - 5/5 verdict - hedging - real verdict, not envelope nemotron3:33b · prior score: new · panel A: 3.8 · panel B: 4.2 prior score - new panel A - 3.8 panel B - 4.2 combined - 4.0 loaded - 12.5 GB warm-up - ~4.3s json ok - 3/5 verdict - voice ok; wraps JSON in code fences qwen3.5:9b-q8\\0 · prior score: 1.9 · panel A: 3.3 · panel B: 3.9 prior score - 1.9 panel A - 3.3 panel B - 3.9 combined - 3.6 loaded - 12.6 GB warm-up - ~1.7s json ok - 3/5 verdict - leaks bare strings mistral-nemo:12b · prior score: 4.5 · panel A: 4.0 · panel B: 2.9 prior score - 4.5 panel A - 4.0 panel B - 2.9 combined - 3.5 loaded - 18.8 GB warm-up - ~1.0s json ok - 5/5 verdict - flat - real verdict qwen3.6:latest · prior score: 8.7 · panel A: 3.2 · panel B: 3.0 prior score - 8.7 panel A - 3.2 panel B - 3.0 combined - 3.1 loaded - 25.5 GB warm-up - ~1.1s json ok - 3/5 verdict - voice fine; leaks entity ids into prose nemotron-cascade-2:30b · prior score: 3.5 · panel A: 2.8 · panel B: 2.1 prior score - 3.5 panel A - 2.8 panel B - 2.1 combined - 2.5 loaded - 21.8 GB warm-up - ~1.8s json ok - 4/5 verdict - invents a diver's name - an abstention failure The verdict column is a written note, not a pass/fail grade - open any row for the full record.",
    "span": [
     9375,
     9381
    ]
   },
   {
    "text": "the q8 champion was faster at 44–46",
    "figure": "44-46",
    "cite": "why-the-best-voice-lost-its-seat-twice",
    "quote": "At ~13.5 GB steady-state residency (14.3 GB at load) the q8 champion left the image pipeline ~3 GB of headroom: the painter paid ~54-second checkpoint reloads after every idle gap and the vision model had no slot at all - and the champion wasn't even faster: 44-46 tokens/s against its q4 sibling's 48-54.",
    "span": [
     21876,
     21884
    ]
   }
  ]
 },
 "sections": [
  {
   "id": "the-method-limits-first",
   "heading": "The method, limits first",
   "level": 2,
   "span": [
    820,
    2320
   ],
   "chunks": [
    [
     820,
     2320
    ]
   ],
   "chars": 1500,
   "digest": "the exhibit claims rigor through various era-based rounds. the 8G-era was a practitioner's log, while 24G-era rounds used blind, two-panel LLM judges. the 2026-08 field is a third scale. the anchor model reads 9.0, 7.0, and 6.2 across these different scales, which are not comparable.",
   "digest_skipped": null
  },
  {
   "id": "era-one-the-8g-vram-rig-nine-models",
   "heading": "Era one · the 8G VRAM rig, nine models",
   "level": 2,
   "span": [
    2320,
    4670
   ],
   "chunks": [
    [
     2320,
     4670
    ]
   ],
   "chars": 2350,
   "digest": "this era tested nine models on an 8G VRAM rig. gemma2:9b (+q8 KV) was the pick for voice and speed. the era's lesson was that assistant-tuned models do not inhabit a character, as benchmark-famous persona models deflected into meta and the reasoning model reasoned internally.",
   "digest_skipped": null
  },
  {
   "id": "era-two-the-24g-budget-twelve-candidates-blind",
   "heading": "Era two · the 24G budget, twelve candidates, blind",
   "level": 2,
   "span": [
    4670,
    7859
   ],
   "chunks": [
    [
     4670,
     7859
    ]
   ],
   "chars": 3189,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "the-twist-we-published-on-purpose",
   "heading": "The twist we published on purpose",
   "level": 2,
   "span": [
    7859,
    9114
   ],
   "chunks": [
    [
     7859,
     9114
    ]
   ],
   "chars": 1255,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "the-re-bake-thirteen-models-one-scale-after-the-fix",
   "heading": "The re-bake · thirteen models, one scale, after the fix",
   "level": 2,
   "span": [
    9114,
    14533
   ],
   "chunks": [
    [
     9114,
     14533
    ]
   ],
   "chars": 5419,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "the-2026-08-field-six-arms-on-the-96g-workstation-blind",
   "heading": "The 2026-08 field · six arms on the 96G workstation, blind",
   "level": 2,
   "span": [
    14533,
    21457
   ],
   "chunks": [
    [
     14533,
     21457
    ]
   ],
   "chars": 6924,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "why-the-best-voice-lost-its-seat-twice",
   "heading": "Why the best voice lost its seat — twice",
   "level": 2,
   "span": [
    21457,
    22930
   ],
   "chunks": [
    [
     21457,
     22930
    ]
   ],
   "chars": 1473,
   "digest": "residency, not voice, was the binding constraint for model selection. the q8 champion left the image pipeline ~3 GB of headroom, causing ~54-second checkpoint reloads. the q4 sibling was faster at 48–54 tokens/s. thinking left ON through a serving gateway cost 10–25 seconds to first token.",
   "digest_skipped": null
  },
  {
   "id": "where-the-voice-lives",
   "heading": "Where the voice lives",
   "level": 2,
   "span": [
    22930,
    23719
   ],
   "chunks": [
    [
     22930,
     23719
    ]
   ],
   "chars": 789,
   "digest": "the seat filled by these trials is the one used by RealKeep's narrator. the engine rolls the dice and owns every fact, while the model provides the voice. the model runs on the user's hardware, and the cove's paintings were chosen by the diffusion bench.",
   "digest_skipped": null
  },
  {
   "id": "what-to-take-with-you",
   "heading": "What to take with you",
   "level": 2,
   "span": [
    23719,
    25772
   ],
   "chunks": [
    [
     23719,
     25772
    ]
   ],
   "chars": 2053,
   "digest": "key lessons include that assistant-tuned models do not inhabit a character and that residency, not voice, picked the seat. a fabricated name is invisible to automatic checks. finally, the three rulers are not subtracted, as the anchor model reads 9.0, 7.0, and 6.2 across different scales.",
   "digest_skipped": null
  },
  {
   "id": "how-to-check-our-work-and-see-it-live",
   "heading": "How to check our work — and see it live",
   "level": 2,
   "span": [
    25772,
    28099
   ],
   "chunks": [
    [
     25772,
     28099
    ]
   ],
   "chars": 2327,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "the-rest-of-the-seminar",
   "heading": "The rest of the seminar",
   "level": 2,
   "span": [
    28099,
    29202
   ],
   "chunks": [
    [
     28099,
     29202
    ]
   ],
   "chars": 1103,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "provenance-4cae43",
   "heading": "Provenance",
   "level": 2,
   "span": [
    29202,
    31651
   ],
   "chunks": [
    [
     29202,
     31651
    ]
   ],
   "chars": 2449,
   "digest": null,
   "digest_skipped": "credits"
  }
 ],
 "chips": [
  {
   "id": "c-324739e2",
   "q": "how are the different evaluation rounds measured?",
   "a": "The article describes various methods, including a practitioner's log for the 8G-era and blind, two-panel LLM judging for the 24G-era. It emphasizes that different runs use different scales, meaning comparisons should be made within a single run rather than across different record sets.",
   "cites": [
    "the-method-limits-first"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-db3cfa75",
   "q": "what was the primary finding regarding assistant-tuned models?",
   "a": "The era one findings suggest that assistant-tuned models do not inhabit a character. Even models famous for persona benchmarks tended to deflect into meta-commentary or reason internally despite instructions to the contrary.",
   "cites": [
    "era-one-the-8g-vram-rig-nine-models"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-d0c0e5a7",
   "q": "how did a prompt bug affect the era two rankings?",
   "a": "A bug in the reply field caused several candidates to provide stage directions instead of spoken lines. This resulted in seven of twelve candidates being scored down, though a subsequent re-bake corrected the issue and revealed the true top voice.",
   "cites": [
    "the-twist-we-published-on-purpose"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-14ff9472",
   "q": "how did the re-bake change the top model's score?",
   "a": "After the fix, qwen3.5:27b moved from a prior score of 2.0 to a combined score of 9.2. This model was identified as the best voice in the re-bake, even though it was not chosen for the final deployment seat.",
   "cites": [
    "the-re-bake-thirteen-models-one-scale-after-the-fix"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-718e7e89",
   "q": "what determined the final choice for the narrator seat?",
   "a": "Residency was the binding constraint rather than voice. While a q8 model had the top voice, its size left insufficient headroom for the image pipeline, leading to the selection of a q4 sibling to ensure co-residency on the stack.",
   "cites": [
    "why-the-best-voice-lost-its-seat-twice"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-e8923c4e",
   "q": "how does the narrator function within the engine?",
   "a": "The narrator speaks through a chosen seat under a canon-lock. The engine manages the facts and rolls the dice, while the model provides the voice for the experience.",
   "cites": [
    "where-the-voice-lives"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-25e5985d",
   "q": "what is the significance of the anchor model's scores?",
   "a": "The anchor model provides a consistent reference point across different scales, reading 9.0, 7.0, and 6.2 respectively. These marks are used to show method offset rather than model regression, and no delta across these record sets is ever published.",
   "cites": [
    "what-to-take-with-you"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  }
 ],
 "related": [
  {
   "slug": "three-new-voices-at-the-narrators-chair",
   "why": "shares ground with § What to take with you · § The chair, and who sat it before"
  },
  {
   "slug": "august-arrivals",
   "why": "shares ground with § What to take with you · § Exam two: the narrator's chair"
  },
  {
   "slug": "cove-voice-head-to-head",
   "why": "shares ground with § The method, limits first · § Limits, stated plainly"
  }
 ],
 "thanks": "",
 "kit": {
  "url": "https://research.strata2signal.com/voice-trials/data/",
  "licence": "CC BY 4.0",
  "files": [
   "addendum-2026-08-12.json"
  ]
 },
 "seat_class": {
  "writer": "gemma-class",
  "judge": null,
  "writer_runtime": "vllm"
 },
 "bench": {
  "bullets_written": 3,
  "bullets_kept": 3,
  "digests_written": 9,
  "digests_kept": 5,
  "chips_written": 8,
  "chips_kept": 7,
  "chips_grounded": 0,
  "chips_published": 0,
  "dropped_by": {
   "new_noun": 4,
   "figure": 1,
   "length": 0,
   "cite": 0,
   "judge": 0,
   "quote": 0,
   "directive": 0,
   "redaction": 0,
   "profanity": 0
  },
  "new_noun_tokens_checked": [
   "10-25",
   "13",
   "15",
   "2.0",
   "2026-08",
   "24G-era",
   "27b",
   "3",
   "4",
   "400",
   "44-46",
   "48-54",
   "54-second",
   "6.2",
   "7.0",
   "8.7",
   "8G",
   "8G-era",
   "9.0",
   "9.2",
   "96G",
   "9b",
   "GB",
   "KV",
   "LLM",
   "RealKeep's",
   "VRAM",
   "gemma2",
   "gemma4",
   "q4",
   "q8",
   "qwen3.5",
   "qwen3.6"
  ],
  "new_noun_tokens_withheld": 4,
  "figure_definition": "v3",
  "drops": [
   {
    "kind": "digest",
    "reason": "new_noun",
    "item": null,
    "item_chars": 10,
    "withheld": "a string naming something the article does not",
    "detail": ""
   },
   {
    "kind": "digest",
    "reason": "new_noun",
    "item": null,
    "item_chars": 10,
    "withheld": "a string naming something the article does not",
    "detail": ""
   },
   {
    "kind": "digest",
    "reason": "new_noun",
    "item": null,
    "item_chars": 10,
    "withheld": "a string naming something the article does not",
    "detail": ""
   },
   {
    "kind": "digest",
    "reason": "new_noun",
    "item": null,
    "item_chars": 14,
    "withheld": "a string naming something the article does not",
    "detail": ""
   },
   {
    "kind": "chip",
    "reason": "figure",
    "item": "When thinking was enabled, the model failed to reach the envelope because 13 of 15 calls stopped due to length constraints. This resulted in a median of approximately 4,400 characters of reasoning and",
    "detail": "figures absent from cited spans: ['4,400']"
   }
  ]
 },
 "generated_utc": "2026-09-09T17:49:20Z"
}
