{
 "schema_version": "1.0",
 "page_kind": "article",
 "pack_kind": "generate",
 "slug": "two-new-frontier-models-at-the-rules-desk",
 "title": "Two New Frontier Models at the Rules Desk",
 "dek": "RuleSage, our board-game rules helper, answers a question from the rulebook’s own words: it finds the passage, cites it by number so a player can check, and says plainly when the book does not answer. GPT-6 Astra and Claude Fable 5.1 sat that job’s exam, frozen in August, judged blind by models from seven families that built neither, beside the small local model that already does the job. Neither new model cleared it, and a blind seven-family panel did not separate the two over thirty-six cases.",
 "published": "2026-09-06",
 "series": [
  "bench"
 ],
 "licence": "CC BY 4.0 for the kit's records; the rulebook passages are the publishers' and are not republished",
 "status": "pending-judge",
 "pack_note": "no judge is seated, so every chip is held",
 "notice": "items with published:false are held and are not this page's published words",
 "source": {
  "url": "https://research.strata2signal.com/two-new-frontier-models-at-the-rules-desk/",
  "md_url": "https://research.strata2signal.com/two-new-frontier-models-at-the-rules-desk/index.md",
  "md_sha": "709057386359b1366d0c9792842a10d94252c92d2a7c579ef09d7024a111860d",
  "html_sha": "b3961dc168234d60326843e7721fea515feec48bef50924e91d1f9b932a46ea5",
  "anchors_sha": "338a7769e2face3116198433917cc3c5f97383a7a5958b5e89ae3dde04ce47e1"
 },
 "short": {
  "paragraph": "RuleSage, our board-game rules helper, answers a question from the rulebook's own words: it finds the passage, cites it by number so a player can check, and says plainly when the book does not answer. GPT-6 Astra and Claude Fable 5.1 sat that job's exam, frozen in August, judged blind by models from seven families that built neither, beside the small local model that already does the job. Neither new model cleared it, and a blind seven-family panel did not separate the two over thirty-six cases.",
  "from_dek": true,
  "counts": {
   "words": 14805,
   "minutes": 67,
   "tables": 17,
   "kit": true
  },
  "bullets": [
   {
    "text": "the pooled preference rate for Claude Fable 5.1 is 0.544",
    "figure": "0.544",
    "cite": "head-to-head",
    "quote": "Over the 36 cases with an observation, the pooled preference rate is 0.544 for Claude Fable 5.1 - read the other way, 0.456 for GPT-6 Astra.",
    "span": [
     57303,
     57311
    ]
   },
   {
    "text": "the total dollars that changed hands come to $23.05",
    "figure": "$23.05",
    "cite": "the-bill",
    "quote": "The dollars that changed hands come to $23.05 (GPT-6 Astra $17.79 + kimi-k3 $5.25 - each rounded for the table; summed before rounding for the total, $17.7931 + $5.2548 = $23.0479), against a registered cap of $60.00 (A $25.00 · C $20.00 · probes $5.00 · reserve $10.00); no cap bound and no call was cut short by one.",
    "span": [
     100657,
     100664
    ]
   }
  ]
 },
 "sections": [
  {
   "id": "two-strangers-at-the-door",
   "heading": "Two strangers at the door",
   "level": 2,
   "span": [
    1361,
    5518
   ],
   "chunks": [
    [
     1361,
     5518
    ]
   ],
   "chars": 4157,
   "digest": "RuleSage is a board-game rules helper that finds passages, quotes them, and says when the book does not answer. This page compares Claude Fable 5.1, GPT-6 Astra, and Gemma 4 (26B) on two jobs: a rules desk and a filing cabinet, using an instrument sealed before the models arrived.",
   "digest_skipped": null
  },
  {
   "id": "who-judged-and-who-wrote-this",
   "heading": "Who judged, and who wrote this",
   "level": 2,
   "span": [
    5518,
    10536
   ],
   "chunks": [
    [
     5518,
     10536
    ]
   ],
   "chars": 5018,
   "digest": "A panel of 7 judging seats from 7 model families read answers blind. The agents that designed, built, ran, audited, and wrote this page run on one of the two vendors' models. Mistral Large 3 also wrote the replacement questions and sits on the judging panel.",
   "digest_skipped": null
  },
  {
   "id": "the-words-this-page-leans-on",
   "heading": "The words this page leans on",
   "level": 2,
   "span": [
    10536,
    18277
   ],
   "chunks": [
    [
     10536,
     18277
    ]
   ],
   "chars": 7741,
   "digest": "This section defines the terminology used throughout the article, including terms like arm, road, transport, seat, the shelf, Leg A, Leg H, Leg C, gate, floor, and the answer fence. It also explains technical states like COLLECTED, NOT-COLLECTED, and the meaning of the brackets.",
   "digest_skipped": null
  },
  {
   "id": "three-weeks-earlier",
   "heading": "Three weeks earlier",
   "level": 2,
   "span": [
    18277,
    19323
   ],
   "chunks": [
    [
     18277,
     19323
    ]
   ],
   "chars": 1046,
   "digest": "The rules desk instrument was sealed on 2026-08-15. GPT-6 Astra's model record appeared on 2026-08-27, Claude Fable 5.1 shipped on 2026-09-01, and GPT-6 Astra shipped on 2026-09-03. The operator co-signed the pre-registration on 2026-09-05.",
   "digest_skipped": null
  },
  {
   "id": "the-rules-desk",
   "heading": "The rules desk",
   "level": 2,
   "span": [
    19323,
    54318
   ],
   "chunks": [
    [
     19323,
     38825
    ],
    [
     38825,
     54318
    ]
   ],
   "chars": 34995,
   "digest": "Sixty questions about board games were put to three models three times. The task is to cite passages, abstain when the book does not answer, and ignore hidden directives. GPT-6 Astra and Claude Fable 5.1 cleared the hidden directives gate, while the local seat missed it.",
   "digest_skipped": null
  },
  {
   "id": "head-to-head",
   "heading": "Head-to-head",
   "level": 2,
   "span": [
    54318,
    64243
   ],
   "chunks": [
    [
     54318,
     64243
    ]
   ],
   "chars": 9925,
   "digest": "Seven outside families read pairs of rules-desk answers blind to choose between the two frontier models. The pooled preference rate is 0.544 for Claude Fable 5.1. A cluster bootstrap over 36 cases puts that rate at [0.444, 0.638], an interval that covers 0.5.",
   "digest_skipped": null
  },
  {
   "id": "the-filing-cabinet",
   "heading": "The filing cabinet",
   "level": 2,
   "span": [
    64243,
    77558
   ],
   "chunks": [
    [
     64243,
     77558
    ]
   ],
   "chars": 13315,
   "digest": "The filing cabinet tests if a model can find a planted sentence in long text or say the text does not say when nothing was planted. GPT-6 Astra and Claude Fable 5.1 achieved 18 of 18 recall and 18 of 18 abstention. The local seat collected 24 items.",
   "digest_skipped": null
  },
  {
   "id": "what-the-numbers-are-allowed-to-mean",
   "heading": "What the numbers are allowed to mean",
   "level": 2,
   "span": [
    77558,
    79328
   ],
   "chunks": [
    [
     77558,
     79328
    ]
   ],
   "chars": 1770,
   "digest": "Four rules apply to the frontier model rows: they are not candidates for products, no ordering is drawn between them, the only comparison is blind preference, and everything is dated because hosted models are not fixed objects.",
   "digest_skipped": null
  },
  {
   "id": "the-ledger-and-what-left-our-machines",
   "heading": "The ledger, and what left our machines",
   "level": 2,
   "span": [
    79328,
    93201
   ],
   "chunks": [
    [
     79328,
     93201
    ]
   ],
   "chars": 13873,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "what-this-page-does-not-say",
   "heading": "What this page does not say",
   "level": 2,
   "span": [
    93201,
    98150
   ],
   "chunks": [
    [
     93201,
     98150
    ]
   ],
   "chars": 4949,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "the-bill",
   "heading": "The bill",
   "level": 2,
   "span": [
    98150,
    104900
   ],
   "chunks": [
    [
     98150,
     104900
    ]
   ],
   "chars": 6750,
   "digest": "The total cost for the metered parts of the round was $23.05. GPT-6 Astra cost $17.79 and Kimi K3 cost $5.25. Claude Fable 5.1's row shows a list-rate estimate of $31.44. The registered cap was $60.00.",
   "digest_skipped": null
  },
  {
   "id": "what-to-take-with-you",
   "heading": "What to take with you",
   "level": 2,
   "span": [
    104900,
    108958
   ],
   "chunks": [
    [
     104900,
     108958
    ]
   ],
   "chars": 4058,
   "digest": "Neither frontier model cleared the whole rules desk. All 7 rival families read both frontier answers blind and did not separate them. On the filing cabinet, both frontier models were at ceiling, while the local seat ran out of room.",
   "digest_skipped": null
  },
  {
   "id": "how-to-check-our-work-and-see-it-live",
   "heading": "How to check our work — and see it live",
   "level": 2,
   "span": [
    108958,
    112261
   ],
   "chunks": [
    [
     108958,
     112261
    ]
   ],
   "chars": 3303,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "who-ran-this-and-thanks",
   "heading": "Who ran this, and thanks",
   "level": 2,
   "span": [
    112261,
    115541
   ],
   "chunks": [
    [
     112261,
     115541
    ]
   ],
   "chars": 3280,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "the-rest-of-the-seminar",
   "heading": "The rest of the seminar",
   "level": 2,
   "span": [
    115541,
    116620
   ],
   "chunks": [
    [
     115541,
     116620
    ]
   ],
   "chars": 1079,
   "digest": null,
   "digest_skipped": "credits"
  }
 ],
 "chips": [
  {
   "id": "c-dd081f64",
   "q": "how are the two frontier models being compared?",
   "a": "The article compares Claude Fable 5.1 and GPT-6 Astra on specific tasks performed by RuleSage, a board-game rules helper, alongside a local Gemma 4 (26B) model. The comparison focuses on how these models handle rules retrieval, citation accuracy, and finding specific information within long documents.",
   "cites": [
    "two-strangers-at-the-door"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-b6e66807",
   "q": "how was the judging panel for this round composed?",
   "a": "The panel consisted of seven judging seats from seven different model families, including Mistral Large 3, NVIDIA's Nemotron 3 Ultra, and Alibaba's Qwen 3.5. The judges performed blind scoring, meaning they saw the answers without knowing which model produced them.",
   "cites": [
    "who-judged-and-who-wrote-this"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-9a4e93af",
   "q": "what was the purpose of the rules desk measurement?",
   "a": "The rules desk measured how well models follow a contract to cite passages correctly, say when a book does not answer a question, and ignore hidden directives within source passages. It tested sixty cases, including questions that the rulebook answers and those it does not.",
   "cites": [
    "the-rules-desk#0"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-36edd049",
   "q": "how did the head-to-head voting result between the two frontier models?",
   "a": "The pooled preference rate was 0.544 for Claude Fable 5.1 and 0.456 for GPT-6 Astra. Because the 95% cluster bootstrap interval of [0.444, 0.638] covers 0.5, the panel did not statistically separate the two arms on the rules desk.",
   "cites": [
    "head-to-head"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-433f6e5c",
   "q": "what tasks were required for the filing cabinet test?",
   "a": "The filing cabinet task required models to recall a single invented sentence planted within a long text of card-game prose and to correctly identify when a sentence was never planted at all. The text tiers ranged from 8k to 32k estimated tokens.",
   "cites": [
    "the-filing-cabinet"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-d7086aea",
   "q": "how were the different model roads reached during the testing?",
   "a": "Claude Fable 5.1 was accessed through a sealed command-line tool, while GPT-6 Astra was reached through the OpenAI API. The local seat was accessed via the workshop's own hardware instance.",
   "cites": [
    "what-this-page-does-not-say"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-60a669cf",
   "q": "what were the total costs for the models that answered or judged?",
   "a": "The total amount that changed hands was $23.05, consisting of $17.79 for GPT-6 Astra and $5.25 for kimi-k3. Claude Fable 5.1's cost was listed as a $31.44 estimate based on a subscription-billed account.",
   "cites": [
    "the-bill"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  }
 ],
 "related": [
  {
   "slug": "three-new-voices-at-the-narrators-chair",
   "why": "shares ground with § Who judged, and who wrote this · § Seven seats, six carried, and where their zeros sit"
  },
  {
   "slug": "the-open-call",
   "why": "shares ground with § The bill · § $2.2826, over 56 metered calls, against $38.49 of caps"
  },
  {
   "slug": "outside-judges",
   "why": "shares ground with § Who judged, and who wrote this · § The audition"
  }
 ],
 "thanks": "## Who ran this, and thanks {#who-ran-this-and-thanks}\n\n**The texts.** Foster's Complete Hoyle (1914), Project Gutenberg #53881, public domain — the filing cabinet's filler, digitised by Project Gutenberg's volunteers, to whom the cabinet owes its million and a half characters. The rulebooks the rules desk cites, thirty-nine titles across the sixty cases, whose publishers wrote the passages every answer was measured against and who owed us nothing: they are held under each publisher's own terms, the corpus is not republished, and the kit carries the scores, the citation counts and each judged answer's sha, never the passages themselves.\n\n**The judges.** The seven model families that built neither contestant, named above, with a version where the shelf pins one, every one of them through the ollama.com shelf — and one of them, Mistral Large 3, also reading the pre-registration for us and writing the thirty-four replacement questions.\n\n**The local seat.** Gemma 4 (26B) on our own hardware, at its production settings — the model that answers our users — and the ollama runtime it runs on.\n\n**The layer beneath.** Python 3 and its standard library, plus RuleSage's own frozen answer assembly and abstention matcher, copied into the harness at a pinned commit so that every arm was asked exactly what the product asks — and nothing else. The harness makes its calls with the standard library's own HTTP client, on purpose, so that every byte on the wire is in a file a reader can open. The Claude Code command-line tool (2.1.261) was the road to Claude Fable 5.1; the OpenAI API was the road to GPT-6 Astra. The harness itself is the descendant of three earlier ones on this shelf — the open call's rig, the offload seat gate, and the filing cabinet — copied in with the sha of every file it took.\n\n**The humans.** An operator read the pre-registration and co-signed it before the first call, ruled on what could leave our machines, and read this page before it went out. The numbers were written by the scorers; three reader passes re-derived them from the kit, and their reports ship under `readers/`. An adversarial reader running on an Anthropic model — the same company as one of the two arms, because no outside model with the context to do that reading sat for it, and the page names the limit rather than call the reader independent — read a late draft in full against the kit, the sealed pre-registration and the scorer output; what it found, what changed, and its pass on the version you are reading are in its receipt under `receipts/`, and its critique is listed in the kit's index as withheld, its sha beside it. The pre-registration had its own outside reader, `mistral-large-3:675b`, whose raw reply (sha 5e274118…) ships in the kit under `receipts/`.\n\n<!-- fill `operator-read` — prereg/receipts/20260905T221500Z-operator-read.json: version_read, utc, found, found[].what / .landed_in -->\n<!-- fill `hostile-reader-and-outside-read` — prereg/receipts/20260905T224559Z-hostile-read.json: verdict, evidence, evidence.<count fields>, version_read, utc, verify_pass, verify_pass.version, verify_pass.new_must_fixes ; prereg/receipts/20260905T043934Z-outside-prereg-read.json: model, reply_sha256 ; results/kit/index.json: files[].file, withheld[].file -->",
 "kit": {
  "url": "https://research.strata2signal.com/two-new-frontier-models-at-the-rules-desk/data/",
  "licence": "CC BY 4.0",
  "files": [
   "legA-gates.json",
   "legA-queries.json",
   "legA-g4.json",
   "pairwise.json",
   "pairwise-key.json",
   "legC-cells.json",
   "legC-cells-think-true.json",
   "legC-needles.json",
   "legC-fixture-manifest.json",
   "counting-rules.json",
   "seats.json",
   "bill.json",
   "seal-manifest.json",
   "prereg.md",
   "prereg-index.md",
   "legC-score.py",
   "score.py",
   "pairwise-score.py",
   "design-floors.py",
   "prereg-integers.py",
   "article-fill.py",
   "fills.json",
   "public-copy.py",
   "receipts/2026-09-04-astra-effort-high.json",
   "receipts/2026-09-04-astra-effort-low.json",
   "receipts/2026-09-04-astra-smoke.json",
   "receipts/2026-09-05T140417Z-cli-claude-fable-5-1-g-quota.json",
   "receipts/2026-09-05T140422Z-cli-claude-fable-5-1-hook-nonce.json",
   "receipts/2026-09-05T140425Z-cli-claude-fable-5-1-session-persistence.json",
   "receipts/2026-09-05T140429Z-cli-claude-fable-5-1-g-tools.json",
   "receipts/2026-09-05T140431Z-openai-gpt-6-astra-g-tools.json",
   "receipts/2026-09-05T140442Z-cli-claude-fable-5-1-g-effort.json",
   "receipts/2026-09-05T140448Z-openai-gpt-6-astra-g-effort.json",
   "receipts/2026-09-05T144532Z-cli-claude-fable-5-1-g-quota.json",
   "receipts/2026-09-05T144819Z-cli-claude-fable-5-1-g-quota.json",
   "receipts/20260905T043729Z-ollama-shelf-read.json",
   "receipts/20260905T043934Z-outside-prereg-read.json",
   "receipts/20260905T140442Z-cli-claude-fable-5-1-double-canary.json",
   "receipts/20260905T140448Z-openai-gpt-6-astra-double-canary.json",
   "receipts/20260905T140451Z-local-gemma4-26b-double-canary.json",
   "receipts/20260905T144854Z-cli-identity-census.json",
   "receipts/20260905T181859Z-g-pen.json",
   "receipts/20260905T181859Z-hostile-read.json",
   "receipts/20260905T195339Z-hostile-read.json",
   "receipts/20260905T195711Z-hostile-read.json",
   "receipts/20260905T203537Z-hostile-read.json",
   "receipts/20260905T210832Z-hostile-read.json",
   "receipts/20260905T221500Z-operator-read.json",
   "receipts/20260905T224559Z-hostile-read.json",
   "receipts/20260905T230427Z-g-pen.json",
   "receipts/20260905T231355Z-g-pen.json",
   "receipts/20260906T001050Z-g-pen.json",
   "receipts/20260906T003223Z-g-pen.json",
   "receipts/20260906T005938Z-g-pen.json",
   "receipts/20260906T010732Z-g-pen.json",
   "receipts/20260906T011522Z-g-pen.json",
   "receipts/20260906T012303Z-g-pen.json",
   "receipts/20260906T012814Z-g-pen.json",
   "receipts/20260906T013439Z-g-pen.json",
   "receipts/20260906T013912Z-g-pen.json",
   "receipts/20260906T014310Z-g-pen.json",
   "receipts/20260906T014743Z-g-pen.json",
   "receipts/20260906T041725Z-g-pen.json",
   "readers/v1-coldscroll.md",
   "readers/v1-hostile.md",
   "readers/v1-numbers.md",
   "readers/v1-stranger.md",
   "readers/v2-numbers.md",
   "readers/v2-stranger.md",
   "readers/v3-coldscroll.md",
   "readers/v4-numbers.md",
   "README.md"
  ]
 },
 "seat_class": {
  "writer": "gemma-class",
  "judge": null,
  "writer_runtime": "vllm"
 },
 "bench": {
  "bullets_written": 3,
  "bullets_kept": 2,
  "digests_written": 12,
  "digests_kept": 10,
  "chips_written": 8,
  "chips_kept": 7,
  "chips_grounded": 0,
  "chips_published": 0,
  "dropped_by": {
   "new_noun": 0,
   "figure": 3,
   "length": 1,
   "cite": 0,
   "judge": 0,
   "quote": 0,
   "directive": 0,
   "redaction": 0,
   "profanity": 0
  },
  "new_noun_tokens_checked": [
   "0.444",
   "0.456",
   "0.5",
   "0.544",
   "0.5442",
   "0.638",
   "17.79",
   "18",
   "2026-08-15",
   "2026-08-27",
   "2026-09-01",
   "2026-09-03",
   "2026-09-05",
   "23.05",
   "24",
   "26B",
   "3",
   "3.5",
   "31.44",
   "32k",
   "36",
   "4",
   "5.1",
   "5.1's",
   "5.25",
   "60.00",
   "7",
   "8k",
   "95%",
   "API",
   "Alibaba's",
   "Astra",
   "Astra's",
   "C",
   "COLLECTED",
   "Claude",
   "Fable",
   "GPT-6",
   "Gemma",
   "H",
   "K3",
   "Kimi",
   "Large",
   "Leg",
   "Mistral",
   "NOT-COLLECTED",
   "NVIDIA's",
   "Nemotron",
   "OpenAI",
   "Qwen",
   "RuleSage",
   "Ultra",
   "kimi-k3"
  ],
  "new_noun_tokens_withheld": 0,
  "figure_definition": "v3",
  "drops": [
   {
    "kind": "digest",
    "reason": "figure",
    "item": "This section provides bookkeeping for text that left the machines, socket samples, and probe results. No personal data appeared in any round artifacts. The total cost for the metered arms and the judg",
    "detail": "figures absent: ['23.05']"
   },
   {
    "kind": "digest",
    "reason": "figure",
    "item": "Claude Fable 5.1 used a sealed command-line tool while GPT-6 Astra used an API. The local seat did not read the whole of a 32k prompt, as the runtime reported a ratio of 0.5442. Both frontier arms ran",
    "detail": "figures absent: ['32', '0.5442']"
   },
   {
    "kind": "bullet",
    "reason": "figure",
    "item": "the local seat's context ratio was 0.5442",
    "detail": "terminal token not verbatim in the cited span"
   },
   {
    "kind": "chip",
    "reason": "length",
    "item": "what specific terms are used to define the models and their testing environments?",
    "detail": "q 81 chars"
   }
  ]
 },
 "generated_utc": "2026-09-09T01:24:58Z"
}
