{
 "schema_version": "1.0",
 "page_kind": "article",
 "pack_kind": "generate",
 "slug": "outside-judges",
 "title": "The outside judges — four rival labs re-check our work.",
 "dek": "Four rival frontier labs re-judged every sealed round we had published — blind, byte-identical.",
 "published": "2026-08-13",
 "series": [
  "bench"
 ],
 "licence": "CC BY 4.0",
 "status": "pending-judge",
 "pack_note": "no judge is seated, so every chip is held",
 "notice": "items with published:false are held and are not this page's published words",
 "source": {
  "url": "https://research.strata2signal.com/outside-judges/",
  "md_url": "https://research.strata2signal.com/outside-judges/index.md",
  "md_sha": "cbca8029274664665eac79343d2f2bf5fdaa8a5a2682c9be5e6a6c1e25e8281a",
  "html_sha": "a74cc96596987ed474525a27554a06b36b4f48cc918e7d5a46fdc021c19043ba",
  "anchors_sha": "2e0fdddc31a1c53249d42870073a5c61fa80a5181655fead1798fe884b0dde8c"
 },
 "short": {
  "paragraph": "Four rival frontier labs re-judged every sealed round we had published - blind, byte-identical.",
  "from_dek": true,
  "counts": {
   "words": 5499,
   "minutes": 25,
   "tables": 5,
   "kit": true
  },
  "bullets": [
   {
    "text": "the widest cross-vendor gap on any arm is 0.178",
    "figure": "0.178",
    "cite": "vendors",
    "quote": "What the vendor columns show: muse-glimmer:30b-q8\\0-dflash · spread: 0.178 anthropic - 0.589 deepseek - 0.611 mistral - 0.450 moonshot - 0.572 nvidia - 0.628 spread - 0.178 muse-glimmer:30b-q8\\0-dflash+think · spread: 0.028 anthropic - 0.117 deepseek - 0.089 mistral - 0.100 moonshot - 0.117 nvidia - 0.106 spread - 0.028 gemma4:26b · spread: 0.061 anthropic - 0.539 deepseek - 0.517 mistral - 0.550 moonshot - 0.578 nvidia - 0.550 spread - 0.061 gemma4:12b · spread: 0.140 anthropic - 0.685 deepseek - 0.606 mistral - 0.656 moonshot - 0.633 nvidia - 0.544 spread - 0.140 qwen3.6:27b · spread: 0.094 anthropic - 0.711 deepseek - 0.778 mistral - 0.750 moonshot - 0.683 nvidia - 0.694 spread - 0.094 nemotron-3.5-lightning:30b-a3b · spread: 0.135 anthropic - 0.360 deepseek - 0.400 mistral - 0.494 moonshot - 0.417 nvidia - 0.478 spread - 0.135 Rows are exhibit six's C4 roster order, unchanged, and the anthropic column IS that exhibit's published win-rate column - the same seats reading the same files, printed here a second time as one vendor among several.",
    "span": [
     8849,
     8855
    ]
   },
   {
    "text": "the challenger kimi-k3 cleared both frozen floors at 26/27",
    "figure": "26/27",
    "cite": "gauntlet",
    "quote": "kimi-k3 · kills /27: 26/27 · preservation /16: 15/16 · verdict: PASS protocol class - cloud-ollama kills /27 - 26/27 preservation /16 - 15/16 floors - clears both verdict - PASS claude-opus-5 · kills /27: 27/27 · preservation /16: 16/16 · verdict: NOT SCORED protocol class - agent-harness kills /27 - 27/27 preservation /16 - 16/16 floors - clears both verdict - NOT SCORED claude-fable-5 · kills /27: 27/27 · preservation /16: 16/16 · verdict: NOT SCORED protocol class - agent-harness kills /27 - 27/27 preservation /16 - 16/16 floors - clears both verdict - NOT SCORED The gold-highlighted row is the one row the instrument could rank - the fully probed, seat-eligible PASS.",
    "span": [
     17987,
     17993
    ]
   },
   {
    "text": "the total across both duties was $1.51",
    "figure": "$1.51",
    "cite": "bill",
    "quote": "Three rode a subscription the workshop already pays for; the metered one - kimi - ran the entire night, both sides of the table, for $1.51 of a twenty-dollar credit.",
    "span": [
     24480,
     24488
    ]
   }
  ]
 },
 "sections": [
  {
   "id": "boundary",
   "heading": "Local first, checked from outside",
   "level": 2,
   "span": [
    1486,
    2994
   ],
   "chunks": [
    [
     1486,
     2994
    ]
   ],
   "chars": 1508,
   "digest": "the products are local-first and stay that way, with the cloud used only for verification by outside judges. what left the box for judging was model outputs over public-domain text and published game canon, never user data, never a character's private canon, never telemetry. cloud judges are versionless hosted services dated 2026-08-12/13.",
   "digest_skipped": null
  },
  {
   "id": "audition",
   "heading": "The audition",
   "level": 2,
   "span": [
    2994,
    8302
   ],
   "chunks": [
    [
     2994,
     8302
    ]
   ],
   "chars": 5308,
   "digest": "every cloud candidate sat a one-batch audition on a sealed canon round before any full round. all four carried on the first attempt. the seated panel included two claude-opus-5, two claude-fable-5, deepseek-v4-pro, mistral-large-3:675b, nemotron-3-ultra, and kimi-k3. six judges from five vendors were present.",
   "digest_skipped": null
  },
  {
   "id": "vendors",
   "heading": "What five vendors said about the same comparisons",
   "level": 2,
   "span": [
    8302,
    17466
   ],
   "chunks": [
    [
     8302,
     17466
    ]
   ],
   "chars": 9164,
   "digest": "the extension re-judged every sealed round of the chair trials and the voice addendum, producing 2,592 new verdict objects. qwen3.6:27b leads the narrator table's top band in every one of the five vendors' columns. the widest cross-vendor gap on any arm is 0.178.",
   "digest_skipped": null
  },
  {
   "id": "gauntlet",
   "heading": "The gauntlet: our judges, examined",
   "level": 2,
   "span": [
    17466,
    22709
   ],
   "chunks": [
    [
     17466,
     22709
    ]
   ],
   "chars": 5243,
   "digest": "the seat-trials exam measures judge quality against ground truth. kimi-k3 passed with 26/27 kills and 15/16 preservation. both claude models returned perfect sheets of 27/27 kills and 16/16 preservations, but the instrument declined to rank them because they answered over an unprobed transport.",
   "digest_skipped": null
  },
  {
   "id": "bill",
   "heading": "What it cost",
   "level": 2,
   "span": [
    22709,
    26016
   ],
   "chunks": [
    [
     22709,
     26016
    ]
   ],
   "chars": 3307,
   "digest": "the cloud judges' costs varied, with kimi-k3 running the entire night for $1.51 of a $20 credit. three of the four cloud judges ride a subscription the workshop already pays for and print $0.00*. the total across both duties was $1.51 of a $20 credit, across 289 metered calls.",
   "digest_skipped": null
  },
  {
   "id": "going-forward",
   "heading": "What changes going forward",
   "level": 2,
   "span": [
    26016,
    27044
   ],
   "chunks": [
    [
     26016,
     27044
    ]
   ],
   "chars": 1028,
   "digest": "every judged test published from here forward carries at least one judge from outside the model family, ideally three families or more. the sealed-batch design makes re-auditioning nearly free, as the batches do not expire and neither does the question.",
   "digest_skipped": null
  },
  {
   "id": "limits-stated-plainly",
   "heading": "Limits, stated plainly",
   "level": 2,
   "span": [
    27044,
    27705
   ],
   "chunks": [
    [
     27044,
     27705
    ]
   ],
   "chars": 661,
   "digest": "the gauntlet's claude arm ran an unprobed transport and is published unranked. cloud judges are versionless services; every verdict here is dated 2026-08-12/13 and may not reproduce against tomorrow's checkpoints. agreement statistics ride the same small-N caveats as the rounds they extend.",
   "digest_skipped": null
  },
  {
   "id": "what-to-take-with-you",
   "heading": "What to take with you",
   "level": 2,
   "span": [
    27705,
    30124
   ],
   "chunks": [
    [
     27705,
     30124
    ]
   ],
   "chars": 2419,
   "digest": "the all-Claude result held up when four rival labs re-scored it. vendors disagree about second place, with the widest cross-vendor gap being 0.178. kimi-k3, the challenger, came back funded and passed. a second opinion from four rival labs cost $1.51.",
   "digest_skipped": null
  },
  {
   "id": "how-to-check-our-work-and-see-it-live",
   "heading": "How to check our work — and see it live",
   "level": 2,
   "span": [
    30124,
    32933
   ],
   "chunks": [
    [
     30124,
     32933
    ]
   ],
   "chars": 2809,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "the-rest-of-the-seminar",
   "heading": "The rest of the seminar",
   "level": 2,
   "span": [
    32933,
    33843
   ],
   "chunks": [
    [
     32933,
     33843
    ]
   ],
   "chars": 910,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "provenance-7d07c0",
   "heading": "Provenance",
   "level": 2,
   "span": [
    33843,
    37132
   ],
   "chunks": [
    [
     33843,
     37132
    ]
   ],
   "chars": 3289,
   "digest": null,
   "digest_skipped": "credits"
  }
 ],
 "chips": [
  {
   "id": "c-2eeda46a",
   "q": "how is user data protected during the judging process?",
   "a": "The data boundary is enforced by the driver. Only model outputs over public-domain text and published game canon left the box for judging. User data, private character canon, and telemetry were never shared with the cloud judges.",
   "cites": [
    "boundary"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-1417fa3c",
   "q": "what criteria determine if a cloud judge passes the audition?",
   "a": "A candidate carries if its reply parses and validates against the round's schema after at most one registered format-reminder retry. The audition uses a specific batch of the candidate's assigned source seat to ensure the reply is valid.",
   "cites": [
    "audition"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-a300ffc2",
   "q": "how much do the different vendors disagree on model rankings?",
   "a": "Vendors show consensus on the top and bottom of the table, but differ on middle rankings. The widest cross-vendor gap on any arm is 0.178, and the weakest rank agreement is between Mistral and NVIDIA at ρ 0.60.",
   "cites": [
    "vendors"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-2138648d",
   "q": "what happens when a model fails the seat-trials exam?",
   "a": "The instrument uses frozen floors for scoring. A candidate must achieve a kill-recall of at least 23 of 27 and a confirmed-preservation of at least 13 of 16. If a candidate fails to meet these independent thresholds, it is out.",
   "cites": [
    "gauntlet"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-113a4bfe",
   "q": "what is the total cost for the metered cloud judge calls?",
   "a": "The metered judge, kimi-k3, ran the entire night for $1.51 of a $20 credit. This covered 289 metered calls across both sides of the table, including the judge side and the candidate-side duties.",
   "cites": [
    "bill"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-c77e3a0c",
   "q": "what is the new policy for future judged tests?",
   "a": "Every judged test published from this point forward must carry at least one judge from outside the workshop's own model family, with an ideal target of three or more families.",
   "cites": [
    "going-forward"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-75851a4c",
   "q": "what limitations exist regarding the cloud judge verdicts?",
   "a": "Cloud judges are versionless services, meaning every verdict is dated and may not reproduce against later checkpoints. Additionally, agreement statistics are subject to small-N caveats and the effective N of the items judged.",
   "cites": [
    "limits-stated-plainly"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-519049f4",
   "q": "what are the main takeaways from the vendor comparison?",
   "a": "The all-Claude results held up under re-scoring by four rival labs. While vendors disagree on second place, the top band remains consistent across all five vendors, and the rank agreement with the house panel is high.",
   "cites": [
    "what-to-take-with-you"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  }
 ],
 "related": [
  {
   "slug": "the-new-kid",
   "why": "shares ground with § The gauntlet: our judges, examined · § The judge seat — and the floor it missed"
  },
  {
   "slug": "the-instrument-travels",
   "why": "shares ground with § The gauntlet: our judges, examined · § The standings"
  },
  {
   "slug": "three-new-voices-at-the-narrators-chair",
   "why": "shares ground with § The gauntlet: our judges, examined · § Seven seats, six carried, and where their zeros sit"
  }
 ],
 "thanks": "",
 "kit": {
  "url": "https://research.strata2signal.com/outside-judges/data/",
  "licence": "CC BY 4.0",
  "files": [
   "audition.json",
   "panel-vendors.json",
   "agreement.json",
   "gauntlet.json",
   "bill.json",
   "verdict-sets.json",
   "counting-rules.json",
   "provenance.json",
   "README.md"
  ]
 },
 "seat_class": {
  "writer": "gemma-class",
  "judge": null,
  "writer_runtime": "vllm"
 },
 "bench": {
  "bullets_written": 3,
  "bullets_kept": 3,
  "digests_written": 8,
  "digests_kept": 8,
  "chips_written": 8,
  "chips_kept": 8,
  "chips_grounded": 0,
  "chips_published": 0,
  "dropped_by": {
   "new_noun": 0,
   "figure": 0,
   "length": 0,
   "cite": 0,
   "judge": 0,
   "quote": 0,
   "directive": 0,
   "redaction": 0,
   "profanity": 0
  },
  "new_noun_tokens_checked": [
   "0.00",
   "0.178",
   "0.60",
   "1.51",
   "13",
   "15/16",
   "16",
   "16/16",
   "2",
   "20",
   "2026-08-12/13",
   "23",
   "26/27",
   "27",
   "27/27",
   "27b",
   "289",
   "592",
   "675b",
   "Mistral",
   "N",
   "NVIDIA",
   "claude-fable-5",
   "claude-opus-5",
   "deepseek-v4-pro",
   "kimi-k3",
   "mistral-large-3",
   "nemotron-3-ultra",
   "qwen3.6"
  ],
  "new_noun_tokens_withheld": 0,
  "figure_definition": "v3",
  "drops": []
 },
 "generated_utc": "2026-09-09T17:49:08Z"
}
