# The chair trials — data kit

The machine-readable companions to **exhibit six, "The chair trials"**
(https://research.strata2signal.com/chair-trials/index.html): nine models and eleven arms through five fresh
exams in a single day, on one workstation that kept serving its real users the
whole time.

**Licence: CC BY 4.0.** Take these rows, re-plot them, check our arithmetic, run the
sets against your own models, publish what you find. Attribution:
strata→signal research, research.strata2signal.com. If you find an error in them, we want to
hear about it: hello@strata2signal.com.

## What is in here

Seven table companions, five runnable sets, and the paperwork.

| file | what it holds |
|---|---|
| `roster.json` | the eleven arms: vendor-labelled architecture, quantization read first-hand off the daemon, measured residency in bytes and GiB at the context length every leg ran, and the postures each leg put each arm through |
| `c1-judge.json` | chair one, the judge: ten rows against two floors with Wilson intervals, self-consistency, the protocol-fence polarity probes, the markdown-fence caveat and the reasoning-dial delta |
| `c2-assistant.json` | chair two, the assistant: eight posture rows over twenty code-checked items, with every item's per-repeat verdict and the per-item matrix across all eight rows |
| `c3-toolbench.json` | chair three, the toolbench: nine rows over nineteen tasks, with every task's calls — tool, arguments, disposition — and the pre-probe that decided which runtimes could carry the contract at all |
| `c4-narrator.json` | chair four, the narrator: six arms, 1 080 pairwise comparisons with cluster-bootstrap intervals, the distinctness round, the canon adjudication with per-judge votes and notes, and the per-judge position-bias rates |
| `c5-stopwatch.json` | chair five, the stopwatch: eight arms across three prompt tiers with every raw repeat value, the two unexplained anomalies counted per arm, the drafter byte-compare receipt and the reasoning-strength sweep |
| `c6-serving.json` | the serving probes: twenty-six real asks against the live app's production path, per-ask walls with their overlap flags, the paired quiet baselines, and the burst, pair and cold probes with their receipts |
| `judge-c1.json` | the runnable judge set: all twenty-one claim, span and verdict triples, whole |
| `assistant-c2.json` | the runnable assistant set: all twenty frozen items with their prompts and expectations, plus the two trimmed at freeze, marked rather than deleted |
| `CHECKERS.md` | what each assistant checker actually tests — generated from the scorer module itself, so the published rule is the rule that ran |
| `tools-tasks.json` | the runnable toolbench set: all nineteen tasks with their expected calls and predicates |
| `tools-schemas.json` | the ten tool schemas exactly as every model received them, dumped from the suite that served them |
| `c3-example-trace.json` | one complete toolbench transcript, round by round, so the per-task calls in the C3 companion can be read for what they are |
| `voice-questions.json` | every question put to a character, with its canon note and what it probed — questions only: the persona blocks are withheld and the fence is stated |
| `c6-asks.json` | every ask the serving leg was allowed to send, frozen before the first one went out |
| `counting-rules.json` | every leg's counting rules in one file, for a reader who wants the rules without the rows |
| `provenance.json` | the pre-registration and golden-set hashes: what was pinned, when it was frozen, and what each one governs |
| `README.md` | this file |
| `index.json` | this directory, listed: file names, sizes, sha256s, and what each one answers |

Every table on the page is a projection of one of these files. The page's cells
are read back OUT of the companion by the same builder that wrote it, so page and
data reconcile by construction — if they ever disagree, that is our bug and we
want the report.

## The runnable sets

Four of these files are not records of what happened; they are the exams
themselves, complete enough to re-run against your own models:

- `judge-c1.json` — twenty-one claims, each with the rulebook span it must be
  checked against and the verdict we expect. Built from Foster's Complete Hoyle,
  1914, public domain in the US.
- `assistant-c2.json` + `CHECKERS.md` — twenty items and the code-level rule that
  decides each one. Thirteen exact, seven proxy, stamped per item, never added
  together.
- `tools-tasks.json` + `tools-schemas.json` — nineteen tasks and the ten tool
  schemas exactly as the models received them.
- `voice-questions.json` — every question put to a character, with the canon note
  the judges saw. **Questions only.** The persona blocks are withheld: they carry
  latch-gated canon and spoilers for a game people are playing. The prompt hashes
  are published so you can verify the same prompt ran on every arm without our
  handing over the prompt.

## Receipts

Files as published (sha256 of the bytes in this directory):

    roster.json              a37ad080f3a7186f5ec0c758acc4af1dcc6a65775f9223dd65f08f381997fdab
    c1-judge.json            7bbba26d0581ca241699692759fcc32e641eee952e8fc1932d24de8cd1ebdcc0
    c2-assistant.json        0ffee8f453cfbc838cebdc51e05f2ef38a22b24b14e003ae3a41ca96db729e4d
    c3-toolbench.json        75098737f370310b288f9bdcfba46e5120fd8ce0fcfb31caa306f2f2ef1ae847
    c4-narrator.json         8de910fb384032990e446f008b69045d641d7f120da664b4204fe8e76cd9460e
    c5-stopwatch.json        3a9d8e5a4585f1abc5d5a276471aac1b9d3b719c4e525e02fef12ea419fa3c79
    c6-serving.json          41183bd4fea2b362a67e30862c957c83403496da1b8f3cad12056226050805d8
    judge-c1.json            6695757e19ed0299d0a5b78db7b8c08bdbe95e056b46ccc91437a66b0a855513
    assistant-c2.json        dbf38ec7eeb3448cf5ba48e4eeb8ad906ba210a5995b972478b3a8ae7680c61c
    CHECKERS.md              95255bcbf04fae6764e1634d8b8c53f8c80966c90997483c1cfacb3c51e19d19
    tools-tasks.json         593636e04340cb3a4599fdb7ddda8835209b69bc466583e84cc52cbcafd7a585
    tools-schemas.json       63caa4edd086b320b91328e6ae1b535f66dc85bf58925c38391910850a12b150
    c3-example-trace.json    c07cda701ff63f1c9b49aa86fc6a394cde6f4d649df59e374084f7e6cdf723e2
    voice-questions.json     f10a83025d0acdf2fb337a6422baaac6345bcd520a5c5e5ff7de28064dd47a20
    c6-asks.json             09e6d04821528672e633bb4010e106dff9504e594894453583621b83d4daa736
    counting-rules.json      8a74239ac354760f94aeae40448548716b51fa07eb01793eca0d31d755e28ce4
    provenance.json          f26ce98b7fcd5591c7b14f1653379c23a2037291f8e8488ec66b13149f587fbe

The pre-registration and golden-set hashes — what was pinned before it could be
scored — are in `provenance.json`, by artifact name. No filesystem paths, no host
names, no addresses: a data kit that leaks its own machine is not a kit, it is a
disclosure.

## How to read a number here

Start from `counting_rules` in whichever file you are reading, or from
`counting-rules.json` if you want them all at once. Every figure on this page is a
count of rows in an append-only record, or a nearest-rank statistic over those
rows, and the rules say exactly which rows were counted, which stayed in the
denominator without producing a verdict, and which never became a verdict at all.

Five words do specific work, in every file:

- **RANKED** — these numbers were produced under the contract and are admissible.
  Admissible, never placed above another row: the descriptive legs carry the same
  word for the same state, because a second word for one state is how a reader
  learns to rank things nobody ranked.
- **EXPLORATORY** — a protocol fence tripped. The row ran and is published, unranked.
- **UNMEASURABLE** — response failures above the registered ceiling. The measured
  cells still print; the score is not admissible as a score.
- **NOT-CARRIED** — zero valid attempts. The raw emission is published beside the
  row, because "could not use the tools" and "could not do the work" are
  different findings.
- **NOT-RUN** — never run, with the reason on the row. Never a blank.

An em dash is a figure we do not hold. It is never a zero.

## Two fences that make these figures mean anything

**Contract failures are not capability failures.** Most of what failed in this arc
failed on the response contract with the capability visibly intact underneath —
valid JSON inside a markdown fence, a correct answer past the length budget, a
dropped envelope under a posture flag. Every state in these files is a statement
about a model, a runtime, a configuration and a contract together. Change any one
of the four and the state can change without the model changing at all.

**One box, one day, one runtime version.** These are our own trials, on our own
hardware, for our own chairs — and the box was serving live production requests
the whole time, in both directions: bench numbers carry contention flags, and the
serving companion measures what the visitors paid.

## Corrected after publication

**2026-08-16**, by a content-accuracy pass that read this kit against the exhibit
page. Two stale sentences, both of them labels: no figure, denominator, verdict,
interval or counting rule moved in either file, and both are re-hashed in the
receipts above rather than quietly swapped.

- `c4-narrator.json` — `judges.canon_and_distinctness_panel` said the distinctness
  and canon rounds ran on **three** judges, and named one Fable judge's four
  verdict batches as missing from them. They had been misfiled, not lost, and were
  recovered before scoring. Every other field in that file already recorded the
  recovered state: four seats in `distinctness_full.per_judge` at 108 assignments
  each (432 in all), four in `canon_full.per_judge` at 36 judgements each (144),
  and an empty `unscored_batches` in all three legs. The field now reads four and
  says what it used to say. The exhibit page carried the same stale sentence and
  was corrected in the same pass.
- `voice-questions.json` — the `voice-c4` set's `purpose` opened "Exhibit five,
  leg C4". Exhibit five is the August arrivals; the chair trials publish as exhibit
  **six**, which this README and the page's own eyebrow have said throughout. The
  neighbouring "exhibit three" reference is the voice trials, and it is correct.

Nothing in this kit is sealed behind a seal manifest — these are published files
with published hashes, so a correction is made in place, said in the corrected
file's own text, and re-hashed here. A kit that silently rewrote a byte would be
worth less than one that never corrected anything at all.

## The contamination caveat

Published 2026-08-13. Models with a later training cutoff may have seen these
sets. We author fresh sets each cycle; this one is not a standard, it is our kit,
yours to reuse.
