{
  "schema": "s2s-bench-v1",
  "exhibit": "chair-trials",
  "published_utc": "2026-08-13",
  "status": "PUBLISHED — live at research.strata2signal.com.",
  "licence": "CC BY 4.0",
  "attribution": "strata→signal research, research.strata2signal.com",
  "hardware": "One 96G VRAM workstation, which also answered live requests for three of our apps throughout these trials — two public, one serving playtests for humans and agents alike. Contention is measured in both directions and published with the figures rather than assumed away.",
  "dataset": "chair-trials-counting-rules",
  "what_this_is": "Every counting rule on the page, in one file. Each one is copied out of the leg companion that uses it rather than restated here, so a rule cannot drift between the two places it appears.",
  "shared_vocabulary": {
    "RANKED": "this row's numbers were produced under the contract and are admissible — admissible, NOT placed above another row. The descriptive legs print this same word for this same state; their scorers emit it under the label DESCRIPTIVE and the pre-registrations say the two are one state with one meaning.",
    "EXPLORATORY": "a protocol fence tripped — the contract probe failed. The row ran and is published, unranked.",
    "UNMEASURABLE": "response failures above the registered ceiling. The measured cells still print; the score is not admissible as a score.",
    "NOT-CARRIED": "zero valid attempts under the contract. The raw emission is published beside the row, because 'could not use the tools' and 'could not do the work' are different findings.",
    "NOT-RUN": "never run, and the reason always rides the row. Never a blank.",
    "TIED": "differences under two items read TIED. Nothing on the page orders by margin.",
    "counts_of_n": "Every figure is a count against its denominator. No percentage is published under N = 30; intervals ride beside counts as labelled companions and gate nothing.",
    "strict_parse": "The strict parse binds every verdict. A leak-stripped or fence-stripped parse is a diagnostic column only and never moves a state.",
    "em_dash": "An em dash is a figure we do not hold. It is never a zero, and a blank cell means the column does not apply to that row."
  },
  "legs": [
    {
      "leg": "the roster and residency census",
      "companion": "roster.json",
      "verdict_discipline": null,
      "primary_metric": null,
      "counting_rules": {
        "residency": "size_vram is read verbatim from ollama's /api/ps for the named tag, in bytes, immediately after a load that generated exactly 1 token at the named num_ctx, with the model still resident and no other candidate resident. GB figures anywhere downstream are size_vram / 1073741824 (GiB, 1024^3), rounded to 2 dp, and the byte count always travels with them. A row is one (tag, num_ctx) pair: size_vram = weights + KV cache, and the KV cache is a function of num_ctx, so a figure quoted without its ctx is not a figure. Protected residents (the live seat, the live embedder) are READ, never loaded: their rows carry the ctx the daemon reports, which is whatever production last set, not a value we chose.",
        "row_is_an_arm": "One row is one ARM — a model tag as these legs ran it. Three of the eleven are the same weights under different serving choices (two quantizations, one served without its draft model), which is why the roster has eleven rows and the arc has nine models.",
        "ctx": "Every figure here is at num_ctx 32768, the length every scored C-leg ran at. size_vram is weights plus KV cache and the KV cache is a function of the context length, so a residency figure without its context length is not a figure. The census also holds rows at 16 384 for the four arms that sat the frozen seat exam; those belong to the previous exhibit and are published there.",
        "postures": "`postures_run` is derived from the legs' own scored artifacts — the postures each leg actually recorded for that arm — never declared by hand."
      },
      "limits": [
        "size_vram is what this daemon allocated under this configuration — not what a model 'needs', and not what another runtime would allocate.",
        "The three drafter rows carry an open finding raised by the previous exhibit's census and unresolved at publication: a model served with a draft model beside it reports a tenth of what the same weights report without one, and two different quantizations of it report a byte-identical allocation. Read that finding before quoting those three residency figures anywhere.",
        "Quantization differs across rows and is printed on every one. Comparing across quantizations is a confound we name, not one we equalize."
      ]
    },
    {
      "leg": "chair one — the judge",
      "companion": "c1-judge.json",
      "verdict_discipline": "RANKED — against floors ONLY, no ordering by margin",
      "primary_metric": null,
      "counting_rules": {
        "floors": "Both floors bind independently and were derived by the published rule from exhibit two's ratios, rounded up on this population: kill ≥ ceil(12 × 0.852) = 11 of 12, preservation ≥ ceil(9 × 0.8125) = 8 of 9. Pass or fail only — two rows that both cleared are TIED whatever the margin, and nothing on this page orders models by margin.",
        "aggregation": "Three repeats per case at temperature zero; each repeat maps to catch/preserve FIRST, then strict majority (the registered map-then-majority convention). A case whose repeats never produced a verdict stays in its denominator rather than being discounted.",
        "binding_parse": "STRICT parse binds every verdict; the leak-stripped parse is a diagnostic column only (PLAN v2 line 17, PREREG-C1 §Protocol)",
        "wilson": "Wilson 95% intervals ride beside every rate as labelled companions and gate nothing. Counts bind; every denominator in this leg is under 30, so no rate is published as a bare percentage.",
        "outcome_states": "RANKED = the contract probe carried and response failures are under the 10% ceiling. EXPLORATORY = a protocol fence tripped: the row ran and is published, unranked. UNMEASURABLE = response failures above the ceiling. NOT-CARRIED = zero valid attempts, raw emission published beside the row.",
        "response_failure_denominator": "ok + response_failure. Length-truncated and transport failures sit OUTSIDE the ceiling by pre-registration, which is why one row shows 0/29 beside 63 calls."
      },
      "limits": [
        "A synthetic entailment set over public-domain text is a weaker instrument than exhibit two's human-verified field key, and it is audit-ready precisely because every case ships whole in this kit.",
        "The floors are inherited ratios applied to a new population, not a new calibration.",
        "One roster (nine models, ten rows), one runtime version, one day."
      ]
    },
    {
      "leg": "chair two — the assistant",
      "companion": "c2-assistant.json",
      "verdict_discipline": "DESCRIPTIVE - no floors, no winner, no ordering by margin (PLAN v2)",
      "primary_metric": "conjunctive pass over 2 repeats, out of 20 frozen items",
      "counting_rules": {
        "conjunctive": "Two repeats per item at temperature zero; an item scores 1 only if BOTH repeats pass. The disagreement column is the count of items whose two repeats disagreed — published rather than smoothed.",
        "two_halves_never_added": "Thirteen items are checked EXACTLY (a deterministic predicate over the answer: a sentence count, a parsed field, a number within a tolerance). Seven are checked by PROXY (marker- and length-based heuristics). The /20 is the registered CONJUNCTIVE ITEM COUNT and the two subtotals are its composition, not two scores summed: what the registration bars is blending them into one QUALITY figure, because a proxy check is evidence about shape and an exact check is evidence about correctness, and averaging those two would claim a precision neither has. The subtotals are published side by side so a reader can weight them differently than we did.",
        "no_floors": "Descriptive: no floors, no winner, no ordering by margin. Differences under two items read TIED by pre-registration.",
        "unmeasurable": "A row whose response-failure rate exceeds 10% of its calls is UNMEASURABLE and its score is not admissible as a score, however the arithmetic came out.",
        "long_context": "The two long-context items vary the PROMPT length (8k and 24k tokens) while num_ctx stays 32 768 everywhere. Observed prompt_eval_count rides on every repeat so a reader can check the tier label against what the model actually read."
      },
      "limits": [
        "Twenty items is a probe, not a suite: a single item's pass or fail moves a row by a twentieth, which is why the tie rule is two items wide and the item matrix ships.",
        "Proxy checkers measure shape. Seven of the twenty scores are therefore weaker evidence than the other thirteen, and the two subtotals are published apart so a reader can weight them differently than we did.",
        "Thinking postures were run for two arms only; every other row is think:false, and the stopwatch leg carries the thinking story for the rest."
      ]
    },
    {
      "leg": "chair three — the toolbench",
      "companion": "c3-toolbench.json",
      "verdict_discipline": "DESCRIPTIVE",
      "primary_metric": "task success, conjunctive over 2 repeats, out of 19",
      "counting_rules": {
        "task_success": "Conjunctive over two repeats: a task scores only if BOTH attempts pass its registered predicate. 19 tasks, so the denominator is 19 and never a percentage.",
        "grounded": "task_success_grounded counts tasks that passed AND made every pre-registered call. The gap between the two fields is the tasks a model got right without using the tool it was supposed to use — right answer, wrong method, and the two fields exist so that difference cannot hide.",
        "tool_selection": "Of the 14 tasks with a single unambiguous expected first tool, how many the model chose correctly.",
        "argument_fidelity": "Of the 66 registered argument predicates across the suite, how many the model's actual call arguments satisfied.",
        "spurious_calls": "Calls the task did not need, over ALL tool calls the arm made across all 19 tasks and both repeats. The denominator is tool calls, not tasks, so a chatty arm and a quiet one are compared on the same axis. Wilson 95% intervals ride beside the rate.",
        "honesty_traps": "Five tasks whose honest answer needs no tool at all. A pass means the model declined to invent a figure; reaching for a tool is not itself a failure, asserting an unfounded answer is.",
        "outcome_states": "RANKED = the pre-probe carried and response failures are under the 10% ceiling. UNMEASURABLE = response failures above it — the numbers stand in the table and are not admissible as a score. NOT-CARRIED = the pre-probe produced no parseable tool_calls in any posture; the row's task attempts are published anyway, with the raw emission, because 'could not use the tools' and 'could not do the work' are different findings.",
        "preprobe_economics": "Two calls per arm per posture before any task ran. The REGISTERED budget is four — two postures — and the MEASURED cost of this arc's one NOT-CARRIED finding was TWO, because that arm registered one posture. Two calls and an honest row, instead of eighty scored calls and a table of zeroes that reads like a capability finding.",
        "tie_language": "Differences under two tasks read TIED (PREREG-C3). On task success that puts six arms in one top band, 16 to 15 of 19, and no ordering inside the band is claimable — which is why no row in this table carries the gold highlight."
      },
      "limits": [
        "Nineteen tasks over one fixture corpus. A tool suite is a shape, not a census of agentic work, and the tasks ship whole so you can see which shape.",
        "The web tool is a local rate-limited fixture endpoint, not the internet.",
        "Spurious-call rate counts calls, not intent: a model exploring the corpus and a model flailing look the same to a counter, which is why the traces carry the calls.",
        "NOT-CARRIED is a statement about this runtime and this tag on this day. It is not a claim that the model cannot call tools anywhere."
      ]
    },
    {
      "leg": "chair four — the narrator",
      "companion": "c4-narrator.json",
      "verdict_discipline": "DESCRIPTIVE",
      "primary_metric": "overall pairwise win-rate, cluster-bootstrap CI over items",
      "counting_rules": {
        "pairwise": "Sample-matched pairings across 15 arm pairs; every comparison judged in BOTH positions by the same judge, and a swap-discordant verdict scores 0.5/0.5 rather than being resolved. Forced choice on one dimension per verdict — overall — with in-character and warmth tallied as secondaries inside the same verdict object.",
        "intervals": "Cluster bootstrap over ITEMS (n = 9), because two samples of the same question are not independent evidence. The effective item count is printed beside every rate — and it is not 9 everywhere: the gate-passing view of the thinking-on Glimmer arm rests on 3 item clusters, which is why its interval is nearly the whole unit line and why the number of clusters travels with the rate. A Wilson interval over 360 comparisons would be a much prettier and much more wrong number.",
        "tie_threshold": "The effective N is 9 items, so a difference smaller than 2 items of 9 reads TIED on the primary: |Δ win-rate| < 2/9 = 0.2222. Distinctness: differences under 2 tasks. Canon: differences under 2 replies. Registered before the first judging batch was built, and printed on the page, because a threshold a reader cannot see is a threshold a reader cannot apply.",
        "gate_passing": "The same comparisons restricted to replies that passed the engine's envelope gate — the JSON contract a reply has to hold for our engine to use it at all. An arm with no gate-passing replies has no gate-passing rate: an em dash and a reason, never a zero.",
        "distinctness": "Three replies by the same arm, one per character, shown together and unlabelled: which character said which? Chance is 1/3. Accuracy over 324 scored assignments with a Wilson interval per arm and, on the headline, a cluster-bootstrap interval over tasks beside the Wilson one; the questions are withheld from the judges so topic cannot substitute for register. The registered per-arm denominator is 72 assignments (four judges) and every arm scored all of them; an earlier draft reported a four-batch shortfall — those batches had been misfiled, not lost, were recovered before scoring, and `unscored_batches` is empty. The cluster bootstrap exists on the aggregate only: the scorer emits it there and this file does not compute statistics the scorer did not.",
        "canon": "Three adversarial items, one per character, each carrying a fabricated premise no canon supports. Every reply judged into one of three categories; the reply's majority is published and a tie is its own outcome, never rounded. Counts, not rates — the population is 36 replies. The registered per-arm denominator is 24 judgements (six replies × four judges) and 18 were scored, by the same three-judge panel; both numbers print.",
        "in_voice": "A separate judgement inside the canon round: did the character speak at all? A reply that is empty, or that is raw JSON, or that refuses out of character, is not in voice however clean its canon."
      },
      "limits": [
        "Judges are LLMs from two families, one of which wrote the questions, the canon notes and this article. The per-family split and the inter-panel agreement are published rather than averaged away.",
        "Nine standard items and three adversarial ones across three characters: a register probe, not a campaign, and not a canon audit.",
        "The distinctness judges also judged these same replies in the pairwise round.",
        "The canon judges see the canon note, so that round measures adjudication against a STATED canon, not a judge's own knowledge of the world.",
        "First-attempt only: production's corrective-retry ladder is not simulated, and it would recover several of the envelope failures counted here."
      ]
    },
    {
      "leg": "chair five — the stopwatch",
      "companion": "c5-stopwatch.json",
      "verdict_discipline": "DESCRIPTIVE — no floors, no winner, no ordering by margin",
      "primary_metric": "decode tok/s at the 1k tier = eval_count / (eval_duration / 1e9), per call, from ollama's own counters",
      "counting_rules": {
        "decode": "decode tok/s at the 1k tier = eval_count / (eval_duration / 1e9), per call, from ollama's own counters",
        "ttft_proxy": "ttft_proxy_ms = total_duration - eval_duration. stream=False, so there is no first-token event: this is an UPPER BOUND, not TTFT, and is never placed beside a vendor's streamed figure",
        "cold_load": "Labelled warm-page-cache/cold-VRAM: no sudo means no page-cache drop, so these are cold-VRAM loads over a warm file cache and are not comparable to a vendor's cold-start figure or to a reboot. First-load-after-pull is excluded.",
        "statistics": "min / median / max / n only. No p95 anywhere in this leg: the largest repeat set is 10, and a p95 over ten samples is the second-largest sample wearing a statistics hat",
        "reading_min_median_max": "On a lightly-shared box the MAX approximates the uncontended ceiling, the median is real life, and a dragged min is the moment a live user shared the card. That is why all three are printed and no p95 appears anywhere: the largest repeat set in this leg is 10.",
        "tie_rule": "two arms read TIED on decode tok/s when their repeat-set min-max ranges overlap; on the DFlash count, differences under 2 prompts read TIED",
        "outcome_rule": "c5_speed.arm_outcome — the same function the live runner uses, so a rebuilt state and a run state cannot diverge",
        "rule_4_null_counters": "A call that returns HTTP 200 with every counter and done_reason null is excluded from every statistic. It is not a slow call and it is not a failure: it is a call we cannot time, and its cause is unexplained — transport_error is null and nothing in the run logs accounts for it.",
        "rule_5_response_failures": "A call that returns valid counters with EMPTY response text and done_reason=length is a response failure. Three arms exceed the registered 10% ceiling on this rule and are UNMEASURABLE; their measured cells still print, because a state is not a reason to hide the numbers that produced it.",
        "re_blocking": "Four arms ran a block twice — the original pass and tonight's completion run. Both values publish. Where both occurrences measured nothing, the row says so rather than reporting one absence."
      },
      "limits": [
        "One box, one runtime version, one day, one batch configuration. The runner's argv and batch size were captured for all eleven tags and matched none of them, so the registered `-b 1024 / -c 32768` bracket has no captured value and is not published.",
        "Quantizations differ per row and are printed on the roster. Cross-quantization speed comparison is a confound we name, not one we correct for.",
        "TTFT-proxy is an upper bound, not TTFT: these calls were not streamed.",
        "The GIN-overlap column is UNMEASURED, not clean — see the contention disclosure.",
        "Two arms' cells are absent for a reason that is itself unexplained: calls that returned text with every counter null. We publish the absence and the count rather than a figure derived from the calls that happened to work."
      ]
    },
    {
      "leg": "the serving probes",
      "companion": "c6-serving.json",
      "verdict_discipline": "DESCRIPTIVE — no floors, no winner, no ordering by margin",
      "primary_metric": null,
      "counting_rules": {
        "wall_s": "POST leaving this process to the ruling page's bytes in hand, including the app's queue. It INCLUDES the app's queue, deliberately: when ten people ask inside a hundred seconds and the box answers one at a time, nine of them wait, and a figure that hid the queue would be measuring the app's politeness.",
        "poll_interval": "Every transition is observed at the app's own fastest self-refresh cadence, 1 s, and is accurate to within one interval. None of these is a server timestamp, and no figure here can resolve anything faster than the poll.",
        "queue_and_generation": "queue_wait_s = send → the first poll that shows the ask being worked on. generation_s = that poll → the ruling. Both quantised to the poll interval.",
        "second_arrival_penalty": "The later completion's wall_s minus the earlier completion's, both measured from their own POST. The two asks of a trial are the same book at different player counts, so the difference is about arrival order rather than rulebook length.",
        "overlap": "overlap_candidate_generating is computed from TIMESTAMPS, never from intent: the ask's window intersected with the co-resident generation's window, or it did not. An ask in the generating regime whose overlap is false publishes as false, and five of them did.",
        "asks_ahead": "`max_ahead_observed` — the largest queue depth the APP ITSELF displayed to that ask, scraped from its wait page at the one-second poll. It is an observation, not a computed concurrency: the ask currently generating is the holder, not 'ahead', so on a saturated queue this column reads one lower than the number of asks actually in flight when the ask was sent. The burst's fifth ask is the worked example — the column says 3 and four earlier asks were still unfinished at its send. Both facts are in this file: `max_ahead_observed` per record, and `send_epoch` / `end_epoch` on every record so the in-flight count can be recomputed. An em dash means no poll ever showed a queue position.",
        "regime_outcome": "Two regimes are TIED on wall_s when their min–max ranges overlap. All three drip regimes overlap all three ways, so the outcome column reads TIED on every row — including the row whose median is the lowest, which is exactly the row a table-only reader would otherwise screenshot as a finding.",
        "content_rule": "asks are scored on CONTENT (a registered setup-shaped phrase in the served ruling), and the content check is recorded, never a gate — an honest abstain is a served answer",
        "statistics": "min / median / max / n only, per probe and per regime. No p95 anywhere — the largest set in this leg is ten. Counts of N everywhere; no rate in this leg is published as a bare percentage.",
        "tie_rule": "Two regimes read TIED on wall_s when their min–max ranges overlap. All three drip regimes overlap, and the page says so rather than reading a median difference as a finding.",
        "no_checkpoints_in_flight": "no canary and no checkpoint call is made while asks are in flight; a checkpoint fired mid-probe would be contention we caused"
      },
      "limits": [
        "The poll interval is a floor. Nothing in this leg resolves faster than one second, and every transition is quantised to it.",
        "Game is a confound ACROSS probes: each probe uses different books, so a drip wall and a cold wall are not comparable. The three paired quiet baselines are the only within-item comparison in the leg.",
        "The regimes do not separate. All three drip regimes' wall ranges overlap, which by the pre-registered tie rule reads TIED — a finding, and NOT evidence that co-residency is free at every scale.",
        "The co-resident's generation lasted about five seconds, not the fifty the regime spans, so 'actively generating' is a weaker treatment than the design intended. Published as measured, with the overlap flag on every ask.",
        "The cold probe measured a cold QUEUE, not a cold seat: the seat was still resident at both windows' close, because production holds it. A genuinely cold seat would cost a model load on top of everything printed here.",
        "Two trials is not a distribution. The cold rows are two walls with their receipts, and they are printed as two walls."
      ]
    }
  ],
  "contamination_caveat": "Published 2026-08-13. Models with a later training cutoff may have seen these sets. We author fresh sets each cycle; this one is not a standard, it is our kit, yours to reuse.",
  "not_a_standard": "Our own trials, our own hardware, for our own chairs. Not first independent numbers, and not a benchmark."
}
