# The counting rules — read these before you compare anything

Four rules govern every number in this kit and every number on the exhibit page
it belongs to. They were pre-registered before scoring, not chosen after it. If
you re-derive a figure under a different rule you will get a different answer,
and the difference will be the rule, not the measurement.

---

## 1. A count travels across cards. A rate does not.

This is the rule that makes the whole comparison legal, and it is the reason
some numbers in this kit are compared to an earlier run and some never are.

- A **count** — kill-recall, preservation, task success, items right, recall,
  abstentions, fabrications — is a property of **the model**. It is comparable
  across cards. The same frozen items, the same seeds, the same scorer, run on
  different silicon, should produce the same count; when they do not, that is a
  finding about the model or its build, not about the hardware.
- A **rate** — decode tokens a second — is a property of **the model AND the
  silicon**. It is not comparable across cards. A slower figure here than one
  measured on a larger card is expected and says nothing about the model.

So: the correctness counts in this kit may be read against the 2026-08-12 run on
a 96 GB workstation card. **The `c5-recovered.json` decode rates may not.** They
are new measurements of new hardware and are never published as a delta.

The corollary the bench holds itself to: a repro row tests *reproducibility of
the correctness verdict*, and nothing else.

## 2. Past a 10 % measurement-failure rate, an exam is UNMEASURABLE.

Every exam carries a pre-registered `response_failure_ceiling` of **0.1**. When
more than one call in ten fails to produce a scorable response — an unparseable
body, a broken output contract, a transport error — the exam refuses to convert
what remains into a verdict.

**What a truncation counts as depends on WHICH exam you are reading, and the two
exams in this kit differ.** (Corrected 2026-08-29: this rule previously listed a
truncation against the reply budget among the failures that count toward the
ceiling, flat, for every exam. That is the seat exam's rule, not this battery's,
and stating it as a single rule contradicted both the pre-registration and the
scorer files shipped here.)

- **This battery's own judge trial (C1) puts a truncation in its own run-quality
  bucket, OUTSIDE the ceiling, and reduces the failure denominator by it.** The
  clean qwen row in `summary-C1-qwen3.6_27b.json` reads `calls: 63, ok: 60,
  length_truncated: 3, response_failures: 0, response_failure_denominator: 60,
  response_failure_rate: 0.0` — three calls hit the 1,024-token reply budget, and
  the exam scored the other sixty rather than counting the three against a
  ceiling they never tested.
- **The judge-seat exam's own frozen rule counts a truncation as a response
  failure, INSIDE the ceiling.** That exam is an older, separately frozen
  instrument (rule 4), quoted as published rather than re-tuned, and its
  `response_failures` figures in `seat43-summary.json` are what carry two
  candidates past the ceiling.

Where the two exams treat the same failure differently, the difference is the
exam, not the model. The exhibit page states this in the same terms.

Two things follow from the ceiling itself, and both are visible in this kit:

- **The judge-seat exam refuses the counts outright.** `seat43-summary.json`
  records `muse-glimmer:30b` at 18/129 (14 %) and `qwen3.6:27b` at 93/129 (72 %)
  and marks both UNMEASURABLE. The underlying kill/preservation counts exist in
  that same file; they are not published as results anywhere, because the exam
  refused them.
- **The tools trial prints them fenced as inadmissible.** `c3-scores.json` marks
  `muse-glimmer:30b` UNMEASURABLE on **6 of 38 calls failing (15.8 %)** — the
  denominator is *calls*, not tasks. Its task score of 13/19 is printed fenced,
  and all 19 tasks are fenced, not 13 of them.

A verdict of UNMEASURABLE is a statement about the measurement, never about the
model.

**A note on the failure denominator.** Where a leg reduces its denominator for
truncated calls, the scorer says so in the same object. The judge trial's clean
qwen row reads `calls: 63, ok: 60, length_truncated: 3, response_failures: 0,
response_failure_denominator: 60` — sixty calls scored clean, three hit the
1,024-token reply budget and were excluded from the failure denominator, and
none failed the response contract. All four numbers are in
`summary-C1-qwen3.6_27b.json`; read them together.

## 3. A difference under 2 items reads TIED.

Pre-registered before scoring: on any frozen item set in this battery, a
cross-run difference of **fewer than 2 items** is read as TIED, not as a move.
Preservation going 9/9 → 8/9 between the two runs is one item, and is TIED.
Task success going 16/19 → 14/19 is two items, and is **not** inside the band —
it is printed as a real move.

The floors themselves are separate from the tie band and bind independently:

| exam | kill-recall floor | preservation floor |
|---|---|---|
| judge trial (C1), 21 items | **≥ 11 of 12** | **≥ 8 of 9** |
| judge-seat exam, 43 cases × 3 calls | **≥ 23 of 27** | **≥ 13 of 16** |

Both floors must clear. A judge that fails everything aces kill-recall and
flunks preservation, which is exactly why preservation binds too.

## 4. One exam ran at its own settings, and they are not the battery's settings.

**Stated plainly, because it changes how a third of the standings grid reads.**

This battery's house laws are a **32,768-token context** everywhere and an
**explicit `think` boolean** on every call. The judge seat's 43-case exam is an
**older, separately frozen instrument** and its own pre-registration fixes
**`num_ctx` 16384** and **`think` omitted entirely**.

The operator ruled that the seat exam **runs at its own published settings**, and
that the deviation is disclosed rather than eliminated: comparability against
that exam's published floors of 23/27 and 13/16 is its entire value, and a run at
32,768 with an explicit `think` would produce numbers that *look* like the
published ones while being a different instrument. A frozen external exam is
quoted, not rewritten.

So exactly one instrument in this battery runs outside the house laws, and it is
that one. Every other leg holds 32,768 and an explicit boolean.

**The receipt is unaltered and in this kit.** Under `raw/seat43/`, every one of
the 390 harness records that carries an options block reads:

    "options": {"temperature": 0.0, "top_p": 1.0, "num_predict": 1024, "num_ctx": 16384}

and each file's `run-start` record reads `"think": "omit"` and `"num_ctx": 16384`.
The string `32768` appears nowhere under `raw/seat43/`. Check it yourself:

    grep -c '"num_ctx": 16384' raw/seat43/*.jsonl
    grep -c '"think": "omit"'  raw/seat43/*.jsonl
    grep -c '32768'            raw/seat43/*.jsonl      # → 0, three times

These wire options were deliberately left **visible and unaltered** when this kit
was sanitised. Rewriting them to match the battery's house laws would have
destroyed the one thing that makes the disclosure checkable.

---

## What the exams' outcome words mean

| word | meaning |
|---|---|
| **RANKED** | a clean, comparable run |
| **EXPLORATORY** | real numbers, fenced from ranking |
| **DESCRIPTIVE** | numbers reported, no registered floor to pass or fail |
| **UNMEASURABLE** | valid attempts exist, but the measurement-failure rate crossed the 10 % ceiling (rule 2) |
| **NOT-CARRIED** | zero valid attempts came back — nothing to score |
| **NOT-RUN** | never attempted, with the reason named |

A cell never pretends, and there are **four** reasons a cell carries no score —
the tables keep them apart rather than printing one blank for all of them:

- **`NOT-CARRIED`** — valid attempts came back zero; there is nothing to score.
- **`NOT-RUN (spill)`** — the model did not fit the card whole, and this bench
  does not run speed or long-context through a half-CPU model. A rule, not a
  judgment call.
- **`NOT-RUN (stop)`** — the disk halt fired before the leg was reached.
- **`NOT-RUN (trim)`** — the operator's roster time-box.

**One cell in this battery is none of the four**: laguna-xs-2.1:latest's judge
trial was **cut mid-leg**. Its 19 calls landed clean across 7 of the 21 items
before the stop cut the leg; they ship in this kit and they are not a result, and
the cell says `cut mid-leg` rather than pretending the leg never started.

A row's first empty reason propagates rightward.

*(Corrected 2026-08-29: this section listed three NOT-RUN reasons and did not
name `NOT-CARRIED` as a fourth empty-cell reason or the `cut mid-leg` cell at
all, where the exhibit page has carried all five words since draft v5. The kit
now says what the page says.)*
