# Counting rules — exhibit thirty-three

Five rules govern every figure in this kit and on the page it belongs to
(https://research.strata2signal.com/the-same-sixteen/). They are collected here so
a reader does not have to reconstruct them from prose, and each names where it was
registered.

## 1. A reference arm is ranked against nothing

`glm-5.3-flash:cloud` is not a candidate for any chair, seat or product decision
this house makes, and no floor is applied to it. Its row is printed **beside** the
candidate rows, never among them, and carries no ordering, no margin and no
"beats"/"loses to" language. The scorer assigns outcome states mechanically and
stamped the toolbench row `RANKED`; **the class label overrides it**. Registered
before the run in `prereg-reference-arm.md` (operator ruling D-20260828-18), fence
2 of 4. Every published cell for this arm is labelled
**REFERENCE ARM · cloud · dated**.

## 2. The reading is dated, and only dated

A hosted model behind a tag is not a fixed object: the weights, the router and the
serving stack behind `glm-5.3-flash:cloud` can change under the same spelling with
no notice and no digest anyone outside the vendor can pin. So this is a statement
about **2026-08-28, 16:37:59Z–16:49:43Z**, stamped at both ends, and it is never
carried forward as "the ceiling" after that day. `prereg-reference-arm.md`, fence 3.

## 3. Two legs read; the other seven are refused with a reason, never left blank

Only the toolbench (C3) and the field exam could honestly be read on this path.
Four legs are NOT-RUN because the path invalidates them — two bind on a JSON-schema
grammar the cloud path accepts and silently ignores, and two are speed and
long-context readings that mean nothing measured over the public internet against
someone else's fleet. Three never applied to a model that cannot hold a seat. Each
disposition and its reason is in `reference-arm-row.json` and `RESULTS-TABLES.md`,
and the evidence for the two `format` exclusions is in
`REFERENCE-glm-5.3-flash-2026-08-28.gate.json` — a probe taken **after** the last
scored call, so the receipt covers the whole scored window.

## 4. `think:false` is a measurement failure on this path, not a low score

Every leg here scores `message.content`. On this cloud path `think:false` does not
switch the model's reasoning off; it only stops ollama parsing the reasoning out of
the reply, so the raw deliberation lands in the field the scorer reads. Four probes,
four leaks — three closed by a literal `</think>`, one not delimited at all. So
`think:false` was **never run** and every figure here is `think:true`. The frozen
pre-probe (`raw/c3/glm-5.3-flash_cloud/preprobe.json`) captured the same leak
independently of the gate receipt, which is why it is two findings rather than one.

## 5. A count is a floor where a literal-string checker touches it

Two of the nineteen toolbench tasks are scored by checkers that match string
literals and therefore **fail correct answers**: `t15-honesty-draught` accepts only
the word "depth" and rejected *"a measure of how deep she sits in the water"*;
`t04-single-clock-now` wants an ISO date and rejected *"09:00 on Wednesday,
12 August 2026"*. The seat and the arm fail both, for the same lexical reason. A
checker that can only ever reject a correct answer pushes a row down and never up,
so **every row those two tasks touch is a floor, not a reading** — the arm's 16/19
included. The frozen numbers stay as scored: re-scoring a frozen instrument to
flatter a row is how benches stop meaning anything. Widening the checkers is a
pre-registered revision of `tasks.json`, applied to every row equally, and it is
not applied here.

## And the resolution of the ruler itself

Against a 16/19 row this bench separates a comparator only at a gap of seven tasks
or more (Fisher exact, α = 0.05: 16-vs-9 separates; 16-vs-10 does not). Wilson 95%
intervals are printed beside every rate; a Wilson interval is a function of the
count and the item total and nothing else, so two rows on the same count carry the
same interval by construction. Nineteen items buy a span thirty-two points wide.
The field exam is saturated at this level — every local model this house has sat on
it scored 39/40 or better on the comparable forty — so no gap on it is measurable
either. Both facts are findings about the instruments, not about any model on them.
