# The Instrument Travels

*A rules question comes in — a table of friends, mid-game, one of them typing 'can I build on a card I already played this turn?' — and a model we picked by exam answers it. Two weeks ago we picked on a 96 GB workstation card. This time we ran the very same frozen exam — same items, same seeds, same pass marks — on a 24 GB consumer card, to find out whether it would say the same thing in a smaller room. Five models sat it: two we had examined before, re-pulled at the tags as they stood on 2026-08-26; two newcomers; one we had blocked and wanted to re-check. One reproduced its published verdict on a different build and different silicon, which is the whole point of freezing an exam. The other four returned four different verdicts — a seat blocked by a stray byte, a judge that caught every planted error and failed anyway, two models that would not board the card, and a row that owes a re-run. Nothing earned a seat, and the bench added a fifth result by stopping itself at a disk floor rather than gambling on the margin.*

*Published 2026-08-29 (UTC) · a small (human) team and a fleet of AI agents — we set the exams, ruled the calls and signed the numbers; the harness ran the legs itself · main measurement window 2026-08-26, 07:00–15:13Z (UTC): the main run stopped at 09:58Z, a gap-fill lane ran on to 10:14Z when the trigger fired again, and scoring resumed at 13:55Z*

**the short version:** Carry a frozen exam from a 96 GB workstation card to a 24 GB consumer one and the model it ranked two weeks ago ranks again — at the floor, on a different build, with nothing new seated.

4,490 words · about 20 minutes (at 220 words/min) · 5 tables · data kit: yes

https://research.strata2signal.com/the-instrument-travels/

---

## The chair, and who gets to sit in it {#the-chair-and-who-gets-to-sit-in-it}

Somewhere a rules question comes in — a table of friends, mid-game, one of them typing "can
I build on a card I already played this turn?" — and the answer comes from a model sitting
in a chair we gave it after an exam. Every few weeks the chairs get contested: a new model
ships, a familiar one gets a new build under the same name, and someone asks whether the
newcomer would answer that question better. The only fair way to find out is to make it sit
the *same* exam the incumbent sat, and to trust nothing the exam did not see.

The question is never "is it good?" — it is "does it earn a chair, on hardware a normal
person owns, under the same exams everything else passed?" What is different this time is
the room. Two weeks ago we published [our ranked verdicts, measured on a 96 GB workstation
card](/chair-trials/); this battery ran on a 24 GB consumer card of the kind sitting in a
gaming PC today. So it asked two questions at once: do any newcomers earn a chair — and when
models we examined before are re-pulled and re-examined in a smaller room, does the frozen
instrument say the same thing? If it does not, every ranking we publish is a photograph of a
moment. If it does, the instrument travels.

## What we ran, and the rule behind every number {#what-we-ran-and-the-rule-behind-every-number}

**Five candidates on a single 24 GB RTX 3090, one at a time, that card emptied for the
run:**

- **two repro rows** — qwen3.6:27b and nemotron-3.5-lightning:30b, examined on 2026-08-12
  and re-pulled at the tags as they stood on 2026-08-26. Their earlier verdicts were not the
  same: qwen cleared both judge floors and was RANKED, nemotron missed both (9 of 12 kills,
  6 of 9 preservations) and its toolbench did not carry. Only one is a new build — qwen's
  manifest digest changed since 2026-08-12, nemotron's is byte-identical, so a nemotron row
  is a harness-and-hardware check, not a model comparison.
- **two new kids** — gemma4:31b and laguna-xs-2.1:latest.
- **one re-probe** — muse-glimmer:30b, first probed two weeks ago and seat-blocked then. The
  earlier chair trials ranked glimmer's Q8 build; this battery pulled the default tag.

The box's second card kept serving live traffic throughout — a separate card with its own
memory, and the checking is in the kit (the receipts published beside this page) rather than
in our word for it: the lane log records a read-only probe of all three live seats at nine
checkpoints, all five fit receipts carry `prod_untouched: true`, and the teardown receipt
carries the last probe of all.

Candidates run **as shipped**: the default published tag, the vendor's own quantisation — a
fit verdict is a verdict on that shipped artifact, not on the architecture. Every count
below is copied from that leg's own scorer file — a *leg* is one exam run for one model —
and the page computes nothing. **The fit gate loads each model at a 32,768-token context**:
residency is a fact about weights *plus* a working context, and 32k is what our seats run.

One rule governs every comparison here. A **count** — kill-recall, preservation, task
success, items right — is a property of the model, and travels across cards. A **rate** —
tokens a second — is a property of the model *and* the silicon, and does not. So the counts
below are compared to the 96 GB run and the speeds are not, ever.

**The run itself is part of the data, told plainly.** The battery **stopped itself** at
09:58Z, mid-leg, when free space on the bench box's disk fell below the harness's 60 GiB
safety floor. The stop was not idle time, and the page has to say so because three published
cells were measured inside it: with the battery's own units standing down, a gap-fill lane
re-ran the re-probe's outstanding legs — glimmer's schema probe, assistant trial and
toolbench, finishing at 10:00Z, 10:06Z and 10:14Z, and those three runs are the glimmer
cells below. One second later, at 10:14:59Z, the speed leg tried to start, the trigger fired
again, and that was the end of it.

We cleared the disk, restarted the bench process (the lane log records the new process ID,
so a reproducibility claim spanning the halt spans a restart too), re-verified all five
weight digests — 5 of 5 MATCH — resumed at 13:55Z, and time-boxed the remainder: qwen ran to
completion, nemotron ran its fit gate only, which turns out to be the only result it needed,
and we dropped laguna's remaining legs. When the battery sealed we tore its model store down
to 16 KB under four guards: one host and never production, results verified present before
any delete, named tags only and no wildcards, before/after receipts. Weights are re-pullable
by design; measurements are not. That teardown ran at 18:13Z, after the window closed —
bookkeeping rather than measurement.

## The words these tables use, and the floors they score against {#the-words-these-tables-use-and-the-floors-they-score-against}

- **A seat, which we also call a chair** — a model slot our products actually serve from:
  RuleSage, our rules assistant, and its sibling games and tools. Every seat here runs on
  Ollama 0.32.13 at `think:false`. A seat is the prize; the things a model sits to win one
  are exams.
- **Posture** — whether the model's hidden reasoning channel is off (`think:false`, the seat
  contract) or on (`think:true`). "Held both ways" means the probe was run in both.
- **The exams** — a frozen battery: a schema-discipline probe (C0), a judge trial with
  pre-registered floors (C1), an assistant trial (C2), a tool-use trial (C3), speed (C5), a
  long-context filing test (C7), a 60-item field exam in three 20-item tasks, and the judge
  seat's own trial (43 cases × 3 calls = 129). The C-numbers are the harness's own labels,
  kept because they are the filenames in the kit; the numbering has gaps because the harness
  is older than this roster, and C4, the narrator's chair, is excluded by operator ruling.
  Frozen means same items, same seeds, same floors, same reply budgets as the published
  runs.
- **Kill-recall and preservation** — the two counts a judge trial scores. We plant known
  errors in the material, and **kill-recall** is how many the model caught. We also leave
  correct material alone in there, and **preservation** is how much of that it left
  standing. A judge that flags everything scores a perfect kill-recall and a terrible
  preservation, which is exactly why both floors have to clear.
- **RANKED** — a clean, comparable run.
- **EXPLORATORY** — real numbers, **fenced**: printed, but walled off from any ranking.
- **DESCRIPTIVE** — numbers reported, with no registered floor to pass or fail.
- **UNMEASURABLE** — valid attempts exist, but the measurement-failure rate crosses 10% and
  the exam refuses to convert what remains into a verdict. The judge-seat exam refuses the
  counts outright; the tools trial prints them fenced as inadmissible.
- **SEAT-BLOCKED** — the model works, but something in what it returns breaks the contract
  our seats require.
- **Strict parse binds** — a caller who requested a JSON schema and got unparseable bytes
  has a failed contract *whatever the reason*. Diagnosis can be sympathetic; the verdict is
  not.
- **The empty cells** — four reasons leave a cell without a score and the tables keep them
  apart: **NOT-CARRIED** (valid attempts came back zero, nothing to score), **NOT-RUN
  (spill)** — the model did not fit the card whole, and we do not run speed or long-context
  through a half-CPU model, which is a rule rather than a judgment call — **NOT-RUN (stop)**
  for the disk halt and **NOT-RUN (trim)** for the time-box we set. One cell is none of the
  four: laguna's judge trial was **cut mid-leg**, with real calls already made, and it says
  so rather than pretending the leg never started. A row's first empty reason propagates
  rightward.

**The floors, pre-registered before any scoring.** The judge trial passes at kill-recall ≥
11 of 12 AND preservation ≥ 8 of 9; the judge seat's trial at kill-recall ≥ 23 of 27 AND
preservation ≥ 13 of 16. A cross-run difference under 2 items reads TIED. Past a 10%
measurement-failure rate an exam is UNMEASURABLE. The registration fixing all of this ships
in the kit with its sha256, `423ed960…`, and so does the 21-item judge fixture it counts,
`fede4e32…`.

**One exam runs at settings that are not this battery's, and it changes how a column
reads.** The judge seat's 43-case trial is an older, separately frozen exam and runs at
**its own registered settings — a 16,384-token context and no `think` flag at all — where
every other instrument here runs 32,768 with `think:false` explicit.** We ran it as
published rather than re-tuned, because its floors of 23/27 and 13/16 only mean anything
against the run they were set on. Its counting rule differs too: this battery's judge trial
puts a truncation against the reply budget in its own run-quality column, outside the 10%
ceiling, while the seat exam's own frozen rule counts a truncation as a response failure.
Where the two treat the same failure differently below, the difference is the exam, not the
model.

## The standings {#the-standings}

Every cell is copied from that leg's own scorer file. Sizes are the fit receipts' own MiB,
read with a 32,768-token context already loaded — a loaded footprint, not a shipped weight
file.

The schema probe (C0) comes first because it gates the rest: qwen3.6:27b and gemma4:31b held
the schema both ways, muse-glimmer:30b is **SEAT-BLOCKED**, laguna-xs-2.1:latest returned
the battery's one **inverted trap**, and nemotron-3.5-lightning:30b never sat it. Both
findings have their own section below.

**Table A — the door: the fit gate, at a 32,768-token context**

| Candidate | Role | Boards whole | On GPU (%) | Loaded footprint (MiB) | On GPU (MiB) | Spilled to CPU (MiB) |
|---|---|---|---|---|---|---|
| qwen3.6:27b | Repro | Yes | 100 | 16,469 | 16,469 | 0 |
| muse-glimmer:30b | Re-probe | Yes | 100 | 15,831 | 15,831 | 0 |
| laguna-xs-2.1:latest | New kid | Yes | 100 | 19,440 | 19,440 | 0 |
| gemma4:31b | New kid | No | 92.5 | 20,552 | 19,011 | 1,541 |
| nemotron-3.5-lightning:30b | Repro | No | 83.1 | 24,434 | 20,298 | 4,136 |

**Table B — the judge trial (C1): 21 items × 3 repeats = 63 calls**

| Candidate | Posture | Kill-recall (of 12, floor 11) | Preservation (of 9, floor 8) | Calls scored clean (of 63) | Calls truncated | Verdict |
|---|---|---|---|---|---|---|
| qwen3.6:27b | think:false | 11 | 8 | 60 | 3 | RANKED |
| muse-glimmer:30b | Reasoning-strength medium | 11 | 7 | 63 | 0 | EXPLORATORY |
| muse-glimmer:30b | Reasoning-strength high | 12 | 6 | 52 | 10 | EXPLORATORY |
| gemma4:31b | think:false | Not scored | Not scored | 0 | 0 | NOT-CARRIED |
| laguna-xs-2.1:latest | think:false | Not scored | Not scored | 19 | 0 | Cut mid-leg |
| nemotron-3.5-lightning:30b | Not sat | Not scored | Not scored | 0 | 0 | NOT-RUN (trim) |

Glimmer's high row also lost one call to a transport failure (52 + 10 + 1 = 63), and both
its rows are fenced by the C0 verdict whatever the counts say. gemma4's 63 calls were all
response failures. Laguna's 19 calls landed clean across 7 of the 21 items before the stop
cut the leg; they ship in the kit, and they are not a result.

**Table C — the judge seat's own 43-case exam: 129 calls per candidate**

| Candidate | Kills (of 27, floor 23) | Preservation (of 16, floor 13) | Calls that failed to score (of 129) | Replies unwrapped from a code fence | Verdict |
|---|---|---|---|---|---|
| gemma4:31b | 27 | 9 | 12 | 117 | FAIL |
| muse-glimmer:30b | Refused by the exam | Refused by the exam | 18 | 0 | UNMEASURABLE |
| qwen3.6:27b | Refused by the exam | Refused by the exam | 93 | 0 | UNMEASURABLE |

laguna and nemotron never reached this exam — NOT-RUN (stop) and NOT-RUN (trim). The two
UNMEASURABLE rows do have kill and preservation counts in the scorer file; the exam refused
them at the ceiling, so we do not print them as results.

**Table D — the exams with no registered floor: counts only, no pass or fail**

| Candidate | Assistant, C2 (of 20) | Tools task success (of 19) | Tools outcome | Field exam (of 60) | Filing recall (of 18) | Filing abstentions (of 18) | Filing fabrications |
|---|---|---|---|---|---|---|---|
| qwen3.6:27b | 15 | 14 | RANKED | 58 | 16 | 18 | 0 |
| muse-glimmer:30b | 16 | 13 | UNMEASURABLE | 55 | 18 | 18 | 0 |
| gemma4:31b | 16 | 12 | RANKED | 58 | NOT-RUN (spill) | NOT-RUN (spill) | NOT-RUN (spill) |

laguna and nemotron sat none of these — NOT-RUN (stop) and NOT-RUN (trim). The tools trial
runs 19 frozen tasks twice each, and glimmer failed 6 of those 38 attempts (15.8%), over the
10% ceiling: its 13 is fenced as inadmissible, and so are all 19 of its tasks, not 6 of them.
The filing exam asks 36 filings twice each, 72 calls, and neither candidate that sat it
failed a single response.

**Table E — speed (C5): median decode rate, one real prompt loaded at each context size**

| Candidate | Prompt size (tokens) | Median decode (tok/s) | Calls counted | Calls with null counters |
|---|---|---|---|---|
| qwen3.6:27b | 1,000 | 62.0 | 13 | 0 |
| qwen3.6:27b | 8,000 | 60.8 | 10 | 0 |
| qwen3.6:27b | 32,000 | 59.1 | 10 | 0 |
| muse-glimmer:30b | 1,000 | 41.1 | 13 | 1 |
| muse-glimmer:30b | 8,000 | NOT-CARRIED | 0 | 10 |
| muse-glimmer:30b | 32,000 | 39.5 | 9 | 1 |

gemma4 is NOT-RUN (spill) at every tier, laguna NOT-RUN (stop), nemotron NOT-RUN (trim). The
measured prompts run 996 to 31,996 tokens, a 32× spread, so the near-flat curve (62.0 → 59.1
tok/s) is a finding about attention cost on this card, not an unfilled tier.

## What the battery was for: the instrument travels {#what-the-battery-was-for-the-instrument-travels}

On 2026-08-12, two weeks before this battery, qwen3.6:27b earned RANKED on the judge trial —
kill-recall 11/12, preservation 9/9 — on the 96 GB card. This battery re-pulled the same tag
and got a **different build** — a different manifest digest, `9d5803d4…` where 2026-08-12
recorded `a50eda8e…` — on **different silicon**, and the frozen exam read: kill-recall
**11/12**, preservation **8/9** — both floors cleared, **RANKED again**.

The margin earns a plain sentence, because the 8/9 is not a judgment the model got wrong. Of
63 calls, 60 scored clean, none failed the response contract, and three hit the 1,024-token
reply budget — and those three were all three repeats of a single preservation item, so that
item carries no verdict at all. qwen preserved eight of the eight items it was scored on,
and the ninth was unscorable rather than lost. Against 2026-08-12's 63 clean calls and 9 of
9, the preservation delta is one item, under the pre-registered two-item tie band, and reads
TIED — the whole distance between "cleared at the floor" and "clean repeat" is those three
truncations. Self-consistency across repeats moved the same single item: 21 of 21 on
2026-08-12, 20 of 21 here.

Honesty about the claim itself: that is one row, and the second repro row was time-boxed to
its fit gate — and its 2026-08-12 judge row had failed both floors anyway, so the sibling
this claim awaits is a second passing row, not nemotron's. "Travels" is an n=1 result, not
yet a law.

The detail that makes this measurement rather than memory: the kill-recall count is
identical but **the missed item moved** — the earlier build missed `c1-k-bez-c1`, this
battery's build missed `c1-k-bez-h1`: a different case in the same family, landing on the
same score. A frozen instrument catching a moving model in a slightly different place, at
the same reading, is what reproducibility actually looks like.

Two other instruments sat both runs, and counts travel, so both are fair to compare. The
assistant trial — descriptive, but frozen — reproduced **exactly**: 15 of 20, 11 of the 13
items with an exact check, 4 of the 7 scored by the registered proxy rule, zero response
failures, the same in both runs on a different build and a different card. The toolbench
went the other way: **16/19 task success on 2026-08-12, 14/19 in this battery**, grounded —
passed *and* made every pre-registered call — 14/19 → 12/19, both RANKED and admissible, the
same 19 frozen tasks. A two-item move is not under the tie band, so we print it as a real
move rather than noise, and it is a build-to-build move rather than a card-to-card one,
because counts do not care about silicon. One instrument reproduced exactly, one reproduced
at the floor, one slipped. That is what "travels" honestly looks like at n=1.

## Schema-perfect, and still blocked {#schema-perfect-and-still-blocked}

We first wrote muse-glimmer:30b's schema-probe finding down wrong, and the correction earns
its paragraph. On the long-context filing exam glimmer was flawless — 18/18 recall, 18/18
correct abstentions, zero fabrications across 72 calls. And it is **SEAT-BLOCKED**, because
under an enforced JSON schema it emits *schema-perfect JSON* — and then the runtime leaks a
trailing end-of-text sentinel into the bytes, which breaks a strict parser. Our probe first
said "can't hold a schema" while the field exam's own 20-item schema-extraction task scored
it 19/20; two instruments, same model, opposite verdicts — so we read the raw bytes, and the
bytes overruled both summaries. The verdict *stands* (strict parse binds: a caller who asked
for a schema got unparseable bytes), but the diagnosis inverts — a one-line fix at the
caller, not a model to discard. A blocked seat and a broken model are different facts, and
the exam records which one it saw.

## The judge's chair scored one candidate of three {#the-judges-chair-scored-one-candidate-of-three}

Three candidates sat the judge seat's 43-case trial; the exam scored exactly one. gemma4:31b
caught **all 27 planted kills — clearing that floor — and failed anyway**: preservation 9/16
against a floor of 13. A judge that fails almost everything will of course catch every
planted error; the trial exists precisely to price that.

One thing about that verdict belongs on the page rather than only in the kit. 117 of
gemma4's 129 replies arrived wrapped in a code fence and were unwrapped before scoring. That
exam asks for JSON in the prompt rather than through a schema grammar, so unwrapping is what
its own frozen rule does, and strict parse binds where a schema was *sent* — which this exam
never sends. It is still worth printing that the one candidate the exam scored is the only
one of the three that needed it: the other two arrived at zero fenced replies each.

The other two came back **UNMEASURABLE**: glimmer at 18 of 129 calls failing to score (14%),
qwen at 93 of 129 (72%) — truncations against that exam's own 1,024-token reply budget, in
its own 16,384-token context. A verbose model meeting a short cap, refused by the exam
rather than converted into a score. Under this battery's judge trial the same failure mode
sits outside the ceiling entirely; under the seat exam's older rule it crosses it. The
difference is the exam.

The same trade appeared inside glimmer's judge trial, where raising its reasoning-strength
dial from medium to high bought a perfect kill sweep (11/12 → 12/12) and *paid for it* in
preservation (7/9 → 6/9). That dial is a line in the system prompt, not Ollama's `think`
flag: both rows went out at `think:false`, the seat contract, and the kit's per-call records
carry the wire settings.

**Nothing was seated this battery.** qwen3.6:27b boarded the card whole and returned a
ranked or clean result on seven of the eight exams it sat; the eighth — the judge seat — was
UNMEASURABLE under the frozen budget. We intend a re-run, and the frozen-instrument rule
binds us too: the reply budget is part of the exam, gemma4's FAIL was measured under the
same 1,024 tokens, so a changed budget would be a new exam version applied to every
candidate. We would publish such a re-run fenced from the published floors — not as a
seating run.

## What a 24 GB door actually filters {#what-a-24-gb-door-actually-filters}

Three of five board the card whole at 32k — 15,831, 16,469 and 19,440 MiB resident, 100% on
GPU. Two do not: gemma4:31b runs 92.5% on GPU and spills 1,541 MiB, and
nemotron-3.5-lightning spills **4,136 MiB, running 83.1% on GPU**. Nemotron's shipped
weights are **25.43 GB (23.68 GiB)** against a card holding 24.0 GiB, so once a 32,768-token
context is loaded beside them there is nowhere left to sit and 4,136 MiB goes to the CPU.
That shipped size is read from the runtime's own tag listing, and the same runtime reports
different sizes for byte-identical builds across versions — so the digest, not the size,
identifies a build here. The kit records that anomaly as an open flag, unresolved as it
ships.

Both spills are published as the finding, speed and long-context rows skipped rather than
measured through them — and both are verdicts on the artifact as shipped: repacked smaller,
either might board, and that would be a different candidate. gemma4:31b still posted the
joint-best field exam of the battery (58/60, tied with qwen) — a good model this card cannot
hold whole is a *different fact* from a bad model, and the tables keep them apart.

## The row that owes a re-run {#the-row-that-owes-a-re-run}

laguna-xs-2.1:latest fits whole (19,440 MiB) and showed the battery's only **inverted
trap**. Under `think:false` — the exact posture our seats contract — its schema output
parses clean; under `think:true` it drops the schema constraint. That is the inversion:
glimmer, and the documented majority, fail the other way round, so a caller who "fixed" a
schema problem here by turning thinking on would be breaking it. Whether that lives in the
model's schema discipline or simply in the runtime's leak not firing on laguna's template is
exactly what its full battery would have said.

It lost that battery twice, first to the disk stop mid-leg, then to the time-box we set.
What survives is 19 valid calls across 7 of the judge trial's 21 items, made and never
scored. Its row stays honestly empty, and it is first in line for the next window.

## What this page does not say {#what-this-page-does-not-say}

- **No seats were awarded.** qwen's claim covers seven exams and is incomplete by the
  eighth; the re-run rule above says what would make it complete, for every candidate
  equally.
- **One tier is a build-specific anomaly, resolved by a pre-named test:** glimmer's 8k speed
  tier returned empty counters on all ten attempts while qwen's returned ten clean readings
  on the same harness, card and prompt construction. The anomaly belongs to the glimmer
  build, and it is recorded NOT-CARRIED.
- **Both speed rows are recoveries, and the timed work is intact.** Both C5 summary passes
  crashed *after* their timed calls — glimmer on a manifest-shape bug of ours, qwen on a
  hardcoded byte-compare 404 in the harness. The medians here are recomputed from the
  per-call records under the frozen counting rules, and the kit ships the file that says so.
- **The speed numbers were measured on a box that was also serving.** The live seats ran on
  the other card, but the host is shared, so the timed calls could carry contention. The
  check is in the raw: each candidate's 1k tier opens with a three-call clean baseline, and
  it lands within 0.3% of the ten timed calls that follow for qwen (62.1 against 62.0 tok/s)
  and 2.4% for glimmer (41.9 against 40.9). The published 1k figure is the median of all
  thirteen.
- **One battery, one card, build-pinned digests.** Every verdict names the build it
  measured; a tag is not a model — which is exactly why the repro rows exist.
- **Speeds are not compared to the 96 GB run**, by the count-travels rule above; a reader
  who diffs them against the earlier page is measuring silicon, not models.

## What to take with you {#what-to-take-with-you}

*Five things, in the page's own words:*

- **Counts travel; rates don't.** A frozen exam's correctness counts are a property of the
  model and can be compared across cards; tokens-a-second belong to the model *and* the
  silicon and are new measurements every time.
- **The instrument travelled at n=1, and all three readings are printed.** One trial
  reproduced exactly, one reproduced at the floor with the missed item moved to a sibling
  case, and one slipped by two items, which the tie band does not cover.
- **A 24 GB door filters by artifact, not architecture.** Two of five spilled as shipped;
  one of those posted the joint-best field exam. Repacked smaller they might board — and
  that would be a different candidate.
- **Frozen means frozen, including the exam you inherited.** The judge seat's trial runs at
  its own registered settings and its own counting rule, both disclosed here, because
  "fixing" it to match the battery would make it a different exam with the same name.
- **The bench stopped itself rather than fill its disk.** It halted at a conservative floor,
  resumed only after every digest re-verified, and we tore its store down under four guards
  — weights are re-pullable; measurements are not.

## How to check our work {#how-to-check-our-work}

[The kit](/the-instrument-travels/data/) is beside this page, and what it is for is
re-deriving every number printed here from the calls we actually made: the five fit
receipts, every scorer file these tables read, the per-call records (including the judge
seat's rows with their own `num_ctx` and `think` settings left exactly as sent), the build
digests, the stop/resume log, the four-guard teardown receipt, the pre-registration with its
golden hashes — `423ed960…` for the judge trial's registration, `fede4e32…` for the 21 judge
items — and a `counting-rules` note that states the count-travels rule, the 10% ceiling, the
two-item tie band and the judge-seat deviation in one place. Where a scorer computes a rate
it carries a Wilson 95% interval beside it; where it counts, it prints counts of N.

Every timestamp of the run itself in the kit is UTC. There is one exception and the kit's
README names it: the tool-use trial's 19 frozen tasks are set in a fictional harbour with
its own scenario clock, dated 2026-08-01, -08-02 and -08-12 rather than the run day,
published as captured because the published checker matches the clock reading itself. None
of them is a reading of any real clock.

If a number here does not reproduce from the kit, [say so at the contact desk](/contact/),
where a person reads every message.

## The rest of the seminar {#the-rest-of-the-seminar}

The chairs these candidates contested were set in *[the chair trials](/chair-trials/)* on
the 96 GB card, with *[the seat trials](/seat-trials/)* and *[the outside
judges](/outside-judges/)* behind them. The consumer-card side of the workshop is *[A Rig
Your Friend Already Owns](/a-rig-your-friend-already-owns/)*, and what a smaller room costs
in watts rather than chairs is *[What 150 Watts Buys](/what-150-watts-buys/)*. Thanks are
owed to the people behind the five builds measured here; we wrote none of them.

[The whole shelf](/) holds the benches behind the claims we publish, failures included. If
there is a piece of the machinery you want opened next, [say so](/contact/) — the suggestion
box is read.

*Released under the house licence; main measurement window 2026-08-26, 07:00–15:13Z (UTC),
with a stop at 09:58Z and a resume at 13:55Z. One 24 GB consumer card, Ollama 0.32.13, every
candidate at its default published tag and the vendor's own quantisation; the 96 GB card of
the earlier run is named by class throughout.*

<!-- derived 2026-09-26 (UTC) by tools/derive_md.py from the pour source.
     source html sha256: 67b0c320e1bb79eebb3d38e311b40e5e7d3fa89373d2a9768305a9ce07714ffc
     derivation sha256:  bc7b46f69b330d15e3ceefdb28680c63dc0c19120a5e9ebcded24e09c3f8575e
     the {#id} on each heading is the anchor that heading carries on the page. -->
