# The Instrument Travels — data kit

**What this directory holds.** Every receipt behind the exhibit *The Instrument
Travels* (https://research.strata2signal.com/the-instrument-travels/index.html):
the five fit receipts, every scorer file the standings grid reads, the per-call
records behind those scorers, the digests that identify the builds, the lane log
with its stop and its resume, the teardown receipt, and the pre-registration and
frozen-fixture hashes that make the word "frozen" mean something.

Five models, **one 24 GB consumer card**, one at a time, on one day —
**2026-08-26**. The card was emptied of its resident classifier for the run; the
box's *other* card kept serving live traffic throughout, which is what the
teardown receipt is checking at every checkpoint.

Measured **2026-08-26T07:00:15Z – 2026-08-26T15:13:00Z**, with the battery
**stopped at 09:58:39Z** (and again at 10:14:59Z) on its own disk-safety floor
and **resumed at 13:55:19Z**. Both stops and the resume are in `battery.log`, and
the first stop is also its own scorer artifact, `summary-C1-none.json`.

**Licence: CC BY 4.0.** Take these rows, re-count them, check our arithmetic,
publish what you find. Attribution: strata→signal research,
research.strata2signal.com. If a number on the page does not reproduce from this
kit, we want to hear about it: hello@strata2signal.com.

## Start here

1. **`counting-rules.md`** — the four rules every number obeys: what travels
   across cards and what does not, the 10 % measurement-failure ceiling, the
   two-item tie band, and the one exam that ran at its own settings.
2. **`digests.md`** — which builds were measured, and the standing size-vs-digest
   flag that means you should quote the digest and not the size.
3. **`index.json`** — every file with its size, its sha256, and what it answers.
4. **`provenance.json`** — every file again, with both sha256s, the named rules
   applied to it, and how many times each one fired.

## What is in here

**177 files.** `index.json` lists 175 of them with size and sha256 — every file
except this kit's browsable listing (`index.html`) and index.json itself, because
a file cannot carry its own sha256; `provenance.json` lists 174, the same set less
itself, and `index.html` lists the 176 that are not it. (Corrected 2026-08-29:
this sentence read "lists every one of them", which a reader running `jq
'.files | length'` gets 175 for.) This table says what each *kind* of file is,
and which named rules touched it.

| path | files | what it holds | rules applied |
|---|---:|---|---|
| `README.md`, `counting-rules.md`, `digests.md` | 3 | This file, the counting rules, and the build-identity tables. | authored for publication |
| `index.json`, `index.html`, `provenance.json` | 3 | The directory, listed; the kit's door; the transform record. | — |
| `prereg/PREREG-C1.md` | 1 | The judge trial's own pre-registration, as frozen. | none — byte for byte |
| `prereg/GOLDEN-SHAS.md` | 1 | The five fixture sha256s that were verified byte-identical to the frozen originals at clone time, plus the judge seat's external fixture sha. | authored for publication |
| `fit-<tag>.json` | 5 | **The primary result for all five candidates** — the only instrument every one of them completed. Resident size, spill, fraction on GPU, and the verdict, all measured with a 32,768-token context loaded. | `local-timestamps-to-utc`, `private-host-to-role` |
| `probes-<tag>.json` | 4 | The schema-discipline probe (C0): the polarity verdict per candidate, the seat-blocked decision and its reason, the leaked-sentinel diagnostic, and all six per-call records behind them. Two repeats per configuration, six calls per model — enough to see a contract break, not enough to rate one. | `local-timestamps-to-utc`, `private-host-to-role` |
| `summary-C1-<tag>.json` | 3 | The judge trial's scorer output per candidate: per-item verdicts, the kill/preservation tallies with their floors and their miss ids, and the outcome block with its call, truncation and failure counts. | `local-timestamps-to-utc`, `absolute-paths-to-relative`, `box-name-to-role`, `operator-name-to-the-operators` |
| `summary-C1-none.json` | 1 | **The disk stop, as a scorer artifact**: the trigger, the free-space reading, the 60 GiB floor, and the full C1 protocol block including the pre-registered floors, the reply budget and the 10 % ceiling. | `local-timestamps-to-utc`, `box-name-to-role` |
| `summary-C1.json` | 1 | A derived index of the three per-model files. It holds no number of its own and says so. | `local-timestamps-to-utc` |
| `c2-scores.json` | 1 | The assistant trial: per-model score out of 20, split into exact and proxy items, with per-item records and the failure counts. | `utc-offset-to-z-stamp` |
| `c3-scores.json` | 1 | The tool-use trial: task success out of 19, plus grounded success, tool selection, argument fidelity, chain completion, honesty traps and the spurious-call rate — each with its own denominator — and the run-quality block the UNMEASURABLE verdict is read from. | none — byte for byte |
| `c5-recovered.json` | 1 | The decode rates, recomputed from the per-call raw records under the frozen counting rules after both summary passes crashed *after* their timed work. Warm blocks only, null-counter calls excluded, median per prompt tier. | `absolute-paths-to-relative` |
| `c7-<tag>.json` | 2 | The long-context filing exam: recall, correct abstentions and fabrications out of 18 each, across 72 calls, with per-tier and per-depth splits. | `private-host-to-role` |
| `field/<tag>.json` | 3 | The 60-item field exam: three 20-item sections, the wire options it ran at, its scoring patterns, and every item. | `private-host-to-role` |
| `seat43-summary.json` | 1 | The judge seat's 43-case exam, scored: kills, preservation, both floors, latency percentiles, failure counts, and the verdict per candidate. | none — byte for byte |
| `battery.log` | 1 | The lane log, start to teardown. **The stop and resume lines are the point**, and so are the two digest verifications around them. | `local-timestamps-to-utc`, `private-host-to-role`, `port-to-seat-role`, `absolute-paths-to-relative`, `box-name-to-role`, `card-name-to-class`, `operator-name-to-the-operators`, `project-slug-to-neutral`, `gate-collision-label-expanded` |
| `TEARDOWN-RECEIPT.md` | 1 | The four-guard teardown: before and after disk, the tags removed, the tags skipped (none), and the read-only probe of every live seat. | `local-timestamps-to-utc`, `private-host-to-role`, `port-to-seat-role`, `absolute-paths-to-relative`, `box-name-to-role` |
| `raw/C1/<tag>/` | 4 | Every judge-trial call: the request as sent (schema, options, messages), the reply, the thinking channel, wall time, done reason and the scored verdict. | `local-timestamps-to-utc`, `provenance-header-added` |
| `raw/C5/<tag>/` | 3 | Every decode-rate call, including the discarded warmup blocks, labelled as such. | `local-timestamps-to-utc`, `provenance-header-added` |
| `raw/c2/` | 124 | Every assistant-trial call: one JSON file per item per repeat, plus a `calls.jsonl` per model and the run manifest. | `utc-offset-to-z-stamp`, and on the manifest also the host, path and box rules |
| `raw/c3/` | 10 | Every tool-use call and its full round-by-round trace, plus the pre-probe and the run manifest with its pinned task, tool and prompt hashes. | `local-timestamps-to-utc`, `utc-offset-to-z-stamp`, `private-host-to-role`, `provenance-header-added` |
| `raw/seat43/` | 3 | **The judge seat's per-call records, which live outside the bench tree** — 129 scored calls per candidate plus the run-start envelope. The wire options here are deliberately unaltered; see the exception below. | `absolute-paths-to-relative`, `project-slug-to-neutral`, `private-host-to-role`, `port-to-seat-role`, `provenance-header-added` |

## About the timestamps

The bench wrote its records on **two clocks**. The field-exam and assistant-trial
scorers stamped UTC; every other harness record stamped the bench machine's own
local wall clock with an explicit numeric offset. Those two conventions sat side
by side in the same tree, which means the offset was recoverable by subtraction
even if it had merely been de-labelled.

**In this kit there is one clock.** Every harness timestamp is UTC and
Z-stamped: the local ones were **converted** (not de-labelled), and the ones
already at UTC were re-written from `+00:00` to `Z` so nothing here needs a
reader to know two formats. `provenance.json` prints, per file, how many stamps
each rule moved. No row was added, dropped or re-ordered; nothing but the
timestamp text changed.

**One deliberate exception, and it is not the bench's clock.** The tool-use
trial's 19 frozen tasks are set in a fictional harbour, and that scenario has its
own clock: a fixed "now", a set of harbour log-line stamps, and the models'
graded answers quoting them back. Those are **fixture content**, they are dated
2026-08-01, -08-02 and -08-12 rather than the run day, and the published checker
matches the clock reading itself — converting them would have rewritten graded
answers and destroyed the one property that makes those grades re-derivable. They
are published **as captured**, 185 of them, under a standing 2026-08-20 ruling,
and `provenance.json` counts them per file as `fixture-clock-left-as-captured`.
They are the only stamps in this kit that are not UTC, and none of them is a
reading of any real clock.

**Why a grep for the offset returns 197 and not 185** (added 2026-08-29, because
both numbers are right and a reader who counts will find them disagreeing). The
185 is the count of fixture *stamps* — a full
`2026-08-01|02|12T<hh:mm:ss>-06:00` — that the sanitiser saw in the seven files
carrying them and deliberately left alone; that is what `provenance.json`'s
per-file `fixture-clock-left-as-captured` counts, and the seven files total
exactly 185 (16, 16, 1, 43, 61, 24, 24). A plain `grep -o -- '-06:00'` over the
same files returns **197**. The extra **twelve** are not stamps the harness
wrote: they are the models' own graded answers **re-writing** the fixture clock
in a shape of their own — six of `she was moved at 09:26:02-06:00 on 2026-08-02`
in `raw/c3/gemma4_31b/think_false/trace.jsonl`, and six of
`**2026-08-12 09:00:00 -06:00**` in
`raw/c3/muse-glimmer_30b/think_false/trace.jsonl`. They fall under the same
2026-08-20 harbour-fiction ruling for the same reason the stamps do: they are
graded answer text, and editing them would rewrite the answers the grades are
derived from. So — **185 fixture stamps left as captured; 197 occurrences of the
offset string, the extra 12 being the fiction quoted back inside answers.**
Neither number is a reading of any real clock, and no figure on the exhibit page
depends on either.

## Sanitised at the pen, not copied

Nothing here is a raw copy of a file that named a machine. Each file was read and
re-written, and every transform has a NAME that `provenance.json` prints per file
with the number of times it fired — so you can see which rules touched a file and
be sure nothing else did. The rules, in full:

- **`local-timestamps-to-utc`** (782) — a harness timestamp on the bench
  machine's local clock is converted to UTC and Z-stamped. The instant is
  preserved; the offset is not published.
- **`utc-offset-to-z-stamp`** (247) — a harness timestamp already at UTC but
  written `+00:00` is re-written `Z`. Same instant, one format.
- **`fixture-clock-left-as-captured`** (185) — counted, not applied: the frozen
  harbour fixture's own scenario stamps, published unchanged. See above.
- **`private-host-to-role`** (90) — a private network address the harness called
  becomes the role it played: `the bench endpoint`, `the 96 GB box's production
  seat A`, and so on. *What* was served, and how, is unchanged in the same file.
- **`absolute-paths-to-relative`** (88) — a path into a private filesystem
  becomes the file's own path **inside this kit**, so a scorer's `raw_path` still
  resolves: `summary-C1-qwen3.6_27b.json` points at
  `raw/C1/qwen3.6_27b/C1-qwen3.6_27b.jsonl`, and it is there.
- **`box-name-to-role`** (28) — a machine name becomes what it did: `the bench
  box`, `the 96 GB box`, `the bench card`. A card's hardware UUID becomes
  `<card uuid redacted>`.
- **`port-to-seat-role`** (19) — a seat named by a bare port number becomes the
  seat: `the bench endpoint`, `production seat A` / `B`. Kept distinct so the
  teardown's read-only probe of *three separate live seats* still reads as three.
- **`project-slug-to-neutral`** (8) — an internal project slug in a fixture
  version string and a path becomes a neutral name. No version changes.
- **`gate-collision-label-expanded`** (4) — one column label in the lane log, a
  three-letter abbreviation of the word "estimate" used for the estimated call
  count, is spelled out in full, because the abbreviation collides with a
  publish-time check for timezone names. A label; no number, row or column moves.
- **`operator-name-to-the-operators`** (3) — a person's name becomes `the
  operator`.
- **`card-name-to-class`** (1) — the card's model name becomes its class, **a
  24 GB consumer card**. The class is what every figure in this kit is about.
- **`provenance-header-added`** (19) — every `.jsonl` in this kit gains ONE extra
  first line, a record of `"type": "kit-provenance"`, naming the card class, the
  leg's wire settings, the runtime version and the run window. **Skip records
  whose `type` is `kit-provenance` and you have the harness's own stream.**

### The one exception, and why it is an exception

**The judge seat's wire options are published unaltered, and deliberately so.**
Under `raw/seat43/`, every record carrying an options block reads `"num_ctx":
16384` and `"num_predict": 1024`, and each file's run-start record reads
`"think": "omit"`. Those settings are *not* this battery's house laws — every
other instrument here ran at 32,768 with an explicit `think` boolean — and that
difference is disclosed on the page and stated in `counting-rules.md`, rule 4.
Rewriting them to match would have destroyed the only thing that makes the
disclosure checkable.

## What was left out, named rather than silently dropped

- **The bench's night report.** An internal working document. The two tables a
  reader actually needs from it — the staging digests and the repro-row digest
  comparison — are extracted into `digests.md`; nothing else from it ships.
- **The derived results views** (`RESULTS-TABLES.md`, `C3-RESULTS.md`). By their
  own description they compute nothing: every number in them is in the scorer
  JSON that ships here. Omitted rather than shipping a second, drifting copy.
- **A VRAM census taken on a different machine.** One early blocking measurement
  in the bench tree was taken on the 96 GB box before the battery was re-targeted
  to the consumer card. It is not this kit's subject and would have been the one
  genuinely misleading file in it, so it is not here.
- **The clone-provenance note.** Its content that matters — the five verified
  fixture hashes — is in `prereg/GOLDEN-SHAS.md`.

## Two caveats we are stating rather than fixing

1. **The size-vs-digest anomaly is unresolved.** `digests.md` §3: two tags report
   different `size` values across runtime versions on byte-identical manifest
   digests, one of them by 8 %. This kit publishes sizes, so read that section
   before quoting one, and prefer the digest for identity.
2. **The bench's internal night report miscounts one weight.** In its final
   standings section it lists `gemma4:31b`'s weights as 24,434 MiB — that is
   nemotron's figure. The correct number is in this kit:
   `fit-gemma4_31b.json` reads `size_mib: 20552` with 1,541 MiB spilled. Nothing
   in this kit or on the exhibit page carries the wrong figure; it is named here
   in case any part of that table is ever reproduced.

## What was not touched

Every count, verdict, floor, latency, token figure, digest, sha256 and reply text
is published as recorded. No row was dropped, re-ordered or re-scored. No failed
call was excluded — the failures are the finding in three of these rows. The
models' replies are quoted verbatim, including the ones that break.

## Not a standard

Five candidates on one card on one day, under one frozen battery, sized to one
question: does a frozen exam say the same thing about a re-pulled build on
hardware a normal person owns? It is not a leaderboard and not a benchmark suite.
It is n=1 on the row that reproduced, and the kit says so. A different card, a
different quantisation, a different runtime version will land somewhere else, and
nothing here is evidence about any of them.
