# The Same Sixteen — data kit

**What this directory holds.** Every receipt behind the exhibit *The Same Sixteen*
(https://research.strata2signal.com/the-same-sixteen/index.html): the full results
narrative for a hosted reference arm, the tables it is read from, the toolbench
per task, the field exam per item, every raw call the arm made, the two frozen
artifacts the page fingerprints, the local rows the count is set beside, and the
registration that says what a reference arm may and may not be used for.

One model, **`glm-5.3-flash:cloud`** — hosted, 320B total / 18B active per token on
its vendor's figures — read on **two** of this house's nine frozen instruments on
**2026-08-28**. It is a **REFERENCE ARM · cloud · dated**: never a candidate, ranked
against nothing, and no threshold is ever derived from it.

Measured **2026-08-28T16:37:59Z – 2026-08-28T16:49:43Z** (UTC, both ends). The
toolbench ran 16:41:17Z–16:44:06Z; the field exam 16:47:49Z–16:48:57Z. No weights
were loaded on any machine of ours for this reading — the calls went to a hosted
model through a local daemon that forwards `:cloud` tags upstream — so it addressed
no production seat and disturbed nothing.

**Licence: CC BY 4.0.** Take these rows, re-count them, check our arithmetic,
publish what you find. Attribution: strata→signal research,
research.strata2signal.com. If a number on the page does not reproduce from this
kit, we want to hear about it: hello@strata2signal.com.

## Start here

- **The two fingerprints the page prints.** `sha256sum tasks.json` returns
  `d172e9c326735df3848389e24f0a82d67b11d97df62805285b2a0b5d527ee140`, and
  `sha256sum tools-manifest.json` returns
  `37a3d20f99a366413a6be7b69be7d80ff2a7cb642c8abf8dcd90d7f81189941b`. Those are the
  frozen task file and the ten-tool roster, and they are the same two shas the run's
  own manifest recorded before the first scored call
  (`raw/c3/manifest-glm-5.3-flash-cloud-2026-08-28.json`, under `pins`).
- **The count, decomposed.** `C3-RESULTS.md` puts all nineteen tasks against each of
  the four rows, with the checker's own note on every one. The arm's 16/19 and the
  three it missed are there by name.
- **Why seven of nine instruments are blank.** `reference-arm-row.json` carries every
  leg with its disposition and its reason in one place;
  `REFERENCE-glm-5.3-flash-2026-08-28.gate.json` is the evidence behind the two
  `format` exclusions and the `think` posture.
- **The rules.** `counting-rules.md` — five of them, each naming where it was
  registered. Read rule 5 before quoting any count from this instrument.

## What is in here

**The narrative and the tables**

- `REFERENCE-glm-5.3-flash-2026-08-28.md` — the full write-up, made the day the arm
  was read: the three gates, the ledger of what could not be read, the toolbench and
  its five honesty traps, the rows it sits beside, the field exam, the cost signals,
  and what the reading does and does not license. It ends with **two dated
  corrections, both 2026-08-30**, appended rather than folded in, because the house
  convention is that a correction appends and leaves the corrected sentence standing.
  The first fixes a §5 cell that attached 19/20 to schema conformance where the
  receipts say conformance was 20/20 and *values* were 19/20. The second fixes three
  things in §§3 and 7: "four orders of magnitude" for a 12.3× size ratio (it is about
  1.1), a bracket that was the union of two different Wilson intervals and is
  therefore a confidence interval for nothing, and a claim of three literal-string
  checker tasks where the evidence in the document names two.
- `RESULTS-TABLES.md` — the whole battery, per candidate × per instrument, with the
  arm's row among them and set apart. Every cell is copied from that leg's own scorer
  output; this view computes nothing.
- `C3-RESULTS.md` — leg C3 in full, including the per-task detail for all four rows.

**The frozen instrument**

- `tasks.json` — the 19 tasks with their checkers, byte-identical to the pinned file.
- `tools-manifest.json` — the 10 tools as the `tools` array rode the wire, in the
  canonical form the run hashes (sorted keys, no whitespace).
- `prereg-reference-arm.md` — the reference-arm class as registered on 2026-08-28,
  before any scored call: what it is, its four binding fences, the two legs it runs,
  the reason each other leg is NOT-RUN, and the gate it had to clear first.

**The evidence**

- `REFERENCE-glm-5.3-flash-2026-08-28.gate.json` — the gate receipt: the morning's
  401 history verbatim, all three `think` postures with the leak reproduced, the
  tool-emission gate with the `tool_calls` it returned, and the `format` enforcement
  probe taken after the last scored call.
- `field/glm-5.3-flash-cloud.json` — the field exam per item, all 60, including the
  20 excluded ones. The exclusion is the evidence, so its items ship.
- `raw/c3/glm-5.3-flash_cloud/preprobe.json` — the frozen pre-probe, both postures.
- `raw/c3/glm-5.3-flash_cloud/think_true/trace.jsonl` — every scored attempt, 38 of
  them, with each round's wire exchange and the final answer as scored.
- `raw/c3/glm-5.3-flash_cloud/think_true/calls.jsonl` — every individual tool call,
  88 of them. This is the file the 44.3% spurious-call figure counts.
- `raw/c3/manifest-glm-5.3-flash-cloud-2026-08-28.json` — the run's own manifest and
  its pins.

**The rows it sits beside**

- `c3-primary-rows-2026-08-12.json` — the 2026-08-12 local rows on this same frozen
  instrument, as the chair trials published them. `gemma4:12b` and `qwen3.6:27b` are
  the two that land on the same 16/19. Taken sixteen days earlier, on the shared
  production daemon rather than the cloud path, at a posture this arm cannot run —
  the page prints no delta across that line and neither should anyone else.

**This kit's own accounting**

- `counting-rules.md`, `README.md` (this file), `index.json`, `provenance.json`, and
  the browsable `index.html`.

## About the timestamps

Every harness timestamp in this kit is **UTC, Z-stamped**. The harness wrote some of
them on the bench machine's local wall clock; those were converted at the pen, under
a named and counted rule, and the instant is preserved while the offset is not
published.

**One clock is deliberately NOT converted**, and it is the one you will see most: the
frozen tool-use fixture's own fictional harbour clock, dated 2026-08-01, 2026-08-02
and 2026-08-12 and written `-06:00`. It is scenario content, it is quoted inside
graded answers, and the published checker matches the clock reading itself — so
converting it would destroy the re-derivability of exactly the grades this kit exists
to let you check. It is counted in `provenance.json` and never touched. This is the
same fixture, under the same standing 2026-08-20 ruling, that exhibit thirty-one's
kit publishes.

## Sanitised at the pen, not copied

Nothing here was copied raw. `provenance.json` names every rule applied to every
file with the number of times each one fired, and carries a **second sha256** — the
pre-sanitisation one — for every file that is not byte-identical to what the harness
wrote. Three files are byte-identical and say so: `tasks.json` (which has to be, or
its pin means nothing), `C3-RESULTS.md`, and `c3-primary-rows-2026-08-12.json`.

The rules, in short: a machine name becomes the role it played, a private address
becomes the endpoint it was, a path into a private filesystem becomes either the
file's own path inside this kit or the thing it names, a person's name becomes "the
operators", a local-clock timestamp becomes UTC, and every `.jsonl` gains one extra
first line of provenance which you can skip by its `type`. Full text and per-file
counts in `provenance.json`.

## What is not here, named rather than silently dropped

- **The bench's runner and its harness.** `tools/c3_run.py`, `tools/c3_score.py`,
  the field-exam runner and the sandbox are not published. This kit ships receipts.
  The three harness changes this run needed are described, additively and by name, in
  the narrative's §8.
- **`c3-scores.json` and `battery.log`.** The narrative and the tables refer to both.
  They belong to the 2026-08-26 night battery, not to this reading, and they are
  already public: they ship in [exhibit thirty-one's
  kit](https://research.strata2signal.com/the-instrument-travels/data/). Neither was
  modified by this run — the narrative's §8 says why, and the scoring for this arm
  was written to a scratch path for exactly that reason.
- **The bench's plan, apart from one section.** `PLAN.md` is an internal working
  document. The one part of it that binds anything here — the reference-arm
  registration — ships whole as `prereg-reference-arm.md`; the narrative and the
  tables cite it under its original name.
- **A currency figure.** None is quoted anywhere, because none was taken. What the
  API's own counters reported is published verbatim instead, in
  `reference-arm-row.json` and in the narrative's §6.
- **The model.** It is hosted, cloud-only on ollama, and there is no blob to pull, no
  digest to pin and no `ollama show --license` receipt to take. The architecture
  figures and the licence in this kit are **vendor-stated**, read off the vendor's
  own library page on 2026-08-28 and labelled as vendor figures everywhere they
  appear. We cannot weigh a model we cannot hold, and we do not pretend to.

## One open item, stated rather than fixed

The narrative's **second correction, item 3** is unresolved on purpose: the evidence
in this document names **two** tasks with literal-string checkers (`t15`, a trap, and
`t04`, single-call). Whether a third exists has not been re-read item by item, and
the bench lane owes that recount before the claim is cited at "three". It is left
open here rather than guessed, and the page states the same thing.

## What was not touched

Every count, verdict, interval, token figure, sha256 and reply text is published as
recorded. No row was dropped, reordered or rescored. No failed call was excluded, and
the three tasks the arm failed are in `C3-RESULTS.md` beside the sixteen it passed.
The model's replies are quoted verbatim, including the deliberation it leaked under
the posture we refused to score.

## Not a standard

One hosted model, on one day, on two of this house's own frozen instruments, sized to
one question: how far is our local fleet from a model that size, measured with the
ruler we already trust. It is not a leaderboard, not a benchmark suite, and not a
verdict about any model. The answer it returned was mostly about the ruler.
