# The narrator's open call — data kit

The machine-readable companions to **exhibit eleven, "A kid, an elder, and a tired
parent walk into the cove"** (https://research.strata2signal.com/the-open-call/index.html): twenty language models handed the same
frozen moment of one small fishing town, seven judges from six model families
reading every reply blind, and the whole round's arithmetic.

**Licence: CC BY 4.0.** Take these rows, re-derive our figures, disagree with
them in public. Attribution: strata→signal research, research.strata2signal.com.
If you find an error in them, we want to hear about it: hello@strata2signal.com.

## What is in here

| file | what it holds |
|---|---|
| `scores.json` | the scoring artifact entire: every arm, every act, every one of the 738 scoring cells with its seat, its two dimension scores, its canon verdict and the judge's own note — plus all 102 recused cells kept with their verdicts, the anchor calibration, the agreement matrix, the tie bands, and the canon tallies. Every figure in the SEALED round is a projection of this file; the addendum's thirty cells are a projection of addendum-scores.json beside it |
| `seal-manifest.json` | the seal: for every frozen bundle, the FULL sha256 of the bundle, its event log and its boot log, the full wire sha256 the arms answered, the byte counts, and the verbatim user line each arm was handed |
| `report.md` | the complete internal report, built by machine from the round's own artifacts — the per-arm and per-ask tables, the anchor calibration, the agreement matrix, the canon findings including the convergent-invented-span cell table with every judge's deciding words, the mirror table, the fabrication classes, every SPLIT in full, the bill and the curated highlights |
| `known-world-issues.md` | the fence: the known-world-issues list that rode every judge sheet — as published, 39 table rows naming 64 distinct issue ids — so a reader can check the rows a judge was allowed to reach for, and check that everything off the fence stayed the model's own |
| `counting-rules.json` | every rule that governs a figure, in one file: the counting rules, the cell arithmetic spelled out, the curation rule as registered AND as implemented, the registered extraction tiers, the mirror column's rule from both of its sources, the fence counted, the tie band with its chaining warning, the forbidden-claims list verbatim — and a note saying why the pre-registration is extracted rather than published whole |
| `seats.json` | the seven judge seats, NAMED: each seat's mean on the shared calibration reply and its offset from the panel, its canon strictness with its own denominator, the pairwise agreement matrix, and the local-seat axis |
| `bill.json` | the bill reconciled: three metered lines, each receipt with the token counts it was billed on, its cited per-token rate and that multiplication performed, the registered caps, the four cost states, the nine discarded cold-start warmups that were billed and never scored, and the difference between the round's carriage artifact and the joined total, named |
| `addendum-scores.json` | the addendum's scoring artifact: the twenty-first arm (qwen3.8:27b, run after the round was sealed) over 30 cells and 5 families, with every cell's seat, dimension scores, canon verdict and judge note, its own anchor calibration — the bridge that makes it readable beside a round it cannot share blinding letters with — and its floor-gate check. A SEPARATE round: nothing in it re-scores the sealed twenty |
| `addendum-samples.json` | the addendum arm's six judged records, whole: each reply's spoken line, its client-wall clock and the endpoint's own counters, the sha256 of the reply, and the wire_sha256 it answered — which is the sealed round's own, byte for byte, so a reader can prove the late arm sat the same exam. Published where the sealed round's records are withheld, and the README says why |
| `addendum-runtime-pins.json` | the addendum's runtime provenance: the ollama version it ran on, the full digest, size and quantisation of the late arm's weights, and the sealed round's three local tags re-resolved today with the one fact that qualifies them — whether each tag's file predates the seal. It also records, in its own words, that the sealed round's runtime was never written down: this file is the receipt for the caveat the page cannot cure, including the part of it that is an inference rather than a reading |
| `addendum-probe.json` | the release re-probe, all 16 calls: four request configurations against both the late arm and the local judge seat, run on the new runtime BEFORE any reply was captured, testing whether asking for JSON with reasoning switched off silently drops the constraint. Every call's verdict, its latency and the first 200 characters it returned — so the claim that the trap did not reproduce is checkable rather than asserted |
| `addendum-bill.json` | the addendum round's own bill, in bill.json's shape: a line per judge seat with the token counts it was billed on, the rate they were priced at and that multiplication performed, the three seats that bill nothing published with their counts too, and the sealed round's total beside this one with a note on why the two add rather than overlap. It exists because the addendum published with $0.3579 on the page and no token counts behind it — this is the file that makes that figure re-derive |
| `README.md` | this file |
| `index.json` | this directory, listed: file names, sizes, sha256s, what each one answers, what was scrubbed from it, and every correction annotated into it |

`schema` on every derived file reads `s2s-bench-v1`.

## How to check the page against these files

Every figure on the exhibit is a projection of one of these companions, and the
builder that writes the page reads them rather than being told them. The three
checks worth running first:

1. **The cell arithmetic.** `scores.json` → `cells_filed` 868, `recused_count`
   102, 28 calibration anchors, **738 scoring cells**. Every sheet was answered
   blind, so a judge filed on every letter in front of it; recusal is applied
   afterwards by key-join, and those 102 cells are kept, printed, and enter no
   figure. The page leads with 738 and decomposes 868 in the same breath.
2. **The canon tallies.** `scores.json` → `canon_tally.overall` gives
   235 not-clean of 738, split 127 / 49 / 45 / 5 / 9. `canon_tally.by_seat`
   gives the same figures per named seat, which is where the panel's own
   disagreement lives.
3. **The bill.** `bill.json` → three metered lines. Every receipt under every
   line carries the token counts it was billed on, the rate they were priced at,
   and `usd_re_derived_from_these_tokens` — the multiplication performed — so
   each line is a check rather than a figure to trust. It also names two things
   a reader would otherwise have to guess at: the nine **discarded cold-start
   warmups** (one per metered arm per leg, registered in prereg §8.2, billed,
   receipted, never scored — they are the whole of the gap between a metered
   arm's call count and its published replies), and a gap in the round's own
   artifact — the carriage table bills by judge-seat key, so one metered arm
   that seats no judge chair had receipted dollars and no carriage line. The
   page publishes the joined total. The post-publication addendum is a separate
   round and has its own bill in the same shape, `addendum-bill.json` → two
   metered lines summing to `$0.3579`, plus the three seats that bill nothing
   with their counts. The two files add and never overlap.

## What is NOT in here, and why

**The sealed bundles themselves.** `seal-manifest.json` carries their full
sha256s, byte counts and the verbatim user line every arm answered — so a holder
of those bytes can prove they hold the bytes this round used. The bundles are
not published, because each one carries an unpublished persona block: the
interior of a living game's characters, which is the product and not the
exhibit. That is a real limit on what "sealed by hash" buys a stranger, and it
is stated here rather than left to be discovered: **the manifest proves
integrity, it does not let you reproduce the round.**

**The judge letter map.** The sheets were read blind against shuffled letters
from a recorded seed. The scored verdicts are published against arm names in
`scores.json`; the letter map is not, because publishing it would retire the
sealed sheets for any future re-judging.

**The sealed per-reply records** (`results/samples/*`). One file per arm per act,
holding each reply's spoken line, its extraction tier and both clocks. They are
the sheets a future round would re-judge, and a published record is a retired
one — so they stay sealed. Every line they hold is printed verbatim on the
exhibit page, in the reply cards and in each act's fold.

The **addendum** arm's records are the one exception, and they publish
(`addendum-samples.json`). The principle has not changed: at six records the
page already prints every line verbatim, so withholding the file would withhold
nothing a reader cannot already read — and publishing it lets them check those
lines against the clocks and the wire hashes that produced them. The sealed
twenty are a different matter and stay back.

## The post-publication addendum (2026-08-16)

Two files in this kit belong to an addendum added after the exhibit was
published: `addendum-scores.json` and `addendum-samples.json`. They cover a
twenty-first arm — `qwen3.8:27b`, whose weights shipped the day the round was
sealed — run afterwards on the same four frozen asks.

It is a **separate round** (`opencall-addendum-qwen38`), not an extension of
this one, and the distinction is load-bearing rather than pedantic:

* The sealed twenty are **unchanged**. Every file above keeps the bytes and the
  sha256 it was first published with; the addendum is an addition to this kit,
  never a new edition of it.
* Blinding letters cannot be shared between a 20-arm build and a 21-arm one —
  they come from a seeded shuffle over the whole reply list — so the addendum's
  figure is **not** comparable letter-for-letter with the twenty. What bridges
  the two rounds is the **anchor**: one calibration reply, the same text in
  both, scored by the same seats. `addendum-scores.json → anchor_calibration`
  holds its side; `scores.json` holds the round's. The difference between them
  is each seat's zero moving, and the page prints the arm's figure both raw and
  with that drift subtracted.
* The addendum panel seated **five of the seven** judge seats, so its figure is
  read against this report's leave-one-family-out *minus anthropic* column and
  never against the headline one.
* **The blindness figure is this round's own, and it is not a zero.** The sealed
  round's "0 of 868 cells claimed to recognise a system" is a statement about the
  sealed round. The addendum sheets carried the same question and
  `addendum-scores.json → self_disclosure` records **4 of 50** cells claimed —
  all four from one seat, all four landing on the **anchor**, which scores no
  arm, and none of them naming a family. The arm's own thirty scored cells are
  clean. The page prints that beside the sealed round's zero rather than letting
  the zero read as a claim about both.
* **This addendum has its own bill, and it was owed for a few hours.** The page
  prints `$0.3579` for the two seats that metered on this round, and
  `addendum-bill.json` is what makes that re-derive: a line per judge seat with
  the token counts it was billed on, the rate they were priced at, and
  `usd_re_derived_from_these_tokens` — the same shape `bill.json` publishes for
  the sealed round. The history is worth keeping, because it is the shape of the
  fault rather than the fault itself: the addendum published on 2026-08-16 with
  that figure on the page and no token counts behind it anywhere in this kit,
  since `bill.json` is scoped to `opencall-r2` and was never reopened. A
  content-accuracy pass that same day marked it as the one figure a reader could
  not check rather than quietly leaving it, and the file landed the same evening.
  It is a **join, not a re-run** — every count in it was recorded when the round
  ran, and no model was called to produce it.
* **The two bills add; they do not overlap.** `bill.json` covers the sealed
  round's $2.2826 and is unchanged by any of this; `addendum-bill.json` covers
  this round's $0.3579. No call appears on both, and the addendum bill publishes
  the sum so a reader who wants the whole exhibit's cost — $2.6405 — does not
  have to guess whether adding them double-counts.
* **The three seats that bill nothing are on that file too**, with their token
  counts. Two are plan-included and one is local, which is what the page's own
  sentence claims; a bill covering only the metered pair would have left the
  other three asserted rather than shown.

**The harness ledgers** (`LEDGER-*.json`) and **the judge run files**
(`RUN-*.json`). The round's per-leg token counts, in files that also carry
internal filesystem paths — they are working artifacts of the harness rather
than of the round. The counts themselves are not withheld: the metered arms' are
joined into `bill.json` and the addendum round's judge seats into
`addendum-bill.json`, both with every dollar re-derived from them. It is the
files that stay back, never the figures they carry.

Together those two are where the files behind the replies, both clocks and most
arms' token counts sit. This kit publishes the figures and withholds the files,
and that is the largest gap between what this page prints and what a reader can
recompute from the kit alone.

## The scrub, declared

Seven of these files carry **path-scrub** substitutions: filesystem paths, one
daemon configuration flag, and a human operator's name — which this hub prints
as a role, never a person. They are `scores.json`, `report.md`, `known-world-issues.md`, `seats.json`, `addendum-scores.json`, `addendum-runtime-pins.json` and `addendum-probe.json`. Every substitution is listed in
`index.json` per file, with the reason it exists, and a file whose bytes were not
touched carries `"scrub": null`. A hash here is a hash of the bytes **in this
directory**, which is why the difference is declared rather than quietly
absorbed.

Nothing measured was changed by the scrub. No figure, no verdict, no note, no
denominator.

## The corrections, annotated in place

Two of the copied files carry a bracketed **correction** inserted where a
claim in them is wrong, with the original line left standing beside it:
`report.md`'s summary line subtracts the recused cells and forgets the 28
calibration anchors, so it reads *766 scoring cells* where every table under it
reads 738; and `known-world-issues.md` inherits a *58 open rows* headline from a
source worklist that is not published, where the list it actually prints holds a
different, countable number of rows. Each annotation is listed in `index.json`
against the file it touched, with the reason it exists, and each one is a count a
reader can redo from the same bytes. A kit that quietly disagreed with the page
it backs would be worse than either figure being wrong.

## Limits, in one place

- **n is small and one-world.** Three scenarios, four judged asks, n=2 on the
  kid's act and n=1 elsewhere, from a single draw of a single world.
- **The author is a contestant.** These scenarios were written by a model that
  sits in the round as an arm and holds one of the seven judge seats. Recusal
  cures the scoring — the two Anthropic seats scored no Claude arm — and cures
  nothing about who chose the questions.
- **Decoding parameters were not pinned.** Each transport's defaults were used
  and are not recorded in the artifacts. Temperature, top-p and sampling seed
  are therefore not reproducible from this kit, and no figure here should be
  read as a settings-controlled comparison.
- **Judges are LLMs, not people**, and one of the seven ran locally. Their zeros
  sit in seven different places, which `seats.json` measures rather than
  assumes.
- **Not a standard.** Our own round, our own world, our own judges. Not first
  independent numbers, and not a benchmark.

## Contamination caveat

Published 2026-08. The scenario prose was written fresh for this round and
sealed before any arm ran, but this world's earlier public exhibits exist, and
models with a later training cutoff may have read them — or may read these
files. We author fresh sets each cycle; this one is not a standard, it is our
kit, yours to reuse.

## Receipts

Files as published (sha256 of the bytes in this directory):

    scores.json             722e04db3ae473e2fdffb9129eb1d7bd2089a59e3548317d1dad25eabec350ab
    seal-manifest.json      05cdb3d5deb6a3f82e0c188b3be45688fb15fd7635d6ea3277589847f61a78b5
    report.md               50fce08f56fa027f8cb2cedc46b496be539e03728862efdd60bf0a3f1472a340
    known-world-issues.md   dfb49b79a8ed758ab3e2ffe127d8ed8f165fb384450ba75ce9019467a9a5001a
    counting-rules.json     4c394f718e7b8c6069703e6cd26feb4ba17c62ae574f25aeeff387f8c8feaae6
    seats.json              4c14d999cdee42ac0edefd8eea4738b6969c82ab9218dff2acbc50c4ddaec6d4
    bill.json               dd5acbdaf2b5c60fdbe6ce519429207bcb1943002e3c0e3b2ad628ebba980767
    addendum-scores.json    c68a603bd6b8c90d14a97610681f5fbc174123184718361c4ec8550f99d6678e
    addendum-samples.json   eedf665f7f317f73b10c3d169de0300ec640261d3965713e95a3b76d7db98830
    addendum-runtime-pins.json  4ef9f335ef4ad17aacda84dc1ed380afb6e0c35427c07d7699eca95c1e192fcb
    addendum-probe.json     c86923cd8148465576a8fec9c3c17d06b9639102ead5c352d2e25d907ecdf1d8
    addendum-bill.json      7815061fc9cc7152505c8307d235a8e69cc735916c973e1b622ce0eb0eff2bcf
