# The August arrivals — data kit

The machine-readable companions to **exhibit five, "The August arrivals"**
(research.strata2signal.com/august-arrivals/index.html): the fresh class of
August 2026 sitting two house instruments — the judge seat trial, frozen in July
and run verbatim, and the narrator's voice probe, which was **rebuilt in the
original's shape** because the original's five questions were never persisted —
plus the residency census that measures what every arm on the roster costs to
keep loaded. (Corrected 2026-08-16: this sentence read "two frozen house
instruments", which is true of one of them. `voice-rows.json` has recorded the
rebuild since it was written, under `scale.label_registered_PREREG_B_7` and
`outcome`, and the exhibit page's body and limits both say so; the page's own
opening line carried the same error and was corrected in the same pass.)

**Licence: CC BY 4.0.** Take these rows, re-plot them, check our arithmetic,
publish what you find. Attribution: strata→signal research,
research.strata2signal.com. If you find an error in them, we want to hear about
it: hello@strata2signal.com.

## What is in here

| file | what it holds |
|---|---|
| `seat-rows.json` | the four fresh-class rows on the judge seat trial: kills, keeps, verdicts, warm latency, the per-class kill ledger, the response-failure ledger by kind, self-consistency, floors, counting rules, deviations, and the run-level provenance hashes |
| `voice-rows.json` | the six arms of the narrator's fourth table: per-panel and combined scores on all six dimensions, per-sample spread, the envelope/JSON ledger, the abstention adjudication with per-judge votes, the blind letter map, judge model ids, inter-panel agreement, and the prompt hash |
| `residency.json` | all seventeen registered rows of the VRAM residency census: per (model, context length) the loaded `size_vram` in bytes and GiB, the load seconds, the outcome, the daemon configuration the figures were allocated under, the counting rules, the open finding on the three drafter rows, and the provenance hashes |
| `index.json` | this directory, listed: file names, sizes, sha256s, and what each one answers |

Both *row* files come from the **same scored records** that back the addendum
tables on exhibits two and three — one scored record, three surfaces, no
re-derivation in between. `voice-rows.json` is byte-identical to exhibit three's
addendum. `seat-rows.json` is not, and the difference is additive: it holds the
same four fresh-class rows byte for byte, and then carries exhibit two's twelve
July rows after them for side-by-side reading (see `carried_rows` in the file,
which says so itself). The seat-trials copy carries the four only. Siblings from
one record, in other words, rather than twins. `residency.json` is this exhibit's
own, and has no counterpart elsewhere: the census ran once, for this page.
`schema` on every file reads `s2s-bench-v1`.

## Receipts

Files as published (sha256 of the bytes in this directory):

    seat-rows.json    aed07dd3174150ae0bd97a4128f878661c9924acadfd20918eb4c8b65ccfbb26
    voice-rows.json   6fa1aeae38bc8f1b9fc116bc31e7feda632b70861c8ee8cfb7f01266af7b8835
    residency.json    1a6c7acf5790f59bfe21796ad2c5ba6dee3e3f9358682cadb664d8bb5e819617
These are the bytes served today. They are not the hashes this kit published on
2026-08-12: `seat-rows.json` grew the carried July field later that day, and all
three files had a `status` line rewritten from STAGED to PUBLISHED on 2026-08-13
when the kit went live. Both edits were real and neither was re-stamped, so the
receipts above went stale while the files themselves stayed honest. Re-stamped
from disk 2026-08-15. Nothing measured changed in any of the three.

Re-stamped again 2026-08-16, and for a smaller reason: all three files carried
`licence` and none carried `attribution`, so a reader who took a file on its own
had the terms without the name they attach to. One key was added to each, beside
the licence it belongs to, and the hashes above follow. `voice-rows.json` is
byte-identical to exhibit three's addendum and that file was changed in the same
pass, so the two are still twins — the claim is checkable rather than asserted:
both hash to `6fa1aeae…` above. Nothing measured changed in any of the three.

Instruments, pinned before the first scored call and re-derived from disk
(artifact names only — no filesystem paths, by house rule):

    judge-cases-v1.json (answer key, 43 cases)      6a226b62dceb005dbd658bc13cf229e38f6f0a2f07ae51091f444ec3e96b56d1
    THRESHOLDS.md (both floors)                     37af75ffa18484a969528e894ef16f7fdcd10eb364e0669ac4f9333ce5279ba0
    judge-prompt-packs-v2-prompted.md               92af6030dddface2c863fe9d672bec6a6bb1017122c0028e435214acd9fb298d
    PREREG-A-addendum.md (seat leg)                 305dce4679f6663c09c955396cf6d9fef1b4c59ba10ee6f4b961775afae6aeb3
    PREREG-B.md (voice leg)                         3773885ef02d82e63f1145df1b6e7e485520de149dfef137a4e7664eae68ae3f
    brisa-5.json (voice item golden, 5 items)       fee89c103c4f50f5169df0577a37435492ad196f49f39c96a17a7f824d62b55f
    voice_score.py (voice scorer)                   11cdafcb95ccd6551761eb4c210554862ca951c4b964efdc00a2dfb1734618e4
    judge_fan.py (blind judging fan)                c006d6430d309501dc17689c2818c511974cd6e1ba638ef051141e49ee2d4b46
    voice_gates.py (envelope gate)                  dc2aafb85011d9719a7849cd8b79bf04fede671fe9df71c9bb851e8337be84aa
    RECONCILIATION-CHECKLIST.md                     a2768ee8d645960918fa20a4e31db2afbd9f0ca951e1240d70de8a59044ca7d8
    PREREG-C7-census.md (residency leg)             3ea13e0676300032c6ca31849c84eace65108e74b1f11be3b49ed6123350afc0
    c7-roster.json (the 17 registered rows)         d6f67db93acf55f10b941aab3a64ee46d97df2881bb032891daac779be586844
    vram_census.py (the census)                     0c87af2760c790f0fd8b71e3da7d77845c669e414863466508fd5b1ba378bcde
    vram-census.json (the census record)            8c53028e8b22a6e1b76c70aa7c208939f3f1e0f8974eda852db2920a55ebf130

Each row file carries the same hashes inline under `provenance`, plus its
`counting_rules`, its `deviations`, and its `limits`. The page's tables are a
projection of these files and nothing else: the seat table renders `rows[].model`,
`kills_27`, `keeps_16`, `verdict`, `warm_p50_ms`, and the response-failure ledger
under `rows[].detail.response_failures` (the `truncated` column is `done_reason`;
the `other kinds` column is `envelope_json + content_json + schema` plus
`detail.transport_failures.total`); the voice table renders `rows[].page_cells`
plus `rows[].envelope.outcome`; the residency table renders `rows[].model`,
`params_and_quant`, `ctx`, `loaded_vram_gib`, `load_s` and `note`, and its six
columns are listed under `columns` in that file. Regenerating the page from these
files is meant to reproduce it exactly — if it does not, that is our bug.

One caveat that belongs at the top of `residency.json` and not buried in it:
three rows carry a figure we do not yet trust, and the file says so under
`open_findings` rather than dropping them. The two drafter tags report a tenth of
what the same weights cost without a drafter, and two different quantizations of
them report a byte-identical allocation — read that finding before quoting those
three rows anywhere. Every other row is an independent load and is unaffected.

## The contamination caveat

Published 2026-08-12. Models with a later training cutoff may have seen these
sets. We author fresh sets each cycle; this one is not a standard, it is our kit,
yours to reuse.

## How to reuse this

Start from `counting_rules` in whichever file you are reading — every figure in
it is a count of rows in an append-only record or a nearest-rank percentile over
those rows, and the rules say exactly which rows were counted, which were left in
the denominator, and which never became a verdict. Read `limits` before quoting a
number, and honour the two fences that make these figures mean anything: the seat
figures are UNMEASURABLE verdicts about a serving stack at one configuration, not
rankings of model judgment, and the voice figures live on their own scale — the
anchor model reads 9.0, 7.0 and 6.2 on three different rounds of the same probe,
unchanged, so no two numbers from different tables may be subtracted. Within a
single table, compare rows freely; that is what the table is for. These are our
own trials, on our own hardware, for our own chairs — not first independent
numbers, and not a benchmark anyone should treat as a standard.
