# Data kit — Ten Minutes with Living Artists

Everything the article claims, with the artifact behind it. Published 2026-09-03 (UTC).
**The records in this directory are CC BY 4.0** — the hub's usual licence for a data kit.
**The audio is not**: every clip is share-alike, because the corpus behind the adapter was,
and the clips are therefore filed beside the article at `../audio/` with their own
`MANIFEST.sha256` rather than in here. Two licences, two directories, on purpose.

The article's own promise is the contents list:

> the full fingerprints of every clip and the map from each player to its file, the request
> payloads verbatim, the training logs' pass lines for both runs, the intake manifest with
> its 208 checksums and the 41 licence records copied from the archive's own metadata, the
> bandwidth audit with its flagged rows, the loudness, peak and flat-factor records for
> every clip in all four rounds, the sealed pre-registration of the coherence bench with its
> five filed amendments — two of which correct the registration itself — and its results down
> to every embedding, and the Round 4 records and blind-sheet layout (the key stays sealed
> until a verdict is recorded)

All of it is here except the blind key, which is the one thing that is deliberately absent —
read "What is NOT here, and why" before you read anything else.

## The map from player to file

The clips play from `../audio/`, one directory up. Each row is the exact file the player
streams; the page prints the first sixteen characters of the digest beside it, and the
digests below were re-read off the staged bytes rather than copied from a note.

| player | file | full sha256 | what it is |
|---|---|---|---|
| Listen 1 | `mnml-prose-1-base.wav` | `d6b62483885e02af9bc76ad964605e46903d32e45cfdb9c715c6b373b2b21222` | request A, nothing added |
| Listen 2 | `mnml-prose-1-adapter.wav` | `b1627a2f09727e89bb040fbfe069626d3078be974ecb61dca293264496c275bd` | request A, the adapter at 1.0 |
| Listen 3 | `mnml-shaped-1-base.wav` | `e71b8c65156cb10c4e64c55dc50cac56e8d0bf8925a8c2839c54ca7338da137c` | Schwerelos (original mix), nothing added |
| Listen 4 | `mnml-shaped-1-adapter.wav` | `0327d06999713dce12739e85936c6492bc91aacac2025ca7ad7af66057458ab0` | Schwerelos, the adapter at 1.0 |
| Listen 5 | `mnml-shaped-2-base.wav` | `5889ca518f8b726e74176d831975eff712f05b878f351d74b70a13a8268068ef` | Basement Loop 7, nothing added |
| Listen 6 | `mnml-shaped-2-adapter.wav` | `0cecc9a342b5bc04aac2e5a56ca4a820d0bffc63816f8a2cf8124cb7862bb415` | Basement Loop 7, the adapter at 1.0 |
| Listen 7 | `mnml-shaped-1-control.wav` | `5db77e0345b032c5c8c5dcf83c4449b4383748d3c8c3121e077ff4d906533cbd` | Schwerelos, the one-artist adapter at 1.0 |
| Listen 8 | `mnml-shaped-2-control.wav` | `f0ebb8d258039c5c233b9879ef4e52af948736b92a4bbd5044fd69730c895594` | Basement Loop 7, the one-artist adapter at 1.0 |
| Listen 9 | `r4-dub-techno-base.wav` | `6f0b2d81c610b7e9f849bef18ae57ab784d4984b9f6ab84064bb21ad95b40b65` | dub techno, nothing added |
| Listen 10 | `r4-dub-techno-adapter.wav` | `31f288228c7aab86498526b35ced0390f66a71e4186f97730ec551e7607e2787` | dub techno, our adapter at 1.0 |

To check them yourself, from `../audio/`:

    sha256sum -c MANIFEST.sha256

Every line should print `OK`. `clips.json` carries the same digests together with the
request each clip answered — caption, seed, key, tempo, step count, guidance, the
normalisation setting — and the loudness and raw-peak measurements taken on it.

## The files

| file | what it is |
|---|---|
| `MANIFEST.sha256` | Every kit file's sha256 in `sha256sum -c` form, over the bytes published here — provenance.json included. It does not list itself, index.json or index.html; index.json carries its digest instead. |
| `PREREG-COHERENCE.md` | The coherence bench's pre-registration, sealed before the first embedding, with its five filed amendments — two of which correct the registration itself (A-C3 measures a registered disjointness claim FALSE; A-C4 corrects the seal's own clock). |
| `PREREG-ROUND4.md` | Round 4's pre-registration, written before the first clip of it existed: the twelve captions, the four arms, one seed per prompt, the settings pinned, and what would count as a result. |
| `README.md` | What each file is, the player-to-file map with full fingerprints, what was projected out of every file and why, what is NOT here, and the licence split between the records and the audio. |
| `bandwidth-audit.json` | The bandwidth instrument's output for all 208 files: measured ceiling, roll-off, cliff and steepness, and the cutoff verdict — including the 20 rows flagged PREMATURE LOSSY CUTOFF. |
| `clip-peaks-raw.tsv` | The clipping diagnostic, one row per RAW render across all four rounds (65): sample peak, ffmpeg astats flat factor, peak count, integrated loudness, and the first sixteen characters of each render's sha256. The join key is (round, clip_id), because one clip id was rendered in two rounds. |
| `clips.json` | The ten clips the article's players stream: full sha256 re-read off the staged bytes, byte count, the round and render each came from, the request as the engine received it (caption, seed, key, tempo, steps, guidance, normalisation), the loudness the harness recorded, and the raw sample-peak and flat-factor row measured for D-20260903-08. |
| `coherence/RESULTS.json` | The same results as data. |
| `coherence/RESULTS.md` | The coherence bench's results in prose and tables: both metrics, the intervals, the predictions scored against them, and what the run does not license anyone to say. |
| `coherence/RESULTS.sha256` | The bench's OWN receipt over its twelve output files, computed on the run's unprojected bytes. Files this kit projected will not match it; MANIFEST.sha256 covers the published bytes, and the README names which files differ. |
| `coherence/RUN.log` | The run's transcript, step by step, with the timings and the two boxes' cross-checked crop digests. |
| `coherence/crop_windows.sh` | The cropping tool the bench shells out to. |
| `coherence/crops.csv` | Every crop the bench took, with its source file, window, offset, digest and duration. |
| `coherence/embeddings-index.csv` | The row order of embeddings.npz — which crop is which vector. |
| `coherence/embeddings.npz` | Every embedding the bench computed, as written. The results are re-derivable from this file and the index beside it. |
| `coherence/metric1-corpus-spread.csv` | Metric 1: the per-corpus spread estimates at equal N. |
| `coherence/metric1-draws.csv` | Every subsample draw behind Metric 1's intervals. |
| `coherence/metric2-pairs.csv` | Metric 2: each base/adapter render pair and its distance to the corpus centroid. |
| `coherence/metric2-render-centroid.csv` | Metric 2's per-render distances, one row per render. |
| `coherence/run_bench.py` | The bench itself, verbatim: the corpora it reads, the crop policy, the embedder pin, both metrics and the interval procedure. |
| `corpus-manifest.json` | The intake manifest, one row per acquired file (208): source URL, the licence URL recorded on the archive item, tier, creator string, measured format, duration, loudness and bandwidth ceiling. |
| `corpus-manifest.sha256` | The 208 checksums, in `sha256sum -c` form, relative to the corpus audio directory. |
| `corpus-provenance.md` | How the corpus was assembled and what was measured on it: the two search pools, the licence-path match, the integrity check, the bandwidth audit and its verdicts, the exclusions and the sentence written before the result. |
| `credits.md` | The attribution owed, release by release — the 41 releases, their archive items, their creator strings as recorded, and their licence tiers. |
| `cross-arm-check-round1.json` | Round 1's byte-distinctness checks: every arm pair at the same caption and seed, with each file's digest prefix. |
| `cross-arm-check-round2.json` | Round 2's byte-distinctness checks, over every pair of the round's seven clips. |
| `exclusions.json` | The eight tracks excluded before training and why: a container-rate exclusion decided per item, with the ruling's own words and every excluded track named. |
| `freeze-receipt.json` | The freeze's own receipt: 208 in, 8 excluded, 200 post-exclusion, 41 held out, 159 trained, and the self-sha256 of both manifests it wrote. |
| `intake-receipt.md` | The download receipt, frozen at acquisition time: what was fetched, when, with which user agent, and what each response carried. |
| `records-round1.json` | Round 1's six renders, one record each, exactly as the render harness wrote them: the request, the adapter and its strength, the seed used, the output sha256, timings, VRAM and the loudness measurement. |
| `records-round2.json` | Round 2's seven renders, same shape. |
| `records-round3.json` | Round 3's four renders — the one-artist control — plus the cross-round determinism check that reproduced two Round 2 renders byte-for-byte. |
| `records-round4.json` | Round 4's forty-nine renders (48 clips plus the determinism repeat): twelve captions, four arms each, with the prompt id, style, arm, dose and the caption as sent. |
| `round4-blind-sheet.json` | The blind sheet's LAYOUT — pair ids, styles, contests and the alpha/beta labels. It carries no arm names. The key that maps label to arm is NOT in this kit and stays sealed until a verdict is recorded. |
| `round4-captions-as-rendered.txt` | Every Round 4 caption exactly as it was sent, one line per (prompt, arm) — including the trigger word the preprocessor prepends, which is the difference between the arms. |
| `s-per-step.md` | The two instruments that disagree about seconds-per-step, both readings, and the counting rule for each — which denominator, what it includes, and which one the lab's ledger marks authoritative. |
| `split-fmctl0.json` | The control run's split, written by the same script under the same rule. |
| `split-mnml0.json` | The wide run's dataset split as the freeze script wrote it: the holdout rule verbatim, the caption rule and its PLACEHOLDER status, the per-tier walk that chose the held-back items, and the 159 training and 41 held-out work keys by name. |
| `split-rule.py` | The holdout rule as an executable pure function — the one implementation the freeze script and the dry run share, so the split cannot disagree with itself. |
| `train-fmctl0.log` | The one-artist control's training run, the same trainer, the same recipe, its own log start to finish — 1m 33s. Same projections. |
| `train-mnml0.log` | The wide corpus's training run, the trainer's own log start to finish: the configuration it printed for itself, ten per-pass lines, both checkpoint writes, and the Training Complete banner carrying 10m 08s. Filesystem paths projected; the card named by class. |
| `unload-restore-control.json` | The control that found the contamination: does attaching and then unloading an adapter restore the base decoder bit-exactly? It does not, and the record says which clip that ruined. |
| `index.json` · `provenance.json` · `index.html` | The kit's own manifest, its projection record, and the browsable listing generated from this directory. None of the three lists itself. |

## What is NOT here, and why

**The Round 4 blind key.** `round4-blind-sheet.json` gives the sheet's layout — twelve
pairs, their styles, which contest each pair runs, the shuffle seed 4104, and the
`alpha` / `beta` labels. The file that says *which label is which arm* is not in this kit
and will not be until a verdict is recorded. The article says so in its own text. A blind
sheet published with its key is not a blind sheet, and this is the one absence here that
exists to protect a measurement rather than a machine.

**The adapter weights.** Two 84 MiB LoRA files, one per run, are not published. Both are
identified by sha256 inside the round records, so a reader can tell which adapter rendered
which clip even without the bytes.

**The corpus audio.** 208 recordings by other people are not redistributed here. What is
published instead is complete: every source URL, the licence URL recorded on each archive
item, the creator string as recorded, and a sha256 per file — `corpus-manifest.json` and
`corpus-manifest.sha256`, 208 rows each.

**The clip audio.** Share-alike, and therefore beside the article rather than in this
CC BY 4.0 directory. Every fingerprint is in `clips.json` and in `../audio/MANIFEST.sha256`.

**The model, the trainer and the embedder.** Named, linked and licence-checked in the
article's colophon; none of them is vendored here.

**A listening result for the one-artist control.** None exists as this is written. The
article says so twice, and no number in this kit stands in for it.

## What was projected out, and why

Every file here is published as recorded, with one class of change: **this house does not
publish the shape of its own machines.** Nothing measured was touched — every loss,
duration, digest, byte count, loudness figure, peak, flat factor and interval is as the
instrument wrote it. `provenance.json` lists, per file, which rules fired and how many
times.

- **Filesystem paths → role-relative paths.** An absolute path becomes the path it names
  (`music/…`, `estate/…`). This is the rule that fires most: the render records alone carry
  one output path per clip.
- **Machine names → the role each played.** *The training box* is where both adapters
  trained and all 65 clips rendered. *The evaluation box* is where the coherence bench
  computed its embeddings. Also *the public host* and *the laptop*.
- **The training card → its class.** It publishes as *a 96 GB workstation-class card*,
  never by model name. The two consumer cards in the evaluation box publish by name — as
  earlier exhibits on this shelf do — because they are hardware a reader might own and the
  page's argument is partly about that.
- **GPU serial numbers withheld.** The runbooks pin a card by UUID. That a card was pinned
  is published; the UUID is not — it lets nobody check anything and it fingerprints a
  machine.
- **One timestamp restated in UTC.** Amendment A-C4 of the coherence pre-registration
  recorded a commit time with a local offset. The instant is published in UTC, which is the
  only clock this house publishes.

Two notes so nothing reads as sleight of hand.

**The bare timestamps in both training logs are UTC.** The training box runs `Etc/UTC` and
the trainer writes stamps with no zone marker. They are left exactly as written, because
editing a timestamp inside a verbatim log is worse than explaining it.

**`coherence/RESULTS.sha256` is the bench's receipt, not this kit's.** It was computed over
the run's own unprojected bytes. Six of the files it lists were projected for publication —
`RESULTS.md`, `RUN.log`, `run_bench.py`, `crops.csv`, `embeddings-index.csv` and
`metric2-render-centroid.csv` — and will not match it. The other six will. Running

    cd coherence && sha256sum -c RESULTS.sha256

prints exactly those six as `FAILED` and the other six as `OK`; that is the disclosure, not
a defect. `MANIFEST.sha256` in this directory covers what is actually published, and
`provenance.json` says, per file, which rules fired and how many times.

**The coherence pre-registration's seal is in the same position.** `PREREG-COHERENCE.md`
§0.4 carries `sha256 = 4bbe9548945d1bd849b355a080b49c481c067e21fd105e98d8d6d30c9f2fe5f2`,
computed over the registered file with the seal line itself still reading its placeholder.
The copy published here is projected and does not hash to it. What the seal buys a reader
is that the registration preceded the run: the seal commit is stamped 2026-09-03T03:08:03Z
and the first embedding began at 03:19:28Z — eleven minutes later, and `coherence/RUN.log`
carries that second stamp independently.

## One thing in this kit refuses to be tidy, on purpose

**The four rounds spell their loudness fields two different ways, and both spellings are
published as written.** Rounds 1 and 2 write `lufs_integrated` / `lufs_range` /
`true_peak_dbtp` / `threshold`; Round 3 writes `input_i_lufs` / `input_lra` /
`input_tp_dbtp` / `input_thresh_lufs`. They are the same four measurements from the same
tool called two ways. Normalising them here would be a small kindness that hid a real
trap — a fold that reads one spelling silently drops clips instead of failing — so the
records keep the harness's own words and `clips.json` says so in its header.

The same principle governs `clip-peaks-raw.tsv`. Its join key is `(round, clip_id)` and not
the id alone, because `base__C__seed777` was rendered in Round 2 and again in Round 3 and
came back **byte-identical** both times. That is the determinism control, recorded in
`records-round3.json`; it also means one digest maps to two records, and a table keyed on
the id would quietly lose one.

## The clipping diagnostic, in one line

`clip-peaks-raw.tsv`: **0 of 65 exceed the engine's −1 dBFS ceiling; 18 of 65 carry
flat-topped stretches (flat factor > 0), the model's own ceiling scaled down by the engine's
normalisation, untouched renders included.** Every raw render's sample peak reads
−0.999589 dBFS — the ceiling, to the sample — and integrated loudness across the 65 spans
−19.6 to −13.2 LUFS. Three of the eighteen flat-topped renders are base renders with no
adapter attached at all, which is why the finding is about the model and not about the
adapter.
