# PRE-REGISTRATION — THE CORPUS-COHERENCE BENCH

**Status: DRAFT. Not registered. Nothing has been embedded.**
Written 2026-09-03T02:5xZ by the `coherence-prereg-draft` recon lane, under ruling **D-20260903-02**.
Ratification is the orchestrator's; every open point is collected in **§12** and marked `⧖ RATIFY` where it appears.

**Registration becomes real when**: (1) the orchestrator resolves §12, (2) this file is sha256-sealed and the seal written into §0.4, (3) row 25 of `PREREG-LEDGER-OF-LEDGERS.md` is added **in the same stroke** (§11), and (4) only then does the first embedding run.

---

## §0 — WHAT THIS FILE IS, FOR A READER WHO ARRIVED COLD

### 0.1 The one-paragraph version

Five music corpora were used to train five LoRA style adapters for a text-to-music model. One of them — a 159-track "minimal techno" set assembled from 41 Creative-Commons releases by dozens of artists — produced an adapter the operator's blind ear **rejected**: the un-adapted base model won both A/B pairs. A one-artist control corpus was then trained under the identical recipe to test whether **corpus coherence** was the reason. This document registers, *before any measurement is taken*, how coherence will be measured, what is predicted, and what the null would mean — so that the numbers cannot be chosen after their answer is known.

### 0.2 The governing ruling, verbatim

From `estate/DECISIONS.md`, **D-20260903-02**:

> THE CORPUS-COHERENCE BENCH IS PRE-REGISTERED BEFORE IT RUNS — CLAP-embedding spread per corpus (Chopin · Sousa · Bach · the 41-release mnml slice · the one-artist control), equal-N subsamples, metric + prediction + null written at bench/ace-mnml-2026-09/PREREG-COHERENCE.md before the first embedding; a second pre-registered arm = render-to-corpus-centroid distance (base vs adapter arms of the existing Round 1–3 clips: did the adapter move its output toward its corpus at all); the numbers ride the article beside the ear's verdict, MEASURED never gating (D-20260831-17); a null publishes as a null; the embedder's licence read from its LICENSE file before adoption; runs on whichever box holds the corpora, never on a live seat's card

### 0.3 The standing constraint that outranks every number below

**MEASURED, NEVER GATING** (D-20260831-17). No result in this bench can pass, fail, promote, or retire an adapter. The instrument of record for whether an adapter is any good is **the operator's blind ear**. These numbers ride *beside* that verdict and never in front of it.

### 0.4 Seal

`sha256(this file at registration) = 4bbe9548945d1bd849b355a080b49c481c067e21fd105e98d8d6d30c9f2fe5f2` — computed with this line still reading `⧖ TO BE FILLED AT SEAL` (see §13); sealed 2026-09-03 ~03:5xZ by the orchestrator
Timebase: UTC throughout.

### 0.5 What in this file is DRAFT

Everything is draft until §12 is resolved. But the parts that are **substantive proposals rather than recon findings** — the ones an orchestrator should actually weigh — are marked inline as **`⧖ RATIFY-n`** and listed in §12. Measured facts (counts, hashes, versions, licences) are not proposals and carry their source instead.

---

## §1 — PURPOSE, AND THE TWO QUESTIONS

### 1.1 Why this bench exists

The minimal-techno corpus was assembled by matching an **archive.org subject tag**. A tag is a folksonomy, not a sound: it records what uploaders across twenty years of netlabel culture chose to call their releases. The corpus that results may be internally incoherent — many artists, many eras, many production chains, many mastering decisions — in a way a single-artist corpus is not. If a LoRA cannot learn a style from an incoherent corpus, that is a **useful negative result about corpus construction**, not merely a failed run.

The control already exists and was built for exactly this. From `music/corpus/fm-control/PROVENANCE.md`:

> It exists to be a **CONTROL**. … one artist, the same pipeline, the same recipe, the same era, the same production and the same licence tier as the mnml v1 adapter — so that **corpus coherence is the only variable that moves**. If v1 and this control sound different, coherence is why; that is the whole design.

What has been missing is a **number for coherence**. This bench supplies one, in one named embedding space, with its limits stated.

### 1.2 Question 1 — is the wide corpus actually less coherent?

*Do the training tracks of the minimal-techno corpus sit further apart, in CLAP embedding space, than the training tracks of the one-artist control — by more than the uncertainty of the estimate, at equal sample size?*

If the answer is no, the coherence explanation for the ear's rejection loses its central premise, and the article says so.

### 1.3 Question 2 — did the adapter move its output toward its corpus at all?

*For every base/adapter render pair at the same caption and seed, is the adapter's render closer to its own training corpus's centroid than the base render is?*

This is the sharper question, and it is the one that can be answered with the clips that already exist. An adapter that changed the audio (they are not byte-identical; the LUFS differ) but did **not** move it toward its corpus learned something other than its corpus. That is a publishable finding either way.

---

## §2 — THE FIVE CORPORA, FROZEN

### 2.1 The unit of analysis: the TRAINING set, not the corpus directory

**`⧖ RATIFY-1`.** The registered primary unit is the **frozen training set** of each run — the `samples` array of that run's `dataset.json` — because the question is what each adapter *saw*. The corpus directory is a superset (it includes held-out and excluded tracks). A secondary, clearly-labelled arm repeats METRIC 1 over the full corpus directories as a sensitivity check.

This choice has a consequence the orchestrator must accept knowingly: **N = 19**, from Chopin's training set, not 24.

### 2.2 The table, measured

| corpus | run id | **train n** | corpus files | distinct releases (train / corpus) | licence tiers |
|---|---|---:|---:|---|---|
| Chopin (`pd-shakedown`) | `20260829T162506Z-stage0` | **19** | 24 | — / — (one collection) | CC0 (Musopen dedication) |
| one-artist control (`fm-control`) | `20260831T193904Z-fmctl0` | **24** | 30 | 8 / 10 | BY-SA 3.0 ×18, BY-SA 4.0 ×12 |
| Bach (`bach-shakedown`) | `20260830T104030Z-bach0` | **63** | 79 | 2 / 2 | CC0 ×31, PD-mark ×48 |
| Sousa (`sousa-shakedown`) | `20260830T004943Z-sousa0` | **84** | 106 | — / 106 | PD-mark ×106 |
| minimal techno (`mnml-shakedown`) | `20260831T023450Z-mnml0` | **159** | 208 | **27 / 41** | CC0 51, BY 58, BY-SA 50 (train) |

Total tracks to embed: 19 + 24 + 63 + 84 + 159 = **349**.

### 2.3 The input hashes — what "frozen" means here

The bench reads these files and no others. If any hash differs at run time, the run aborts.

```
d26114c0d3828f1b2aaab947db6e676fc5a617cfbe3def62508c1924b8ae8333  mnml-shakedown/20260831T023450Z-mnml0/dataset.json
13b16b5314807d0c4d45d61e37ded51e5cd4e176a761e2e043f1303d58d358ac  fm-control/20260831T193904Z-fmctl0/dataset.json
baf091ea1d251a84c6b96652874d706d78d0da54edc63d14357f76633b112d7d  pd-shakedown/20260829T162506Z-stage0/dataset.json
5de3991c18e356be0cb82a9b08c0f8699b24306b3acea0a5285664245d077054  sousa-shakedown/20260830T004943Z-sousa0/dataset.json
5460a13eaf487e0bd2689cec495e0883735b598adddc7e3b86cb265ebd37cbcf  bach-shakedown/20260830T104030Z-bach0/dataset.json
73040ef2e45694efef61e308ecb333c10f7aa21580fc086826a42245bd64e41f  corpus/mnml-shakedown/manifest.json
11a12a746a929b8046d09905bf46ed98abefa3ae160859ee334760769563754c  corpus/fm-control/manifest.json
```

Each corpus also carries its own `manifest.sha256` covering the audio bytes; the run re-verifies the specific files it reads (not the whole corpus — that is 10 GB of hashing for no added assurance).

### 2.4 Three properties of these corpora that the analysis must not ignore

1. **The control's tracks are a subset of the wide corpus's bytes.** `fm-control` was carved from `mnml-shakedown` by hardlink; no audio was re-downloaded. The two corpora are **not independent samples**. However, none of the control's 24 training tracks is in the wide corpus's 159 training tracks (its Monokrak items were held out or fall outside the wide training split), so the two *training sets* are disjoint. **Both facts are stated wherever a between-corpus comparison is reported.**
2. **Sample rates run 16 kHz → 96 kHz across the five corpora**, and codecs run mp3 / ALAC / FLAC / PCM. Without a single resampling target the metric would partly measure the mastering and delivery chain. §4 fixes one target.
3. **Bach is two works, Sousa is 106 separate items, Chopin is one collection.** Their internal structure is not comparable to a netlabel corpus's release structure. H2 is therefore **exploratory only** (§7.2).

### 2.5 The number the article must get right

**The wide corpus spans 41 releases and 208 tracks. The adapter trained on 159 tracks from 27 releases.** The 20 % holdout was stratified *by release* (an item's tracks are all-in or all-out) and two further releases were excluded outright for telephone-rate source, so 14 of the 41 releases contributed no training track: 41 − 12 held − 2 excluded = **27**. Both numbers are true of their own subject; the phrase *"159 tracks from 41 releases"* is true of neither.

---

## §3 — THE EMBEDDER

### 3.1 The instrument, and why no installation is needed

The bench uses the **already-pinned, already-hash-verified evaluation environment** built 2026-08-29 and registered as **Amendment A1** to `PREREG-ACE-LORA-FIRST-TRAIN-2026-08-27.md`. It is reused, not rebuilt, so this bench inherits A1's discipline rather than opening a second instrument.

**Instrument:** `CLAP-laion-music` — `laion_clap` **1.1.7**, checkpoint `music_audioset_epoch_15_esc_90.14.pt`.
**Interpreter:** CPython **3.12.14**. **Build tool:** `uv` **0.12.7**.

| package | version |
|---|---|
| `laion-clap` | **1.1.7** |
| `torch` | **2.8.0+cpu** |
| `torchaudio` | **2.8.0+cpu** |
| `numpy` | **1.26.4** |
| `scipy` | **1.17.1** |
| `soundfile` | **0.14.0** |
| `librosa` | **0.10.2.post1** |
| `transformers` | **4.57.6** |

Full freeze: `freeze-A1.txt`, 105 lines, `sha256 = 258a10ee9966ed60eb7839ce9ae9695c431277e4e324564c9fe2bf4e2550ebb1` — **re-verified byte-identical by this lane on 2026-09-03**, so the instrument has not drifted since A1 was filed.

Weights, verified present with A1's sizes and hashes:

```
2352471003  music_audioset_epoch_15_esc_90.14.pt
            sha256 fae3e9c087f2909c28a09dc31c8dfcdacbc42ba44c70e972b58c1bd1caf6dedd
```

**`torch` is a CPU-only build.** It has no CUDA runtime linked and cannot allocate on a card. D-20260903-02's *"never on a live seat's card"* is therefore satisfied by construction, not by scheduling.

`audiobox-aesthetics`, `fadtk` and `mir_eval` are present in the same environment and are **not used by this bench**. In particular no Fréchet-family distance is computed: the estate has already ruled that FAD/KAD are excluded at small N for the equal-N reason (`the evaluation box's ACE-LoRA plan (2026-08-27):762`), and this bench's smallest cell is N = 19.

### 3.2 Two loading traps, pinned so they cannot bite

**(a) The checkpoint must be passed explicitly.** From A1:

> loading a 630k checkpoint would silently satisfy the import and quietly answer a different question, so the music checkpoint is passed explicitly as `ckpt=` and pinned by hash here.

**(b) The backbone default is wrong for this checkpoint.** `laion_clap/hook.py:21` reads:

```
def __init__(self, enable_fusion=False, device=None, amodel= 'HTSAT-tiny', tmodel='roberta') -> None:
```

The music checkpoint is an **HTSAT-base** model. The bench constructs the module as `CLAP_Module(enable_fusion=False, amodel='HTSAT-base')` and passes `ckpt=` explicitly. **`⧖ RATIFY-2`: this is inferred from the default and the checkpoint's identity, not from a load — STEP 0 (§10.2) exercises it before any corpus is touched, and the run aborts if it does not reproduce A1's receipt exactly.**

### 3.3 THE LICENCE, read from the LICENSE file

D-20260903-02 requires the embedder's licence be read from its LICENSE file before adoption. Done, from the file on disk in the pinned environment.

**Code — `laion_clap 1.1.7`.** Bundled `LICENSE`, 7,048 bytes / 121 lines, opens verbatim:

```
Creative Commons Legal Code

CC0 1.0 Universal
```

and runs to the standard CC0 closing clause. **Reading: CC0 1.0 Universal.**

**⚠ The package's own metadata contradicts its own LICENSE file, and both readings are recorded rather than reconciled:**

```
License: Creative Commons Legal Code            ← the CC0 file's first line
Classifier: License :: OSI Approved :: Apache Software License
License-File: LICENSE                            ← points at the CC0 text
```

The `License:` field carries the opening line of the bundled CC0 text (a setuptools artefact); the classifier declares Apache-2.0. **The estate rule is that the LICENSE file is the reading of record**, so the adopted reading is **CC0-1.0**. Nothing turns on the disagreement — both candidates are permissive, neither is NC or ND — but the article discloses it rather than being caught omitting it.

**Weights — `lukewys/laion_clap`.** The model repository declares its licence as **`cc0-1.0`** (read 2026-09-03; the repository README is empty and carries no separate licence prose). The checkpoint was fetched 2026-08-29T15:15–15:17Z from `https://huggingface.co/lukewys/laion_clap/resolve/main/music_audioset_epoch_15_esc_90.14.pt` and pinned by sha256 in A1.

**Verdict: adoptable.** Code CC0-1.0 (classifier says Apache-2.0), weights cc0-1.0. This also matches why the estate chose CLAP in the first place — *"because it is permissively licensed and therefore publishable beside the result it gates."*

**Credit owed on any published page** (per the hub's THANKS convention): LAION's CLAP, the `laion_clap` package, and the `lukewys/laion_clap` checkpoint host, each named with its licence as read above.

---

## §4 — CROP AND LOUDNESS POLICY

### 4.1 Why crop at all

CLAP's audio encoder consumes a **10-second window at 48 kHz mono**. The corpus tracks average 133–240 s; the renders are 60 s. Something must choose which seconds are measured, and that choice must be registered rather than made at the keyboard.

### 4.2 The registered policy `⧖ RATIFY-3`

**Three 10-second windows per item, centred at 25 %, 50 % and 75 % of the item's true duration**, where "true duration" is read from `ffprobe` on the **source file** (never from `dataset.json`, whose `duration` is clamped at the 240 s preprocessing cap).

Each window is cut and conditioned in one deterministic ffmpeg pass:

- downmix to **mono**
- resample to **48 000 Hz**
- loudness-normalise to the estate's existing operating point: **−16 LUFS integrated, −1 dBTP** (the same point at which A1's live verification was measured, so this bench's embeddings are directly comparable to A1's receipt)
- write **16-bit PCM WAV**, and record each crop's `sha256`

**Justification, point by point.**

- **Three, not one:** a single window can land on an intro, a breakdown or a fade and mistake structure for style. Three cheaply covers early / middle / late.
- **Three, not ten:** cost scales linearly and the marginal window adds little; the sensitivity arm (§4.4) tests whether the count mattered at all.
- **25/50/75, not 0/50/100:** the endpoints catch silent lead-ins and fade-outs, which are the least style-bearing seconds of any track and the most variable across mastering conventions.
- **Deterministic, not random:** no seed is needed for crop placement, so the crop set is reproducible from the corpus alone. (Seeds are still fixed for the subsampling — §5.3.)
- **Mono 48 kHz:** CLAP's native input. Resampling from 16 kHz–96 kHz sources to one rate removes the delivery chain as a confound.
- **Loudness normalisation:** this is the trap Round 2 already sprang. Corpus loudness varies enormously (one track measures **−31.26 LUFS integrated**) while every render is peak-normalised to −1 dBTP. Un-normalised, "corpus spread" would partly be a spread of mastering decisions. Normalising to a fixed point removes level as a variable.

**A track's embedding** is the **L2-normalised mean of its three crop embeddings**, renormalised to unit length. CLAP's outputs are already unit-norm (A1's receipt: `shape (1, 512), L2 norm 1.0000`), so cosine distance is `1 − dot(a, b)` with no further scaling.

### 4.3 Cross-box crop equivalence, checked rather than assumed

Two of the five corpora live on a different box from the instrument (§10.1). Both boxes run the **identical** ffmpeg build, `8.0.1-3ubuntu2`, so crops cut on either box should be byte-identical — but "should be" is not a receipt. **STEP 1 cuts the same three crops of one nominated track on both boxes and compares sha256.** If they match, cross-box cropping is proven for this run and the crops travel. If they do not, the run falls back to copying source audio and cutting every crop on the instrument's box, and the fallback is recorded in the results rather than hidden.

### 4.4 Pre-registered sensitivity arms

Reported alongside the primary, not instead of it, and **not** used to select the headline:

- **S1 — no loudness normalisation.** The whole analysis recomputed on crops that are mono/48 kHz but level-untouched. If S1 and the primary disagree in direction, the bench reports that the coherence signal is entangled with loudness and draws no coherence conclusion.
- **S2 — single centre crop.** One 10 s window at 50 %. Tests whether the three-window policy mattered.
- **S3 — full corpus directories** instead of training sets (§2.1).

---

## §5 — METRIC 1: CORPUS SPREAD

### 5.1 The two statistics

For a set of track embeddings `E = {e₁ … e_n}` (each unit-norm):

- **mean pairwise cosine distance**
  `D_pair(E) = (2 / (n(n−1))) · Σ_{i<j} (1 − eᵢ · e_j)`
- **mean distance to the centroid**
  `c = normalise( (1/n) Σ eᵢ )`, then `D_cent(E) = (1/n) Σ (1 − eᵢ · c)`

Both are reported for every corpus. `D_pair` is the headline (it has no dependence on a centroid that itself moves with N); `D_cent` is reported because METRIC 2 uses the same centroid and the two must be read together.

**Higher = more spread out = less coherent**, in this embedding space and no other.

### 5.2 Equal-N, and why it is not optional

`D_pair` is a mean over pairs and is comparatively stable in N, but `D_cent` is not, and any Fréchet-style quantity is badly N-biased — which is exactly why the estate already excludes FAD/KAD at small N (`the evaluation box's ACE-LoRA plan (2026-08-27):762`: *"FAD/KAD stay excluded at N = 3–5 for the equal-N reason"*). Comparing a 159-track corpus to a 19-track corpus without equalising N would let sample size masquerade as coherence.

**N = 19**, the smallest training-set count (Chopin). `⧖ RATIFY-4` — the brief assumed 24; 19 is the measured value under §2.1's unit of analysis.

### 5.3 The registered estimator, with its known degeneracy stated up front

**Point estimate.** For each corpus: **200 draws of N = 19 tracks without replacement**; compute `D_pair` and `D_cent` within each draw; the point estimate is the mean over the 200 draws.
*Known property, stated before the run:* for a corpus with exactly 19 training tracks, all 200 draws are the same set, so the point estimate equals the full-set value and the draw-to-draw variance is exactly zero. **This is a property of the design, not a precision.**

**Interval.** Because the without-replacement distribution is degenerate at the smallest corpus, the reported **95 % CI comes from a separate, comparable resampling**: **200 nonparametric bootstrap draws of 19 tracks with replacement**, and the interval is the 2.5th / 97.5th percentile of the bootstrap statistic.
⚠ **The duplicate-pair rule, registered because it is the easy way to get this wrong:** a with-replacement draw contains repeated tracks whose pairwise distance is exactly 0, which deflates spread by an amount that depends on N. **`D_pair` within a bootstrap draw is computed over pairs of DISTINCT track ids only**, with the pair count adjusted to the number of distinct tracks actually drawn.

**Seeds.** `numpy.random.default_rng(20260903)` for the without-replacement draws; `default_rng(20260904)` for the bootstrap. Fixed here, before the run.

### 5.4 The decision rule for H1 `⧖ RATIFY-5`

D-20260903-02 asks for equal-N subsamples with `mean ± 95 % CI`. Non-overlapping CIs are registered as the **primary** decision rule, and a second, more appropriate test is registered alongside it because CI-overlap is a weak instrument for a two-group difference:

**Primary (as ruled):** H1 is supported if the wide corpus's `D_pair` 95 % CI lies entirely **above** the control's.
**Secondary (registered, not a substitute):** a **two-sample permutation test** — pool the 159 + 24 track embeddings, randomly relabel into groups of 159 and 24, recompute the difference in equal-N `D_pair` at N = 19, **10 000 permutations**, one-sided p for "wide > control". Seed `default_rng(20260905)`.
Both are reported. If they disagree, the disagreement is the finding and is published as such.

---

## §6 — METRIC 2: RENDER-TO-CENTROID

### 6.1 The construction

For each render, compute its embedding under the identical §4 policy (three 10 s windows at 25/50/75 % of its 60 s duration → 15 s, 30 s, 45 s).

For each adapter, compute its **own training corpus's centroid** from the same track embeddings used in METRIC 1 (full training set, not a subsample — the centroid is a property of the corpus, not of a draw).

For each render `r`: `d(r, C) = 1 − ê_r · c_C`.

### 6.2 The pairs

**Primary set — the six scale-1.0, final-checkpoint pairs**, each a base and an adapter render at the same caption and seed:

| # | adapter | caption | seed | base sha256 (12) | adapter sha256 (12) |
|---|---|---|---:|---|---|
| 1 | mnml 1.0 | A | 42 | `d6b62483885e` | `b1627a2f0972` |
| 2 | mnml 1.0 | B | 4242 | `b3f55e8eb66c` | `2cef5f056c7b` |
| 3 | mnml 1.0 | C | 777 | `e71b8c65156c` | `0327d0699971` |
| 4 | mnml 1.0 | D | 7777 | `5889ca518f8b` | `0cecc9a342b5` |
| 5 | fm-control 1.0 | C | 777 | `e71b8c65156c` | `5db77e0345b0` |
| 6 | fm-control 1.0 | D | 7777 | `5889ca518f8b` | `f0ebb8d25803` |

**The base arm of pairs 3/5 and 4/6 is the same audio** — Round 3 reused Round 2's base renders byte-for-byte (identical sha256). That is a strength: the two adapters are compared against one fixed origin. It is also a de-duplication hazard for the implementation, which must key on `output_sha256`, not `clip_id` (the clip_ids collide too).

**Secondary set — the adapter-strength ladder**, reported as a curve, never pooled into the sign test (the renders share seeds and are not independent): mnml at scale 0.7 (captions A, B), 0.5 and 0.35 (caption C), and the epoch-5 checkpoint at 1.0 (caption C).

**One honest caveat on "same caption".** The adapter arms prepend the trigger token — `mnml-shakedown, minimal techno, …` and `fm-control, minimal techno, …` — so the captions are identical *modulo the trigger*. That is by design, but it means the base/adapter difference is "adapter + trigger token", not "adapter alone". Stated wherever the pairs are reported.

### 6.3 The test, and its power, computed before the run

**Exact one-sided sign test** on the paired differences `d(base, C) − d(adapter, C)`, positive meaning the adapter moved *toward* its corpus.

⚠ **Registered power arithmetic.** With **n = 4** (the mnml-only pairs), the minimum attainable one-sided p is 1/2⁴ = **0.0625** — the test *cannot* reach 0.05 even if all four pairs agree. With the pooled **n = 6** it is 1/2⁶ = **0.015625**, which can.

Therefore:
- the **mnml-only** sign test is registered as **DESCRIPTIVE by construction** and will be reported without a significance claim;
- the **pooled six-pair** sign test is the single registered inferential test;
- the **per-pair direction table with the raw distances** is the primary deliverable of METRIC 2, because with six pairs the directions are more informative than any p-value.

`⧖ RATIFY-6`: accept that METRIC 2 is descriptive-to-weak by construction, or expand the clip set (which would require fresh renders and a fresh go, and is out of scope for this article).

### 6.4 The cross-centroid control, registered

For every render, also compute the distance to the **other** adapter's corpus centroid. This separates two very different findings:

- adapter renders move toward **their own** corpus → the adapter learned its corpus;
- adapter renders move toward **both** centroids → the adapter learned "minimal techno in general", or merely moved in a direction both corpora happen to share;
- adapter renders move toward **neither** → the adapter changed the audio without moving it toward anything the corpus defines.

The base renders are shared between the two comparisons, so `d(base, C_mnml)` and `d(base, C_fm)` are measured on identical audio and are directly comparable.

---

## §7 — PREDICTIONS, REGISTERED BEFORE THE RUN

*These are stated now, in advance, so they cannot be adjusted to whatever the numbers say.*

### 7.1 H1 — the wide corpus is more spread out than the one-artist control

**Predicted: yes.** `D_pair(mnml_train) > D_pair(fm_train)` at equal N = 19, with non-overlapping 95 % CIs and a one-sided permutation p < 0.05.

**Reasoning stated in advance:** the wide corpus is dozens of artists across three licence tiers, four sample rates, four codecs and roughly twenty years; the control is one artist on one netlabel in one production chain at one sample rate in one codec. CLAP embeddings are timbre- and texture-dominated, which is precisely the axis on which those two sets should differ most.

**Confidence: high** — this is close to a sanity check on the instrument. If H1 fails, the first suspicion is the instrument or the crop/loudness policy, not the corpora, and §8.1 says what happens then.

### 7.2 H2 — the three composer corpora (EXPLORATORY)

**No directional prediction is registered.** H2 is exploratory and is reported as a ranking with intervals, not as a test.

Weak expectations recorded for the record, so that any post-hoc storytelling is visibly post-hoc: Bach (two works, one performer, one instrument, one session chain) should be the tightest of all five; Chopin (one collection, one instrument, one commissioning project) close behind; Sousa (106 separately-sourced band recordings) looser than both but tighter than the wide netlabel corpus. **If any of these lands differently, it is a finding to report, not a hypothesis to have had.**

The composer corpora are **not** matched to the techno corpora on anything but N, and their internal structure differs (Bach = 2 items, Sousa = 106). H2 cannot support a claim of the form "classical is more coherent than techno"; at most "these three classical training sets, as assembled here, sit closer together in CLAP space than this techno training set does."

### 7.3 H3 — the adapters moved their output toward their corpora

**Predicted for the control: yes** — the fm-control renders sit closer to the fm-control centroid than the base renders do, in both pairs.
**Predicted for the wide adapter: uncertain, leaning yes-but-smaller** — the mnml renders move toward the mnml centroid in a majority of the four pairs, by a smaller margin than the control's.

**Reasoning stated in advance:** the mnml adapter trained to a bouncing, non-monotonic loss (epoch means 0.6603 · 0.6432 · 0.6418 · 0.6507 · 0.6334 · 0.6514) while the control descended monotonically for all ten epochs (1.2104 → 0.7659). If corpus coherence is the lever, the control should show the cleaner pull.

**The interesting outcome is the one that is not predicted:** if the mnml renders move *away* from the mnml centroid, or do not move at all, then the adapter the ear rejected did not merely learn its corpus badly — it did not move toward its corpus at all, which is a stronger and more publishable statement of the negative result.

---

## §8 — THE NULL, AND WHAT IT MEANS FOR THE ARTICLE

**A null publishes as a null** (D-20260903-02). Each of the three is written out here, in advance, with the sentence the article would carry.

### 8.1 If H1 is null — the corpora are not measurably different in spread

The coherence explanation loses its central premise *in this embedding space*. The article says: the wide corpus and the one-artist corpus do not sit measurably differently in CLAP space at equal sample size, so whatever made the adapter fail the ear, this instrument cannot see it. **The ear's verdict is unaffected** — it was never gated on this number.

Before publishing that, the run must state whether the sensitivity arms (S1 no-normalisation, S2 single-crop) also come back null. A null in the primary that flips under S1 means the metric was measuring loudness; that is reported as an instrument finding, not a corpus finding.

### 8.2 If H3 is null — the adapters did not move their renders toward their corpora

Two readings, and the bench cannot choose between them:

1. the adapters genuinely did not learn their corpora; or
2. CLAP cannot see the axis along which they did.

The article carries both. The second is not a hedge: CLAP-class embeddings are timbre-dominated, and a LoRA at scale 0.35–1.0 over 8 inference steps may move an axis this instrument compresses. **A null here is fully consistent with the ear's rejection and is the more interesting of the two outcomes** — it says the negative result is not "the adapter learned the wrong thing" but "the adapter did not measurably move toward the thing it was trained on".

### 8.3 If everything is null

Then the bench's contribution is a **method with its limits measured**: a pre-registered, licence-clean, CPU-only way to ask "is this corpus coherent, and did the adapter move toward it", which on this material returned nothing. That publishes. It also stands as a receipt that the coherence story was tested rather than asserted — which is worth more to the article than a number that happened to agree with the narrative.

### 8.4 What no outcome can do

No outcome promotes, retires, or re-ships an adapter. No outcome overrides the ear. No outcome changes ruling **D-20260903-03** (the operator's pending Round 3 verdict), which is the operator's alone and must not be anticipated, hinted at, or pre-empted by any number this bench produces.

---

## §9 — WHAT IS NOT CLAIMED

1. **CLAP distance is not musical quality.** Nothing here measures whether audio is good, pleasant, danceable, or correct. The estate has already ruled which instrument does: `PREREG-ACE-LORA-FIRST-TRAIN-2026-08-27.md:260` — *"CLAP/MuQ-class embeddings are texture- and timbre-dominated — which is precisely what a LoRA legitimately learns, so a legitimate style match and a memorised melody can sit at the same cosine"* — and the estate's answer is that **the embedding screens and the listen confirms**.
2. **Spread is not disorder.** A wide corpus is not "worse". A style may legitimately be broad. The bench measures dispersion, and dispersion is not a defect.
3. **No causal claim.** Even if H1 and H3 both land as predicted, this bench cannot establish that corpus incoherence *caused* the ear's rejection. The ear's evidence is two blind pairs — itself not a powered experiment. Two under-powered instruments agreeing is a consistent story, not a proof.
4. **No generalisation to genres.** "Minimal techno" here is one archive.org subject tag sampled on one date under three licence tiers. It is not the genre.
5. **No generalisation to models or embedders.** One text-to-music model, one adapter recipe, one embedding space. A different embedder may order these corpora differently; nothing here says it would not.
6. **Not a memorisation check.** METRIC 2 measures distance to a *centroid*, not to nearest training tracks. It cannot detect memorisation and is not offered as a screen for it (that is A2's job, and A2 remains unfiled).
7. **No claim about the held-out tracks.** The holdout was never rendered against; it plays no part in this bench.
8. **Not an equivalence test.** A null is "we did not detect a difference at this N with this instrument", never "the corpora are the same".

---

## §10 — THE PROCEDURE

### 10.1 The box, with its receipts `⧖ RATIFY-7`

D-20260903-02 says *"runs on whichever box holds the corpora"*. **The corpora are on two boxes**, and the pinned instrument is on the one holding three of the five. **This pre-registration proposes running the whole bench on the instrument's box** — the corpora travel to the instrument, as 10 s crops — because rebuilding the environment elsewhere would create a second instrument and require a fresh numbered amendment under A1's own discipline:

> **Instrument discipline, restated from §6.1:** both arms are scored by this one pinned instrument in one session. Any change to the table above requires a **new numbered amendment**, filed before the session it governs — never an in-place edit.

**This is a deviation from the letter of the ruling and is flagged, not slipped through.**

**The receipts, taken 2026-09-03T02:4xZ, verbatim.**

Instrument box (2 × consumer 24 G cards, both occupied by live seats; the bench uses neither):

```
index, uuid, name, memory.used [MiB], memory.total [MiB]
0, GPU-<uuid withheld>, NVIDIA GeForce RTX 3090, 23402 MiB, 24576 MiB
1, GPU-<uuid withheld>, NVIDIA GeForce RTX 3090, 19326 MiB, 24576 MiB

Filesystem                         Size  Used Avail Use% Mounted on
/dev/mapper/ubuntu--vg-ubuntu--lv  295G  123G  160G  44% /

CPU: AMD Ryzen 7 3700X 8-Core (nproc 16) · RAM 30 G total, 10 G available · ffmpeg 8.0.1-3ubuntu2
load average: 0.08, 0.06, 0.01
```

Corpus box (one workstation-class card, occupied; the bench uses none of it — it only cuts crops on CPU):

```
index, uuid, name, memory.used [MiB], memory.total [MiB]
0, GPU-<uuid withheld>, <96 G workstation card>, 63978 MiB, 97887 MiB

Filesystem      Size  Used Avail Use% Mounted on
/dev/nvme0n1p2  1.8T  688G  1.1T  40% /

CPU: AMD EPYC 7532 32-Core (nproc 64) · RAM 123 G total, 104 G available · ffmpeg 8.0.1-3ubuntu2
```

**Pre-flight aborts, registered:** the run refuses to start if the instrument box shows **< 8 GiB available RAM** (`free -g`), if free disk on either box is **< 5 GiB**, or if any input hash in §2.3 differs. A bench that swaps is a bench whose timings are fiction.

### 10.2 The steps

**STEP 0 — instrument check (no corpus touched).** Verify `freeze-A1.txt` sha, verify both weight shas against `weights.sha256`, construct `CLAP_Module(enable_fusion=False, amodel='HTSAT-base')`, `load_ckpt(ckpt=<pinned path>)`, embed A1's own 25 s Chopin verification clip, and **require the printed receipt to reproduce A1 exactly**:

```
CLAP-laion-music embedding: shape (1, 512), L2 norm 1.0000
```

Any deviation aborts before a single corpus file is read.

**STEP 1 — cross-box crop equivalence (§4.3).** Cut the three registered crops of one nominated track on both boxes; compare sha256. Record match or fall back.

**STEP 2 — crop.** Cut 349 × 3 = **1,047** corpus crops and 15 × 3 = **45** render crops under §4.2. Write `crops.csv` with `(corpus, track_id, source_path, source_sha256, window_pct, crop_sha256, source_duration_s)`.

**STEP 3 — transfer.** Ship the 549 crops from the corpus box to the instrument box (see §10.4 for the size arithmetic). Re-verify every crop sha after transfer.

**STEP 4 — embed.** 1,092 forward passes. Write `embeddings.npz` (float32, `(n, 512)`, unit-norm) plus `embeddings-index.csv` mapping row → `(corpus | render, id, window_pct)`.

**STEP 5 — METRIC 1.** Per §5, all five corpora, primary + S1 + S2 + S3.

**STEP 6 — METRIC 2.** Per §6, six primary pairs + the ladder + the cross-centroid control.

**STEP 7 — outputs and hashes.** Write the artefacts in §10.5 and `sha256sum` every one of them into `RESULTS.sha256`.

### 10.3 The runnable skeleton (illustrative; the run's script is committed before it runs)

```python
# run under the pinned instrument only:
#   music/venvs/eval/.venv/bin/python  (CPython 3.12.14, torch 2.8.0+cpu)
import json, hashlib, numpy as np, laion_clap, soundfile as sf

RNG_SUBSAMPLE = np.random.default_rng(20260903)
RNG_BOOTSTRAP = np.random.default_rng(20260904)
RNG_PERMUTE   = np.random.default_rng(20260905)
N_EQUAL, N_DRAWS, N_PERM = 19, 200, 10_000
CKPT = "<eval-venv>/weights/music_audioset_epoch_15_esc_90.14.pt"

model = laion_clap.CLAP_Module(enable_fusion=False, amodel="HTSAT-base")
model.load_ckpt(ckpt=CKPT)                      # explicit: never the 630k default

def embed_track(crop_paths):                     # 3 crops -> 1 unit vector
    x = np.stack([sf.read(p, dtype="float32")[0] for p in crop_paths])
    e = model.get_audio_embedding_from_data(x=x, use_tensor=False)   # (3, 512)
    v = e.mean(axis=0)
    return v / np.linalg.norm(v)

def d_pair(E):                                   # mean pairwise cosine distance
    G = E @ E.T
    iu = np.triu_indices(len(E), k=1)
    return float((1.0 - G[iu]).mean())

def d_cent(E):
    c = E.mean(axis=0); c /= np.linalg.norm(c)
    return float((1.0 - E @ c).mean())

def spread_equal_n(E, rng, n=N_EQUAL, draws=N_DRAWS):
    idx = np.arange(len(E))
    return np.array([d_pair(E[rng.choice(idx, n, replace=False)]) for _ in range(draws)])

def spread_bootstrap(E, rng, n=N_EQUAL, draws=N_DRAWS):
    out = []
    for _ in range(draws):
        pick = np.unique(rng.choice(len(E), n, replace=True))   # DISTINCT ids only
        out.append(d_pair(E[pick]))                             # duplicate-pair rule
    return np.array(out)
```

**ffmpeg crop (one window), the exact registered form:**

```
ffmpeg -nostdin -v error -ss <START> -t 10 -i "<SRC>" \
  -ac 1 -ar 48000 \
  -af loudnorm=I=-16:TP=-1.0:LRA=11:print_format=summary \
  -c:a pcm_s16le "<OUT>.wav"
```

`<START> = max(0, pct * true_duration − 5)`, clamped so the window stays inside the file; `true_duration` from `ffprobe -show_entries format=duration` on the source.

### 10.4 Transfer arithmetic

- corpus tracks on the instrument's box already: 19 + 84 + 63 = 166
- corpus tracks that must travel: 159 + 24 = **183** → 183 × 3 = **549 crops**
- one crop as mono 48 kHz 16-bit PCM: 48 000 × 10 × 2 = 960 000 B ≈ **0.96 MB**
- transfer total: 549 × 0.96 MB = **527 MB**
- versus copying source audio: 8.0 G + 2.7 G = **10.7 G** — a **20×** saving

### 10.5 Outputs, all sha256'd

| file | content |
|---|---|
| `crops.csv` | every crop: corpus, track id, source path + sha256, window %, crop sha256, true duration |
| `embeddings.npz` | float32 `(1092, 512)` unit-norm, plus the index array |
| `embeddings-index.csv` | row → (corpus\|render, id, window %) |
| `metric1-corpus-spread.csv` | per corpus: n, `D_pair`, `D_cent`, equal-N point estimate, bootstrap 95 % CI lo/hi, per arm (primary, S1, S2, S3) |
| `metric1-draws.csv` | all 200 × 5 draw statistics, so the intervals are re-derivable |
| `metric2-render-centroid.csv` | per render: arm, caption, seed, scale, distance to own centroid, distance to other centroid, sha256 of the wav |
| `metric2-pairs.csv` | the six pairs: paired difference, direction, and the ladder rows |
| `RESULTS.json` | every headline number with its inputs, the seeds, the instrument versions and weight hashes |
| `RESULTS.sha256` | sha256 of every file above |
| `RUN.log` | stdout/stderr, with STEP 0's receipt at the top |

### 10.6 Runtime estimate `[ESTIMATE — to be replaced by measured]`

| stage | estimate | basis |
|---|---|---|
| STEP 0 instrument load + check | 1–3 min | 2.35 GB checkpoint from disk, CPU |
| STEP 1 equivalence check | < 1 min | 6 crops |
| STEP 2 cropping, 1,092 crops | 5–15 min | ffmpeg two-pass loudnorm on 10 s, 8-way parallel |
| STEP 3 transfer 527 MB | **unmeasured** | the private network throughput not measured by this lane |
| STEP 4 embedding, 1,092 forwards | 10–30 min | ~0.5–1.5 s per 10 s clip, HTSAT-base on 16 threads |
| STEPS 5–7 statistics + writeout | < 2 min | 171 000 pair distances is arithmetically free |
| **total** | **~20–55 min + transfer** | |

The bench is **not** GPU-bound, **not** disk-bound, and does not compete with any seat. The one real resource risk is RAM on the instrument box (§10.1's abort).

---

## §11 — THE LEDGER ROW OWED, IN THE SAME STROKE

`PREREG-LEDGER-OF-LEDGERS.md` LAW 1:

> Every pre-registration on this estate lives under a location NAMED IN THIS FILE. Opening a new prereg location = adding its row here in the same stroke (the lane brief carries the duty).

This is a **new location**. The file's last row is **24**. Proposed **row 25**, in that file's column shape, for the orchestrator to paste at ratification (this lane did not edit the ledger):

> `| 25 | estate/bench/ace-mnml-2026-09/PREREG-COHERENCE.md (body) + estate/bench/ace-mnml-2026-09/coherence/ (runs) | the laptop (repo; the body) + the evaluation box (the run artifacts) | **THE CORPUS-COHERENCE BENCH** (D-20260903-02) — CLAP-embedding spread across five ACE-Step LoRA training corpora at equal N=19 with 200-draw intervals, plus render-to-corpus-centroid distances for the six base/adapter render pairs of Rounds 1–3. Instrument reused, not rebuilt: Amendment A1 of row 24's prereg (laion-clap 1.1.7, music_audioset_epoch_15_esc_90.14.pt sha fae3e9c0…, torch 2.8.0+cpu). Predictions H1/H2/H3, the null, and the non-claims registered before the first embedding. MEASURED NEVER GATING per D-20260831-17. sha256:⧖ at seal. A rule-2 hunt reads this row AND row 24 — the instrument pin lives there. |`

---

## §12 — THE OPEN RATIFICATION POINTS, COLLECTED

| id | the question | this draft's recommendation | cost of the other choice |
|---|---|---|---|
| **⧖ RATIFY-1** | Unit of analysis: **training sets** or corpus directories? | Training sets (what the adapter saw); corpus directories as sensitivity arm S3 | Corpus directories change N to 24 and change the question from "what it learned from" to "what was collected" |
| **⧖ RATIFY-2** | Accept `amodel='HTSAT-base'` as inferred, gated by STEP 0? | Yes — STEP 0 aborts if A1's receipt is not reproduced | Nothing; the alternative is exercising it now, which needs a 2.4 GB load |
| **⧖ RATIFY-3** | Crop policy: 3 windows at 25/50/75 %, mono 48 kHz, −16 LUFS / −1 dBTP | As stated, with S1 (no normalisation) and S2 (single crop) registered | More windows cost linearly and add little; no normalisation risks measuring mastering (Round 2's trap) |
| **⧖ RATIFY-4** | **N = 19**, not 24 | 19 — the measured smallest training count (Chopin) | 24 requires switching to corpus files (see RATIFY-1) |
| **⧖ RATIFY-5** | H1 decision rule: CI-overlap (as ruled) plus a permutation test | Report both; publish any disagreement as the finding | CI-overlap alone is a weak instrument for a two-group difference |
| **⧖ RATIFY-6** | METRIC 2 is underpowered by construction (min p = 0.0625 at n = 4) | Accept: mnml-only test is descriptive, pooled n=6 is the one inferential test, the direction table is the deliverable | Expanding the clip set needs fresh renders and a fresh operator go — out of scope for this article |
| **⧖ RATIFY-7** | Run on the **instrument's** box (corpora travel as crops), deviating from "whichever box holds the corpora" | Yes — rebuilding the instrument elsewhere costs a new A1-style amendment | Rebuilding is ~15 min of CPU saved against a second instrument to defend |
| **⧖ RATIFY-8** | Fix the "159 tracks from 41 releases" phrasing before the title is ruled | Corpus = 41 releases / 208 tracks; training set = 159 tracks / **27** releases | The candidate title "Forty-One Releases" describes the corpus, not the adapter — usable only if framed as the corpus |
| **⧖ RATIFY-9** | Does this bench's output publish in the article, or only in the pack? | D-20260903-02 implies the article ("the numbers ride the article"); a figure/table menu has not been drawn | — |

---

*End of draft. Nothing in this file has been run. The first embedding waits on §12 and the seal.*

---

## §13 — RATIFICATION (the orchestrator, 2026-09-03 ~03:5xZ, BEFORE the first embedding)

Every §12 point is ruled here. The seal in §0.4 is computed over this file WITH the placeholder text
`⧖ TO BE FILLED AT SEAL` still in place, then written in; verifying the seal therefore means restoring
that placeholder and hashing. Any later change is an AMENDMENT block appended below §13, never an edit.

| id | ruling | note |
|---|---|---|
| RATIFY-1 | **Training sets** are the unit; corpus directories ride as S3 | what each adapter saw is the question |
| RATIFY-2 | **Accepted**, gated by STEP 0 | the run aborts if A1's receipt (`shape (1, 512), L2 norm 1.0000`) is not reproduced |
| RATIFY-3 | **As stated**: 3 windows at 25/50/75 %, mono 48 kHz, −16 LUFS / −1 dBTP, 16-bit PCM; S1 and S2 registered | Round 2's loudness trap is the reason normalisation is primary |
| RATIFY-4 | **N = 19** | the measured smallest training count |
| RATIFY-5 | **Both** rules reported; disagreement publishes as the finding | CI-overlap primary as ruled, permutation secondary |
| RATIFY-6 | **Accepted for Set A** (Rounds 1–3: mnml-only descriptive, pooled n = 6 inferential) — AND expanded by AMENDMENT A-R4 below, which registers Round 4's fresh pairs as Set B before they exist | fresh renders carry a fresh go: D-20260903-04 |
| RATIFY-7 | **Run on the instrument's box, CPU only; crops travel** | satisfies the ruling's intent ("never on a live seat's card") by construction; the deviation from its letter is recorded here |
| RATIFY-8 | **The article's phrasing is fixed**: the corpus = 41 releases / 208 tracks; the adapter trained on 159 tracks from **27** releases (12 releases held out by the split, 2 excluded for telephone-rate source) | "159 tracks from 41 releases" is true of neither and does not publish |
| RATIFY-9 | **The numbers publish in the article**: Metric 1 as one table (five corpora, `D_pair` with CI, primary arm; S1–S3 in the kit), Metric 2 as a direction table (Set A) plus the Set B dose curve; every CSV rides the kit | MEASURED, never gating — stated beside the table |

### AMENDMENT A-R4 — Round 4's renders join METRIC 2 as Set B (registered 2026-09-03 ~03:5xZ, before any Round 4 clip exists)

Round 4 (`PREREG-ROUND4.md`, D-20260903-04) renders twelve captions × four arms at 30 s: `base`,
`mnml-0.5`, `mnml-1.0`, `fm-control-1.0`, one seed per prompt shared across its arms.

- **Crops:** the §4.2 policy applied to a 30 s render → windows centred at 7.5 s, 15 s, 22.5 s.
- **Set B pairs, keyed on `output_sha256`:** B-mnml = 12 pairs (`base` vs `mnml-1.0`); B-fm = 12 pairs
  (`base` vs `fm-control-1.0`); B-half = 12 pairs (`base` vs `mnml-0.5`), reported as the ladder's middle
  rung. Sets A and B are reported SEPARATELY (60 s vs 30 s renders are different compositions); a pooled
  figure is reported last and labelled pooled.
- **Test per set:** the exact one-sided sign test of §6.3 on `d(base, C) − d(adapter, C)`, n = 12 per
  set (min attainable one-sided p = 1/4096) — Set B is therefore the bench's best-powered arm and is named
  as such in the article.
- **Cross-centroid control (§6.4)** applies to every Set B render.
- **H3b (registered):** the control's renders sit closer to the control centroid than base in a majority of
  the 12 (prediction: yes, ≥ 9/12); the wide adapter's renders sit closer to the mnml centroid than base in
  a majority of the 12 (prediction: uncertain, leaning yes, 7–9/12).
- **H5, the dose-response (registered, descriptive):** for each prompt, `d(·, C_mnml)` is non-increasing
  across the ladder `base → 0.5 → 1.0`. Prediction: monotone in a majority of the 12 prompts. A
  non-monotone majority publishes as "the dial moves the audio but not along the corpus axis".
- **H4 (from PREREG-ROUND4 §1, exploratory):** any adapter effect on `d` is larger on P01–P06 than
  P07–P12. Reported as two group medians with no test.
- **Ordering:** Metric 1 and Set A run immediately; Set B runs when `round-4/records.json` exists on the
  training box (the bench lane polls, ≤ 90 min); if Round 4 has not landed by then, RESULTS.md publishes
  Sets A only and says Set B is owed.

*Ratified by the orchestrator under the operator's standing "roll with your recs" of 2026-09-03 02:2xZ (D-20260903-01/02).*

---

### AMENDMENT A-C1 — the S3 crop budget (filed 2026-09-03 by the bench lane, BEFORE the first crop)

**What.** §4.4 registers sensitivity arm **S3 — full corpus directories instead of training sets**,
and §2.1 registers the same arm a second time (*"a secondary, clearly-labelled arm repeats METRIC 1
over the full corpus directories"*). §10.5 registers a `metric1-corpus-spread.csv` carrying a column
*"per arm (primary, S1, S2, S3)"*. But **§10.2 STEP 2 budgets `349 × 3 = 1,047` corpus crops** and
§10.4's transfer arithmetic counts `159 + 24 = 183` travelling tracks — both of which count the
**training sets only**. S3 cannot be computed from crops that were never cut, so the registered arm
and the registered crop count contradict each other.

**Why this is an amendment and not an edit.** The contradiction is internal to the sealed file and was
found at run time. §13's rule is that any later change is an amendment appended below, never an
in-place edit of a registered section.

**The resolution.** **S3 is run as §4.4 and §2.1 register it** — METRIC 1 over the full corpus
directories — because dropping a twice-registered arm would leave a registered output column empty,
while cutting extra crops changes no metric, no seed, no N, no prediction, and no set definition.
§10.2's `1,047` is corrected to the crop set the registered analysis actually requires:

| crop set | tracks | windows | crops | serves |
|---|---:|---:|---:|---|
| corpus, `norm` (−16 LUFS), full corpus directories | 447 | 3 | 1,341 | primary (the 349 training rows), S2 (its 50 % rows), S3 (all rows) |
| corpus, `raw` (level untouched), training sets | 349 | 3 | 1,047 | S1 |
| renders, `norm`, Set A | 15 | 3 | 45 | METRIC 2 Set A |
| **total before Set B** | | | **2,433** | |
| renders, `norm`, Set B (Round 4, if it lands) | 48 | 3 | 144 | METRIC 2 Set B (A-R4) |

The 447 full-corpus tracks are §2.2's own "corpus files" column (24 + 30 + 79 + 106 + 208), measured
and confirmed on disk before this amendment was filed. `N_EQUAL` stays **19** in every arm, as §5.2
registers it; S3 does not re-derive N from its own smallest corpus.

**S1's scope.** §4.4 says S1 recomputes *"the whole analysis"*. The arm column exists only in §10.5's
METRIC 1 outputs (`metric1-corpus-spread.csv`, `metric1-draws.csv`); the METRIC 2 outputs carry no arm
column. S1/S2/S3 are therefore run as **METRIC 1 arms**, and METRIC 2 is reported on the primary
conditioning only. This is recorded so the reading is visible rather than assumed.

### AMENDMENT A-C2 — path resolution for the §2.3 inputs and the Round 4 landing site (filed 2026-09-03 by the bench lane, before the first crop)

**What.** §2.3 names its five frozen inputs by run-relative path (`mnml-shakedown/20260831T023450Z-mnml0/dataset.json`
and siblings). Those names do not resolve under `music/out/`; the files live under **`music/datasets/`**.
**Every one of the seven §2.3 sha256 values matches byte-for-byte at the resolved path**, so the frozen
inputs are the intended files and nothing registered has moved — only the shorthand needed expanding. The
resolved paths are recorded in `RESULTS.json` and `RUN.log`.

Second, the lane brief named Round 4's landing site as `music/out/mnml-round4/round-4/`; `PREREG-ROUND4.md`
§5 registers **`music/out/mnml-shakedown/round-4/`** on the training box. The prereg governs; the bench
polls **both** paths so either lands, and A-R4's ≤ 90 min window is unchanged.

**Neither item changes a metric, a seed, a crop policy, an N, a prediction, or a set definition.**

### AMENDMENT A-C3 — §2.4's disjointness claim is MEASURED FALSE (filed 2026-09-03T03:3xZ by the bench lane, after cropping, before any metric was computed)

**What §2.4 registered**, as a stated property the analysis must not ignore:

> **The control's tracks are a subset of the wide corpus's bytes.** … However, none of the control's
> 24 training tracks is in the wide corpus's 159 training tracks (its Monokrak items were held out or
> fall outside the wide training split), so the two *training sets* are disjoint. **Both facts are
> stated wherever a between-corpus comparison is reported.**

**What is measured.** Comparing the two frozen `samples` arrays by the sha256 of every source file:

| quantity | measured |
|---|---:|
| fm-control training tracks | 24 |
| of those, byte-identical to a track in the mnml **training** set | **21** |
| of those, byte-identical to a file anywhere in the mnml **corpus** directory | 24 |
| fm-control corpus files whose bytes appear in the mnml corpus directory | 30 of 30 |
| mnml training tracks also in the fm-control training set | 21 of 159 |

The overlap is release-shaped, exactly as §2.5's by-release stratification implies: the control's
training set is 8 Monokrak releases × 3 tracks, and **7 of those 8 releases landed in the wide
corpus's training split**, not its holdout — `Monokrak200`, `monoKraK204`, `Monokrak205`,
`Monokrak209`, `Monokrak214`, `Monokrak216`, `Monokrak217`. Only `Monokrak203` (3 tracks:
*Drink Your Caotika*, *Les Mauvaises Wagnieres*, *Melody Transformer*) fell outside it.

**The two training sets are therefore NOT disjoint: they share 21 of the control's 24 tracks (87.5 %).**

**What changes, and what does not.** No metric, seed, crop policy, N, prediction, or set definition
changes — the bench runs exactly as registered, and H1/H2/H3/H3b/H5 stand as written. What changes is
the sentence §2.4 obliges the run to state wherever a between-corpus comparison is reported. The
corrected statement, which RESULTS.md carries at every such point:

> The one-artist control's training set is very nearly a **subset** of the wide corpus's training set
> — 21 of its 24 tracks are the same bytes. The H1 comparison is therefore not "two independent
> corpora" but "21 shared tracks, with and without 138 others". A difference in spread between them is
> attributable to the 138 additional tracks, which is a *sharper* statement of the coherence question
> than the registered framing; a null is correspondingly weaker evidence, because the two sets are far
> more alike than §2.4 assumed.

The same correction applies to METRIC 2's cross-centroid control (§6.4): `C_mnml` and `C_fm` are not
centroids of disjoint material, and the "moved toward both centroids" reading of §6.4 must be read
with that in mind.

**A disjoint-subset re-analysis is NOT part of this registration.** RESULTS.md may report one only as
an explicitly labelled POST-HOC, NOT-PRE-REGISTERED descriptive, never as the headline, and never in
place of the registered numbers.

### AMENDMENT A-C4 — the seal's clock, stated exactly (orchestrator, 2026-09-03 ~05:5xZ)

§13 says the ratification was written "~03:5xZ"; that was an estimate typed before checking a
clock, and it is wrong. The seal commit (estate `071cd03`) carries committer time
**2026-09-03T03:08:03Z**; the run scripts were committed at
03:18:37Z (`ec37c02`), and the bench's first embedding began at 03:19:28Z (RESULTS.md §8). The
seal therefore preceded the run by eleven minutes, as the registration requires. Nothing in
§0.4–§13 is edited (the seal covers them); this note is the correction.
