# RESULTS — THE CORPUS-COHERENCE BENCH

**What this measured.** Five music corpora were used to train five LoRA style adapters for a
text-to-music model. One of them — a "minimal techno" set assembled by matching an archive.org
subject tag across 41 Creative-Commons releases by dozens of artists — produced an adapter the
operator's blind ear **rejected**. A one-artist control corpus was trained under the identical
recipe to test whether **corpus coherence** was the reason. This bench asked two questions, both
registered in writing before the first number existed: **(1)** do the wide corpus's training
tracks sit further apart in CLAP embedding space than the one-artist control's, at equal sample
size? and **(2)** for every base/adapter render pair at the same caption and seed, did the
adapter's output move *toward* its own training corpus's centre?

**Measurement window: 2026-09-03T03:19:28Z → 2026-09-03T03:40:50Z (UTC), 21 min 22 s wall.**
Every number below was produced inside that window by one pinned CPU-only instrument, from a
pre-registration sealed at 2026-09-03T03:5xZ *the previous evening* — i.e. before any of it ran.

**Against the pre-registered predictions, in one sentence each.**

- **H1 (predicted yes, high confidence): SUPPORTED**, decisively and in all four analysis arms —
  the wide corpus's mean pairwise cosine distance is 0.4489 against the control's 0.2378 at
  equal N = 19, with non-overlapping 95 % intervals and a permutation p of 0.0001.
- **H2 (exploratory, no directional prediction): the ranking landed Bach → Sousa → Chopin →
  control → wide**, which puts Sousa where the recorded weak expectation had put Chopin.
- **H3 (predicted: control yes in both pairs, wide adapter uncertain-leaning-yes): NOT SUPPORTED** —
  4 of 6 pooled pairs moved toward their corpus (p = 0.3438), and the control, the arm predicted
  most confidently, moved toward in only 1 of its 2.
- **H3b (predicted: control ≥ 9/12, wide adapter 7–9/12): NOT SUPPORTED** — in the bench's
  best-powered arm the control scored 5/12 and the wide adapter 6/12, both indistinguishable
  from a coin.
- **H5 (predicted: monotone in a majority of 12): NOT SUPPORTED** — 4 of 12, which publishes in
  the pre-registration's own registered words as *"the dial moves the audio but not along the
  corpus axis"*.
- **H4 (exploratory, no test): the direction held** — the adapter's median pull toward the corpus
  was +0.0316 on the six techno-family prompts and −0.0143 on the six house/breaks/trance ones.

**MEASURED, NEVER GATING** (D-20260831-17). Nothing here passes, fails, promotes or retires an
adapter. The instrument of record for whether an adapter is any good is the operator's blind ear;
these numbers ride beside that verdict and never in front of it.

---

## §1 — The seal, and the instrument's receipts

### 1.1 The seal

`PREREG-COHERENCE.md` §0.4 carries `sha256 = 4bbe9548945d1bd849b355a080b49c481c067e21fd105e98d8d6d30c9f2fe5f2`,
computed with that line still reading its placeholder. Verifying it means restoring the placeholder
and hashing. Done, before anything was read:

```
method     restore the §0.4 line to `sha256(this file at registration) = ⧖ TO BE FILLED AT SEAL`
           in a temp copy, then sha256 that copy
computed   4bbe9548945d1bd849b355a080b49c481c067e21fd105e98d8d6d30c9f2fe5f2
§0.4 seal  4bbe9548945d1bd849b355a080b49c481c067e21fd105e98d8d6d30c9f2fe5f2
VERDICT    MATCH — the registration this run executed is the sealed one.
```

### 1.2 STEP 0 — the instrument check, which had to reproduce Amendment A1 exactly

The bench reuses the evaluation environment pinned as **Amendment A1** of
`PREREG-ACE-LORA-FIRST-TRAIN-2026-08-27.md` (built 2026-08-29) rather than building a second
instrument. STEP 0 ran before a single corpus file was read, and aborts on any deviation:

```
[03:19:28Z] freeze-A1.txt sha256 = 258a10ee9966ed60eb7839ce9ae9695c431277e4e324564c9fe2bf4e2550ebb1
[03:19:28Z] freeze-A1.txt lines  = 105  MATCHES A1
[03:19:31Z] music_audioset_epoch_15_esc_90.14.pt bytes = 2352471003
            sha256 = fae3e9c087f2909c28a09dc31c8dfcdacbc42ba44c70e972b58c1bd1caf6dedd  MATCHES A1 pin
[03:19:37Z] python 3.12.14  torch 2.8.0+cpu  numpy 1.26.4
[03:19:37Z] torch.cuda.is_available() = False   (CPU-only by construction, §3.1)
[03:19:44Z] CLAP_Module(enable_fusion=False, amodel='HTSAT-base') + explicit ckpt loaded in 6.7s
[03:19:45Z] RECEIPT:     CLAP-laion-music embedding: shape (1, 512), L2 norm 1.0000
[03:19:45Z] A1 REQUIRES: CLAP-laion-music embedding: shape (1, 512), L2 norm 1.0000
[03:19:45Z] PASS — A1's receipt reproduced exactly; the corpora may now be read.
```

| pinned item | value | source |
|---|---|---|
| `laion-clap` | 1.1.7 | `freeze-A1.txt`, sha verified above |
| `torch` / `torchaudio` | 2.8.0+cpu | ditto |
| `numpy` / `scipy` | 1.26.4 / 1.17.1 | ditto |
| `soundfile` / `librosa` / `transformers` | 0.14.0 / 0.10.2.post1 / 4.57.6 | ditto |
| interpreter | CPython 3.12.14 | measured at run time |
| checkpoint | `music_audioset_epoch_15_esc_90.14.pt`, 2,352,471,003 B | sha verified above |
| backbone | `amodel='HTSAT-base'`, `enable_fusion=False`, explicit `ckpt=` | §3.2, exercised by STEP 0 |

`RATIFY-2` is discharged: `HTSAT-base` was inferred from the default (`HTSAT-tiny`) and the
checkpoint's identity, and STEP 0 exercised it rather than assuming it.

**The CPU-only guarantee is structural, not scheduled.** The pinned `torch` is a CPU build with no
CUDA runtime linked; `torch.cuda.is_available()` returned `False`. D-20260903-02's *"never on a live
seat's card"* therefore held by construction. No process on either box was touched, evicted or
restarted; both boxes' live seats and the Round 4 render lane ran throughout.

**A1's verification clip.** A1 records *"a 25 s clip cut from the pd-shakedown corpus and
loudness-matched to the §1 operating point"* without naming the file, and pins a **shape-and-norm**
receipt rather than a value. This run named its own and recorded it:
`Ballade no. 1 - Op. 23.m4a`, duration 613.589 s, 25 s from 294.294573 s, conditioned to
−16 LUFS / −1 dBTP / 48 kHz stereo, clip
`sha256 2fe2595d7d31067cb0597a1b5acb00502083a05a2cd036163e9ed298669ce6c2`.

### 1.3 The embedder's licence, read from the LICENSE file

Required by D-20260903-02 before adoption, and read from the file on disk in the pinned
environment: the bundled `LICENSE` of `laion_clap 1.1.7` is **CC0 1.0 Universal** (7,048 B / 121
lines). The package's own metadata contradicts its own LICENSE file — the `License:` field carries
the CC0 text's opening line and the classifier declares Apache-2.0 — and **both readings are
recorded rather than reconciled**; the estate's rule is that the LICENSE file is the reading of
record, so the adopted reading is **CC0-1.0**. The weights (`lukewys/laion_clap`) declare
**cc0-1.0**. Both candidates are permissive; neither is NC or ND. Credit is owed on any published
page to LAION's CLAP, the `laion_clap` package, and the `lukewys/laion_clap` checkpoint host, each
with the licence as read here.

---

## §2 — The corpora as embedded

### 2.1 What was read, and the hash receipts

The bench reads the **frozen training set** of each run — the `samples` array of that run's
`dataset.json` (`RATIFY-1`: what each adapter actually saw) — plus the full corpus directories for
sensitivity arm S3. All seven §2.3 input hashes were re-verified at run time and **all seven
matched**:

| input | sha256 | verdict |
|---|---|---|
| `mnml-shakedown/20260831T023450Z-mnml0/dataset.json` | `d26114c0…e41f` | MATCH |
| `fm-control/20260831T193904Z-fmctl0/dataset.json` | `13b16b53…358c` | MATCH |
| `pd-shakedown/20260829T162506Z-stage0/dataset.json` | `baf091ea…2d7b` | MATCH |
| `sousa-shakedown/20260830T004943Z-sousa0/dataset.json` | `5de3991c…7054` | MATCH |
| `bach-shakedown/20260830T104030Z-bach0/dataset.json` | `5460a13e…54cf` | MATCH |
| `corpus/mnml-shakedown/manifest.json` | `73040ef2…4e41` | MATCH |
| `corpus/fm-control/manifest.json` | `11a12a74…3754` | MATCH |

(The §2.3 shorthand paths resolve under `music/datasets/`, not `music/out/` — recorded as
**AMENDMENT A-C2**. Nothing moved; only the shorthand needed expanding.)

### 2.2 Counts, every one matching the registered table

| corpus | run id | train n | corpus files | box |
|---|---|---:|---:|---|
| Chopin (`pd-shakedown`) | `20260829T162506Z-stage0` | 19 | 24 | instrument box |
| one-artist control (`fm-control`) | `20260831T193904Z-fmctl0` | 24 | 30 | corpus box |
| Bach (`bach-shakedown`) | `20260830T104030Z-bach0` | 63 | 79 | instrument box |
| Sousa (`sousa-shakedown`) | `20260830T004943Z-sousa0` | 84 | 106 | instrument box |
| minimal techno (`mnml-shakedown`) | `20260831T023450Z-mnml0` | **159** | **208** | corpus box |
| **total** | | **349** | **447** | |

Every count equalled the registered value; a mismatch would have aborted the run.

**The number the article must get right (`RATIFY-8`):** the corpus spans **41 releases / 208
tracks**; the adapter trained on **159 tracks from 27 releases** (12 releases held out by the
by-release split, 2 excluded for telephone-rate source). *"159 tracks from 41 releases"* is true of
neither and does not publish.

### 2.3 ⚠ The registered disjointness claim is measured FALSE

`PREREG-COHERENCE.md` §2.4 registered, as a property the analysis must not ignore, that *"none of
the control's 24 training tracks is in the wide corpus's 159 training tracks … so the two training
sets are disjoint"*, and obliged the run to state that fact wherever a between-corpus comparison is
reported. **Measured by the sha256 of every source file, it is not true.**

| quantity | measured |
|---|---:|
| fm-control training tracks | 24 |
| of those, byte-identical to a track in the mnml **training** set | **21** |
| of those, byte-identical to a file anywhere in the mnml **corpus** directory | 24 |
| fm-control corpus files whose bytes appear in the mnml corpus directory | 30 of 30 |

The overlap is release-shaped, exactly as the by-release stratification implies: the control's
training set is 8 Monokrak releases × 3 tracks, and **7 of those 8 landed in the wide corpus's
training split**, not its holdout. Only `Monokrak203` (3 tracks) fell outside it.

**The corrected sentence, which this document carries at every between-corpus comparison below:**

> The one-artist control's training set is very nearly a **subset** of the wide corpus's training
> set — 21 of its 24 tracks are the same bytes. The H1 comparison is therefore not "two independent
> corpora" but "21 shared tracks, with and without 138 others". A difference in spread between them
> is attributable to the 138 additional tracks, which is a *sharper* statement of the coherence
> question than the registered framing; a null would correspondingly have been weaker evidence,
> because the two sets are far more alike than §2.4 assumed.

Filed as **AMENDMENT A-C3**, appended after §13/A-R4 rather than edited into §2.4. No metric, seed,
crop policy, N, prediction or set definition changed.

### 2.4 Crops

Registered policy (`RATIFY-3`): three 10 s windows centred at 25 %, 50 % and 75 % of each item's
**true duration** read from `ffprobe` on the source file, downmixed to mono, resampled to 48 kHz,
loudness-normalised to −16 LUFS integrated / −1 dBTP, written as 16-bit PCM WAV. A track's
embedding is the L2-normalised mean of its three crop embeddings.

| crop set | tracks | crops | serves |
|---|---:|---:|---|
| corpus, −16 LUFS, full corpus directories | 447 | 1,341 | primary (its 349 training rows), S2 (its 50 % rows), S3 |
| corpus, level untouched, training sets | 349 | 1,047 | S1 |
| renders, −16 LUFS, Set A (Rounds 1–3, 60 s) | 15 | 45 | METRIC 2 Set A |
| renders, −16 LUFS, Set B (Round 4, 30 s) | 48 | 144 | METRIC 2 Set B |
| **total** | | **2,577** | |

All 2,577 crops cut with 0 failures; every crop 48 000 Hz, 1 channel, 10.000000 s (a handful
9.999979 s from input-seek granularity); **no source was shorter than the 10 s window**. §10.2's
`1,047` counted only the training sets and could not have produced the registered S3 arm — recorded
and resolved as **AMENDMENT A-C1**.

### 2.5 Cross-box crop equivalence — checked, not assumed (§4.3)

Two corpora live on a different box from the instrument. Both boxes run ffmpeg `8.0.1-3ubuntu2`, so
crops *should* be byte-identical — but "should be" is not a receipt. The three registered crops of
one nominated track were cut on both boxes:

| window | start (s) | the training box sha256 | the evaluation box sha256 | bytes |
|---|---:|---|---|---:|
| 25 % | 81.573866 | `aed72638…acdb9` | `aed72638…acdb9` | 960,194 |
| 50 % | 168.147732 | `95d5612e…884ae` | `95d5612e…884ae` | 960,194 |
| 75 % | 254.721599 | `ce5b8bf9…a3f167` | `ce5b8bf9…a3f167` | 960,194 |

**3/3 byte-identical.** Cross-box cropping is proven for this run; the crops travelled and §4.3's
fallback (copy source audio, cut everything on the instrument box) was not used. After transfer,
**2,577 of 2,577 crops were re-verified by sha256** — the receipt STEP 3 actually requires.

Every crop vector came back with L2 norm exactly 1.0000 (min = max = mean = 1.0000 over all 2,577),
confirming A1's property that CLAP's outputs are already unit-norm, so cosine distance is `1 − dot`
with no rescaling.

---

## §3 — METRIC 1: corpus spread

**Higher = more spread out = less coherent**, *in this embedding space and no other*.
`D_pair` = mean pairwise cosine distance (the headline; it has no dependence on a centroid that
itself moves with N). `D_cent` = mean distance to the corpus centroid. Equal-N point estimate =
mean over **200 draws of N = 19 without replacement**, seed `default_rng(20260903)`. The 95 % CI is
the 2.5th/97.5th percentile of **200 bootstrap draws of 19 with replacement**, seed
`default_rng(20260904)`, with `D_pair` computed over **distinct track ids only** (the registered
duplicate-pair rule). Corpora consume the RNG streams in the §2.2 table order.

#### primary — training sets, 3 windows, -16 LUFS (§4.2)

| corpus | n | D_pair (full set) | D_cent (full set) | equal-N D_pair | 95 % CI lo | 95 % CI hi | rank |
|---|---:|---:|---:|---:|---:|---:|---:|
| Bach (`bach-shakedown`) | 63 | 0.0862 | 0.0434 | 0.0872 | 0.0672 | 0.1072 | 1 |
| Sousa (`sousa-shakedown`) | 84 | 0.0923 | 0.0467 | 0.0921 | 0.0776 | 0.1083 | 2 |
| Chopin (`pd-shakedown`) | 19 | 0.1555 | 0.0766 | 0.1555 | 0.1316 | 0.1786 | 3 |
| one-artist control (`fm-control`) | 24 | 0.2359 | 0.1203 | 0.2378 | 0.1909 | 0.2773 | 4 |
| minimal techno (`mnml-shakedown`) | 159 | 0.4468 | 0.2544 | 0.4489 | 0.3652 | 0.5239 | 5 |

#### S1 — training sets, 3 windows, NO loudness normalisation (§4.4 S1)

| corpus | n | D_pair (full set) | D_cent (full set) | equal-N D_pair | 95 % CI lo | 95 % CI hi | rank |
|---|---:|---:|---:|---:|---:|---:|---:|
| Sousa (`sousa-shakedown`) | 84 | 0.0878 | 0.0444 | 0.0878 | 0.0738 | 0.1018 | 1 |
| Bach (`bach-shakedown`) | 63 | 0.0876 | 0.0441 | 0.0881 | 0.0706 | 0.1045 | 2 |
| Chopin (`pd-shakedown`) | 19 | 0.1508 | 0.0742 | 0.1508 | 0.1134 | 0.1797 | 3 |
| one-artist control (`fm-control`) | 24 | 0.2408 | 0.1230 | 0.2428 | 0.1965 | 0.2825 | 4 |
| minimal techno (`mnml-shakedown`) | 159 | 0.4752 | 0.2735 | 0.4779 | 0.3943 | 0.5546 | 5 |

#### S2 — training sets, single centre window, -16 LUFS (§4.4 S2)

| corpus | n | D_pair (full set) | D_cent (full set) | equal-N D_pair | 95 % CI lo | 95 % CI hi | rank |
|---|---:|---:|---:|---:|---:|---:|---:|
| Bach (`bach-shakedown`) | 63 | 0.1195 | 0.0606 | 0.1207 | 0.0976 | 0.1428 | 1 |
| Sousa (`sousa-shakedown`) | 84 | 0.2096 | 0.1095 | 0.2090 | 0.1698 | 0.2504 | 2 |
| Chopin (`pd-shakedown`) | 19 | 0.2384 | 0.1201 | 0.2384 | 0.2117 | 0.2616 | 3 |
| one-artist control (`fm-control`) | 24 | 0.3119 | 0.1627 | 0.3136 | 0.2725 | 0.3449 | 4 |
| minimal techno (`mnml-shakedown`) | 159 | 0.5124 | 0.2994 | 0.5154 | 0.4295 | 0.5897 | 5 |

#### S3 — FULL corpus directories, 3 windows, -16 LUFS (§4.4 S3)

| corpus | n | D_pair (full set) | D_cent (full set) | equal-N D_pair | 95 % CI lo | 95 % CI hi | rank |
|---|---:|---:|---:|---:|---:|---:|---:|
| Sousa (`sousa-shakedown`) | 106 | 0.0919 | 0.0466 | 0.0921 | 0.0786 | 0.1105 | 1 |
| Bach (`bach-shakedown`) | 79 | 0.0954 | 0.0482 | 0.0964 | 0.0705 | 0.1240 | 2 |
| Chopin (`pd-shakedown`) | 24 | 0.1473 | 0.0733 | 0.1482 | 0.1219 | 0.1713 | 3 |
| one-artist control (`fm-control`) | 30 | 0.2523 | 0.1305 | 0.2509 | 0.1983 | 0.3078 | 4 |
| minimal techno (`mnml-shakedown`) | 208 | 0.4717 | 0.2716 | 0.4736 | 0.3851 | 0.5631 | 5 |

**The registered degeneracy showed up exactly where §5.3 said it would.** Chopin's training set has
exactly 19 tracks, so every one of its 200 without-replacement draws is the same set: its equal-N
point estimate equals its full-set value to the last digit (0.155487 both ways) and its draw-to-draw
standard deviation is **exactly 0.000000**. That is a property of the design, not a precision — which
is why the intervals come from the separate bootstrap resampling and not from the draw spread. The
other four corpora carry real draw spread (mnml 0.0396, fm-control 0.0116, bach 0.0095, sousa 0.0085).

**A second property, worth stating because it makes the equal-N step look like a no-op and is not:**
`D_pair` is a mean over pairs, and every pair is equally likely in a uniform subsample, so the
equal-N estimate converges on the full-set value — visible above, where the two columns agree to
three decimals for every corpus. The equal-N machinery is doing its real work on `D_cent`, which *is*
N-biased (mnml's full-set 0.2544 against its equal-N 0.2423), and on the bootstrap intervals, which
are what make a 159-track corpus and a 24-track corpus comparable at all.

### 3.1 H1 — is the wide corpus actually less coherent than the one-artist control?

**Predicted: yes**, with non-overlapping 95 % CIs and a one-sided permutation p < 0.05.
**Registered decision rules, both reported (`RATIFY-5`):**

| rule | result |
|---|---|
| **Primary (as ruled):** wide corpus's `D_pair` 95 % CI entirely above the control's | wide **[0.3652, 0.5239]**, control **[0.1909, 0.2773]** — **disjoint, wide above** ✅ |
| **Secondary:** two-sample permutation test, 10 000 permutations, seed `default_rng(20260905)` | observed difference **+0.2111**; **0 of 10 000** permutations reached it; **p = 0.0001** ✅ |
| do the two rules agree? | **yes** |

**H1 VERDICT: SUPPORTED.** The minimal-techno training set sits roughly **1.9× further apart** in
CLAP space than the one-artist control at equal sample size (0.4489 vs 0.2378), the intervals do
not touch, and the permutation null distribution (mean −0.0003, 97.5th percentile +0.0848) does not
come close to the observed +0.2111.

**It survives every sensitivity arm.** The wide corpus ranks last (most spread) in all four, and the
control fourth in all four: primary 0.4489 vs 0.2378 · S1 (no loudness normalisation) 0.4779 vs
0.2428 · S2 (single centre window) 0.5154 vs 0.3136 · S3 (full corpus directories) 0.4736 vs 0.2509.
**S1 is the arm that matters most** — §8.1 required the run to report whether a primary result flips
without loudness normalisation, because that would mean the metric was measuring mastering rather
than music. It does not flip; it gets slightly *larger*. The coherence signal is not a loudness
artefact.

**Stated wherever a between-corpus comparison is reported (§2.3 / A-C3):** these two training sets
are not independent — 21 of the control's 24 tracks are byte-identical to tracks inside the wide
corpus's 159. H1 is therefore the statement *"the same 21 tracks plus 138 others are far more spread
out than the 21 tracks alone"*, which is a sharper form of the coherence question, not a weaker one.

### 3.2 H2 — the three composer corpora (EXPLORATORY, no directional prediction registered)

**Measured ranking, primary arm, tightest first:** Bach **0.0872** → Sousa **0.0921** → Chopin
**0.1555** → one-artist control **0.2378** → minimal techno **0.4489**.

The pre-registration recorded weak expectations *"so that any post-hoc storytelling is visibly
post-hoc"*: Bach tightest, **Chopin close behind**, Sousa looser than both. **Bach landed first as
expected, but Sousa and Chopin swapped** — 106 separately-sourced band recordings sit almost as
tightly as two Goldberg-Variations sessions (0.0921 vs 0.0872, overlapping intervals), while one
Chopin collection on one instrument sits 70 % further apart than either. §7.2 registered in advance
that a differently-landing corpus *"is a finding to report, not a hypothesis to have had"*, so it is
reported as one and no explanation is offered here. The ordering is stable in S1 and S3; in S2
(single window) Sousa loosens toward Chopin, which is the one place the three-window policy visibly
mattered.

H2 cannot support a claim of the form "classical is more coherent than techno". At most: *these
three classical training sets, as assembled here, sit closer together in CLAP space than this techno
training set does.*

### 3.3 POST-HOC, NOT PRE-REGISTERED (A-C3)

Because §2.4's disjointness claim failed, a reader will want to know how much of H1 rides on the 21
shared tracks. Stripping them leaves 138 mnml training tracks: `D_pair` **0.4667** (full set),
equal-N **0.4667**, 95 % CI **[0.3674, 0.5334]**. The spread goes **up**, not down. **This figure is
post-hoc, is not a registered arm, and does not replace the headline** — it exists only so the
overlap cannot be mistaken for the cause.

---

## §4 — METRIC 2, Set A: did the adapters move their renders toward their corpora?

For each render, distance to its adapter's **own** training-corpus centroid (computed from the full
training set, primary arm) and to the **other** adapter's centroid. Positive difference = the
adapter moved *toward* its corpus.

**One honest caveat, registered in §6.2 and restated here:** the adapter arms prepend the adapter's
trigger token, so the captions are identical *modulo the trigger*. The base/adapter difference is
therefore **"adapter + trigger token"**, not "adapter alone".

| # | adapter | caption | seed | d(base, own C) | d(adapter, own C) | difference | direction | d(base, other C) | d(adapter, other C) | difference (other) |
|---:|---|---|---:|---:|---:|---:|---|---:|---:|---:|
| 1 | mnml 1.0 | A | 42 | 0.2303 | 0.2204 | +0.0099 | toward | 0.2579 | 0.2334 | +0.0246 |
| 2 | mnml 1.0 | B | 4242 | 0.3870 | 0.4218 | -0.0348 | away | 0.4043 | 0.4371 | -0.0328 |
| 3 | mnml 1.0 | C | 777 | 0.2871 | 0.2497 | +0.0375 | toward | 0.3412 | 0.3031 | +0.0381 |
| 4 | mnml 1.0 | D | 7777 | 0.3696 | 0.3371 | +0.0325 | toward | 0.3904 | 0.3764 | +0.0140 |
| 5 | fm-control 1.0 | C | 777 | 0.3412 | 0.3581 | -0.0169 | away | 0.2871 | 0.2926 | -0.0055 |
| 6 | fm-control 1.0 | D | 7777 | 0.3904 | 0.3510 | +0.0394 | toward | 0.3696 | 0.3348 | +0.0348 |

### 4.1 The sign tests

| test | n | toward | one-sided p | status |
|---|---:|---:|---:|---|
| **pooled six pairs** | 6 | 4 | **0.3438** | the single registered inferential test of Set A |
| mnml only | 4 | 3 | 0.3125 | **DESCRIPTIVE by construction** — min attainable p = 1/2⁴ = 0.0625 |
| fm-control only | 2 | 1 | 0.7500 | descriptive, n = 2 |

The pre-registration computed that arithmetic *before* the run and accepted it (`RATIFY-6`): with
four pairs the test cannot reach 0.05 even if all four agree, so **the direction table above is the
primary deliverable of METRIC 2, not any p-value**.

### 4.2 The adapter-strength ladder (reported as a curve, never pooled into the sign test)

| caption | seed | rung | d(render, C_mnml) | d(render, C_fm-control) |
|---|---:|---|---:|---:|
| A | 42 | base | 0.2303 | 0.2579 |
| A | 42 | mnml 0.7 | 0.2488 | 0.2597 |
| A | 42 | mnml 1.0 | 0.2204 | 0.2334 |
| B | 4242 | base | 0.3870 | 0.4043 |
| B | 4242 | mnml 0.7 | 0.4552 | 0.4566 |
| B | 4242 | mnml 1.0 | 0.4218 | 0.4371 |
| C | 777 | base | 0.2871 | 0.3412 |
| C | 777 | mnml 0.35 | 0.3012 | 0.3483 |
| C | 777 | mnml 0.5 | 0.3030 | 0.3592 |
| C | 777 | mnml 1.0 | 0.2497 | 0.3031 |
| C | 777 | mnml 1.0 ep5 | 0.2499 | 0.3113 |
| D | 7777 | base | 0.3696 | 0.3904 |
| D | 7777 | mnml 1.0 | 0.3371 | 0.3764 |

The ladder is not monotone in dose. On caption C the base sits at 0.2871 and every adapted rung sits
closer to the corpus (0.35 → 0.3012, 0.5 → 0.3030, 1.0 → 0.2497, epoch-5 at 1.0 → 0.2499), but the
0.35 and 0.5 rungs are *further* from the corpus than the 1.0 rung by more than they are from the
base. On caption B the adapter moves *away* at 1.0 (0.3870 → 0.4218) with the 0.7 rung further still
(0.4552). The epoch-5 checkpoint and the shipped checkpoint land within 0.0002 of each other on
caption C.

### 4.3 The cross-centroid control (§6.4)

The base renders are shared between the two comparisons, so `d(base, C_mnml)` and `d(base, C_fm)`
are measured on identical audio and are directly comparable. Reading the last three columns of the
Set A table: where a render moved, it generally moved toward **both** centroids by similar amounts —
pair 3 moved +0.0375 toward `C_mnml` and +0.0381 toward `C_fm`; pair 4 moved +0.0325 and +0.0140;
pair 2 moved away from both. This is §6.4's second reading — *"the adapter learned 'minimal techno in
general', or merely moved in a direction both corpora happen to share"* — and per A-C3 the two
centroids are **not** built from disjoint material (21 shared tracks), so the control separates the
three readings far less cleanly than §6.4 assumed.

### 4.4 H3 verdict

**Predicted: control yes in both pairs; wide adapter uncertain, leaning yes-but-smaller.**
**H3 VERDICT: NOT SUPPORTED.** Pooled, 4 of 6 pairs moved toward their corpus at p = 0.3438 — no
inferential support. The arm predicted most confidently, the one-artist control, moved toward in
only **1 of its 2** pairs, and its one reversal (pair 5, −0.0169) is the opposite of the prediction's
"cleaner pull". The wide adapter moved toward in 3 of 4, with margins (+0.0099, +0.0325, +0.0375)
that are *not* smaller than the control's single positive margin (+0.0394). The prediction's
reasoning — that the control's monotonically descending training loss should show the cleaner pull —
is not borne out by this instrument.

---

## §5 — METRIC 2, Set B: the Round 4 prompt spread (the best-powered arm)

Registered as **AMENDMENT A-R4** before any Round 4 clip existed. Round 4 rendered twelve captions ×
four arms at 30 s, one seed per prompt shared across its arms, on the training box beside its live
seats. Crops at 7.5 / 15 / 22.5 s per the §4.2 policy applied to a 30 s render.

**Round 4's landing receipt.** The bench polled and Round 4 landed inside the window:
`ROUND 4 END 2026-09-03T03:30:03Z`, 49 clips rendered in 1,042 s wall. `sha256sum -c MANIFEST.sha256`
→ **49 lines, 0 FAILED**. `records.json`: 49 clips, **48 distinct `output_sha256`**, 12 prompts × 4
arms plus one determinism repeat; the determinism control reproduced **byte-identical**; **72 of 72
pair checks passed** (every arm pair byte-distinct). Pairs were keyed on `output_sha256`, never on
`clip_id`.

At n = 12 the minimum attainable one-sided p is 1/4096, so **Set B is the bench's best-powered arm**.

**B-mnml**

| prompt | seed | d(base, own C) | d(adapter, own C) | difference | direction | difference (other C) |
|---|---:|---:|---:|---:|---|---:|
| P01 | 4001 | 0.3115 | 0.3485 | -0.0369 | away | -0.0122 |
| P02 | 4002 | 0.4023 | 0.2999 | +0.1024 | toward | +0.1079 |
| P03 | 4003 | 0.3386 | 0.3207 | +0.0179 | toward | -0.0007 |
| P04 | 4004 | 0.5732 | 0.4084 | +0.1648 | toward | +0.1559 |
| P05 | 4005 | 0.3936 | 0.3483 | +0.0453 | toward | +0.0458 |
| P06 | 4006 | 0.1637 | 0.1763 | -0.0127 | away | -0.0453 |
| P07 | 4007 | 0.3696 | 0.3555 | +0.0141 | toward | -0.0347 |
| P08 | 4008 | 0.3787 | 0.3661 | +0.0127 | toward | -0.0126 |
| P09 | 4009 | 0.5754 | 0.6319 | -0.0565 | away | -0.0629 |
| P10 | 4010 | 0.4138 | 0.4937 | -0.0798 | away | -0.0671 |
| P11 | 4011 | 0.2395 | 0.2437 | -0.0042 | away | -0.0073 |
| P12 | 4012 | 0.3416 | 0.3659 | -0.0243 | away | -0.0353 |

**B-fm**

| prompt | seed | d(base, own C) | d(adapter, own C) | difference | direction | difference (other C) |
|---|---:|---:|---:|---:|---|---:|
| P01 | 4001 | 0.3340 | 0.3608 | -0.0269 | away | -0.0474 |
| P02 | 4002 | 0.4339 | 0.3438 | +0.0901 | toward | +0.0562 |
| P03 | 4003 | 0.3972 | 0.3493 | +0.0479 | toward | +0.0356 |
| P04 | 4004 | 0.5660 | 0.5047 | +0.0613 | toward | +0.0779 |
| P05 | 4005 | 0.3775 | 0.3707 | +0.0068 | toward | +0.0021 |
| P06 | 4006 | 0.1955 | 0.2359 | -0.0404 | away | -0.0516 |
| P07 | 4007 | 0.4089 | 0.4216 | -0.0126 | away | -0.0193 |
| P08 | 4008 | 0.4001 | 0.4460 | -0.0458 | away | -0.0334 |
| P09 | 4009 | 0.6255 | 0.6406 | -0.0151 | away | +0.0042 |
| P10 | 4010 | 0.3957 | 0.3994 | -0.0038 | away | -0.0000 |
| P11 | 4011 | 0.3076 | 0.3037 | +0.0039 | toward | -0.0005 |
| P12 | 4012 | 0.4075 | 0.4623 | -0.0548 | away | -0.0423 |

**B-half**

| prompt | seed | d(base, own C) | d(adapter, own C) | difference | direction | difference (other C) |
|---|---:|---:|---:|---:|---|---:|
| P01 | 4001 | 0.3115 | 0.3537 | -0.0422 | away | -0.0286 |
| P02 | 4002 | 0.4023 | 0.3649 | +0.0374 | toward | +0.0548 |
| P03 | 4003 | 0.3386 | 0.3455 | -0.0069 | away | -0.0013 |
| P04 | 4004 | 0.5732 | 0.4854 | +0.0878 | toward | +0.0852 |
| P05 | 4005 | 0.3936 | 0.3720 | +0.0216 | toward | +0.0213 |
| P06 | 4006 | 0.1637 | 0.2167 | -0.0531 | away | -0.0767 |
| P07 | 4007 | 0.3696 | 0.3957 | -0.0262 | away | -0.0452 |
| P08 | 4008 | 0.3787 | 0.3744 | +0.0044 | toward | -0.0043 |
| P09 | 4009 | 0.5754 | 0.5938 | -0.0184 | away | -0.0357 |
| P10 | 4010 | 0.4138 | 0.3896 | +0.0242 | toward | +0.0024 |
| P11 | 4011 | 0.2395 | 0.2460 | -0.0066 | away | -0.0236 |
| P12 | 4012 | 0.3416 | 0.3697 | -0.0281 | away | -0.0321 |

### 5.1 The sign tests

| set | comparison | n | toward | one-sided p |
|---|---|---:|---:|---:|
| **B-mnml** | base vs mnml @ 1.0 | 12 | **6** | 0.6128 |
| **B-fm** | base vs fm-control @ 1.0 | 12 | **5** | 0.8062 |
| **B-half** | base vs mnml @ 0.5 | 12 | **5** | 0.8062 |

### 5.2 H3b verdict

**Predicted:** the control's renders closer to the control centroid than base in **≥ 9/12**; the wide
adapter's closer to the mnml centroid in **7–9/12**.
**H3b VERDICT: NOT SUPPORTED, in the arm best able to have supported it.** The control scored
**5/12** — below chance, and less than half its predicted floor. The wide adapter scored **6/12**,
below its predicted range. Neither is distinguishable from a coin flip, and this arm had the power
to detect a real effect had one been there.

### 5.3 H5 — the dose response (registered, descriptive)

Registered prediction: for each prompt, `d(·, C_mnml)` is **non-increasing** across
`base → 0.5 → 1.0`; monotone in a majority of the 12.

| prompt | d(base, C_mnml) | d(0.5, C_mnml) | d(1.0, C_mnml) | non-increasing? |
|---|---:|---:|---:|---|
| P01 | 0.3115 | 0.3537 | 0.3485 | no |
| P02 | 0.4023 | 0.3649 | 0.2999 | yes |
| P03 | 0.3386 | 0.3455 | 0.3207 | no |
| P04 | 0.5732 | 0.4854 | 0.4084 | yes |
| P05 | 0.3936 | 0.3720 | 0.3483 | yes |
| P06 | 0.1637 | 0.2167 | 0.1763 | no |
| P07 | 0.3696 | 0.3957 | 0.3555 | no |
| P08 | 0.3787 | 0.3744 | 0.3661 | yes |
| P09 | 0.5754 | 0.5938 | 0.6319 | no |
| P10 | 0.4138 | 0.3896 | 0.4937 | no |
| P11 | 0.2395 | 0.2460 | 0.2437 | no |
| P12 | 0.3416 | 0.3697 | 0.3659 | no |

**H5 VERDICT: NOT SUPPORTED — 4 of 12 monotone.** The pre-registration wrote the sentence for this
outcome in advance, and it is used verbatim: **"the dial moves the audio but not along the corpus
axis."** The dial demonstrably moves the audio — all 72 arm pairs are byte-distinct and the 0.5 and
1.0 renders differ from base and from each other — but the direction of that movement, measured
against the corpus centroid, is not ordered by dose. P09 and P10 move steadily *away* as the dose
rises; P04 is the cleanest monotone descent (0.5732 → 0.4854 → 0.4084).

### 5.4 H4 — prompt distance from the corpus (exploratory, no test)

Two group medians of the base-minus-adapter difference at dose 1.0, no test:

| prompt group | style family | n | median difference |
|---|---|---:|---:|
| **P01–P06** | techno family, nearest the corpus | 6 | **+0.0316** |
| **P07–P12** | house / breaks / trance | 6 | **−0.0143** |

**The registered direction held.** Where the caption is already in the corpus's own family, the
adapter pulls the render toward the corpus; where the caption is a neighbouring style the corpus
never contained, it pushes it away. This is the one Round 4 reading that lands as its
pre-registration expected — reported as two medians with no test, exactly as registered.

---

## §6 — The nulls, in the pre-registration's own §8 wording

**H1 is not null**, so §8.1 does not apply. Its conditional obligation was discharged anyway: the
run reports that S1 (no loudness normalisation) and S2 (single crop) agree with the primary in
direction and ranking, so the coherence signal is not entangled with loudness or with the choice of
three windows.

**H3 and H3b are null**, so §8.2 applies, verbatim. Two readings, and the bench cannot choose
between them:

> 1. the adapters genuinely did not learn their corpora; or
> 2. CLAP cannot see the axis along which they did.

The article carries both. The second is not a hedge: CLAP-class embeddings are timbre-dominated, and
a LoRA at scale 0.35–1.0 over 8 inference steps may move an axis this instrument compresses. As
registered: **a null here is fully consistent with the ear's rejection and is the more interesting of
the two outcomes** — it says the negative result is not "the adapter learned the wrong thing" but
"the adapter did not measurably move toward the thing it was trained on".

Two facts sharpen that null rather than softening it. First, the one-artist **control** — the arm
built specifically so coherence would be the only variable, whose training loss descended
monotonically for all ten epochs — is the arm that scored *worst* (5/12, 1/2). Whatever this
instrument is failing to see, it is failing symmetrically. Second, the corpus-side question and the
render-side question separated cleanly: **H1 lands decisively and H3/H3b/H5 do not**. The wide
corpus really is far less coherent; the adapter trained on it really did not measurably move its
output toward it; and the same is true of the coherent control. The two results together say the
coherence difference is real *and* that this instrument cannot connect it to the renders.

**§8.4 stands untouched.** No outcome here promotes, retires or re-ships an adapter, none overrides
the ear, and none anticipates, hints at or pre-empts ruling **D-20260903-03**, the operator's pending
Round 3 verdict.

---

## §7 — What is NOT claimed (§9, restated)

1. **CLAP distance is not musical quality.** Nothing here measures whether audio is good, pleasant,
   danceable or correct. The embedding screens; the listen confirms.
2. **Spread is not disorder.** A wide corpus is not "worse". A style may legitimately be broad.
3. **No causal claim.** Even with H1 landing as predicted, this bench cannot establish that corpus
   incoherence *caused* the ear's rejection. Two under-powered instruments agreeing is a consistent
   story, not a proof — and here they do not even agree on the render side.
4. **No generalisation to genres.** "Minimal techno" here is one archive.org subject tag sampled on
   one date under three licence tiers. It is not the genre.
5. **No generalisation to models or embedders.** One text-to-music model, one adapter recipe, one
   embedding space. A different embedder may order these corpora differently.
6. **Not a memorisation check.** METRIC 2 measures distance to a *centroid*, not to nearest training
   tracks.
7. **No claim about the held-out tracks.** The holdout was never rendered against.
8. **Not an equivalence test.** The H3/H3b nulls are "we did not detect a difference at this n with
   this instrument", never "the adapters changed nothing".

---

## §8 — Measured timings, replacing §10.6's estimates

| stage | §10.6 estimate | **measured** | detail |
|---|---|---|---|
| STEP 0 instrument load + check | 1–3 min | **17 s** | incl. 2.35 GB sha256 and a 6.7 s model load |
| STEP 1 cross-box equivalence | < 1 min | **~3 s** | 6 crops, both boxes |
| crop planning (ffprobe + sha256 every source) | not estimated | **41 s** | corpus box 26.4 s · instrument box 14.2 s |
| STEP 2 cropping | 5–15 min | **87 s** | 1,308 in 27.0 s (16-way) · 1,125 in 55.5 s (6-way) · 144 in 4.5 s |
| STEP 3 transfer 1.26 GB | **unmeasured** | **38 s** | 58.4 MB/s then 109.2 MB/s, two rsync legs |
| STEP 3 post-transfer re-verification | not estimated | **~30 s** | 2,577 / 2,577 crops re-hashed |
| STEP 4 embedding, 2,577 forwards | 10–30 min | **308.6 s** | **8.35 crops/s** on 16 CPU threads |
| STEPS 5–6 statistics | < 2 min | **44 s** | of which the 10,000-permutation test 42.1 s |
| **total, seal check to RESULTS.json** | ~20–55 min + transfer | **21 min 22 s** | 03:19:28Z → 03:40:50Z |

The bench was neither GPU-bound nor disk-bound and competed with no seat. Pre-flight aborts all
passed: instrument box 10 GiB available RAM (threshold 8), 160 G free disk; corpus box 104 GiB
available RAM, 1.1 T free disk (threshold 5 G each).

---

## §9 — GAPS, deviations, and the amendments filed

**Three amendments were filed, each appended after §13/A-R4, none an in-place edit.**

| id | what | why |
|---|---|---|
| **A-C1** | The S3 crop budget | §4.4 and §2.1 both register S3 over the full corpus directories, but §10.2 budgets only 1,047 training-set crops and §10.4 counts only 183 travelling tracks. S3 is run as registered; the corrected crop set is 2,433 before Set B. Changes no metric, seed, N, prediction or set definition. |
| **A-C2** | Path resolution | §2.3's shorthand paths resolve under `music/datasets/`; all seven hashes match. Round 4 landed at `PREREG-ROUND4.md` §5's registered path, not the one the lane brief named; the bench polled both. |
| **A-C3** | §2.4's disjointness claim is measured false | 21 of the control's 24 training tracks are byte-identical to mnml training tracks. The corrected sentence is filed rather than edited in. |

**Reading recorded, not amended.** §4.4 says S1 recomputes *"the whole analysis"*. The arm column
exists only in §10.5's METRIC 1 outputs; the METRIC 2 outputs carry no arm column. S1/S2/S3 are
therefore run as **METRIC 1 arms**, and METRIC 2 is reported on the primary conditioning only.

**Transport substitution, recorded.** the training box and the evaluation box hold no ssh host key for each other, and the
bench changes no box configuration, so the registered rsync ran as two legs via the laptop rather
than one direct hop. STEP 3's actual requirement — re-verify every crop sha after transfer — was met
in full (2,577 / 2,577).

**A1's verification clip is named by this run, not by A1.** A1 pins a shape-and-norm receipt and
describes the clip only as *"a 25 s clip cut from the pd-shakedown corpus"*. This run's choice is
recorded in §1.2 with its sha256 so the check is reproducible; the receipt itself is invariant to
which clip is used.

**Known limits of this run, beyond §9's registered non-claims.**

1. **The H1 comparison is between overlapping sets** (A-C3). The post-hoc figure in §3.3 shows the
   overlap is not driving the result, but it is post-hoc.
2. **METRIC 2's cross-centroid control is blunted** by the same overlap: `C_mnml` and `C_fm` share 21
   of the 24 tracks that define the latter.
3. **`RATIFY-7`'s deviation stands recorded**: D-20260903-02 says *"runs on whichever box holds the
   corpora"*, and the corpora are on two boxes. The bench ran on the instrument's box with the
   crops travelling. The ruling's intent — *"never on a live seat's card"* — held by construction.
4. **A disjoint-subset re-analysis of H1 is not registered** and is not offered as one.
5. **Sets A and B are reported separately and never pooled** (60 s and 30 s renders are different
   compositions), as A-R4 requires.

---

## §10 — The artifacts, and where they live

All under `estate/bench/ace-mnml-2026-09/coherence/`, with `RESULTS.sha256` covering every one:

| file | content |
|---|---|
| `RESULTS.md` | this document |
| `RESULTS.json` | every headline number with its inputs, seeds, instrument versions and weight hashes |
| `RESULTS.sha256` | sha256 of every file below |
| `crops.csv` | all 2,577 crops: corpus, source path + sha256, true duration, window %, mode, start, crop sha256, bytes, duration, sample rate, channels |
| `embeddings.npz` | float32 `(2577, 512)`, unit-norm |
| `embeddings-index.csv` | row → (corpus \| render, id, window %, mode, crop sha256, source path) |
| `metric1-corpus-spread.csv` | per corpus per arm: n, `D_pair`, `D_cent`, equal-N estimates, bootstrap CI, rank |
| `metric1-draws.csv` | all 4,000 draw statistics (4 arms × 5 corpora × 200), so the intervals are re-derivable |
| `metric2-render-centroid.csv` | per render: distance to each centroid, keyed by sha256 |
| `metric2-pairs.csv` | every pair in Sets A and B: raw distances, difference, direction, cross-centroid difference |
| `RUN.log` | the seal receipt, STEP 0's receipt, and every stage log from both boxes |
| `run_bench.py`, `crop_windows.sh` | the run's scripts, **committed before they ran** (§10.3) |

---

## §11 — Who ran this, and thanks

**LAION's CLAP** — the embedding model this entire bench is built on, `laion_clap` 1.1.7, licence
read from its bundled LICENSE file as **CC0 1.0 Universal** (its package metadata separately
classifies Apache-2.0; both readings are recorded above). **`lukewys/laion_clap`**, the host of the
`music_audioset_epoch_15_esc_90.14.pt` checkpoint, declared **cc0-1.0**. **FFmpeg** 8.0.1-3ubuntu2,
which cut all 2,577 crops identically on two different machines. **NumPy** and **PyTorch**, pinned
at 1.26.4 and 2.8.0+cpu. The corpora themselves are Creative-Commons and public-domain works by
their named artists and performers, credited in each corpus's own `PROVENANCE.md` — the one-artist
control's 24 tracks are BY-SA and carry attribution obligations wherever a clip is published.

*Pre-registered under D-20260903-02, sealed before the first embedding, run 2026-09-03T03:19:28Z →
03:40:50Z UTC. MEASURED, NEVER GATING.*
