## The verdict

### What the licence string alone would have bought you

Start with the naive harvest, because it is the one an automated intake performs. Taking every track-bearing CC0 item in the house pool on the strength of its `licenseurl` — the field the search API returns, the field the minimal-techno survey trusted and was right to trust — yields **2,414 tracks / 158.5 hours**. Of those:

- **776 tracks** are commercially released music the uploader had no standing to license — including a major-label catalogue whose own subject field records the streaming service it was ripped from, and an animation studio's library (F-COMMERCIAL).
- **792 tracks** are not music at all — sample packs and sound-effect libraries (kitchen noises, doorbells) that the tag `house` attached itself to (F-NOTMUSIC).
- **286 tracks** matched the *word* house, not the genre: a live set at a venue called House of Blues, a spoken-word item on the House of Commons, a chamber ensemble with *house* in its name (F-NONGENRE).
- **198 tracks** are machine-generated (F-AI).

**These classes overlap, so they do not sum: the union is 1,615 tracks across 57 items — 67% of every CC0 track in the pool.** (A further 1,417 CC0 tracks are fenced on style under D-20260905-84, which is a separate axis from provenance.)

**Had the string been trusted, a house adapter trained on this rung would have learned from a major label's back catalogue, a cartoon studio's library, somebody's kitchen, and a lecture on parliamentary history.** It would have been an artefact the estate could not ship, could not describe honestly in an article, and could not defend if asked where the audio came from. That is the finding of this survey: **not a volume problem, a provenance problem** — and it is invisible to every field the search index exposes.

The minimal-techno survey did not hit this, because its corpus was netlabel-dominated and a netlabel sets the licence on records it owns. The house pool is not netlabel-dominated. That structural difference — not genre, not size — is what separates the two surveys, and it is why the same method yields a different recommendation here.

### Usable tracks, hours and artists per licence rung

| Rung | Provenance | Items | Tracks | Hours | Artists |
|---|---|---:|---:|---:|---:|
| CC0 | jamendo-mirror | 0 | 0 | 0.0 | 0 |
| CC0 | netlabels-curated | 3 | 4 | 0.3 | 1 |
| CC0 | third-party-upload | 108 | 414 | 39.0 | 65 |
| BY + BY-SA | jamendo-mirror | 1,225 | 3,366 | 277.4 | 494 |
| BY + BY-SA | netlabels-curated | 148 | 940 | 81.0 | 93 |
| BY + BY-SA | third-party-upload | 457 | 1,305 | 113.0 | 286 |
| **All open rungs** | jamendo-mirror | 1,225 | 3,366 | 277.4 | 494 |
| **All open rungs** | netlabels-curated | 151 | 944 | 81.3 | 93 |
| **All open rungs** | third-party-upload | 565 | 1,719 | 151.9 | 348 |

### The Jamendo-mirror finding

**1,225 of the 1,941 usable items — 3,366 tracks, 277.4 hours — are Jamendo albums mirrored onto archive.org.** They are the largest single block in the pool, and their provenance is *good*: a Jamendo artist chose the licence on their own recording, which is exactly the provenance story arm 1's corpus already rests on.

That is also the problem. **Arm 1's pool is Jamendo.** Material mirrored from Jamendo is not a new source; it is more of the same source, and some of it is the same music.

Measured overlap, by artist name against arm 1's MTG-Jamendo open house-family set: **49 of 643 mirror artists are artists arm 1 already has, accounting for 1,701 of 4,841 mirror tracks (35%).** This is artist-level overlap, **not** proof of track-level duplication — a mirrored album by an artist arm 1 already has may still contain tracks arm 1 lacks. It must be deduplicated track-by-track before use, and **this survey did not do that**: it would need either the audio or Jamendo track ids, which the archive records do not carry. Read 35% as the share requiring a dedup pass, not as a discount already taken.

Two second-order costs the raw count hides. The mirror re-weights the corpus toward artists arm 1 is already heavy in, which pushes *against* arm 1's stated purpose — arm 1 (WIDE) is the size-and-variety arm; arm 3 is the one-artist control. And duplicated tracks straddling a train/held-out split break the split, which is the D-20260831-03 pattern the workflow already treats as binding.

### Against the arm-1 Jamendo pool

Arm 1 holds **625 tracks / 53.1 h / ~141 artists** of MTG-Jamendo house + deephouse, licence-clean per track via `audio_licenses.txt`. This survey reproduced that pool independently from the same metadata as a check: deephouse open = **89**, matching CONTEXT.md exactly; house open = **560** against CONTEXT's 536 — a 4% difference traceable to which MTG tag file is read. The lane report records it.

| Addition | Items | Tracks added | % on pool | Hours added | Provenance | Dedup risk |
|---|---:|---:|---:|---:|---|---|
| netlabels-curated | 151 | 944 | +151% | 81.3 | good — label set it | none |
| jamendo-mirror | 1,225 | 3,366 | +539% | 277.4 | good — artist set it | **high — 35% same-artist** |
| third-party upload | 565 | 1,719 | +275% | 151.9 | **unverified self-assertion** | unknown |
| everything usable | 1,941 | 6,029 | +965% | 510.6 | mixed | high |

### Recommendation

**Add the netlabels-curated subset — 944 tracks / 81.3 h / 93 artists. Decline the third-party uploads. Treat the Jamendo mirror as a deduplication problem rather than a new source**, and take it only if arm 1 wants Jamendo material it does not already hold — in which case it must be deduplicated against the existing pool first.

The netlabels subset is the one clean win: **+151%** on arm 1's track count and **+153%** on its hours, from 93 artists, with a provenance story as good as the existing pool's and no overlap with it. It is genuinely a different source — netlabel releases are not on Jamendo.

The third-party remainder — **1,719 tracks / 151.9 hours** — is the tempting block and should be declined. Arm 1's corpus is currently 100% licence-clean per track. Mixing in a source whose licence field is unverified self-assertion converts a clean corpus into one that can only be described with an asterisk, in exchange for volume the arm does not need. **Every contaminated item this survey found by hand came from this class.** The trade is bad.

If the orchestrator wants that volume regardless, the minimum defensible path is a per-item human provenance pass over the third-party rows before any download — not a spot check, since the contaminated items are among the *largest* in the pool and several evaded every mechanical screen. The per-item table and the JSONL both carry a `provenance` value, so any subset can be taken without re-running the survey.

**On the licence rung the brief asks about:** the ladder is not the constraint here. CC0 is the *dirtiest* rung in this pool, not the cleanest — it is the string an uploader reaches for when putting someone else's record online, while BY and BY-SA are used disproportionately by netlabels and Jamendo artists who actually own what they license. The minimal-techno survey reached a compatible conclusion by a different route. **"Enough" does not arrive at a licence rung. It arrives at a provenance class.**

