The field guide — a series that opens up one piece of the machinery at a time

Ten Minutes with Living Artists

exhibit thirty-six The field guide
Published 2026-09-03 (UTC)
A small (human) team and a fleet of AI agents.

We taught an open-weights music generator 159 tracks of minimal techno in ten minutes — the fourth adapter of this series, the genre part two promised — and the one person who listened preferred the model untouched, both times. This is that failure, receipted like the successes: the corpus verified licence by licence, the training accounted to the second, the diagnosis typed thirty seconds after the verdict, the control built to test it, and a sealed bench that measured the diagnosis true and could not connect it to anything in the audio. One ear, unblinded, two requests: that is the whole listening evidence, and the page says so. Every clip from every round plays on the companion page, Listen for Yourself.

Where part two left off

If you read part one and part two, you know the machinery: an open-weights music generator (ACE-Step 1.5), a small detachable adapter trained with LoRA in minutes, a strength dial from zero to full, and a name tag — a made-up first word in the request that tells the model which adapter you mean. Three composers, three adapters, each provably changing the sound as the dial opened, none yet proven by anyone's ears to sound like its composer. Part two also told you, in advance, that the fourth adapter's best training pass was never saved to disk. That adapter is this one.

A genre, not a composer

The composers were a convenient place to start: one instrument, one hand, recordings already given away. A genre is where the people who make music actually live. One of us runs the machines and DJs — a human operator, and the one ear in this story — and minimal techno is that DJ's floor: music built from very little, a kick drum and a handful of sounds that change slowly enough that the changes are the point. If an adapter could learn that, the series turns from a parlour trick into a tool.

But a genre has no single hand. The corpus that taught this adapter came from the Internet Archive by matching a subject tag — "minimal techno" — across twenty years of netlabel culture. Forty-one releases matched under the licences we allow: 208 tracks, twenty-three and a half hours. Their metadata carries 28 distinct creator strings, and "creator string" is the honest phrase: one of the 28 is "VA" — various artists — on a 24-track compilation, one is a label, two name six people between them, and seventeen tracks across two releases name nobody at all. Nobody chose these tracks because they sound alike. An uploader, at some point, called each one minimal techno. One thing the DJ in the house did decide first: 63 of the files the survey found ran twenty minutes or longer — 70.7 hours, 63 percent of all the audio surveyed — and every one was thrown out, because a DJ mix is not a track.

The ear

The listening was not a test. It was a human operator — one of us — on 2026-08-31 between 03:08 and 03:19 UTC, with the clips labelled in plain sight — base, adapter at full strength, adapter at 0.7 — and the loudness and fingerprint printed under every player. Two requests, three renders each, sixty seconds a render, the same seed within a request; the 0.7 renders play on the companion page and were never ranked. Everything else pinned: 128 beats per minute, A minor, four-four, instrumental, eight rendering steps of the model's fast variant, and the engine's own output normalisation at −1 dB, which caps every clip's peak at the same ceiling — so what the clips differ by is what the adapter did inside that ceiling. One listener, unblinded, first impression: the only ear that has heard this adapter, and what it said decided the rest of the night.

One receipt before the clips, because every fingerprint on this page rests on it. The first attempt at this round loaded and unloaded the adapter inside one process, and a control found the engine does not restore itself exactly after an unload: a "no adapter" render made afterwards was not byte-identical to a clean one. The round was thrown out and re-rendered one clip per fresh process; the first replacement clip was on disk seventy-seven seconds after that check. Five of the six fingerprints reproduced exactly; the one that moved was the contaminated one. That is why a byte comparison anywhere below means anything.

The first request, verbatim. Every clip made with an adapter puts that adapter's own name tag at the front of the request — mnml-shakedown for the wide one, fm-control for the control you will meet later — and the untouched clips carry none.

minimal techno, hypnotic rolling groove, analog drum machine, deep sub bass, sparse percussion, late night

🔊 Listen 1 — nothing added · sha256 (first 16) d6b62483885e02af 🔊 Listen 2 — the same request and seed, the adapter at full strength · b1627a2f09727e89

🔊 Listen 1 — nothing added · sha256 (first 16) d6b62483885e02af
🔊 Listen 2 — the same request and seed, the adapter at full strength · b1627a2f09727e89

Each code is the first 16 characters of that file's sha256; every clip on this page is AI-generated.

The second request — minimal techno, dubby stabs, four on the floor, clicks and glitch percussion, warehouse haze — was judged the same way, and plays on the companion page rather than here. Revisiting it three nights later, the same ear found it, exactly as typed, "sounds terrible at any level (no adapter, or any adapter at any level)" — a request the model cannot render well with or without help, which makes it a demonstration of the model's limits and of nothing else.

Each clip took about eleven seconds inside the engine, twenty-one end to end. The verdict, quoted exactly as typed at 03:19 UTC: "we delayed the article releases. and for the new listens? they didn't turn out that well... both of the non-adapter verions actually sound better from this one". Thirty seconds later, from the same keyboard, the sentence this article is built on: "maybe it's because the 'minimal techno' honestly could have been anything and is all over the place?" And three nights later, asked whether this failure should be written up at all: "minimal techno article is PERFECT just because the music that came out is so bad haha, and like we say, we DO publish our failures, so let's write it up! great idea". So here we are.

Forty-one releases, licence by licence

Everything before this was public domain, by dedication or by age. This corpus is not. It is Creative Commons, in three tiers, and every licence was verified at intake against the release's own metadata record on the archive — 41 for 41, the deed's URL read off the item, not a flag on a file. (The deeds themselves were not fetched and hashed; the claim is exactly that strong and no stronger.)

licence tierreleasestrackswhat it asks of us
CC0961nothing — a public-domain dedication
CC BY974credit, by name, on anything derived
CC BY-SA2373credit, and share-alike: derived work carries the same licence

Nothing non-commercial, nothing no-derivatives, nothing with an empty licence field: those never enter a training corpus here. But 108 of the 159 tracks that eventually trained carry an attribution obligation, 50 of those share-alike, so the house took the strictest reading before the first byte was fetched. We did not decide whether copyright law makes an adapter a derivative of the songs it trained on; we decided we would not need to find out. The adapter carries the share-alike condition of the strictest licence in its corpus, and so does every clip on this page. Part two's closing line was that none of its corpora demanded credit, "which is exactly why we give it." Thirty-two of these forty-one releases do demand it, so this time the credits at the bottom of the page are the licence.

What the tier table hides. Two releases declared lossless formats at bitrates impossible for CD-quality audio; the intake predicted transcodes, and was wrong about the mechanism and right about the alarm: they are genuine, bit-exact lossless files of a 16 kHz source — telephone rate — with one track at 8. Eight tracks were kept on disk and out of training, the only exclusion the intake ever applied. The split held back twenty percent of the tracks, whole releases at a time — twelve of the forty-one, so no release straddles the line — for a test not yet run: 41 tracks out, 159 in, from 27 releases. The rung was chosen for its even spread of creators across releases, but by track it is lumpy: five large releases carry 85 of those 159 tracks, 53.5 percent, and the split rule keeps all five on the training side by construction. And nothing was levelled. The training path has no loudness stage, so the 159 tracks the model saw span 22 LU of integrated loudness — loudness measured over a whole track, on a scale where 0 is the most a digital file can hold — from −20.9 to a track above zero; the wider working set of 200 runs 32 LU, from −31.3, with peaks to +5.2 dBTP.

Two more numbers. The 159 tracks total 18.75 hours of source audio; the trainer caps every sample at four minutes, and netlabel techno runs long, so what the model saw was 10.39 hours — the cap alone discarded 8.36. Each track was captioned mechanically, as in part one: the words minimal techno, then the track's own title; the name tag is not in the caption, the trainer puts it in front. The intake's bandwidth audit also flagged 15 of the 159 training tracks as lossy-damaged and, by design, did not act on the flag — a deliberately dark mix and a codec-damaged file look identical to that instrument. On the facing table, the five dullest files it kept on purpose, the audit wrote: "if the adapter comes out dull, this is the first table to re-read." The ear's verdict is above; the tables are in the kit.

The question in court

Five days before this piece was drafted, on 2026-08-29, Sony Music Publishing and Warner Chappell, with other music publishers, sued Anthropic and two of its co-founders in the U.S. District Court for the Northern District of California. The complaint, as TechCrunch reported it the same day, alleges a "brazen campaign of illegally torrenting, scraping, and downloading copyrighted works" to train the company's models, including "millions of copies of books" carrying lyrics and sheet music. Anthropic's reply, quoted in the same report: "We disagree with the publishers' claims and we intend to defend ourselves robustly in court." We have read the report, not the complaint; the claims are the publishers', the denial is the company's, and neither is ours.

We mention it for two reasons. The first is that it is the question this page walked around. Whether a model, or an adapter like the one above, is a derivative of the music it trained on is what a court will now be asked to decide at the scale of catalogues; this arc decided it would not need to know. Every track that taught this adapter was given away in advance — by dedication, by age, or by a licence whose one demand was credit and, for share-alike, that anything derived carry the same terms — and the credits below are paid in full, whether or not the adapter worked. That is not a legal opinion about anyone else's training data. It is the only posture a small lab can take and still publish its receipts.

The second reason is disclosure. The fleet of AI agents that did most of the work on this page — the recon, the panels, the bench, and the drafting under a human's direction — runs on Anthropic's models, the defendant in that suit. We say so because a reader weighing this page should know who wrote it, and because the house rule that every clip here is labelled AI-generated would be hollow if the prose were not.

Ten minutes on the big card

The machine. Not the five-year-old card this time. This run went on the lab's workstation-class card — 96 gigabytes of memory, four times what parts one and two had (that was a desktop card in a 2021-era box) — with the day's requests routed to other boxes but the day's models still resident: about 48 gigabytes of other work sat on the card before the run began. The card is not the point; the recipe was held exactly where parts one and two left it, so the runs stay comparable: rank 64, alpha 128, dropout 0.1 on the attention projections, batch size 1 with 4-step gradient accumulation, bf16, learning rate 1e-4 on a cosine schedule, samples capped at four minutes, the fast variant, seed 42.

The training. 2026-08-31, ten passes over the 159 tracks, four at a time: forty weight updates per pass, 400 in all, against the composers' fifty to two hundred and ten. The trainer's closing banner reads 10 minutes 8 seconds and includes loading the model and two checkpoint writes; the ten passes themselves, summed from the trainer's own per-pass clock, come to 10 minutes 6 seconds — 1.515 seconds per update, that sum divided by 400. Both numbers are in the training ledger with their counting rules beside them, because "seconds per step" is only a fact once you say what a step includes. Peak memory: 5.7 gigabytes inside the trainer's own allocator, 6.7 by subtracting the card's occupancy before the run from its peak during it — two instruments, two quantities, and under seven percent of the card either way.

Here is what the ten minutes bought. The per-pass training error fell from 0.99 to 0.65, but not smoothly: down through pass seven, up at eight, down to the run's lowest at nine — 0.633 — and up again at ten. Pass nine was never saved: snapshots were written every five passes, and the file that shipped is pass ten, at 0.651, finishing above four of its own earlier passes. This is the sequel part two warned you about. Its second corpus — John Philip Sousa's marches — taught this lab to ship the checkpoint with the best measured error rather than the last one written, and the lesson was applied one run too late. The training ledger now carries a column whose only job is to record how the shipped checkpoint was chosen; this row reads, verbatim: "lowest epoch_mean among SAVED checkpoints; epoch 9 (0.6334) was the true best but save_every=5 never saved it." The best of what exists, not the best that happened.

Two smaller receipts, both unflattering. The launch asked for a 100-step learning-rate warm-up and the trainer wrote 100 into its own log; the learning-rate curve says otherwise — warm-up ended at step 40 — because one line of the trainer's source clamps warm-up to a tenth of the run, and has done so to every run in this series. And the trainer's closing banner prints a "best loss" figure that appears nowhere else — not in the per-step log, not in the per-pass means, not in the tensorboard record — so we do not quote it. The adapter came out at 88 megabytes — 88,130,280 bytes — with not one of its 44 million adjustable weights left sitting at zero, the check that catches an adapter that trained on nothing. Whatever the ear said, the adapter is not broken. It trained.

Was it the adapter, or the words?

Seven more clips, the same night, 03:23 to 03:27 UTC, tested the two cheapest alternatives to "the corpus was the problem": the adapter turned down to 0.5 and 0.35, the pass-five checkpoint instead of pass ten, and — the one that mattered — the request rewritten in the shape the adapter was trained on. Every training caption looked like minimal techno, <track title>; the two requests above look nothing like that. So two requests were written in the corpus's own shape, with invented titles nothing could have memorised: minimal techno, Schwerelos (original mix) and minimal techno, Basement Loop 7.

Here is the meter that moved. A change of one LUFS is small, about a third of a notch on a mixer's fader. On the two descriptive requests, the adapter made the audio quieter: 1.79 and 1.84 LUFS below the untouched model. What that sounds like, nobody wrote down. On the two corpus-shaped requests, the same adapter, same strength, same seed discipline, moved loudness by −0.22 and +0.49. The drop tracked the shape of the caption, not the adapter; the strongest honest reading is that round one may have been measuring the prompt.

🔊 Listen 3 — Schwerelos (original mix), nothing added · e71b8c65156cb10c 🔊 Listen 4 — the same request and seed, the adapter at full strength · 0327d06999713dce 🔊 Listen 5 — Basement Loop 7, nothing added · 5889ca518f8b726e 🔊 Listen 6 — the same request and seed, the adapter at full strength · 0cecc9a342b5bc04

🔊 Listen 3 — Schwerelos (original mix), nothing added · e71b8c65156cb10c
🔊 Listen 4 — the same request and seed, the adapter at full strength · 0327d06999713dce
🔊 Listen 5 — Basement Loop 7, nothing added · 5889ca518f8b726e
🔊 Listen 6 — the same request and seed, the adapter at full strength · 0cecc9a342b5bc04

One other meter moved, and more interestingly. Loudness range is the gap between a clip's quiet stretches and its loud ones; at full strength the adapter narrowed it on three of the four requests and doubled it on the fourth — Listen 6, 7.4 to 14.7 LU, the widest of the seventeen clips in rounds one to three. Something in it comes and goes. Two limits, stated once and governing every meter below: loudness is not quality, and the dose ladder on the corpus-shaped request drew no curve — 0.35 slightly louder than the baseline, 0.5 quieter, full strength in between, the pass-five checkpoint quietest of all — so the dial is not a slider on loudness either. What the round proves is narrower and more useful: the caption shape is a variable you must pin before you blame the adapter.

One artist instead of a tag

So, a control. If the corpus was the problem, the cheapest test is a corpus that is not — one artist, one label, one production chain, one licence tier, one sample rate — trained by the identical recipe, so that coherence is the only thing that moves. Floating Mind has ten releases on monoKraK, a netlabel that has been giving this music away under share-alike for years, and is the largest single voice in the corpus — thirty tracks, more than any other creator — and the tightest: one format, one rate, and a loudness spread a quarter of the whole corpus's. Twenty-four of the thirty trained, 1.6 hours after the cap, after the same by-release holdout. It is not a separate corpus: the thirty tracks are the same bytes, hard-linked out of the wide one, and twenty-one of the control's twenty-four training tracks were also among the wide run's 159 — the other three the wide run had held out. So the two runs share twenty-one tracks, and the wide adapter saw 138 more, from twenty other releases. The configuration diff between the two trainers is two lines, the input directory and the output directory; every hyperparameter is byte-identical, the trainer is the same commit, and the two adapters' configuration files differ in one more way that is not a difference — the order of four module names in a list.

The control trained in 1 minute 33 seconds by the banner — 60 updates over ten passes, and by the same rule as above, the ten passes' own clock divided by the updates, 1.53 seconds per update, within one percent of the wide run's pace. Its error fell monotonically, every pass lower than the last, against the wide run's bounce; the two runs' loss values are not comparable, only the shape is. It was rendered on 2026-08-31 at 19:48 UTC on the two corpus-shaped requests, against the very same untouched clips as Listens 3 and 5, byte-identical, which is what makes the comparison fair:

🔊 Listen 7 — Schwerelos (original mix), the one-artist adapter at full strength · 5db77e0345b032c5 — against Listen 3. 🔊 Listen 8 — Basement Loop 7, the one-artist adapter at full strength · f0ebb8d258039c5c — against Listen 5.

🔊 Listen 7 — Schwerelos (original mix), the one-artist adapter at full strength · 5db77e0345b032c5 — against Listen 3.
🔊 Listen 8 — Basement Loop 7, the one-artist adapter at full strength · f0ebb8d258039c5c — against Listen 5.

On Basement Loop 7 the meter moved +1.91 LUFS, four times the wide adapter's +0.49 on the same request and the same way; on Schwerelos, −0.31.

The ear ruled three days after the meters, with the clips labelled as before. On Schwerelos, exactly as typed: "those two actually sound pretty similar", and two minutes later, "the adapter one might actually be slightly better though" — a lean, hedged as typed, and the first time in this arc an ear has leaned toward any adapter. On Basement Loop 7: "the non-adapter version sounds way better, the adapter version adds in weird clippy hi hats" — the untouched clip again, on the request where the control's meter had moved most. The same sitting produced the first words on round two's own pair, Listens 3 and 4, our adapter on Schwerelos: "they're similar", then "but also very different", then "the adapter version adds weird rythms that the non-adapter version doesn't" — a difference heard and named, no preference stated. So the control survived its falsification test on the meters and the loss curve, and only half on the ear: one pair similar with a slight lean to the control, one pair to the untouched model. The "clippy" hats are the ear's word, not the meter's — that render's flat-factor row, in the receipt below, reads zero.

Then a number for coherence. On 2026-09-03, between 03:19 and 03:41 UTC — eleven minutes after the file saying what would be measured, what was predicted and what a null would mean was sealed, so the numbers could not be chosen after the answer was known — a bench ran. The instrument is CLAP, a free, openly licensed model that listens to audio and turns it into a list of numbers arranged so that two clips that sound alike land near each other; it ran on CPU. Every training track of all five corpora in this series was cut into three ten-second windows, loudness-matched, embedded, and the spread of each corpus measured as the mean distance between its tracks at equal sample size — nineteen, the smallest training set — with the whole measurement re-drawn from the same tracks two hundred times over, which is where the ranges come from. Higher means further apart, less coherent, in this space and no other.

corpustraining tracksspread at equal N95 % range
Bach630.08720.0672 – 0.1072
Sousa840.09210.0776 – 0.1083
Chopin190.15550.1316 – 0.1786
one artist — the control240.23780.1909 – 0.2773
the tag — minimal techno1590.44890.3652 – 0.5239

Prediction one held, decisively. The tag-assembled corpus sits about 1.9 times further apart than the one-artist corpus, the ranges do not touch, and none of ten thousand random relabelings of the two sets produced a gap that large. It survives every variation registered in advance — without loudness matching, with a single window per track, over the whole corpus directories instead of the training sets. Could have been anything is, in this space, simply true. The composer corpora, for which nothing was predicted, landed far below both techno corpora — even one artist's techno sits half again as far apart as Chopin's nineteen recordings. And the overlap above was a claim the sealed file got wrong — it called the two training sets disjoint — so the correction is filed beside the result, with one number computed afterwards, not pre-registered: strip the 21 shared tracks from the wide corpus and its spread rises to 0.4667. The overlap does not explain the gap.

Then the sharper question. For every base-and-adapter pair the first three rounds rendered — six pairs, five of them on this page and one on the companion — the bench measured whether the adapter's render sits closer to its own corpus's centre than the untouched render does. Prediction three was that the control's renders would, cleanly, and the wide adapter's would, less so. Neither happened. Across those six pairs, four moved toward their corpus and two away, a coin's worth. Across the forty-eight renders of the request spread below — twelve pairs per comparison, the bench's best-powered test — the wide adapter at full strength moved toward its corpus on six requests of twelve, the control on five, the wide adapter at half strength on five; and the dial, from 0 to 0.5 to 1.0, moved a render steadily toward the corpus on four requests of twelve. In the pre-registration's own words for exactly this outcome: the dial moves the audio, but not along the corpus axis.

Two readings, written down before the run, and the bench cannot choose between them: the adapters genuinely did not learn their corpora, or this instrument cannot see the axis along which they did. The second is not a hedge — embeddings of this kind are dominated by timbre and texture, and a small adapter over eight rendering steps may move something the instrument compresses away. Together they say something exact: the coherence difference is real, by about a factor of two, and this instrument cannot connect it to anything in the renders — the coherent control scored no better than the incoherent tag. None of it measures musical quality. A corpus can be wide and good.

Twelve requests, four ways

The ear's verdict rested on two requests, and the request makes a huge difference. So before publication twelve were written down in advance, from the corpus's own genre outward — six techno, four house, electro breaks, uplifting trance — each with its own tempo, and each rendered four ways at thirty seconds: nothing added, the wide adapter at half strength, the wide adapter at full, and the one-artist adapter at full, one seed per request shared across the four. The first request is Listen 1's, verbatim.

requestbase, LUFSwide @ 0.5wide @ 1.0one artist @ 1.0
minimal techno (Listen 1's)−18.40+0.95+1.16+1.45
dub techno−17.36−2.29−1.92−1.93
Detroit techno−13.80+0.21+0.28−0.82
melodic techno−18.41+0.43+1.90+1.60
hard techno−17.52−0.26+2.35+2.56
acid techno−15.85+1.26+1.40+1.57
deep house−16.03−0.89−0.72−0.59
tech house−16.44−0.15−0.23+0.57
disco house−14.70+0.80+0.58+1.37
progressive house−13.44−0.30−0.83−0.79
electro breaks−15.98+1.29+1.02−0.49
uplifting trance−14.13+0.84−0.33+0.10

Adapter columns are the change in integrated loudness against that request's own base render. Rendered 2026-09-03 between 03:13 and 03:30 UTC, plus one repeat that proved the process deterministic: 30 seconds, 48 kHz stereo, the fast variant at eight steps, A minor throughout, tempos from 122 to 145, output normalised at −1 dB, about eleven seconds inside the engine — the same eleven a sixty-second clip takes, because the engine's cost is mostly fixed overhead, not music.

Read the first row against Listens 1 and 2. The same words, a different seed and half the length: the untouched render alone moved 3.66 LUFS between the two rounds, bigger than anything the adapter did in either, and the adapter that made those words quieter at sixty seconds made them louder at thirty. A thirty-second render is a different composition, not a shorter one. Across the twelve, four requests came out louder under every adapter and three quieter under every adapter; on the other five they disagree, and on two of those the dose itself flips the sign — hard techno goes quieter at half strength and 2.35 louder at full, trance the reverse. The meter does move about twice as far on the six techno requests nearest the corpus as on the six farther away — 1.35 LUFS against 0.66, averaged over the three adapter renderings — the round's one pre-stated expectation, and the only thing on this table that behaved. Whatever these adapters do to a render, the loudness meter will not be the instrument that names it. One pair from this round to hear beside Listens 1 and 2 — the request one door down from the corpus's own genre, and the one whose four renders all agree on direction:

dub techno, deep chord stabs drenched in delay, soft muffled kick, tape hiss, slow underwater swing

🔊 Listen 9 — nothing added · 6f0b2d81c610b7e9 🔊 Listen 10 — the same request and seed, our adapter at full strength · 31f288228c7aab86

🔊 Listen 9 — nothing added · 6f0b2d81c610b7e9
🔊 Listen 10 — the same request and seed, our adapter at full strength · 31f288228c7aab86

One more receipt from this round, published rather than tidied. Measured after the fact on all sixty-five raw renders of the arc: none exceeds the engine's −1 dB ceiling, but eighteen carry flat-topped stretches — runs of samples pinned at the peak, the model's own ceiling, scaled down afterwards by the engine's normalisation — untouched renders included, the hard-techno request worst of all. Whatever distortion you hear is the model's, not the adapter's; the peak and flat-factor rows are in the kit. All forty-eight play on the companion page, Listen for Yourself, beside every clip from the three rounds before — sixty-five in all, each offered raw and brought to a common loudness for fairer listening. The sounds are all over the place; that is the finding, and it is yours to hear.

And the overarching theme, from the one ear after all sixty-five, exactly as typed: "it's just not very good at making good edm (yet, this model / version at least)""most of these samples, with or without an adapter, are not pieces of music i would ever listen to for pleasure haha." One listener's taste, stated as such. But it turns the open question of this piece around. We asked whether an adapter could teach this model minimal techno; the ear's answer is that the model, at this version, is the weak instrument, with or without help — which makes the next question not how to teach it, but whether it is a floor worth building dance music on at all.

What we did NOT measure

Whether any adapter in this piece sounds like its corpus — one listener, labels visible, two requests, is not a listening test, and the one-artist control has been heard by that listener once, on two requests. What any of it was played back on — no monitors, room or level were recorded. Whether the untouched model is any good at this genre — never tested by an instrument; the one ear, above, says it is not, with or without help. Whether an unlevelled corpus taught the adapter a level rather than a style — the survey that assembled it asked for a normalisation pass, and none ran. Whether the adapter that finished at pass nine would have fared better — it does not exist. What a DJ-curated subset of the same 208 tracks would teach — the obvious next run, and one of us will pick it. Whether the 41 held-back tracks sit any closer to what the adapter renders — never scored. Whether an embedding space that sees corpus spread this clearly can see an adapter's effect at all — nothing here calibrates that instrument against any ear. And the ledger the lab now keeps holds the wide run end to end and the control not at all; its receipts are on disk, awaiting the ingest they are owed.

What to take with you

  • Same recipe, same trainer, same ten passes: three composer corpora produced adapters that measurably did something; a genre corpus of 41 releases and 28 creator strings, 159 tracks of which trained, produced one the only listener preferred to switch off. The recipe did not change; the corpus did. That is the difference we can point at, not the cause we proved — and the ear's overarching verdict, after every clip, is that the model itself is the weak instrument here.
  • The ear's hypothesis — "could have been anything" — cost ninety-three seconds of training to put on the bench, and measured true by about a factor of two. Whether it is the lever is still open: this instrument saw no pull toward the corpus from either adapter, and the one ear that has heard the control gave it a slight lean on one request and a loss on the other.
  • The caption is a variable. A model trained on genre, title and asked in prose is being asked in a language it never saw, and the loudness meter caught it. Pin the request before you blame the weights.
  • A best pass that is not saved is not a best pass; the ledger now says how every shipped checkpoint was chosen. And the credits are owed whether or not the adapter worked: thirty-two releases require them by licence, and the adapter and every clip here are share-alike because the corpus was.
  • One person, one night, labels in plain sight, two requests. That is the whole listening evidence for the word "failure" on this page, and your ears are as good as ours.

How to check our work

Every number on this page traces to the training log's own per-pass lines, the render service's records, the intake receipt read at acquisition time, or the research ledger the lab now keeps. The data kit beside this article carries the full fingerprints of every clip and the map from each player to its file, the request payloads verbatim, the training logs' pass lines for both runs, the intake manifest with its 208 checksums and the 41 licence records copied from the archive's own metadata, the bandwidth audit with its flagged rows, the loudness, peak and flat-factor records for every clip in all four rounds, the sealed pre-registration of the coherence bench with its five filed amendments — two of which correct the registration itself — and its results down to every embedding, and the Round 4 records and blind-sheet layout (the key stays sealed until a verdict is recorded); the companion page carries every clip with its own manifest. The kit's records carry the hub's usual CC BY 4.0; the audio beside this page is share-alike, and is filed apart from the kit for that reason. Where two honest instruments disagree — the trainer's banner and its own per-pass clock — the kit carries both readings and the rule for each.

Who ran this, and thanks

The corpus first, because this time it is owed. Forty-one releases from the Internet Archive's netlabel collections under CC0, CC BY and CC BY-SA, listed in full after this section: each row links the archive item it came from and names the licence recorded on that item, and the licence instruments are spelled out, deed by deed, above the tables. Two rows read "creator not stated in the record," because an invented artist name in a credits list would be the worst possible bug. One netlabel's name contains an expletive; it is the legally correct attribution string and it is printed as recorded. The model: ACE-Step 1.5 (MIT — its row on our licence ledger); its model card asks that AI involvement be disclosed, and every clip on this page is AI-generated and says so. The trainer: Side-Step, the corrected-timestep LoRA trainer vendored in ACE-Step 1.5; its upstream licence is CC BY-NC-SA 4.0, though the vendored copy says it follows ACE-Step's MIT — we proceed on the stricter reading, and wrote to ask. The coherence instrument: LAION's CLAP — the laion_clap package at 1.1.7 and the lukewys/laion_clap checkpoint host that carries the music weights the bench ran on; the licence file on disk reads CC0 1.0 for the code and the checkpoint repository declares the same, though the package's own index metadata says Apache 2.0 — both permissive, nothing turns on it.

And beneath all of it, the open tools this work stood on without modifying: PyTorch and torchaudio, Hugging Face's PEFT library, NumPy, FFmpeg — which decoded every training file, cut every crop and measured every loudness figure — and the Internet Archive, which carried all forty-one releases and every licence record we read, fetched under a user agent that names us and says how to reach us: strata2signal-research/1.0 (private research; contact via strata2signal.com). None of them owed us anything. A small team and a fleet of AI agents did the work; the humans signed the numbers — how that works is next door.

The licence instruments. Each label in the tables stands for the deed at the address beside it; the items record each address with an http or https scheme, and the dedication appears under both, which is why ten spellings cover nine deeds.

labeldeed
CC0creativecommons.org/publicdomain/zero/1.0/
BY 1.0creativecommons.org/licenses/by/1.0/
BY 3.0creativecommons.org/licenses/by/3.0/
BY 4.0creativecommons.org/licenses/by/4.0/
BY-SA 2.5creativecommons.org/licenses/by-sa/2.5/
BY-SA 3.0creativecommons.org/licenses/by-sa/3.0/
BY-SA 3.0 CHcreativecommons.org/licenses/by-sa/3.0/ch/
BY-SA 3.0 UScreativecommons.org/licenses/by-sa/3.0/us/
BY-SA 4.0creativecommons.org/licenses/by-sa/4.0/

Each release's identifier links to its archive item.

CC0 — nine releases, nothing owed, credited anyway.

releasecreatorfilesarchive itemnotes
RW-Techordings presents [RWT-012] 30-hcir & Bertha James - Split OneRichard Wilmer230-hcir-berthaJamesSplit-RWT-012in the corpus
Black SaturdayArchivOne1archivone-black-saturdayin the corpus
Flourish // PerishBraids10braids-flurish-perishin the corpus
RW-Techordings presents [RWT-010] ElectRICHual - Predictable (Mixes)ElectRICHual2ElectRICHual-Predictable-Mixesin the corpus
Hardware TechnoDavid Murillo Diaz24hardware-technoin the corpus
In NovationKλпξiðλ1in-novationin the corpus
knolios_momentsknolios6knolios_momentsin the corpus
Mr.Dee_D1Mr.Dee3Mr.Dee_D1in the corpus
DawnShinji Wakasa12shinji-wakasa-dawnin the corpus

CC BY — nine releases, attribution required.

releasecreatorlicencefilesarchive itemnotes
Aaron Goldbody Vs. Luke The Wizard - Z Veseljem Exported EP [kahvi019]Aaron goldbody vs. Luke the wizardBY 1.02kahvi019in the corpus
VA-Let the bass ruin your speakers [ONMP215]VABY 3.024onmp215ain the corpus
[shoki005g] - Various - Grey EPShoki RecordingsBY 3.04shoki005gin the corpus
STE34. Jyolrstion (kind of placeless interlude)kraiBY 3.05STE34in the corpus
Nisiru Remixed EPstroboskopBY 3.06stroboskop-label013in the corpus
Reho RemixedstroboskopBY 4.06Stroboskop033in the corpus
[Tranz023] Holocaos- Metamorfose Computador EPCaue MirandaBY 4.08tranz023Holocaos-MetamorfoseComputadorEp_140in the corpus
[unfound88] happy in novi sadcreator not stated in the recordBY 4.015unfound88in the corpus
[unfound91] jukka-pekka kervinen - coffee beansjukka-pekka kervinenBY 4.04unfound91in the corpus

CC BY-SA — twenty-three releases, attribution required and share-alike. Ten are the one-artist control's corpus; two were kept on disk and excluded from training for their 16 kHz source; the notes column says which.

releasecreatorlicencefilesarchive itemnotes
Chuänchö - Profilaxis [inoquo001]ChuänchöBY-SA 2.54inoquo001in the corpus
[inoquo018] Ol - liturgyOlBY-SA 2.54inoquo018in the corpus
Tomorrow EPDerek ScottBY-SA 2.53RAR005_Derek_Scott_Tomorrow_EPin the corpus
Sequentialwork - Sequentialwork EP [bump188]Kiyoshi TomeharaBY-SA 3.03bump188in the corpus
[inoQuo070] v.a. - we are aliveinoQuoBY-SA 3.07inoQuo070in the corpus
[monoKraK 134] Sin Amigos "Ningun Amigo"Sin AmigosBY-SA 3.0 CH2monokrak134SinAmigosnignAmigoin the corpus
(monoKraK173) Yann Detroit VS Floating Mind "Dust"creator not stated in the recordBY-SA 3.02MonoKraK173YannDetroitVSFloatingMind_Dustin the corpus
(monoKraK197) Albert Negredo "Lithium Serendipity"Albert NegredoBY-SA 3.02MonoKraK197AlbertNegredoLithiumSerendipityin the corpus
[monokrak199] Floating Mind "Schöni"Floating MindBY-SA 3.03Monokrak199FloatingMind_Schnicontrol corpus
[monoKraK200] Floating Mind "Birthday Accelerated"Floating MindBY-SA 3.03Monokrak200FloatingMind_Birthday_Acceleratecontrol corpus
[monokrak 203] Floating Mind "A Mind Is Floating"Floating MindBY-SA 3.03Monokrak203FloatingMind_AMindIsFloatingcontrol corpus
[monoKraK 204] Floating Mind "Disko Chill"Floating MindBY-SA 3.03monoKraK204FloatingMind_DiskoChillcontrol corpus
(monoKraK205) Floating Mind "Voyager EP"Floating MindBY-SA 3.03Monokrak205FloatingMind_VoyagerEPcontrol corpus
[monoKraK208] Floating Mind "et si ..."Floating MindBY-SA 3.03MonoKraK208FloatingMind_et_sicontrol corpus
[monoKraK84] Various Artists "Smoked mono vol.7"Zaid Edghaim, Kimo-S, Alicia Hush, Floating MindBY-SA 3.03monokrak84VariousArtistssmokedMonoVol.7in the corpus
[TFN110] Lik-o - TRASHFUCK NET EPLik-oBY-SA 3.0 US3TFN110in the corpus
[TFN211] Graffiti Mechanism - ParkGraffiti MechanismBY-SA 3.0 US2TFN211in the corpus
[TFN223] Graffiti Mechanism - RENukeFamRMXS-EPGraffiti MechanismBY-SA 3.0 US4TFN223kept on disk, excluded from training (16 kHz source)
[TFN247] Graffiti Mechanism - RE-RMXD-RENukeFamRMXS-EPGraffiti MechanismBY-SA 3.0 US4TFN247kept on disk, excluded from training (16 kHz source)
[monokrak 209] Floating Mind "Ready For Flying"Floating MindBY-SA 4.03Monokrak209FloatingMind_Ready_For_Flyingcontrol corpus
[monoKraK214] Floating Mind "Hidden Passion"Floating MindBY-SA 4.03Monokrak214FloatingMind_Hidden_Passioncontrol corpus
[monokrak 216] Floating Mind "Thru Lines"Floating MindBY-SA 4.03Monokrak216FloatingMind_Thru_Linescontrol corpus
[monokrak 217] Floating Mind "Spatial Moments"Floating MindBY-SA 4.03Monokrak217FloatingMind_Spatial_Momentscontrol corpus

The rest of the seminar

Part one — Teaching a Music Model Chopin in Five Minutes — is the field guide; part two — Half an Hour with Dead Composers — is the three composers and the algebra of blending them. Coming next: what a corpus actually costs to teach, by the second of audio rather than the track, which the 10.39 hours above make concrete; the memory question part one left open; and the curated retrain — the same 208 tracks, chosen by a DJ instead of a tag. One piece at a time.

elsewhere in the workshop

a strata→signal property · hello@strata2signal.com · say hello