The field guide — a series that opens up one piece of the machinery at a time
It Was Already There
exhibit twenty-one The field guide
Published 2026-08-24 (UTC)
a small (human) team and a fleet of AI agents
Models don’t conjure knowledge out of nothing. They mine two places that were waiting long before the machines arrived: patterns sitting unread in archives people spent decades building, and regions of rule-space nobody had walked. Then they hand what they find to the one authority that has ever settled a question about the world — the world itself, put to the question by people patient enough to ask properly.
Since 1994 a contest has graded protein-structure predictions the only way that settles anything: blind. Entrants get sequences whose real shapes are settled in a laboratory but sealed until the scoring is done. In late 2020 the people who run it looked at one entrant’s scores and said that for single proteins, in a qualified sense, the problem was solved. Structural biologists spent the same month listing what that left out — how proteins pair up, how they move, what one mutation does — and they were right too.
The story gets told as a machine making knowledge out of nowhere. It is close to the opposite. Nearly everything AlphaFold needed already existed, free to download, and none of it by accident: the sequence databases, accumulating since the 1980s, and the Protein Data Bank, half a century of structures solved one at a time in laboratories. The second is what the model trained against. The experimentalists were its first teachers before they were ever its examiners.
Here is the signal it read. A protein is a chain that folds, and two links far apart along it can end up pressed together; evolution then keeps the pair compatible, so a mutation at one position that would break the contact tends to get answered by a compensating mutation at the other. Line the protein up against every related sequence anyone has recorded — most never assigned to a named species — and those answered mutations show up as positions that move together. Biologists had called that co-evolution for decades; what began in 1994 — the same year, unrelatedly — was reading it off an alignment as evidence that two positions touch. It half-worked for years, because correlation spreads and a pair that never touches can look exactly like a pair that does. By the early 2010s several groups had statistics sharp enough to separate them, and the fingerprints became usable contact predictions — real, published, public.
What was missing was a reader. The signal at any single pair is faint; it sharpens only weighed against every other pair at once, across millions of sequences. The 2020 system, AlphaFold 2, weighed a whole alignment in one go, every pair of positions against every other, letting sequence statistics and candidate geometry revise each other pass after pass until the disagreement stopped shrinking.
Then the check, which is the part that matters. The program scored a median of 92.4 across every structure it was scored on, on the main 0-to-100 scale the contest uses: lay the prediction over the answer, score what share landed within each of four distance tolerances, and average them. Around ninety is where a prediction stops being a guess about a structure and starts being usable as one.
And note the tell that argues for the archive rather than the magic: where the archive is thin, so is the answer. Proteins with few known relatives — a shallow alignment, too few sequences to stack up — come back less reliable on average, and the program’s confidence drops on them. That score counts for something because it was built as a prediction and benchmarked against real structures: a checked estimate, not a tone of voice. (It also drops where a stretch has no single shape to predict. A low score can mean I could not read this, and it can mean there is nothing here to read.)
That is the same principle the piece next door lays out. A model is a compression of what it was trained on — the training objective, not a figure of speech — so it can only be as detailed as its source was: thin archive, smooth invention; thick archive, the real detail. The thinness is in the evidence the model gets for this protein rather than in what it was trained on — a different variable, the same shape of failure. Nothing was conjured. Something was finally read.
The piece next door taught us to say it this way: almost nothing is true chaos — when you look closer. Co-evolution looked like noise until somebody could hold enough of it in view at once.
The first engine: the archive nobody had read
Call it present in the data, unmined.
In 2020 an MIT-led team published a model trained on 2,335 molecules, each labelled by whether it stopped E. coli growing in a dish. They pointed it at compound libraries — catalogues of molecules that already exist and can be ordered — and asked it to rank them. Near the top of the first came a molecule that had sat on that shelf for years, worked up earlier as a diabetes candidate and dropped before it reached a clinic, structurally unlike the antibiotics the field was working through. They named it halicin. In the dish it killed bacteria that shrug off existing drugs; in mouse models of particular infections it worked too. On the published record it has never been through the human trials that would decide whether it is a drug.
That first library held 6,111 shelved compounds. The hundred-million-compound catalogue the same paper screened afterwards is the number everyone remembers, and it surfaced a different, smaller handful — not halicin. Six thousand molecules nobody had thought to re-rank is the sharper indictment of the missing reader anyway.
Nothing there is alchemy: the structure-activity pattern — which molecular shapes actually do the job — was latent in the training labels. What the model added was a willingness to rank whatever it was handed, without a medicinal chemist’s prior about what an antibiotic should look like. That prior is one of the things that had stood between the field and this compound, along with the plainer fact that nobody had thought to put that molecule in that dish. The model did not escape human priors either; it inherited a different set, learning from tests people chose to run on a library assembled for something else. A trained eye is a prior, and a prior pays for itself right up to the moment the answer is somewhere else.
Halicin is not a drug you can be prescribed. What it has is receipts from dishes and mice.
The second engine: the rooms the rules imply
The second engine doesn’t need the pattern to be in the data at all. It needs a rule that holds past the rows it was fit to.
Kepler is the cleanest case. He squeezed decades of Tycho Brahe’s observations into three laws plus a short list of numbers per planet, and the shortness was the discovery. He thought he was reading God’s geometry, and the shortness only looks like the discovery from here. A compression that captures the generating rule emits entries nobody has recorded — though emitting them was not new: the Alfonsine and Prutenic tables had issued future positions for centuries, from a rule that was wrong. Kepler’s Rudolphine Tables of 1627 placed planets at dates that had not happened yet, among them a transit of Mercury across the sun in November 1631. Pierre Gassendi watched it the year after Kepler died — the first time any human being had seen a planet cross the sun.
The prediction held, imperfectly, in the way that counts. Mercury crossed on the day Kepler named, about five hours off the hour he named, and its disc came up far smaller than anyone expected. A rule good enough to put an unrecorded event on the calendar was not yet good enough to time it. The event was in no row of Tycho’s notebooks. It was implied by them.
Dirac walked the same kind of room. In 1928 he wrote an equation for the electron that finally sat right with both quantum mechanics and special relativity — and that handed back the electron’s spin unasked, a property nobody had put in. It also came with solutions of negative energy nobody wanted. He spent three years explaining them away, for a while guessing they were protons in disguise, before proposing in 1931 a genuinely new particle: the electron’s mass, the opposite charge. Carl Anderson photographed such a track in 1932, and by his own account was not hunting Dirac’s particle: the theory “played no part whatsoever in the discovery of the positron.” Looking back in 1961 he put it harder — “the discovery of the positron was wholly accidental,” and a sagacious person in a well-equipped laboratory who “had taken the Dirac theory at face value ... could have discovered the positron in a single afternoon.” Nobody knew. The equation had the consequence in it, unread, and it took three years, one wrong answer and an accident to read it out.
AlphaGo’s move 37 is the same move in a smaller universe. Game two against Lee Sedol, Seoul, March 2016: a shoulder hit on the fifth line — a stone laid diagonally against one of Lee’s, a line higher up the board than the tradition puts it. Commentators first read it as a mistake. The team had also trained a separate model on human games, purely to predict what a person would play; by that model’s reckoning the move had roughly one chance in ten thousand of being chosen by a human.
Go permits something like a 2 followed by 170 zeros’ worth of legal positions on a 19×19 board — arrangements of stones the rules allow, not games, counted exactly for the first time in 2016. Every game anyone has written down has visited a corridor so thin the comparison stops meaning anything. Move 37 was legal all along, and the shape was not unknown to Go theory; what the corridor had almost never visited was that shape there, at that moment.
What models add is not the walk — people have always taken it — but its speed, and a learned sense of what tends to work that lets a system try unvisited regions without trying them at random, and with none of the taste a tradition hands down.
Two kinds of there-all-along
Laid out plainly, the taxonomy fits in a pocket:
- Present in the data, unmined — the archive nobody had read. The pattern was already there; what was missing was a reader wide enough to hold it all at once. Co-evolution fingerprints. A rule about which shapes kill bacteria, hiding in two thousand dish results.
- Implied by the rules, unvisited — the rooms the rules imply. No row of the corpus contains it; a rule captured well enough emits it anyway. A planet’s position in 1631. An anti-electron. A shoulder hit on the fifth line.
Neither is creation from nothing — one is an archive finally read, the other a possibility space finally walked. That is no demotion, since reading and walking are, we would argue, most of what discovery has ever consisted of. It just moves the mystery to where it lives.
The seam leaks, and it is worth knowing where. A system that plays itself a hundred million times has written a new archive and then mined it — engine one running on data engine two produced — and a machine that goes and makes new data, running experiments nobody had run, is doing neither. That third one is the check itself getting faster; it turns up in the next section.
One caution, since these five exemplars are most of the piece’s evidence: they are famous, and famous because the check came back yes. Wins tell you what the engines can do and nothing about how often either proposes something the world then refuses. We know of no published base rate, which is why the next section is the long one.
What would count against this piece: a discovery whose reconstruction requires information that was in nobody’s archive and no instrument’s reach at the time — genuine arrival from outside the library. We looked for one; the seam cases below are as close as we got, and the honest reading is that they are seams, not exceptions.
What the record shows
Both engines only ever produce a candidate. What turns a candidate into knowledge is external.
AlphaFold became knowledge one structure at a time, because the experimentalists kept agreeing — because structures determined in laboratories, by people bouncing X-rays off crystals and firing electrons through frozen samples, went on matching what the model had said for the proteins it was good at. Not every one: proteins caught mid-motion, in complex, or holding a drug came back wrong often enough that nobody sensible treats a prediction as a structure. The people who had spent half a century filling the Protein Data Bank were now marking what they had taught. Halicin became a finding because bacteria died. Both times the proposal was the cheap half.
Where the claim is about the world, the world referees — though never on its own initiative. Somebody has to grow the crystal, put the molecule in the dish, wait for the transit. The judge is the world plus the institutions that keep putting questions to it, which is why the disposing is the slow half — and why a claim checked only against a simulation, one model marking another model’s homework, is not yet checked against anything outside.
Take the check away and both stories become something else wearing the same confidence. A model that surfaces a pattern nobody can verify is, from outside, indistinguishable from the same model inventing a plausible citation — a hallucination, as people have taken to calling it. Both are credible texture across a region where the answer has not been established: an unchecked discovery is a thin archive coming back smooth, in a lab coat. That is no contradiction of the confidence score in the opening. That one was built as a prediction and benchmarked against experiments; fluency was never built to predict anything, and nobody checked it. Confidence was never the tell. When a model hands you something new, the question is never how sure it sounds. It is what would have happened if the check had come back no.
The counter-receipts are in the public record too, and better said plainly here than left to a critic. In one week of November 2023, two papers landed in the same issue of Nature. Merchant and colleagues’ GNoME search reported 2.2 million candidate crystal structures, about 380,000 of them predicted stable — predicted meaning a calculation puts the composition at or near the energy floor, not that anyone had made it. Szymanski and colleagues’ A-Lab, a bench that plans its own experiments and runs them with robot arms, reported — as its abstract then read — 41 new compounds from 58 targets in 17 days. Both drew published critiques within months: Leeman, Palgrave and co-authors re-read the lab’s diffraction data in PRX Energy and argued its products had been misidentified, while Cheetham and Seshadri sampled the predicted structures in Chemistry of Materials and reported scant evidence for compounds meeting novelty, credibility and utility at once. Only one team answered in print, and that answer is the next paragraph.
Then the slow half finished a piece of its work. In January 2026 Nature published an author correction to the A-Lab paper and the word novel came out of its own title: an autonomous laboratory for the accelerated synthesis of novel materials became inorganic materials, the authors clarifying they had meant new to the prediction platform rather than new to science. The count came down with it: a re-analysis of the diffraction patterns, peer-reviewed after publication, stood behind 36 of the 40 reported successes and left four inconclusive, with a further compound dropped for having been in the training data all along. The abstract now reads 36 compounds from 57 targets. The prediction paper stands uncorrected and still disputed. Read that as the loop working rather than as a scandal — then notice the exchange rate. The lab’s proposing took seventeen days. The disposing took two years, and two years of disposing is what a published record is for.
It is also the shape of the work on this shelf. A bench claim here ships with its counting rule beside the number, and refusals get published on the same page as passes — the seat trials is what that looks like when twenty-three models sat the same exam, five earned a chair, and the four newest sat it and did not. When our own panel had graded every judged number here, we handed the same sealed comparisons to judge models from four other labs and published what came back. This page was drafted with the machines it describes, then read, corrected and signed by the people in the byline — how that works is next door. A colleague is not the same kind of outside as a bacterium; that is accountability, not a world-check. The stronger half is publishing all of it in a form that lets somebody who is not us go and disagree.
Where that leaves us
The loop is old. Notice something, guess at the rule, go and check. What changed in the past decade is not the loop but the throughput of one half of it. The proposing got very fast and very strange — fast enough to read archives no career could finish, strange enough to walk into rooms no tradition had entered. The deciding stayed exactly where it has always been: outside, slow, unimpressed, and entirely uninterested in how confident the proposal sounded.
That is the good news, not the disappointment. A proposer that ratifies its own claims about the world has nothing left to be surprised by; it is checking the compression against the compression. Mathematics is the exception that shows the shape of the rule: there a machine really can check a machine, and a proof assistant will refuse a bad proof all day, because the claim is about the derivation and nothing outside it. Ask whether a molecule kills a bacterium, and no amount of checking inside the model will tell you.
The pattern was already there. So was the rule. Whether either one is true has never been the finder’s call.
What to take with you
Six things, the explanations in the piece’s own words:
- These discoveries were found, not conjured. Nearly everything AlphaFold needed already existed, free to download, and none of it existed by accident — the sequence databases accumulating since the 1980s, and the Protein Data Bank, half a century of structures solved one at a time in laboratories. What was missing was a reader wide enough to hold them all at once.
- Engine one: present in the data, unmined. The pattern is in the archive already. Halicin came out of a public repurposing library of 6,111 shelved compounds — worked up earlier as a diabetes candidate and dropped before it reached a clinic — and what the model added was a willingness to rank whatever it was handed, without a chemist’s prior about what an antibiotic should look like.
- Engine two: implied by the rules, unvisited. A compression that captures the generating rule emits entries no row of the data contains — a planet’s position in 1631, an anti-electron, a shoulder hit on the fifth line that the team’s own model of human play rated at roughly one chance in ten thousand of being chosen by a person.
- Where the archive is thin, the answer is thin. Proteins with few known relatives come back less reliable on average, and the program’s own confidence drops on them. Same principle as hallucination: thin archive, smooth invention; thick archive, the real detail. (It drops for a second reason too — some stretches have no single shape to predict.)
- Confidence is not the tell. Before the check, an unchecked discovery and an invented citation are indistinguishable from outside — credible texture across a region where the answer has not been established — and nothing in the model’s own confidence tells them apart. This is the practical one: when a model hands you something new, the question is never how sure it sounds. It is what would have happened if the check had come back no.
- The check is the knowledge. AlphaFold became knowledge because the experimentalists kept agreeing; halicin because bacteria died — and halicin is still not a drug you can be prescribed. Both engines only ever produce a candidate.
How to check our work — and see it live
Read the blind scoresheet yourself. The contest publishes every entrant’s per-target scores after each round, on the same 0-to-100 scale this page quotes: lay the prediction over the answer, score what share landed within each of four distance tolerances, and average them. The median we print is across every structure that entrant was scored on. Look up the shallow-alignment targets in the same table and the tell is right there — thin archive, thinner answer.
Read the counter-receipts without a machine. Everything in the record section is public and named on purpose. Merchant and colleagues’ GNoME paper and Szymanski and colleagues’ A-Lab paper landed in the same November 2023 issue of Nature; the two published critiques are Leeman, Palgrave and co-authors in PRX Energy and Cheetham and Seshadri in Chemistry of Materials; the author correction that took novel out of the A-Lab title is in Nature, January 2026. Read the correction beside the original title and the two-year exchange rate is on the page in front of you.
See a check that publishes its refusals. Put a real rules question to RuleSage — free, no account — and watch it either attach the rulebook page it read or say plainly that it could not. That is the same distinction this page is about, running in a product. This page ships no data kit, by design: it is a field guide reading the published record, not a bench with rows, and the exhibit data index says so beside its entry.
The rest of the seminar
This page assumes one sentence from its neighbour and spends the rest of its length on the consequence. It’s Not a Metaphor is that sentence — the weights are the compression — with the receipts under it. Three librarians and a careful reader is the check built into a product: how the right page gets found so an answer can arrive with the page attached. The seat trials is the check built into a hiring decision — how a model earns a job here, and what it looks like when one doesn’t. The outside judges is the check turned on us: our own panel graded every sealed round first, and then judge models from four other labs re-scored the same batches, because a proposer grading its own proposals is the failure mode this whole article is about. The whole shelf holds the benches behind the claims we publish, failures included. If there is a piece of the machinery you want opened next, say so — the suggestion box is read.
Licence: CC BY 4.0, the whole page — name the source and link to it.