The field guide — a series that opens up one piece of the machinery at a time

It’s Not a Metaphor

exhibit twenty The field guide
Published 2026-08-23 (UTC)
a small (human) team and a fleet of AI agents

A language model is a compression of its corpus. That isn’t a figure of speech — it’s the training objective. Hold that one fact and most of what people find mysterious about these systems goes mechanical.

Five words this page leans on. A corpus is the body of text a model was trained on. A token is the unit a model predicts next — roughly a word, or a piece of one. Compression here is the strict kind: a short description that regenerates a long thing. The codebook is whatever the reader must already hold to decode it — for a language model, the weights. Replay is asking the same question at the same settings and checking that the same bytes come back.

On 2026-08-21, one of our bench probes went looking for determinism and found something better.

The setup was simple: ask a locally-run model the same question at the same settings, sampling turned all the way down, and confirm the same bytes come back — a receipt every measurement we run quietly depends on. Short answers replayed perfectly, identical across all five runs. Long, deliberative answers didn’t replay at all: no repeat matched the first byte-for-byte, and our two model families diverged at characters 372 and 464 of roughly a thousand.

That looks like noise. It wasn’t. The probe pinned it: the output depended on what the machine had run immediately before. One prior workload, one stable answer; another, a different one — each reproducible on its own. Two patterns, not zero. What looked like chaos was structure with a hidden variable, and the variable was history.

Years before this company had a name, its founders talked late about that distinction. True randomness, the quantum kind, has — as far as physics can establish — no usable pattern underneath; some readings keep determinism deep down; none let you predict with it. Almost everything else that looks random wears chaos as a costume. The company that grew from that conversation is called strata → signal: dig through the layers, find the pattern pretending not to be there. We build local-first tools now — a rules assistant, a walking-tour maker, benches keeping them honest — and the founding idea turns out to explain the machines they run on.

Three people, one idea

In the early 1600s, Johannes Kepler came into possession of decades of Tycho Brahe’s planetary observations, the sharpest sky measurements yet made without a telescope. Kepler compressed them into three sentences plus a short column of numbers, a few orbital figures per planet — two laws in 1609, the third in 1619 — and when he ran the description back out into fresh tables in 1627, it held. Rule plus parameters, far shorter than the tables they regenerate: the shortness itself was the discovery. Kepler had demonstrated the thing this article is about: understanding is compression. A law is what you call a summary that stops leaking. And remember: the parameters counted; that comes back.

Some three centuries later, in 1948, Claude Shannon gave information a unit: the bit, one yes-or-no answer’s worth. Then, in the early 1960s, three people reached the same idea from three directions, independently. Ray Solomonoff came first, after a theory of induction, in a 1960 report almost nobody read. Andrei Kolmogorov — who in the 1930s had written the short rulebook that made probability real mathematics — arrived in 1965, defining information without presupposing probability. Gregory Chaitin, a teenager when he wrote his paper, published in 1966, chasing the randomness of finite strings. The idea now goes by Kolmogorov complexity — not quite fair, as Kolmogorov himself said, crediting Solomonoff’s priority in print.

What they found: the information in a thing is the length of its shortest description, and a thing is truly random exactly when no recipe for it is meaningfully shorter than itself. Pattern is compressibility. An idea so true it got discovered three times is an idea that was simply there, waiting.

Chaos, real chaos, is the stuff that won’t compress — and right here the mathematics turns on our founding motto. Count the descriptions: short recipes are far scarcer than strings to produce, so almost every possible string has none shorter than itself. Incompressibility isn’t rare; in the space of all that could be written, it’s nearly universal. What’s rare is meeting any of it. Orbits, grammars, tides, faces — what the world hands us keeps coming from the thin compressible sliver; why is its own deep question. “Almost nothing is true chaos” is false of string-space and true of nearly every dataset anyone has cared to collect, and that gap is why science works at all. We live in the exception.

The idea carried a sting the same people proved within a decade: the shortest description, the true optimum, is uncomputable. It exists, well-defined once you fix the describing machine (change machines, it shifts by a bounded constant), yet no algorithm computes it — and past a ceiling set by its own size, no formal system can even prove a description shortest. Every compressor ever built approximates a floor that mathematics put permanently out of reach and out of sight. Put that in your pocket; it matters in a moment.

From that root grew minimum description length — statistics’ rule that the best theory is the shortest one written out beside whatever it fails to explain — and eventually a blunt thesis decades ahead of fashion: compression and intelligence might be two names for one ability.

The punchline

Then someone industrialized it.

Training a large language model means minimizing how surprised it is by the next token of its training text — roughly a word, or a piece of one — across trillions of them for the largest publicly described models. Minimizing surprise and minimizing description length are formally the same job: you only write down what surprised you, so the better the prediction, the shorter the record. A model putting probability p on the next token buys a code about log₂(1/p) bits long — arithmetic coding, exactly the number the training loss reports.

One honest caveat, which returns in the receipts: the loss counts only the bits to describe the text given the model. The model’s own bits — the weights — sit on a separate ledger line, and nothing in the next-token training loss charges for them.

The weights are not like a compressed archive of the corpus; they’re the codebook for one — pair them with a coder and they’ll squeeze that corpus smaller than the general-purpose compressors anyone has measured them against. What they hold is lossy and interpolating: not the sentences, but the shape the sentences made, drawn from a curated slice of what a crawler could reach, at one moment in history. Now the pocketed sting comes due: the optimum is uncomputable, so even the best model anyone has built is a reach toward a limit nothing can certify contact with — always headroom, and nobody can say how much.

So: the weights are the compression. That is the sentence this whole piece exists to hand you, because once you hold it, the mysteries fall in a row.

Bias. A compressor carries the tilts of what it compressed, chosen and unchosen alike. Asking why a model has its corpus’s tilts is close to asking why a zip file holds what you zipped — close, not identical: a shipped model gets tuned and steered afterward, which pushes some tilts around but can’t restore what compression never kept. And mind the cure: inferring from context is thinking; a mind with no priors — no assumptions carried ahead of the evidence — can’t cross a street. What matters is whether the guess stays revisable. A person’s assumption dissolves on meeting the person; an archive’s meets no one unless the system around it insists.

Hallucination. Lossy compression discards detail; decompression papers over the gaps by interpolation. A model inventing a plausible citation isn’t lying — it’s rendering the texture of the region where the real detail used to be. The archive is thin there, and thin archives come back smooth. The rest is incentive: a confident sentence outscores an abstention at training and preference-tuning alike. Compression makes the hole; the scoring rule decides whether the model paints over it or points at it.

Scale. Trained the same way on the same material, bigger models compress with less loss, so the rare stuff — the long tail of facts, languages, styles — survives where a smaller model had to blur it. Measured, not argued — the scaling-law papers (Kaplan et al., 2020, arXiv:2001.08361) chart it: loss against size is one of the field’s best-charted curves, and a large part of why capability arrives with size.

The frozen moment. A compression is a snapshot. Today’s models are lossy cultural archives of roughly the internet-writing decades — dense in the middle, thin at the edges. Unlike a story passed hand to hand, the snapshot doesn’t drift in the retelling; it recirculates. Print froze texts too, but a book waits on its shelf; this copy talks back, in the voice of the moment it was taken. The risk isn’t censorship but fluency asymmetry: the snapshotted consensus stated effortlessly, at a volume no dissent can match, everything outside it arguing with an opponent of infinite rhetorical stamina. Dogma is what you call a prior that no longer has to win arguments to survive. What decades of that do to a culture’s priors is genuinely open, and pretending we know would be a claim past our receipts.

Style collapse. When many writers lean on the same archive, their sentences decompress from the same neighborhoods. The sameness people keep reporting in model-assisted writing — not something we’ve measured — isn’t conspiracy; it’s shared source material at work.

The antidotes: the page, and the byline. If the failure is interpolation over thin archives, the first fix is reintroducing the uncompressed original. Retrieval with citation helps exactly this way, putting the uncompressed page back within the reader’s reach. Our rules assistant is built to answer with the rulebook’s page number attached — and to say so plainly, in place of a citation, when it cannot point at the page. A design rule with failure modes, not a guarantee; but the rule is the point. “Show the page” isn’t a UX flourish; it’s decompression with the original stratum held open beside it. The second fix is older: authorship. Where a surface must carry a prior — a voice, a stance — write it yourself, sign it, let a person bless it before it ships. Every model is bias at scale; the honest alternative isn’t no bias, it’s bias with a byline. That includes this page — drafted with the machines it describes, then read, corrected, and signed by the people behind the byline (how that works): the piece’s own prescription, applied to itself.

What we can show you

Three receipts — two ours, run 2026-08-21 (UTC) on the same bench, a single workstation-class GPU on one runtime build (ollama 0.32.13); one from the published record. For these two receipts we report the measurement rather than the model names — the finding is about the mechanism, not the vendor: both candidates were open-weight models we run locally — one about 24 billion parameters, one about 12 billion — put to identical prompts at identical settings.

The replay probe, in numbers: five runs per condition, two model families, sampling off (“greedy”: take the likeliest token every time), seed pinned, outputs SHA-256-hashed against run one. Both families agreed: short completions (a bare answer, a few tokens) matched run one on four of four repeats; long deliberative ones (the same question reasoned out, roughly a thousand characters), zero of four — the divergence tracked to run order, the cache state whatever ran before left behind, each branch reproducible once controlled, first divergence logged by position. One dial proved dead: two very different seeds, byte-identical output. The label’s knobs are sampling and seed; the unlabeled third is what ran before. We’re not first; people who serve these models for a living know the effect. It lands differently on your own bench.

From the same day (certification window 07:52–08:41 UTC): the exam for a seat, our word for a named job with one named model in it, and a measured reason why. One candidate wrapped 387 of 387 answers in a code fence (the triple-backtick markers that tell a screen “this part is computer text”) that nothing had asked for; a sibling model on the same prompts fenced zero of 387. Different training, different compression, different reflex — we can’t see either model’s diet, only what it made a habit of. (Counting rule: fenced = the harness’s pre-registered one-fence strip changed the response’s bytes before it parsed as the requested object; 43 cases × 9 repeats = 387, every scored generation.) The job went to the fencer, counterintuitively: a habit showing up 387 times out of 387 strips out in one line of code; one that shows up sometimes ambushes you in production. The stripper is part of the published measurement, not a footnote.

The third receipt is external, and it cuts against the usual telling. A trained language model driving an arithmetic coder squeezes English text far past the general-purpose compressors: the 2023 paper “Language Modeling Is Compression” (Delétang et al., arXiv:2309.10668) reported a 70-billion-parameter model squeezing the standard 1 GB Wikipedia-text benchmark (enwik9) to under a tenth of its original size, against roughly a third for gzip. But that scoreboard never charges the model for its own weights — never amortizes the codebook into the score — and the weights outweigh the file by more than a hundred times: 70 billion parameters at 16 bits each is about 140 GB against a 1 GB file. A small archive mailed with a huge decoder is not a small archive. The Hutter Prize — its namesake is an author on that very paper — charges for everything (self-extracting entries, every decompressor byte counted, fixed time and memory, no outside data), so a frontier model can’t even enter — and the standing champions, as of 2026-08 (UTC), are context-mixing compressors whose small neural components train from scratch on the very file they compress. Kepler’s parameters counted; so do these. Which ledger you keep decides who wins. Best found; never certifiable as best possible.

When you look closer

Kolmogorov’s line splits what “random” smears together. True chaos — incompressible, patternless — may have just one clean citizen in physics: quantum measurement. Everything else is the second side: deterministic chaos, fully caused yet unpredictable, its three-line rule hiding all the surprise in an infinitely-detailed initial condition; pseudo-randomness, a seed and a function in a trench coat; noise with a hidden variable, structure whose key you haven’t found yet. Nearly everything we meet lives there — the machines this piece is about included. What looks like spontaneity, examined closely, keeps resolving into history, corpus, and seed.

If you keep one thing: when a model surprises you, ask what its archive looked like at that spot. Thin archive, smooth invention; thick archive, the real detail. That’s most of it. Whether the resolution flatters or unsettles depends on the day you ask us. But the late-night argument this company grew from — two founders on what “random” actually means — had it right about the only ground it claimed, the world as we find it rather than the space of every string that could be written:

True chaos has no pattern. Almost nothing is true chaos — when you look closer.

What to take with you

Seven things, the explanations in the piece’s own words:

  • The weights are the compression. Not like a compressed archive of the corpus — the codebook for one. What they hold is lossy and interpolating: not the sentences, but the shape the sentences made, drawn from a curated slice of what a crawler could reach, at one moment in history.
  • Bias starts as the tilt of what was compressed. A compressor carries the tilts of what it compressed, chosen and unchosen alike. Tuning after training pushes some tilts around — but it can’t restore what compression never kept. Inferring from context is thinking; a mind with no priors can’t cross a street. What matters is whether the guess stays revisable — a person’s assumption dissolves on meeting the person; an archive’s meets no one unless the system around it insists.
  • Hallucination is a thin archive coming back smooth. Lossy compression discards detail; decompression papers over the gaps by interpolation, so an invented citation is the texture of the region where the real detail used to be. Compression makes the hole; the scoring rule decides whether the model paints over it or points at it.
  • Scale buys the long tail. Trained the same way on the same material, bigger models compress with less loss, so the rare stuff — the long tail of facts, languages, styles — survives where a smaller model had to blur it. Measured across the field for years, not argued — this one is the published record’s, not our bench’s.
  • Which ledger you keep decides who wins. The loss counts only the bits to describe the text given the model; the weights sit on a separate line, and nothing in the next-token training loss charges for them. A small archive mailed with a huge decoder is not a small archive: on the scoreboard that charges for the decompressor, the frontier model cannot even enter, and small compressors trained on the file itself hold the record. And the true optimum is uncomputable, so even the best model anyone has built is a reach toward a limit nothing can certify contact with — always headroom, and nobody can say how much.
  • Your bench is less deterministic than the label says. Sampling off and the seed pinned, short answers replayed byte-for-byte, four of four; long deliberative ones, zero of four — the divergence tracked to run order, the cache state whatever ran before left behind. The label’s knobs are sampling and seed; the unlabeled third is what ran before. And from the seat exam: a habit showing up 387 times out of 387 strips out in one line of code; one that shows up sometimes ambushes you in production.
  • Ask what the archive looked like there. When a model surprises you, ask what its archive looked like at that spot. Thin archive, smooth invention; thick archive, the real detail. That’s most of it.

How to check our work — and see it live

Run the replay probe on your own bench. Any locally-run model, sampling off, seed pinned: ask for an answer long enough to need a paragraph of reasoning, then ask four more times and hash what comes back. With ollama (ours was 0.32.13):

for i in 1 2 3 4 5; do
  curl -s http://localhost:11434/api/generate -d '{
    "model": "<your-model>", "stream": false,
    "prompt": "<a question that needs a paragraph to answer>",
    "options": {"temperature": 0, "seed": 42, "num_predict": 400}}' \
  | python3 -c 'import json,sys,hashlib; print(hashlib.sha256(json.load(sys.stdin)["response"].encode()).hexdigest())'
done

Five identical hashes is the label keeping its promise. If the long answers disagree, run a different model in between and go again — that is the unlabeled third dial, the cache state whatever ran before left behind. Our figures above came from five runs per condition on two model families and one runtime build; yours will differ, which is the point of running it on yours.

Read the third receipt without a machine. “Language Modeling Is Compression” (Delétang et al., 2023, arXiv:2309.10668) is where the enwik9 rates come from, and the Hutter Prize publishes its entry rules in full — the conditions that keep a frontier model out of the contest. Read the two side by side and the whole argument about which ledger you keep is there on the page.

See it live. Put a real rules question to RuleSage — free, no account — and watch it either attach the rulebook page it read or say plainly that it could not. This page ships no data kit, by design: it is a field guide, not a bench with rows, and the exhibit data index says so beside its entry; the bench rows behind the two house receipts are available on request at hello@strata2signal.com.

The rest of the seminar

This page is the field guide’s argument in one sentence — the weights are the compression — and four of its neighbors carry the parts it only names. The compressed photograph is that sentence with a ruler on it: what a model’s own bits actually weigh once you open the file, and what the compression cost when one question was put to four builds of the same model — less, it turned out, than changing how the question was phrased. Three librarians and a careful reader is the first antidote built: how the right page gets found so an answer can arrive with the page attached, which is this page’s “show the page” made mechanical. How we work is the second antidote applied to us — who writes these pages, and how a human blesses one before it ships, which is the byline this page argues for. Reading is fast, writing is slow is the physics of running one — why these machines read in bulk and answer a word at a time. The whole shelf holds the benches behind the claims we publish, failures included. If there is a piece of the machinery you want opened next, say so — the suggestion box is read.

Licence: CC BY 4.0, the whole page — name the source and link to it.

elsewhere in the workshop

a strata→signal property · hello@strata2signal.com · say hello