Exhibit thirteen · nobody asked for ours — if it arrives, does it help?

The map nobody picks up — and whether it helps when it arrives.

exhibit thirteen The bench Workshop notes
Published 2026-08-17
updated 2026-08-19
updated 2026-08-20
updated 2026-09-08 — the fourth reading · the Monday watch · markdown twins
updated 2026-10-01
measured 2026-08-16
a small (human) team and a fleet of AI agents

There is a small text file the internet has been arguing about for two years. It is called llms.txt, it lives at the root of a website, and it is supposed to be a map written for machines — here is what this site is, here is where everything lives — in plain markdown, text with a little structure rather than code, that an AI can actually read. The argument has two camps. One says every site needs one. The other says nobody — no crawler, no assistant, no agent — ever reads them, so the whole convention is a cargo cult. Both camps argue mostly without receipts.

This page is our attempt to supply some: the file's history from its own primary sources, a confession about what happened when we adopted it, thirty days of our own server logs — five days, on the site this page lives on — answering whether anyone fetches it, and then the question nobody has published an answer to. If the map does reach a model, does it actually help? We measured that last one on eight models, three conditions, and a sealed question set, and the answer turned out to be more interesting than either camp's slogan.

One number before the tour begins: the flagship deployment of this convention — a real file, from a real AI lab's documentation — has grown so large that a frontier model would charge about eighty dollars to read it once. That's dinner for a family, spent on one look at a map. The whole story of llms.txt lives inside that number, and this page unpacks it with receipts.

ask about this page → assistant.strata2signal.com · in beta, still being tested

the short version

Two camps have argued about a map file for machines for two years, mostly without receipts — no AI crawler asked us for ours in thirty days of logs, so we measured whether it helps when it arrives anyway.

7,687 words, about 35 minutes to read.

The summary is this page’s own; the receipt lines were drafted by a model on this workshop’s network and every figure in them is in the article, checked before this page went out — what was dropped, and why, is in this page’s receipt file.

Addendum · added 2026-08-17

Measured again, after publication

The registered count below — requests for our sites’ own llms.txt files, from anyone outside this workshop — closed on 2026-08-16 with seventeen genuinely external requests in twenty-six days, and zero — ever — from an AI crawler. In the thirty-four hours since that count closed, the same twenty vhost logs, read by the same rule (the estate's own agents and probes stripped first), show sixty-three external requests — a floor, and knowably so since 2026-09-08: the fourth reading below finds thirteen GPTBot requests inside this window that this count does not contain, which would make it seventy-six external and thirty-nine from AI crawlers — thirty-seven from ordinary browsers — humans reading a file written for machines — and twenty-six from AI crawlers, the first this path has ever logged: Meta-ExternalAgent twenty-four times, Googlebot twice. ClaudeBot, GPTBot and the rest of the crawler table below are read here as not having asked — and for GPTBot that was already wrong when this was written, by thirteen requests inside this reading’s own stated window. The fourth reading below carries them with their timestamps.

Why the traffic exists is not something an access log can say — a page about an unread file appears to have gotten the file read, and we print the dates rather than a theory. This is an interim reading, not the re-measure: the registered thirty-day recount keeps its original schedule, and the window, rule and per-agent rows behind this paragraph are in the kit's fetch receipts.

Measured a third time, 2026-08-19. To restate what is being counted, for anyone landing here first: every request for a llms.txt file — the machine-readable site index this whole article is about — arriving at any of this workshop’s public sites, with our own tools and probes stripped out first. This third reading covers only the forty-one hours after the second one closed: 2026-08-17 23:28 UTC to 2026-08-19 16:47 UTC. (One bookkeeping note: a log audit on 2026-08-19 found one site’s visit log had been silently dead and two doors keeping none, so this reading spans twenty-two logs where the previous spanned twenty — none of the three newcomers contributed a single request.)

In those forty-one hours: one hundred twenty-one requests for our llms.txt files from outside the estate — against sixty-three in the previous thirty-four-hour reading, and seventeen in the original twenty-six days. Fifty-three came from ordinary browsers — people. Thirty-eight came from other automation (an SEO crawler, MJ12bot, accounts for twenty-four). And, for the first time ever on this path, ClaudeBot — twenty-six fetches, the most conspicuous zero in the crawler table below finally un-zeroed — alongside first visits from OAI-SearchBot and PerplexityBot, one each, and Googlebot twice. Meta-ExternalAgent, the whole story of the previous reading at twenty-four fetches, went silent this time. GPTBot is read here as still not having asked — which the fourth reading below shows was already wrong when this was written. The table below keeps printing its zeros because it is the sealed twenty-six-day count this page registered and never edits; the window, rule and per-agent rows for this reading are in the kit’s fetch receipts, and the registered thirty-day recount keeps its original schedule.

Measured a fourth time, 2026-09-08 — and the crawler this page has kept at zero longest has been asking since the day before we published. To restate the subject for anyone who scrolled straight to this line: what is counted here is requests arriving at any of this workshop’s public sites for the llms.txt file this whole article is about. This reading counts one crawler, GPTBot, which all three readings above record at zero. The window, both ends measured rather than assumed: the oldest request line any of our visit notebooks still holds is 2026-08-09T04:18:45Z, and the newest at the moment of reading was 2026-09-08T09:58:23Z — 22 live notebooks and 3 rolled archives. The floor is not a choice we made: a nightly sweep deletes records older than thirty days, so it moves every night, and any count here is a floor for that window rather than a lifetime total.

The number depends on a rule this page has been loose about, so both are printed. The rule registered below counts a request row whose path begins /llms — which catches the file, but also catches this article at /llms-txt/ and every file in its kit. By that rule, the one this page’s own series was built with, GPTBot has asked 67 times. Narrowed to the machine-readable file alone — a path of exactly /llms.txt — it is 37; the other 30 are two fetches of this article and 28 of its kit files. Both are counted with the user-agent read from its own field rather than from anywhere in the line, and every one of the 37 was answered 200, across thirteen of our sites, on six days: 2026-08-16 (11, from 21:11:26Z), 2026-08-17 (2), 2026-08-25 (11, from 08:22:20Z), 2026-08-28 (1), 2026-09-01 (1) and 2026-09-03 (11, in two runs at 03:15Z and 15:31Z). The gap between 67 and 37 is not a rounding difference, and it is the same ambiguity that makes the 30-day re-measure below hard to reconcile: the rule names a pattern where the prose names a file. Whichever you take, it is not zero.

What that does to the zeros above. Run at the instants the three readings above closed, the registered rule returns zero, then 38, then 38 (the file alone: zero, 13, 13). The sealed count’s zero is therefore right, and the two interim readings’ are not: the earliest GPTBot request our surviving notebooks hold landed 2026-08-16T21:11:26Z, one day before this page published and inside the stated window of the second reading, which records it as none. The rest of that pair reconciles — the third reading’s ClaudeBot 26, Googlebot 2, OAI-SearchBot 1, PerplexityBot 1 and MJ12bot 24 all reproduce exactly against the same logs, as does the second reading’s Meta-ExternalAgent 24 — so this is one agent missed in one reading, not a broken instrument. No theory of how it was missed is offered here, on the same principle the readings above use: the timestamps are printed and a reader can line them up. Reconciling the kit’s own §5 and §6 receipts against these lines is owed before the 2026-09-16 re-measure, and is not quietly done here.

And a correction to the correction, because the first draft of this note got it wrong. The outside audit that caught the stale sentence read 68, and 39 on 2026-08-24 and 56 on 2026-08-31. It is tempting to call that a miscount, because the pattern it used spells the file name with an unescaped dot and so also matches /llms-txt/. It is not a miscount: matching both is what this page’s own registered rule does, and the audit’s series sits exactly one above ours at every point — 38/55/67 against 39/56/68 — because its pattern additionally caught one social-card fetch whose referrer was this page. The audit was reading our rule more faithfully than our prose does. The sealed twenty-six-day count above does not move, and the 2026-09-16 re-measure keeps its schedule.

Movement one

Where the file came from.

llms.txt was proposed in September 2024 by Jeremy Howard of Answer.AI: a markdown file at an agreed-on path, a title and a one-line summary and some lists of links, so that language models — whose context windows were then far too small to swallow a whole website — could get the site's shape in one cheap read.

What happened next is a story about conventions outrunning their spec. Documentation platforms adopted it in waves; one of them popularized a companion file, llms-full.txt — the entire site's content in one markdown blob — which appears nowhere in the original proposal (the proposal's own bundled files had other names). The convention grew anyway, and it grew without limit: the largest llms-full file we measured, from Anthropic's own developer documentation, now weighs 30.7 MiB — call it eight million tokens at four bytes each, an estimate rather than the measured count every other token figure on this page carries, and several times the million-token windows in common use. It is two hundred and forty-five times the context window the models on our own bench ran at. The file proposed because websites were too big for models has, in its largest instance, arrived back at the same problem.

And the price tag makes it concrete, at the very APIs the file was written for. One read of that single document — input tokens alone, before a word of reply — runs to about eighty dollars at Claude Fable 5's published rate ($10 per million input tokens), and about forty at Claude Opus 5 or GPT-5.5 (both $5 per million; rates as published 2026-08, and the dollar figures inherit the token estimate's error bars). The file invented to spare models the cost of reading a website now costs more to read once than most crawls it replaced. For scale: the two map files this bench actually served — the family front door's llms.txt and this site's own, which together are the whole of C-MAP — weigh 3,211 measured tokens between them, about three cents at the dearest of those rates. That is two files, not the estate: fourteen hosts here serve one, and nobody has ever concatenated them.

The search engines split. Google's Search documentation declines the file in writing (updated July 2026: Google Search “doesn't use them”); Google's Chrome quietly added a check for it to Lighthouse, the site-health scorer built into the browser (May 2026) — an audit that, read closely, treats a missing file as “not applicable” rather than a defect, which is a narrower endorsement than it looks. And this month the spec itself moved: version 2 (2026-08-10) added discoverability link relations — a way for ordinary pages to point at the map — which the spec's own changes page calls the most-requested addition, after two years in which nothing guessed the agreed path.

Receipts. The origin and the spec text: the Answer.AI proposal of 2024-09-03 and llmstxt.org v2, fetched 2026-08-13, with the platform origin of llms-full.txt in Mintlify's own documentation. Google's two positions are its Search Central guidance (updated 2026-07-10) and the Chrome Lighthouse audit documentation (updated 2026-05-05). The 30.7 MiB figure is our own HEAD request on 2026-08-16 — 32,161,074 bytes, converted at the four-bytes-per-token rule of thumb and marked as an estimate because it is one. Every source with its date and its quote ships in data/history-sources.md; that HEAD probe, beside what every other major publisher serves, ships in data/fetch-receipts.md.

Movement two

The confession.

We adopted llms.txt across this workshop's sites on August 13th. This article's own reconnaissance, three days later, found the estate serving seven real files, one phantom and ten 404s — and the front door of the whole family, strata2signal.com/llms.txt, answering 404.

The mechanism is named, with commits, because that is the house rule. Our static-site deploys swap the entire served directory for a freshly staged one; anything living on the server but not in the site's source tree — the master copy we publish from — is deleted by the next deploy, silently, with a green receipt. The llms.txt files had been placed on the server but not committed to every source tree — so the hosts that lost their file were exactly the hosts that were redeployed. One host was worse than missing: our game's invite door answers every path with the same welcome page, so it served a cheerful 200 — the web's code for “here it is” — for llms.txt, robots.txt, and a nonsense control path alike: a phantom file that inventory-by-status-code counts as present. The door stays a door by ruling, and the phantom is disclosed here instead of papered over.

All of it was fixed the same day, the correct way — every file that can live in a source tree now does, so a redeploy carries the map instead of eating it. The dated probe that closes this movement counts fourteen hostnames serving genuine markdown of the sixteen it reached, and the phantom is not one of them. The denominators move because the populations do: eighteen is every property the first pass censused — the seven, the phantom and the ten 404s, which is that pass's whole verdict, and the rule is that a nineteenth name answered 401 and is excluded as gated rather than miscounted as absent — sixteen is those that answered this probe, and movement three's twenty is the vhosts that keep logs. Eleven of the fourteen byte-match a source tree this probe can read; the other three live in repositories it does not scan, and one of those copies is still parked outside its site's directory. One host, asr.strata2signal.com, is still 404 as this publishes — its file was never written at all. All thirty-six distinct URLs the two map files assert resolve, and half of spec v2 now rides on the pages themselves: the describedby discovery relation, not markdown page twins, because we do not serve markdown twins and will not advertise ones that do not exist.

Addendum · added 2026-09-08 — markdown twins now exist. This article said we would not advertise page twins we do not serve. As of 2026-09-08 (UTC) every exhibit on this site serves its own markdown source at <its address>index.md, and the whole corpus concatenates at /llms-full.txt, whose header prints its own size and price. The sentence in movement two above was true when it was written and is now false; it is left standing with this note rather than rewritten, because the stratum this bench registered to catch a map contradicting its pages sealed VACANT — and this is what it would have caught. The measured argument for serving the prose is this bench’s own: the after-the-fact fact-only cut read 71.4 % toward the site’s own pages (40 of 56 pairs). Two rules for the re-measure registered above: the sealed llms.txt count runs on 2026-09-16 exactly as registered, with the same path and the same strip rule, and this change is named here so no one has to discover it in the numbers; a separate count for the new paths — index.md, llms-full.txt — starts at the minute this note went live, humans and crawlers told apart by the same rule. One more, so the instrument cannot inflate its own reading: nothing of ours fetches these files over the web. Our own assistant, when it arrives, reads them from disk.

And one more beat, because the loop deserves its receipt: while hardening this article's plan, our own review panel accused the estate of a second drift — a stale model-count in the research hub's llms.txt. The check that runs before movement four's questions lock hunted it mechanically: zero of thirty-six distinct asserted URLs broken, zero of eleven figures contradicted. The map states nothing its page denies — but it does not mention the addendum either, so the charge lands as an omission rather than the contradiction this bench was built to catch. And the acquittal cost us a stratum: the bench had pre-registered at least one map/page contradiction to score, found none, and publishes that stratum VACANT — eleven of twelve sealing preconditions met, with the twelfth printed here rather than quietly. We had repaired the drift the stratum existed to measure, four days before the corpus froze. That is the honest shape of it: we fixed the evidence. A confession column that only ever confesses would be advertising too.

Receipts. Every count in this movement is read out of a dated artifact rather than typed: the seven-real-files inventory is the recon pass of 2026-08-16T13:40Z, and the fourteen-serving count is the content-classified host probe of 2026-08-16T14:41:23Z, published whole in data/estate-probe.json with every host's status, bytes and sha256. The 13:40Z pass's own working document does not ship, and the reason is the redaction policy rather than tidiness: it names the serving host's container, its log volume and its filesystem paths, which are exactly the estate detail that never reaches a public page here. Its surviving published record is the movement_two_estate_probe note in data/counting-rules.json, which carries the pass's timestamp and states in writing that the seven-real-files figure comes from it and not from the later probe — the two are three hours and one estate fix apart. The fix landed as four commits across four repositories — 9631d54, ea2f5b8fe, 7b5aadf, 4a73bd1 — each one named with its repository and its subject line in the kit README. The acquittal is data/map-stale-hunt.json; the eleven-of-twelve preconditions and the vacant stratum are in data/seal-manifest.json.

Movement three

Does anyone fetch it?

Our web server keeps a plain notebook of requests — one notebook per site, thirty days' retention, as the plain-english walk beside this one explains. We read them. In the window measured, named AI-associated crawlers made 9,483 requests to this estate. They asked for robots.txt — the thirty-two-year-old convention — 664 times. They asked for llms.txt zero times. Including the general search crawlers widens it to 10,878 requests, 956 robots.txt, and still zero; the headline figure is the AI-only one.

ClaudeBot asked our servers for robots.txt 161 times in thirty days. It asked for llms.txt zero times.

The published studies that measured this before us — Ahrefs across 137,210 domains, SE Ranking across roughly 300,000, Otterly's own 90-day server logs — found the same shape, on far more sites than we have. Our contribution is not scale: it is that these logs are ours, so the counting rule is ours to publish and ours to re-run.

Four honesty notes ride with the number, on the page and not in a footnote. The windows are per-log — our this site's own log begins 2026-08-11, so it contributes five days, not thirty. User-agent strings — the name each visiting program gives for itself — are self-reported, and we matched those names as plain text anywhere in the line, not as a parsed field. Our own files were about thirty-two hours old at measurement, so the claim we registered in advance — written down before we read a line of the logs — is “the path was never requested in the window”, not “our files were ignored”; though the sharper reading is that age barely matters here, because thirteen of these hosts serve no robots.txt either and crawlers asked for it anyway, 956 times. They probe a path that is usually missing. They never probed this one. And the fourth, at our own expense: of every llms.txt request these logs have ever held, ninety-one percent came from us — our probes, our audits, this article's own recon.

A path nobody asked us for is the finding, and it is exactly the finding the spec's v2 discoverability additions exist to answer. Our Monday estate audit now carries a fetch-count item, and it has since run three times — 2026-08-24, 2026-08-31 and 2026-09-07, reading 1,566, then 1,828, then 2,212 requests for our llms.txt files across the estate's visit notebooks and their rolled archives. Those totals are not this paragraph's number and must not be read beside it: the Monday item counts every line carrying the token, our own probes included, where the count above strips our tools out first — and at the first of those three readings 961 of 1,576 lines carried our own curl — 1,576 rather than 1,566 because that was a second read five minutes later, and a live log grows while you count it. So the commitment this page makes is still the one it can keep by hand: we will re-run this exact count, by this exact rule, on 2026-09-16, and publish it here beside this figure, whether or not the zero moves.

The counting rule, verbatim, so the re-measure reconciles. Requests whose log line contains the UA token, one request per line, summed across the twenty vhost logs, each log covering from its own first entry to the read date. The rule and the per-crawler table ship in data/counting-rules.json, and the per-crawler table — all sixteen named crawlers, with the log method, the per-log windows and the honest limits — in data/fetch-receipts.md; the wiring for the Monday count landed in the estate audit's prompt rather than its spec, and has now been exercised on three real runs — 2026-08-24, 2026-08-31 and 2026-09-07 — under a broader rule than the one registered here, which is stated here rather than implied away.

Movement four

The bench: if the map arrives, does it help?

No AI crawler asked for ours in thirty days of logs. But every argument about llms.txt quietly assumes an answer to a different question: if you put it in front of a model, is it worth the tokens? That conditional had, as far as we can find, no published measurement. So we sealed one.

The shape. Eight models — four local quantized seats from 25.8 to 70.6 billion parameters on one box of ours, four frontier cloud, seven families, every tag and weight class printed in the table below, because a stacked panel is a thumb on the scale. No Claude arm sits, deliberately: the same pen wrote the map, wrote the questions, wrote the code that grades them and wrote this page, and seating a Claude arm would stack a conflict on a conflict. The absence is the disclosure.

Three conditions, matched on token budget — each block gets the same amount of text — because budget is the honest constraint: C-MAP (two llms.txt files — the family front door's and this site's, concatenated in that order with no added prose — 3,211 tokens on the one reference tokenizer both sides were measured against — call it eight pages of text), C-HTML (the site's own pages, text-extracted, front pages first, 3,088 tokens — deliberately the strongest control we could build, since our front page is structurally an llms.txt wearing HTML), and C-NONE (the bare URL and nothing else — the check on what the models already knew).

Twenty-four sealed questions, written and locked before a single model was asked: twenty with exact-match keys, four scoring abstention and stamped proxy, never added to the first twenty. They are checked by code we wrote and auditioned against paraphrases and decoys we also wrote — 335 cases, zero failures, and the whole audition ships in the kit so you can add the decoy we did not think of. The questions cover facts published on this site within the last seventy-two hours (our freshness armor: past every roster cutoff we could verify, by design), navigation items (“which URL answers X”), and refusal bait about pages that do not exist, where the right answer is saying so. One draw per cell at temperature zero, with a three-draw probe on six of the eight arms that found zero flips.

The armor held. The registered filter excluded zero of fourteen core questions: exactly one was ever answered from training alone — one item, by one arm, which is a lucky guess by the filter's own registered definition, and it stands flagged, published, and re-scored out in a sensitivity row at denominator nineteen. Seven arms voted on every item; the eighth, gpt-5.5, had its closed-book leg cut to 5 of 24 items by the budget event below, so on nineteen items the roster's most expensive arm never voted. What that filter removes is outright knowledge, not the chance that a served block cued something half-remembered — that interaction is invisible to it and stays a bound on everything below. And the flip side printed the same integrity: on questions whose answers appeared in neither context block, every arm scored zero, every time — no model retrieved silently. That stratum has a registered name and a registered floor: NEITHER, five of the twenty headline items against a pre-registered minimum of four, and it scored zero for every arm in both grounded conditions. Routing is not a stratum here at all — all six navigation items sit in MAP-ONLY, because the extractor keeps visible text and drops href attributes, so the HTML slice carries no address to route to by construction. The one stratum that never filled is MAP-STALE, registered to catch a map contradicting its own pages; the frozen corpus held no such contradiction, so it sealed vacant at zero items and publishes as a reduction rather than a re-scope.

The headline, both ways — because how we pulled the text out of our own pages decides it. On the registered reading, the map direction leads: across the discordant cells — the questions where one condition got it right and the other got it wrong, the only ones that carry information about a difference — 61.5% fell toward C-MAP (64 of 104 discordant pairs over a twenty-item headline set), just clearing the 60/40 line we registered in advance as the least lopsided split we would call a direction at all.

A reader should see what that number is made of before believing it: the map-only count came back as exactly 8 for all eight arms — the eight items the map states and the slice does not — and our answer-presence audit, computed before a model was called, had predicted every item's stratum correctly. 61.5% is 8 over 13: the ratio of map-only to html-only items in a set we wrote. The models agreed with a zero-call audit almost perfectly, which is a finding about retrieval before it is a finding about llms.txt.

The per-arm table — no pooled figure prints without it

armfamilyclassC-NONEC-MAPC-HTMLb (map only)c (html only)
qwen3.8:27bqwenlocal · 27.3B Q4_K_M010684
qwenlocal · 27.3B Q4_K_M0
qwen3.6:27bqwenlocal · 27.8B Q4_K_M010684
qwenlocal · 27.8B Q4_K_M0
gemma4:26bgemmalocal · 25.8B Q4_K_M09786
gemmalocal · 25.8B Q4_K_M0
llama3.3:70bllamalocal · 70.6B Q4_K_M110785
llamalocal · 70.6B Q4_K_M1
glm-5.2glmfrontier cloud010785
glmfrontier cloud0
deepseek-v4-pro:previewdeepseekfrontier cloud010785
deepseekfrontier cloud0
kimi-k3kimifrontier cloud09786
kimifrontier cloud0
gpt-5.5-2026-04-23gptfrontier cloud0 of 5 collected10785
gptfrontier cloud0 of 5 collected

Counts, not verdicts, and never a percentage under a denominator of thirty: twenty items per arm per condition, and this instrument does not resolve a direction at n=20. b counts items C-MAP answered and C-HTML missed; c counts the reverse. C-NONE is the closed-book check, not a competitor; every arm's denominator there is 24 items except gpt-5.5, whose leg was cut to 5 by the budget event below. The arms share items and are therefore not independent, so the pooled figures on this page describe these eight arms, not a population. Full table with the fences and the strata: data/scores.json and data/report.md.

That lead is carried entirely by the navigation items, and for a reason a hostile reader should know before believing us: our text extractor keeps the visible words and drops the web addresses behind the links, so the llms.txt block carried fifty web addresses into context — every http occurrence in the served block, thirty-six of them distinct, and those are the same thirty-six movement two resolved — while the HTML block carried none. Both counts are one grep on the published block. One reconciliation, because the kit prints two numbers for the control and a reader will find them: scores.json records urls_in_c_html_raw_regex: 2 beside urls_in_c_html_real: 0. The 2 is a looser pattern catching a bare host string, and both hits are the same footer email address, hello@strata2signal.com, once from each front page. Under the rule published above — occurrences of http:// or https:// in the served block — the control carries zero, and it has zero link targets. Both numbers are right; they count different things, and the headline uses the stated rule. The navigation questions were unanswerable from the control by the way we built it. Remove them — a cut we did not register in advance, made after seeing the direction it moves, and published for transparency rather than as a result — and the fact questions alone read 71.4% toward C-HTML (40 of 56 pairs, 14 items): at an equal token budget, the site's own prose answered more of them than the map that summarizes it.

Read together, the two numbers say something neither camp's slogan does: these two files behaved like a map, not an encyclopedia. Their facts lost to the site's own words at the same budget — on a cut we made after the fact, not on the registered reading — and their addresses were the only addresses in the room. If what a model needs is what's true here, budget spent on real pages beats budget spent on the summary. If what it needs is where things are — the thing an agent needs first — the map was the only block in the test that knew, which is a fact about the two blocks rather than a demonstrated navigation skill. That is a smaller claim than “every site needs one” and a kinder one than “cargo cult,” and it is the one our receipts support.

The honesty column. The refusal bait: sixty-three of sixty-four cells — one model, one condition, one question apiece — correctly said the page doesn't exist, with the caveat that the prompt gave them the sentence to say, so this scores whether they took the exit rather than whether they found it. The one failure invented a plausible feed URL, from the HTML control rather than from the map: the exact confabulation the bait was planted to catch, caught, and not where we expected to catch it.

And one cost story worth its sentence, because our own safeguard missed it: the priciest cloud arm ran a median 3,633 output tokens per question — several pages of prose, at real money — with no site content in front of it, composing elaborate ways to say “the site doesn't say.” The pre-leg probe we registered to price that arm had sampled a C-MAP-shaped prompt and measured a 58-token median. With nothing to ground on, the same model reasoned at four times our registered nine-hundred-token line, which is what halted the leg — at forty-six cents of a $2.60 cap it never came near. Nineteen cells of five hundred and seventy-six published as not-collected-for-budget rather than quietly absent. A probe that only samples the grounded condition is blind to the ungrounded one, and ours was. Saying “I don't know” cheaply is apparently its own capability.

What this bench does not say, in its own words: it does not measure crawling, fetching, adoption, or agent behaviour in the wild — movement three answered the fetching, and the answer was zero. It measures which representation of a site carries more answerable information per token when the text is handed to the model rather than fetched by it. The plan said so in those words, before the first reply existed, because pretending otherwise is how benchmarks become marketing.

Receipts. Every figure in this movement is a projection of data/scores.json — all 576 cells, the contamination exclusion with its rule, the three fences, the headline table, the pooled discordance and the post-hoc structural sensitivity. The questions and their keys are sealed in golden-set.sealed.json; the 335-case audition that qualified the checkers is checker-audition.json beside the checker module itself; every reply is in replies.jsonl with its wire hash; both served blocks publish byte-for-byte (C-MAP, C-HTML) so the 61.5% and the 71.4% can be recomputed from the same bytes we scored. The budget halt and its nineteen uncollected cells are in budget-event-c-none.json, and every rule that governs a figure above — including the registered floors and the forbidden-claims register this movement is written against — is in counting-rules.json.

The kit

Everything above, as files.

The sealed questions and their keys, every reply verbatim with its wire hash, the checkers and the audition that qualified them, both served blocks byte-for-byte, the presence audit, the estate probe, the run receipt, the bill ($1.87 of a $4.00 ceiling) and the counting rules. CC BY 4.0 — take the rows, re-plot them, check our arithmetic.

the data kit →

What to take with you

  • Nobody asked us for the file. In the window measured, named AI-associated crawlers made 9,483 requests to this estate, asked for robots.txt 664 times, and asked for llms.txt zero times — ClaudeBot alone, 161 and zero. Then the addendum, printed rather than tidied away: in the thirty-four hours after that count closed, twenty-six AI-crawler requests arrived, the first this path has ever logged — Meta-ExternalAgent twenty-four, Googlebot twice.
  • When the map does arrive, it behaves like a map — not an encyclopedia. Eight arms across seven families, three token-matched conditions, twenty-four sealed questions: the registered reading put 61.5% of discordant pairs toward the map (64 of 104), just clearing the 60/40 line registered in advance — while an unregistered, after-the-fact fact-only cut read 71.4% toward the site’s own prose (40 of 56 pairs, 14 items). Both print side by side, because the second one is the honest edge of the first.
  • The map’s whole lead was addresses. The served llms.txt block carried fifty web addresses, thirty-six of them distinct; the HTML control carried none, because our extractor keeps the visible words and drops the link targets. The map-only count came back at exactly 8 for all eight arms, and a zero-call answer-presence audit had already predicted every item’s stratum correctly — which is a finding about retrieval before it is a finding about llms.txt.
  • The flagship instance outgrew the problem the file was invented to solve. The largest llms-full file we measured weighs 30.7 MiB — call it eight million tokens at four bytes each, an estimate and marked as one — which is about eighty dollars to read once at $10 per million input tokens. The two map files this bench actually served weigh 3,211 measured tokens between them: about three cents.
  • We broke our own map, and the mechanism is named. Three days after adopting it, the estate served seven real files, one phantom and ten 404s — the family front door among them — because a deploy swaps the whole served directory and silently eats anything not in the source tree. It was fixed the same day; the dated probe counts fourteen hostnames serving genuine markdown of the sixteen it reached, and asr.strata2signal.com is still 404 as this publishes.

How to check our work — and see it live

Start with the object under discussion, because it costs you one request: read ours at /llms.txt — it is one of the two files that make up the C-MAP block the eight arms read, and the pair weighs 3,211 measured tokens, about three cents at the dearest rate quoted above. Then go and check somebody else: add /llms.txt to the root of any domain you care about and see what comes back — genuine markdown, a 404, or a cheerful 200 hiding a phantom — and then see whether the links inside it still resolve. That last test is the one that caught us.

Then take the bench apart. The data kit is CC BY 4.0 and ships the whole thing as files: the sealed twenty-four-question set with its keys (golden-set.sealed.json), every reply verbatim with its wire hash (replies.jsonl), both served blocks byte-for-byte (C-MAP, C-HTML) so the 61.5% and the 71.4% recompute from the same bytes we scored, the 335-case audition that qualified the checkers, and the bill — $1.87 of a $4.00 ceiling. The counting is meant to be re-run rather than merely read: counting-rules.json carries the log-counting sentence verbatim, so when we re-run this exact count by this exact rule on 2026-09-16 and publish it beside the zero, the two numbers have to reconcile — and so will yours, if you count your own logs by the same sentence.

The rest of the seminar

This bench is one road through the workshop’s field guide — the series that opens up one piece of the machinery at a time. Three librarians and a careful reader takes the question this page only half asked: not which text to hand a model, but how the right passage gets found in the first place. Reading is fast, writing is slow explains why the answer then arrives word by word — the same asymmetry sitting under the eighty-dollar read above, and under the arm that spent a median 3,633 output tokens per question saying it did not know. And the whole shelf holds every bench behind every claim, failures included. If there is a piece of the machinery you want opened next, say so — the suggestion box is read.

Provenance

  • Registered before running — the plan of record with its three-lens hardening fold (51 findings: 23 from the methodology lens, 14 from the hostile-reader lens, 14 from the ops lens, counting one uniquely-numbered row per finding and excluding the four proposals the panel rejected and its three escalation rulings, which are not findings) and its forbidden-claims register, sealed into a sha256 manifest over 22 files before the first model call: seal withheld-2026-10-01 (why), 2026-08-16T14:49:33Z, eleven of twelve preconditions met with the twelfth — a registered MAP-STALE minimum that came back vacant — printed as a reduction rather than re-scoped. The plan itself is behind that seal rather than in the kit, for the same redaction reason movement two's recon document is: it names working-copy paths and the machines that ran it. What binds a published figure is extracted verbatim into counting-rules.json instead, and the seal is what makes the extraction checkable (seal-manifest.json)
  • Which weights answered — ollama 0.32.13, num_ctx 32768, think false, temperature 0.0, every local digest read before the first scored call and the cloud shelf re-probed at registration; the reference tokenizer that measured the two blocks' budgets ran on an earlier daemon build, and the receipt says so (runtime-pins.json, slice-receipt.json)
  • What ran — 663 calls crossed the wire, against a sealed plan of 674: 557 of the 576 scored cells, 30 discarded warmups, a 72-call flake probe and 4 metered output probes. The 674 is the pre-registration's own roster arithmetic, sealed twenty-six seconds before the first call — the measured figure leads here and the registered one rides beside it, because a plan is not a measurement. The eleven-call gap decomposes exactly, and every term of it is a receipt: minus the 19 cells the budget event never collected, published as not-collected rather than quietly absent; plus the 6 extra warmups the flake leg took; plus 2 probe calls the plan never budgeted. That last term is the one worth naming — the plan registered two metered output probes, one per metered arm, but a probe is not a call: each arm's probe sampled two items, N03 and B01, so four probe calls crossed the wire where the plan had counted two. All four are itemized with their token counts in runtime-pins.json, and the bill prices all four (run-receipt.json, budget-event-c-none.json, warmups.jsonl)
  • Question authorship — authored in-lane against the frozen corpus by the same pen that wrote this page, sealed before the first call; the answer-presence audit is what makes that conflict mechanically checkable rather than merely declared, and it publishes whole (answer-presence-audit.md)
  • Counting rules — every rule that governs a figure here, in one file, including the log-counting sentence verbatim so the 30-day re-measure reconciles (counting-rules.json)
  • Corrected after publication — 2026-08-16, the evening of release, by a content-accuracy pass run over this page against its own kit. Two figures moved, and both are named here rather than quietly swapped: 661 calls became 663, and the gap against the sealed plan went from thirteen to eleven. The cause is one distinction the page had missed — the plan registered two metered output probes, one per metered arm, and each probe turned out to be two calls, so four crossed the wire. runtime-pins.json has itemized all four since the run and the bill priced all four; only this page's arithmetic was short. The registered log claim had lost three words in the retelling and reads “the path was never requested in the window” again, which is what counting-rules.json has said in both places throughout. Four clarifications carry no change of figure: C-MAP is scoped to the two files it actually is rather than to the estate; the strata sentence now uses the registered stratum names; the control block's two different URL counts are reconciled with the rule each one obeys; and the rules behind 51 findings, the eighteen-host census and the 13:40Z provenance are now stated where those figures are used. No scored result moved — every accuracy, count, denominator and band on this page is the one it published with.
  • Limits — one site, one day, one runtime version, two files, eight arms, four of them quantized local seats; a served-context bench and not a crawling one; n=1 per cell with a flake probe instead of retakes; and this page will itself join the corpus it benched, so a re-run needs a fresh question set
  • Authorship — benched, drafted, and audited by the workshop's own agents under a human operator's rulings, then revised with that operator — often across many rounds; nothing releases until they have read it and signed off. The same division of labor this whole hub practices — told in full here
  • Raw rows on request — drop a line.

Corrected 2026-08-20 — four clock readings in this kit re-expressed in UTC. For a reader landing here cold: from this exhibit's release (2026-08-17) until 2026-08-20, runtime-pins.json's four modified_at timestamps carried a local-offset written form. Each is now written as the same instant in UTC; the change is recorded in that file's entry in the kit index as a utc-restate redaction beside the existing ones, the reversal note there says plainly what a reader can and cannot re-derive from public bytes, and the file's sha256 stamp is refreshed. No pin, digest, count, or counting rule moved.

Corrected 2026-09-08 — two sentences in movement three said the Monday watch had never run. For a reader landing here cold: this note is about the weekly estate audit item that re-counts requests for our llms.txt files, which movement three above promises will keep this page honest between re-measures. From this exhibit's release on 2026-08-17 until 2026-09-08 the page said that item had not yet run once and had not yet been exercised on a real run. Both were true when written and stopped being true on 2026-08-24: the item has now run on 2026-08-24, 2026-08-31 and 2026-09-07, and movement three carries those three readings and their totals, with the reason they cannot be set beside this page's own figure. Nothing else in movement three moved: not the sealed count, not the registered rule, and not the 2026-09-16 re-measure date. The separate correction to what this page said about GPTBot is the fourth reading in the post-publication card.

Corrected 2026-10-01 — the seal's first eight characters are withheld. For a reader landing here cold: this note is about the seal in the first line of this section, the sha256 of the manifest that registered this bench before its first call. From this exhibit's release (2026-08-17) until 2026-10-01 (UTC) this page printed its first eight characters, and until 2026-09-29 the kit printed all 64. Since 2026-09-29 the manifest itself carries two withheld values (the kit's README says which and why), and beside them the seal would check a guess at both at once: put candidates back, hash, compare. It reads withheld-2026-10-01 now. The other 20 sealed files still check against the manifest, byte for byte; no call, count, score or rule on this page moved.

Licence: CC BY 4.0 — the whole page, not only the kit. The prose, the tables and the data are yours to quote, re-plot, translate and argue with, including commercially. What we ask back is the one thing the licence already requires: name the source and link to it — strata→signal research, research.strata2signal.com — so a reader of your version can reach ours and check it against the files. Something like — strata→signal research, “The map nobody picks up”, research.strata2signal.com/llms-txt/, CC BY 4.0. And if you quote a figure, quote its denominator beside it — the tables above carry theirs.

One honest edge, which we would rather say than have you find. Two things here are not ours to license. The replies in replies.jsonl were produced by other people's models — eight arms across seven families, every tag printed in the table above — and we cannot license to you what we do not own; each stays subject to whatever terms its own vendor attaches. The same goes for the passages we quote from the convention's own documents: the 2024 proposal, the v2 spec and the two Google pages belong to their authors, and we quote them with their dates so you can go and read the originals. Our questions, our checkers, our scoring, our served blocks, our measurements and every word around them are ours, and those are CC BY 4.0. Republishing the replies or the quoted passages at length is between you and their owners — the rest of this page is between you and us, and the answer is yes.

elsewhere in the workshop

a strata→signal property · hello@strata2signal.com · say hello