# The map nobody picks up — data kit

The machine-readable companions to **exhibit thirteen, "The map nobody picks up"**
(https://research.strata2signal.com/llms-txt/index.html): the history of
`llms.txt` from its primary sources, a confession about our own estate serving
it wrong, thirty days of our server logs answering whether anything fetches it,
and a sealed eight-model bench on the question nobody had published — *if the
map does reach a model, does it help?*

**Licence: CC BY 4.0.** Take these rows, re-derive our figures, disagree with
them in public. Attribution: strata→signal research, research.strata2signal.com.
If you find an error in them, we want to hear about it: hello@strata2signal.com.

## What this measures, and what it does not

A **served-context** bench. At a matched token budget, which representation of
one site — its `llms.txt` files, or an equal budget of its own pages
text-extracted — carries more answerable information when it is placed in a
model's context. It does **not** measure crawling, fetching, adoption, or agent
behaviour in the wild. Movement three of the exhibit measures the fetching,
separately, from server logs, and the answer there was zero.

## What is in here

`index.json` lists every file with its size, its sha256, the role it plays, and
the exact redactions applied to it. In reading order:

| file | what it holds |
|---|---|
| `golden-set.sealed.json` | the sealed question set entire — all 30 authored items (24 run, 6 held as a pre-sealed reserve), each with its category, its exact key, its accepted variants, the decoys the audition rejected, its evidence span, its `exact`/`proxy` stamp and the published line it was authored against |
| `seal-manifest.json` | the seal: 22 files hashed before the first model call, the twelve sealing preconditions each with its evidence, and the one shortfall recorded with its cause rather than re-scoped |
| `answer-presence-audit.md` · `answer-presence-audit.json` | the audit that makes this bench falsifiable: for every question, whether its answer is present in the map block, the slice block, both, or neither — computed before the first call, at zero model cost, with the author's prediction printed beside the computed answer |
| `checkers.md` · `checkers.py` | one rule per item, generated from the checker functions' own docstrings and carrying the module's sha256 — and the module itself, so the sha can be recomputed and every scoring decision read as code |
| `checker-audition.json` · `checker-audition.md` | the audition run before sealing: 335 cases — every canonical key, every hand-written paraphrase, every decoy, and the shape cases — with each case's verdict |
| `block-c-map.txt` · `block-c-html.txt` | the two served blocks **verbatim**, byte for byte as eight models were handed them: 12,080 B / 3,211 tokens and 13,120 B / 3,088 tokens |
| `slice-receipt.json` | how the HTML block was cut: page order, the reference tokenizer and its settings, the budget band, the byte offset of the cut, the last block included and the first block excluded |
| `extract.py` | the hand-written text extractor — published because it is the thumb on the scale the exhibit confesses: it keeps visible words and drops the link targets, which is why the control could not answer a navigation question |
| `scores.json` | the scoring artifact entire: all 576 cells read, the contamination exclusion with its rule, the three fences, the headline table, the pooled discordance, the post-hoc structural sensitivity, every stratum, the refusal-bait rows and the flake probe |
| `replies.jsonl` | every reply verbatim, one JSON object per cell: the model's text, its reply and wire sha256, the endpoint's own counters, the client-wall latency and the runtime deltas it answered under |
| `warmups.jsonl` | the 30 discarded warmups — one per arm per condition plus the six the flake leg took, billed where metered, never scored — published so the cold-start calls are not an unexplained gap between the call count and the scored cells. The pre-registration budgeted 24; the run made 30, and the wire-call arithmetic counts the 30 that happened |
| `run-receipt.json` · `budget-event-c-none.json` | what actually ran, per leg, with the pause rule and its probes (including the one that was UNAVAILABLE and is recorded as a gap rather than invented); and the budget halt in full |
| `runtime-pins.json` | which weights answered: the daemon version, each local tag's digest, parameter count and quantisation, read *before* the first scored call, plus the cloud shelf re-probe at registration |
| `estate-probe.json` | the confession's counts, dated and content-classified: every host probed, its status, bytes and sha256, and whether the served file byte-matches a source tree |
| `map-stale-hunt.json` | the acquittal: every URL the two map files assert, resolved; every numeric claim they make, checked against the page it describes |
| `counting-rules.json` | every rule that governs a figure on the page, in one file — including the log-counting sentence verbatim, the registered bands, the fences, and the forbidden-claims register |
| `report.md` | the complete internal report, written by machine from the artifacts above; every figure on the exhibit is a projection of it |
| `history-sources.md` | movement one's sources, one bullet each: every claim about the convention's history with the URL it came from, the date it was fetched and a quote — gathered by live fetch on 2026-08-13, never from model memory, and carrying its own UNVERIFIED marks |
| `fetch-receipts.md` | movement one's and movement three's measured fetch evidence: what the major publishers serve (with the 32,161,074-byte HEAD receipt behind the thirty-megabyte figure), the log method and per-log windows, every `llms*` request in the window by user agent, and the per-crawler request table with its honest limits |
| `index.json` | this directory, listed: names, sizes, sha256s, roles, and the redactions applied to each |

## The three checks worth running first

1. **The seal.** `seal-manifest.json` carries the sha256 of 22 files taken
   before the first model call. Twenty of them publish here byte-identical, and
   `index.json` marks each `matches_seal: true` — recompute and compare. Two do
   not, and say so: `runtime-pins.json` and `estate-probe.json` carry the
   redactions listed below, so their published hash differs from the sealed one,
   which is printed beside it. The manifest also records that this bench sealed
   on **eleven of twelve** preconditions, with the twelfth — a registered
   MAP-STALE stratum that came back vacant — published as a reduction.
2. **The headline, both ways.** `scores.json` → `headline.pooled_discordance`
   gives b = 64, c = 40, 104 discordant pairs, **61.5%** toward C-MAP against a
   registered 60/40 band. Then read `structural_sensitivity`: remove the six
   navigation items and it is b = 16, c = 40, 56 pairs, **71.4%** toward C-HTML
   — a cut that was **not** pre-registered and is not this bench's result. And
   look at the per-arm rows while you are there: `b` is exactly **8 for all
   eight arms**, which is the count of MAP-ONLY items in a set we wrote, and
   `answer-presence-audit.json` had predicted every item's stratum before a
   model was called.
3. **The bill.** `run-receipt.json` → **$1.8688 of a $4.00 ceiling**, priced at
   the uncached rates, which is an upper bound by construction. The one budget
   event is in `budget-event-c-none.json`: the closed-book leg of the priciest
   arm halted on a registered 900-output-token line at a measured median of
   3,633 — not on its dollar cap, which it never neared — and 19 cells publish
   NOT-COLLECTED rather than quietly absent.

## Limits, stated plainly

- **One site, one day, one runtime version, two files.** Eight arms, four of
  them quantised local seats between 25.8B and 70.6B, four frontier cloud, on a
  single box. Nothing here generalises to *websites*; it is what these two files
  did for these eight arms on this site.
- **n = 1 per cell** at temperature zero. The gap is closed by measurement
  rather than argument: a three-draw flake probe on the six zero-cost arms over
  six items found zero flips. The two metered arms carry no flake figure, and
  the reason is cost.
- **The contamination filter removed outright knowledge, not contamination.**
  A served block can cue half-remembered knowledge; that interaction is
  invisible to a closed-book filter and is a standing bound on what the
  difference between conditions can mean.
- **The navigation result is structural.** Our extractor drops link targets, so
  the control carried no URLs and could not answer a navigation question. The
  exhibit leads with the number that carries, states the thumb, and prints the
  reversal.
- **We wrote the file, the questions, the checkers and the page.** No Claude arm
  sits on the roster, deliberately. The answer-presence audit exists to make
  that conflict mechanically checkable rather than merely declared.
- **This page will join the corpus it benched.** A re-run of this instrument
  needs a fresh question set; models with a later training cutoff may have read
  this kit.
- **Not a standard and not independent numbers.** Our own file, our own bench,
  our kit — yours to reuse and to disagree with.

## What was redacted, and what is not here

**Redactions.** Every file is a verbatim copy of the sealed workspace artifact
except for the substitutions `index.json` names per file: the local inference
host's internal name becomes `the-gpu-box`, the rented server's becomes
`the-vps`, private addresses and the daemon port are marked redacted in place,
and working-copy paths collapse to a marker that says what was cut. Nothing was
deleted, no measurement was altered, and no figure on the page depends on a
redacted string. Two sealed files carry redactions and therefore publish with a
hash that differs from the seal's; both say so in their index entry, and
reversing the substitutions returns the sealed hash.

**Not here, and why:**

- **The raw vhost access logs.** They hold visitors' IP addresses. The counting
  rule publishes verbatim in `counting-rules.json`, and the per-crawler totals
  publish on the page, so the figure is reproducible against your own logs even
  though ours stay ours.
- **The frozen corpus mirror and its manifest.** The bench froze fourteen pages
  of the live site and hashed them; the manifest is a list of working-copy paths
  on the machines that ran it. What the models actually saw is not the mirror
  but the two blocks cut from it, and those publish here whole, with their
  sealed hashes and the receipt that describes the cut.
- **The plan of record.** It names working-copy paths and internal repo slugs
  throughout. Rather than ship a document scrubbed until a reader could not tell
  what had been cut, the parts that bind a published figure — the counting
  rules, the registered bands, the fences, the forbidden-claims register and the
  titling constraint — are extracted verbatim into `counting-rules.json`.
- **`results/c-none.log`.** The closed-book leg halted mid-run before its runner
  wrote a log, so the file is empty. The leg is reconstructed from the journal
  and the budget event, and `run-receipt.json` labels that row as rebuilt.

One field in `replies.jsonl` deserves a note before it confuses somebody:
`envelope_check` is vestigial. The record schema is lifted from an earlier
bench that asked models for a JSON envelope; this bench asks for one line of
plain text, so `envelope_ok` reads false on every row and means nothing here.
The scoring lives in `scores.json` and the checkers that produced it.

## Two numbering notes, added 2026-08-16 after publication

**`map-stale-hunt.json` numbers the corpus by working copy, not by exhibit.**
Its `claim_rows` carry paths like `01-the-open-call`, `05-chair-trials`,
`08-diffusion`. Those two-digit prefixes are the frozen corpus mirror's own file
ordering — the order the pages were cut in — and they are **not** this site's
exhibit numbers, which they contradict. The mapping, so nobody reconciles the
wrong pair of numbers:

| corpus path in the file | the exhibit it actually is |
| --- | --- |
| `01-the-open-call` | exhibit **11** |
| `05-chair-trials` | exhibit **6** |
| `06-cove-voice-head-to-head` | exhibit **7** |
| `07-outside-judges` | exhibit **8** |
| `08-diffusion` | exhibit **1** |
| `09-seat-trials` | exhibit **2** |
| `10-voice-trials` | exhibit **3** |

The public numbers are publication-order identity and never renumber;
`/data/index.json` is authoritative for them. The file itself is **sealed** —
its sha256 is in the seal manifest and `index.json` records `matches_seal: true`
— so it is disclosed here rather than edited. A sealed artifact that turns out
to carry a confusing internal convention gets a note beside it, never a quiet
correction inside it, because the whole value of the seal is that the bytes
under it did not move.

**The estate fix's four commits**, cited on the page and named here with their
repositories, since one SHA alone is not a receipt. All four landed 2026-08-16
inside a forty-second window:

| SHA | repository | subject line |
| --- | --- | --- |
| `9631d54` | the estate web repo | *llms.txt gets a source of record, estate-wide* |
| `7b5aadf` | the walking-tour guide | *llms.txt, and the deploy learns to carry it* |
| `4a73bd1` | this research hub | *the pages point at llms.txt the v2 way* |
| `ea2f5b8fe` | the RPG engine repo (RealKeep site) | *realkeep site: llms.txt comes home* |

The repositories are private, so the commits are not clickable; the subject
lines and the SHAs are given so the claim is specific enough to be wrong.
