{
 "schema_version": "1.0",
 "page_kind": "article",
 "pack_kind": "generate",
 "slug": "llms-txt",
 "title": "The map nobody picks up — and whether it helps when it arrives.",
 "dek": "Two camps have argued about a map file for machines for two years, mostly without receipts — no AI crawler asked us for ours in thirty days of logs, so we measured whether it helps when it arrives anyway.",
 "published": "2026-08-17",
 "series": [
  "bench",
  "notes"
 ],
 "licence": "CC BY 4.0",
 "status": "pending-judge",
 "pack_note": "no judge is seated, so every chip is held",
 "notice": "items with published:false are held and are not this page's published words",
 "source": {
  "url": "https://research.strata2signal.com/llms-txt/",
  "md_url": "https://research.strata2signal.com/llms-txt/index.md",
  "md_sha": "baa4ed961342bf97444e59d05d63c962616ffabe951da26447863277a5f863cc",
  "html_sha": "b21399c10555c130a428af7937765c4b906f96115c403a6a772d290f0633c4eb",
  "anchors_sha": "10f0ef6b1a0ab48f588e135e05c00241e71a989aa4ab75c8b0089b1e0aab9904"
 },
 "short": {
  "paragraph": "Two camps have argued about a map file for machines for two years, mostly without receipts - no AI crawler asked us for ours in thirty days of logs, so we measured whether it helps when it arrives anyway.",
  "from_dek": true,
  "counts": {
   "words": 7687,
   "minutes": 35,
   "tables": 1,
   "kit": true
  },
  "bullets": [
   {
    "text": "the largest llms-full file measured weighs 30.7 MiB",
    "figure": "30.7 MiB",
    "cite": "where-the-file-came-from",
    "quote": "The convention grew anyway, and it grew without limit: the largest llms-full file we measured, from Anthropic's own developer documentation, now weighs 30.7 MiB - call it eight million tokens at four bytes each, an estimate rather than the measured count every other token figure on this page carries, and several times the million-token windows in common use.",
    "span": [
     9636,
     9647
    ]
   }
  ]
 },
 "sections": [
  {
   "id": "measured-again-after-publication",
   "heading": "Measured again, after publication",
   "level": 2,
   "span": [
    1358,
    8786
   ],
   "chunks": [
    [
     1358,
     8786
    ]
   ],
   "chars": 7428,
   "digest": "the article tracks requests for llms.txt files across several readings. it notes that while initial counts showed zero AI crawler requests, subsequent readings identified GPTBot, ClaudeBot, Googlebot, and others. the text explains discrepancies between different counting rules and provides timestamps to reconcile the data.",
   "digest_skipped": null
  },
  {
   "id": "where-the-file-came-from",
   "heading": "Where the file came from.",
   "level": 2,
   "span": [
    8786,
    12406
   ],
   "chunks": [
    [
     8786,
     12406
    ]
   ],
   "chars": 3620,
   "digest": "llms.txt was proposed in September 2024 to help language models understand site shapes. the text describes how the convention grew, noting that some files like llms-full.txt have become very large and expensive to read. it also mentions Google's varying positions on the file.",
   "digest_skipped": null
  },
  {
   "id": "the-confession",
   "heading": "The confession.",
   "level": 2,
   "span": [
    12406,
    18648
   ],
   "chunks": [
    [
     12406,
     18648
    ]
   ],
   "chars": 6242,
   "digest": "the workshop adopted llms.txt but initially faced deployment issues where files were missing or phantom files were served. the text details how these errors were fixed by committing files to source trees. it also notes that some hosts still return 404 errors.",
   "digest_skipped": null
  },
  {
   "id": "does-anyone-fetch-it",
   "heading": "Does anyone fetch it?",
   "level": 2,
   "span": [
    18648,
    22530
   ],
   "chunks": [
    [
     18648,
     22530
    ]
   ],
   "chars": 3882,
   "digest": "the article reports that no AI-associated crawlers asked for llms.txt in a thirty-day window, despite many requests for robots.txt. it clarifies that the findings are based on the workshop's own logs and specific counting rules, noting that most requests were from the workshop's own probes.",
   "digest_skipped": null
  },
  {
   "id": "the-bench-if-the-map-arrives-does-it-help",
   "heading": "The bench: if the map arrives, does it help?",
   "level": 2,
   "span": [
    22530,
    27447
   ],
   "chunks": [
    [
     22530,
     27447
    ]
   ],
   "chars": 4917,
   "digest": "a benchmark was conducted to see if providing an llms.txt map helps models. testing eight arms against three conditions showed that the map helped with navigation, while the site's own prose was better for facts. the results show the map behaves like a map, not an encyclopedia.",
   "digest_skipped": null
  },
  {
   "id": "the-per-arm-table-no-pooled-figure-prints-without-it",
   "heading": "The per-arm table — no pooled figure prints without it",
   "level": 2,
   "span": [
    27447,
    34187
   ],
   "chunks": [
    [
     27447,
     34187
    ]
   ],
   "chars": 6740,
   "digest": "the table presents counts for eight arms across different conditions. it notes that the map's advantage lies in providing web addresses that the HTML control lacks. the text explains that while the map helps with navigation, the site's own words are more effective for factual questions.",
   "digest_skipped": null
  },
  {
   "id": "everything-above-as-files",
   "heading": "Everything above, as files.",
   "level": 2,
   "span": [
    34187,
    34629
   ],
   "chunks": [
    [
     34187,
     34629
    ]
   ],
   "chars": 442,
   "digest": "this section states that all data, including questions, replies, served blocks, and counting rules, is available as files in the data kit under a CC BY 4.0 licence.",
   "digest_skipped": null
  },
  {
   "id": "what-to-take-with-you",
   "heading": "What to take with you",
   "level": 2,
   "span": [
    34629,
    37043
   ],
   "chunks": [
    [
     34629,
     37043
    ]
   ],
   "chars": 2414,
   "digest": "the article summarizes key findings: no AI crawlers asked for the file, the map behaves like a map rather than an encyclopedia, the map's lead was in providing addresses, and the largest llms-full file is very expensive to read.",
   "digest_skipped": null
  },
  {
   "id": "how-to-check-our-work-and-see-it-live",
   "heading": "How to check our work — and see it live",
   "level": 2,
   "span": [
    37043,
    38608
   ],
   "chunks": [
    [
     37043,
     38608
    ]
   ],
   "chars": 1565,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "the-rest-of-the-seminar",
   "heading": "The rest of the seminar",
   "level": 2,
   "span": [
    38608,
    39468
   ],
   "chunks": [
    [
     38608,
     39468
    ]
   ],
   "chars": 860,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "provenance",
   "heading": "Provenance",
   "level": 2,
   "span": [
    39468,
    48708
   ],
   "chunks": [
    [
     39468,
     48708
    ]
   ],
   "chars": 9240,
   "digest": null,
   "digest_skipped": "credits"
  }
 ],
 "chips": [
  {
   "id": "c-003c100c",
   "q": "how many external requests for llms.txt files were logged?",
   "a": "The article reports various counts across different windows, including seventeen genuinely external requests in twenty-six days and sixty-three external requests in a subsequent thirty-four-hour window.",
   "cites": [
    "measured-again-after-publication"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-ff693886",
   "q": "what is the origin of the llms.txt proposal?",
   "a": "The llms.txt file was proposed in September 2024 by Jeremy Howard of Answer.AI as a markdown file containing a site's shape for language models.",
   "cites": [
    "where-the-file-came-from"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-ad7cdb6e",
   "q": "what caused the disappearance of files during site redeploys?",
   "a": "Static-site deploys swapped the entire served directory for a freshly staged one, deleting anything in the server that was not in the site's source tree.",
   "cites": [
    "the-confession"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-718f00c9",
   "q": "how many requests did ai-associated crawlers make to the estate?",
   "a": "In the measured window, AI-associated crawlers made 9,483 requests to the estate, while requests for robots.txt totaled 664.",
   "cites": [
    "does-anyone-fetch-it"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-d64df01f",
   "q": "what is the primary advantage of the map over the html block?",
   "a": "The map's lead was in addresses; the llms.txt block carried fifty web addresses, thirty-six of them distinct, while the HTML block carried none.",
   "cites": [
    "the-per-arm-table-no-pooled-figure-prints-without-it"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-2fc7feab",
   "q": "what files are included in the provided data kit?",
   "a": "The kit contains the sealed questions and keys, every reply with its wire hash, the served blocks, the presence audit, the estate probe, the run receipt, the bill, and the counting rules.",
   "cites": [
    "everything-above-as-files"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-465fb18a",
   "q": "what are the main findings regarding crawler behavior?",
   "a": "AI-associated crawlers made 9,483 requests but asked for llms.txt zero times. However, a later reading logged twenty-six AI-crawler requests, including twenty-four from Meta-ExternalAgent and two from Googlebot.",
   "cites": [
    "what-to-take-with-you"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  }
 ],
 "related": [
  {
   "slug": "reading-the-answer",
   "why": "shares ground with § The bench: if the map arrives, does it help? · § What we asked it"
  },
  {
   "slug": "three-new-voices-at-the-narrators-chair",
   "why": "shares ground with § Everything above, as files. · § Every published file behind every number on this page"
  },
  {
   "slug": "the-new-kid",
   "why": "shares ground with § The bench: if the map arrives, does it help? · § One more bench it happened to sit — the site-reading"
  }
 ],
 "thanks": "",
 "kit": {
  "url": "https://research.strata2signal.com/llms-txt/data/",
  "licence": "CC BY 4.0",
  "files": [
   "golden-set.sealed.json",
   "seal-manifest.json",
   "answer-presence-audit.md",
   "answer-presence-audit.json",
   "checkers.md",
   "checkers.py",
   "checker-audition.json",
   "checker-audition.md",
   "block-c-map.txt",
   "block-c-html.txt",
   "slice-receipt.json",
   "extract.py",
   "scores.json",
   "replies.jsonl",
   "warmups.jsonl",
   "run-receipt.json",
   "budget-event-c-none.json",
   "runtime-pins.json",
   "estate-probe.json",
   "map-stale-hunt.json",
   "report.md",
   "history-sources.md",
   "fetch-receipts.md",
   "counting-rules.json",
   "README.md"
  ]
 },
 "seat_class": {
  "writer": "gemma-class",
  "judge": null,
  "writer_runtime": "vllm"
 },
 "bench": {
  "bullets_written": 3,
  "bullets_kept": 1,
  "digests_written": 8,
  "digests_kept": 8,
  "chips_written": 8,
  "chips_kept": 7,
  "chips_grounded": 0,
  "chips_published": 0,
  "dropped_by": {
   "new_noun": 0,
   "figure": 2,
   "length": 0,
   "cite": 0,
   "judge": 0,
   "quote": 0,
   "directive": 0,
   "redaction": 0,
   "profanity": 0,
   "identifier": 0,
   "bare_figure": 1,
   "house_voice": 0
  },
  "new_noun_tokens_checked": [
   "2024",
   "211",
   "3",
   "30.7",
   "4.0",
   "404",
   "483",
   "61.5",
   "61.5%",
   "664",
   "71.4%",
   "9",
   "AI",
   "AI-associated",
   "AI-crawler",
   "Answer.AI",
   "C-HTML",
   "C-MAP",
   "C-NONE",
   "CC",
   "ClaudeBot",
   "GPTBot",
   "Google's",
   "Googlebot",
   "HTML",
   "Howard",
   "Jeremy",
   "Meta-ExternalAgent",
   "MiB",
   "September"
  ],
  "new_noun_tokens_withheld": 0,
  "figure_definition": "v3",
  "figure_boundary": "guarded",
  "drops": [
   {
    "kind": "bullet",
    "reason": "figure",
    "item": "the map direction leads toward discordant pairs at 61.5 %",
    "detail": "terminal token not verbatim in the cited span"
   },
   {
    "kind": "bullet",
    "reason": "bare_figure",
    "item": "the workshop's own files weigh 3,211",
    "detail": "the cited section writes this figure as '3,211 measured tokens between' and the bullet prints '3,211' alone"
   },
   {
    "kind": "chip",
    "reason": "figure",
    "item": "The bench compared C-MAP, C-HTML, and C-NONE. The registered reading showed that 61.5% of discordant pairs fell toward C-MAP, while an after-the-fact cut for facts alone read 71.4% toward C-HTML.",
    "detail": "figures absent from cited spans: ['71.4']"
   }
  ]
 },
 "generated_utc": "2026-09-30T20:10:10Z"
}
