{
 "schema": "s2s-bench-v1",
 "kind": "data-kit-index",
 "exhibit": "the-compressed-photograph",
 "exhibit_title": "The compressed photograph — what those Q4_K_M tags actually mean",
 "exhibit_url": "https://research.strata2signal.com/the-compressed-photograph/index.html",
 "published_utc": "2026-08-21",
 "licence": "CC BY 4.0",
 "attribution": "strata→signal research, research.strata2signal.com",
 "contact": "hello@strata2signal.com",
 "what_this_measures": "Two things, both from our own files and cards. (1) The anatomy: the full tensor-by-tensor precision inventory of the same 11.9B model at Q4_K_M and Q8_0, read from the runtime's own /api/show — every tensor name, its storage type, the derived bits-per-weight, and the receipt that the two builds carry identical tensor-name sets and identical full-precision sets — the only difference the inventory can see is storage width. (2) The behavior: one pre-registered four-arm mini-bench (the robber question, two phrasings, n=5 scored calls per arm per phrasing, temperature 0) comparing the same model at both compressions plus the two 27B-class arms from the previous exhibit. The page's accuracy-across-precisions claim re-quotes exhibit fourteen's published precision addendum; nothing there is re-run here. (2) grew on 2026-08-20: the four-build quant ladder of the same 11.9B — every precision the vendor ships — ran under registration #2 and publishes here beside the original pair, attempt #1's refused posture included.",
 "not_a_standard": "One card, one runtime, our own models, short prompts beside live traffic. The robber rows are an illustration under a registered rule, never a ranking, and are NOT comparable to the published 32k-tier census rows. Comparing across quantizations is a confound we name, not one we equalize — the two 12B arms are the same model, the cross-model arms are not. Grew 2026-08-20 with the ladder: four builds of ONE model on two prompts \u2014 its tok/s column orders builds by file size on this card and this runtime and is not a benchmark; it ran beside live product traffic by design, contention measured in both directions.",
 "files": [
  {
   "file": "quant-recipes.json",
   "role": "the anatomy: full per-tensor name-to-precision inventories for gemma4:12b Q4_K_M (338 F32 + 284 Q4_K + 45 Q6_K) and gemma4:12b-it-q8_0 (338 F32 + 329 Q8_0), a flat list of the 45 Q6_K tensor names (1 token embedding, 20 of the 40 attention value projections, 24 of the 48 feed-forward down projections), the same-model receipt (identical tensor-name sets, identical F32 sets), derived bits-per-weight with its rule and basis (whole blob: 5.08 and 8.63; text weights alone for the q4 build: 4.95 — the vision projector sits in the blob but outside the parameter count); the standing-set mixture recipes the page cites are not re-derived here — they live in exhibit seventeen's kit (three-at-the-table/data/active-weights-derivation.md), which this file's header points to",
   "media_type": "application/json",
   "bytes": 154047,
   "sha256": "96b7ccab7b0eda32ef60c40352e9fa2b11c2243dd0079f71eb302b02137e92c2",
   "redactions": []
  },
  {
   "file": "robber-runs.json",
   "role": "the behavior: all 40 scored calls of the four-arm robber mini-bench (4 arms x 2 questions x n=5), runtime counters per call, per-cell medians, the byte-identical determinism receipt (5/5 per cell), and every arm's verbatim reply to both questions — including the q4/q8 answer-length difference (98 vs 64 tokens on the simple ask) and both builds' flip between phrasings",
   "media_type": "application/json",
   "bytes": 20307,
   "sha256": "a6d3e4e2796c5a663ffb282726e0cbed8d33057d1cb6eeede439ed09dabaf2d0",
   "redactions": []
  },
  {
   "file": "ROBBER-PREREG.md",
   "role": "the pre-registration, verbatim as pinned before any scored call: arms, both question specifications, the n=5 rule with decoding parameters, the contention disclosure, and the falsifiable predictions",
   "media_type": "text/markdown",
   "bytes": 1976,
   "sha256": "77b3c4f080cfd4b64bf0fe492122428f6075e6a133f763888a4941b76f89a9e4",
   "redactions": []
  },
  {
   "file": "PREREG-INDEX.txt",
   "role": "the pins, four lines covering all three registrations: ROBBER-PREREG.md (2026-08-19T22:18:55Z, before any scored call), ROBBER-LADDER-PREREG.md pinned twice \u2014 before any pull completed (12:56:18Z) and re-pinned after its appended date-correction note (12:56:37Z), so a reader can verify the amendment changed only what the note says \u2014 and ROBBER-LADDER2-PREREG.md (12:58:31Z, before any scored call). The pin timestamps are the machine record; where the registrations' prose clocks disagree, the pins govern (LADDER-AVAILABILITY.md note 4)",
   "media_type": "text/plain",
   "bytes": 676,
   "sha256": "dbb56c33505d15e2c0876295c0cf0ea4c775213ed058ec721cf10e8d6835ac98",
   "redactions": []
  },
  {
   "file": "ladder-runs.json",
   "role": "the quant ladder, attempt #2 (the publishing run): all scored calls for four builds of the same 11.9B (qat / Q4_K_M / Q8_0 / BF16) on both robber phrasings, n=5 per cell, with per-arm card receipts (driver-level free VRAM + render-queue state before every load), seating verification, the fresh-vs-2026-08-19 byte comparison (all four shared cells IDENTICAL), the byte-identity groups (Q8_0 and BF16 share one sha on the simple ask), verbatim replies, and the blob-bytes receipt behind the derived bits-per-weight column",
   "media_type": "application/json",
   "bytes": 43361,
   "sha256": "8830f82a56fc96ca390deb4d02f57956c7199d78ca3c3b70ed712244c6519a65",
   "redactions": []
  },
  {
   "file": "ladder-runs-attempt1.json",
   "role": "the quant ladder, attempt #1, published as the failure it was: the image pipeline serving the live game re-took the card mid-run, the builds fell to partial residency, and every speed row is REFUSED under the embedded posture ruling — kept because its reply rows still carry the findings: determinism held 5/5 even under partial offload, and TWO cells changed bytes under offloaded arithmetic \u2014 the 8-bit and the bf16 harder-phrasing replies (the bf16 by a single word), both returning to their stable bytes under full residency in attempt #2; the attempt1_vs_attempt2 block in ladder-runs.json holds the machine comparison",
   "media_type": "application/json",
   "bytes": 42250,
   "sha256": "d5f67b06b06a880cb85e11d034d42bfcfffffba44c10e0f0f7737e54793f5c1c",
   "redactions": []
  },
  {
   "file": "ROBBER-LADDER-PREREG.md",
   "role": "ladder registration #1, verbatim as pinned, including its appended pre-run date-correction note: five candidate vendor tags that turned out not to exist — closed unrunnable with no scored calls, no threshold table",
   "media_type": "text/markdown",
   "bytes": 3648,
   "sha256": "b51ebaddd6fb0ed6eba20b30dadc66c6892c5876d0ea3123ef1349ed38298d35",
   "redactions": []
  },
  {
   "file": "ROBBER-LADDER2-PREREG.md",
   "role": "ladder registration #2, verbatim as pinned before any scored call: the four rungs the vendor actually ships, the visitor discipline, and five falsifiable predictions (P1 determinism, P2 cross-day byte-identity, P3 the uncompressed reference's verdict, P4 the two-4-bit-builds question, P5 speed ordering) — all five confirmed by attempt #2",
   "media_type": "text/markdown",
   "bytes": 3548,
   "sha256": "57f2cba33dc1ba6651af65eef85fab9fc9edf4a8b41877c48cddbb34185f2d64",
   "redactions": []
  },
  {
   "file": "LADDER-AVAILABILITY.md",
   "role": "the availability record between the two registrations (the five verbatim 404s and the vendor's real tag list), plus the operator's reverse-contention receipt: the live rules product answering at 35.6 tok/s from a tablet while the heaviest rung decoded, with its own debug-ledger row cited",
   "media_type": "text/markdown",
   "bytes": 4252,
   "sha256": "5b019b78e6d05f2b713b1ae0caab3e94f8239fcdba02f6724c2f8fa2068d8d9c",
   "redactions": []
  }
 ],
 "post_publication_corrections": []
}