{
 "schema_version": "1.0",
 "page_kind": "article",
 "pack_kind": "generate",
 "slug": "the-compressed-photograph",
 "title": "The compressed photograph — what those Q4_K_M tags actually mean",
 "dek": "Nearly every local model wears a tag like Q4_K_M, and most of us read one thing from it: fewer bits, smaller file. We opened all four builds the vendor ships of one 11.9B — 4.80 to 16.13 bits per weight — and asked each the same question: the 8-bit answer is byte-identical to the uncompressed original’s, and the wrong answer survives the whole ladder.",
 "published": "2026-08-21",
 "series": [
  "guide"
 ],
 "licence": "CC BY 4.0",
 "status": "pending-judge",
 "pack_note": "no judge is seated, so every chip is held",
 "notice": "items with published:false are held and are not this page's published words",
 "source": {
  "url": "https://research.strata2signal.com/the-compressed-photograph/",
  "md_url": "https://research.strata2signal.com/the-compressed-photograph/index.md",
  "md_sha": "937f6b53083fa73547682e1c8633f5d8469983cc889473c3a227327d8f3f6a86",
  "html_sha": "6aea7edc4b18b5eb35fde6c69c0e4159b08bea196140135404632ca951642b38",
  "anchors_sha": "12a997ca5d3a949623c042a97deccd4d279c68473e84e7218609165c547a096c"
 },
 "short": {
  "paragraph": "Nearly every local model wears a tag like Q4KM, and most of us read one thing from it: fewer bits, smaller file. We opened all four builds the vendor ships of one 11.9B - 4.80 to 16.13 bits per weight - and asked each the same question: the 8-bit answer is byte-identical to the uncompressed original's, and the wrong answer survives the whole ladder.",
  "from_dek": true,
  "counts": {
   "words": 4948,
   "minutes": 22,
   "tables": 3,
   "kit": true
  },
  "bullets": [
   {
    "text": "the q8 build is 28%",
    "figure": "28%",
    "cite": "the-same-brain-at-two-compressions",
    "quote": "The 8-bit copy is 28% slower in tokens per second - which is the same thing as 1.39× the time per token, the form the arithmetic below wants - and the reason is the whole of part one's physics in one line: every written token pays a trip through the weights - one trip per token here, no guesser on either build - and the q8 copy has 70% more bytes to haul.",
    "span": [
     6732,
     6738
    ]
   },
   {
    "text": "the uncompressed original is slower at 58.1",
    "figure": "58.1",
    "cite": "the-same-question-down-the-whole-ladder",
    "quote": "The same question, down the whole ladder The morning after the pair above ran, we extended it to every precision the vendor ships for this model in a format our runtime loads - four builds of the same 11.9 B brain, from a trained-4-bit build up to the uncompressed 24 GB original - and asked the same two robber questions under the same registered rule (both registrations in the kit, n=5 per build per phrasing, temperature 0, one build loaded at a time beside the live set): build bits/weight (blob rule) file decode (simple ask, tok/s, n=5) the simple ask's verdict qat - trained to be 4-bit 4.80 7.15 GB 134.4 No - with the invented settlement rule, in its own words Q4\\K\\M - compressed to 4-bit 5.08 7.56 GB 129.4 No - same rule, byte-for-byte identical to the 2026-08-19 run Q8\\0 8.63 12.84 GB 92.2 No - byte-identical to the uncompressed original's reply BF16 - the uncompressed original 16.13 24.01 GB 58.1 No - the invented settlement rule, at full precision Measured 2026-08-20, 13:27-13:31 UTC · ollama 0.32.13 · same decoding as the table above · medians of five, byte-identical replies 5/5 in every cell · simple-ask ranges 134.06-135.00, 126.10-129.43, 92.11-92.24, and 48.49-58.12 tok/s (that last low first call is the load-warm artefact the registered warm-up absorbs everywhere else) · bits/weight = file bytes × 8 ÷ parameter count, the whole-blob rule defined below - the blob carries a vision projector the parameter count omits, which is why even bf16 reads above 16 · the q4 rung re-ran a point faster than the 2026-08-19 pair's 128.4 in its own fresh window - same bytes out, different clock · full rows, verbatim replies, load receipts, and all three pre-registrations in the kit.",
    "span": [
     8912,
     8920
    ]
   },
   {
    "text": "the extra bits increased latency by 77%",
    "figure": "77%",
    "cite": "the-receipt-we-bought-three-times-the-bits-and-got-nothing",
    "quote": "And the bill for the extra faithfulness was real: about three times the bits per weight by this page's own division - that seat's q4 runs 5.20 bits, bf16 runs 16; the tags say four times, but the tags always say four - three times the resident memory (17.5 → 29.2 → 53.1 GB), 77% more median latency (406 → 560 → 719 ms).",
    "span": [
     22314,
     22318
    ]
   }
  ]
 },
 "sections": [
  {
   "id": "the-same-brain-at-two-compressions",
   "heading": "The same brain at two compressions",
   "level": 2,
   "span": [
    2418,
    7842
   ],
   "chunks": [
    [
     2418,
     7842
    ]
   ],
   "chars": 5424,
   "digest": "comparing two compressions of an 11.9 billion weight model shows that quantization preserves a confident wrong answer regarding a Catan rule. while the q4 version is faster than the q8 version, the wording of the question determines the verdict rather than the compression level.",
   "digest_skipped": null
  },
  {
   "id": "the-same-question-down-the-whole-ladder",
   "heading": "The same question, down the whole ladder",
   "level": 2,
   "span": [
    7842,
    15127
   ],
   "chunks": [
    [
     7842,
     15127
    ]
   ],
   "chars": 7285,
   "digest": "testing four builds from a trained 4-bit version to the uncompressed original shows that the model's incorrect verdict remains constant across all precisions. decode speed follows file size, with the trained-4-bit build being the fastest at 134.4 tok/s.",
   "digest_skipped": null
  },
  {
   "id": "what-the-tag-actually-says",
   "heading": "What the tag actually says",
   "level": 2,
   "span": [
    15127,
    16592
   ],
   "chunks": [
    [
     15127,
     16592
    ]
   ],
   "chars": 1465,
   "digest": "the Q4_K_M tag describes a specific bargain. Q4 refers to the headline precision, K denotes the K-quants compression format using nested super-blocks, and M indicates the mix of how many tensors are bumped above the headline bits.",
   "digest_skipped": null
  },
  {
   "id": "fewer-bits-is-not-rounding-off",
   "heading": "Fewer bits is not \"rounding off\"",
   "level": 2,
   "span": [
    16592,
    17870
   ],
   "chunks": [
    [
     16592,
     17870
    ]
   ],
   "chars": 1278,
   "digest": "quantization uses blocks and scales rather than simple rounding. weights are grouped into neighborhoods that each carry a higher-precision scale, allowing K-quants to nest these rulers for finer granularity in different ranges.",
   "digest_skipped": null
  },
  {
   "id": "the-reveal-a-4-bit-model-is-not-4-bit",
   "heading": "The reveal: a \"4-bit model\" is not 4-bit",
   "level": 2,
   "span": [
    17870,
    20979
   ],
   "chunks": [
    [
     17870,
     20979
    ]
   ],
   "chars": 3109,
   "digest": "an inventory of the 11.9 B Q4_K_M build reveals it is not strictly 4-bit. while most weights are Q4_K, 45 tensors including token embeddings and parts of the attention and feed-forward projections are stored at Q6_K.",
   "digest_skipped": null
  },
  {
   "id": "the-receipt-we-bought-three-times-the-bits-and-got-nothing",
   "heading": "The receipt: we bought three times the bits and got nothing",
   "level": 2,
   "span": [
    20979,
    23916
   ],
   "chunks": [
    [
     20979,
     23916
    ]
   ],
   "chars": 2937,
   "digest": "a moderation classifier test showed that increasing precision from q4 to bf16 resulted in identical verdicts across 1,122 scored calls. the extra bits increased resident memory and latency without improving accuracy.",
   "digest_skipped": null
  },
  {
   "id": "what-the-freed-memory-buys",
   "heading": "What the freed memory buys",
   "level": 2,
   "span": [
    23916,
    25935
   ],
   "chunks": [
    [
     23916,
     25935
    ]
   ],
   "chars": 2019,
   "digest": "using fewer bytes per weight allows for larger models in the same memory or multiple models to reside on one card. at q4, about 1.7× the parameters can fit in the same weight budget compared to q8.",
   "digest_skipped": null
  },
  {
   "id": "what-this-page-does-not-say",
   "heading": "What this page does not say",
   "level": 2,
   "span": [
    25935,
    27179
   ],
   "chunks": [
    [
     25935,
     27179
    ]
   ],
   "chars": 1244,
   "digest": "this article does not claim q4 is always free, nor does it rank builds or attribute differences to quantization across different models. it also does not measure perplexity.",
   "digest_skipped": null
  },
  {
   "id": "what-to-take-with-you",
   "heading": "What to take with you",
   "level": 2,
   "span": [
    27179,
    28646
   ],
   "chunks": [
    [
     27179,
     28646
    ]
   ],
   "chars": 1467,
   "digest": "the article concludes that tags are recipes rather than exact contents, quantization uses blocks and rulers, and extra precision often buys no additional accuracy for certain tasks.",
   "digest_skipped": null
  },
  {
   "id": "how-to-check-our-work-and-see-it-live",
   "heading": "How to check our work — and see it live",
   "level": 2,
   "span": [
    28646,
    30237
   ],
   "chunks": [
    [
     28646,
     30237
    ]
   ],
   "chars": 1591,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "the-rest-of-the-seminar",
   "heading": "The rest of the seminar",
   "level": 2,
   "span": [
    30237,
    31657
   ],
   "chunks": [
    [
     30237,
     31657
    ]
   ],
   "chars": 1420,
   "digest": null,
   "digest_skipped": "credits"
  }
 ],
 "chips": [
  {
   "id": "c-f2d0a6fc",
   "q": "how do q4 and q8 compressions compare on the same model?",
   "a": "The q4 compression uses about four bits per weight, while the q8 version uses about eight. While they share the same training and tensor sets, the q8 file is 70% larger on disk. In testing, the q4 build was faster, reaching 128.4 tok/s compared to 92.2 tok/s for the q8 build.",
   "cites": [
    "the-same-brain-at-two-compressions"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-1a3060ef",
   "q": "what do the components of a q4_k_m tag signify?",
   "a": "The Q4 indicates the headline precision class of roughly four bits per weight. The K refers to the K-quants compression format using nested super-blocks. The M denotes the mix, which determines how many specific tensors are bumped to a higher precision than the headline bits.",
   "cites": [
    "what-the-tag-actually-says"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-748cb296",
   "q": "how does the block and scale method work in quantization?",
   "a": "Instead of simple rounding, weights are grouped into neighborhoods. Each neighborhood uses a higher-precision scale, or ruler, to read its values. K-quants nest this further by using a 16-bit master ruler for a super-block and smaller, cheaper rulers for its sub-blocks.",
   "cites": [
    "fewer-bits-is-not-rounding-off"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-323d663c",
   "q": "why is a 4-bit model actually higher than four bits per weight?",
   "a": "A 4-bit model uses a recipe that protects certain tensors. In the tested 11.9 B model, 45 tensors—including token embeddings and half of certain projections—are stored at Q6_K (six bits), while 338 structural tensors remain at F32 precision.",
   "cites": [
    "the-reveal-a-4-bit-model-is-not-4-bit"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-e1a0de68",
   "q": "what was the result of testing higher precision on a binary task?",
   "a": "In a moderation classifier exam, testing q4, q8, and bf16 yielded identical results across 1,122 scored calls. The extra bits provided no increase in accuracy, even though the higher precision versions required three times the resident memory and 77% more median latency.",
   "cites": [
    "the-receipt-we-bought-three-times-the-bits-and-got-nothing"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-02a98e2f",
   "q": "what are the practical benefits of using lower bit-rates?",
   "a": "Lower bit-rates allow for more parameters to fit in the same memory budget; about 1.7× the parameters fit at the q4 rate compared to q8. This freed memory also allows multiple models to reside on a single card simultaneously, preventing performance-killing reloads.",
   "cites": [
    "what-the-freed-memory-buys"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-eaf76752",
   "q": "what limitations does this article acknowledge regarding its findings?",
   "a": "The article does not claim q4 is universally superior, nor does it rank builds into a winner. It avoids attributing differences to quantization across different models and does not measure perplexity, focusing instead on verdicts, replies, and clock speeds.",
   "cites": [
    "what-this-page-does-not-say"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  }
 ],
 "related": [
  {
   "slug": "compression-is-the-objective",
   "why": "shares ground with § The same brain at two compressions · § What to take with you"
  },
  {
   "slug": "three-at-the-table",
   "why": "shares ground with § The same brain at two compressions · § The arithmetic that catches a mixture red-handed"
  },
  {
   "slug": "two-hours-on-battery",
   "why": "shares ground with § The same question, down the whole ladder · § The one table to remember"
  }
 ],
 "thanks": "",
 "kit": {
  "url": "https://research.strata2signal.com/the-compressed-photograph/data/",
  "licence": "CC BY 4.0",
  "files": [
   "quant-recipes.json",
   "robber-runs.json",
   "ROBBER-PREREG.md",
   "PREREG-INDEX.txt",
   "ladder-runs.json",
   "ladder-runs-attempt1.json",
   "ROBBER-LADDER-PREREG.md",
   "ROBBER-LADDER2-PREREG.md",
   "LADDER-AVAILABILITY.md"
  ]
 },
 "seat_class": {
  "writer": "gemma-class",
  "judge": null,
  "writer_runtime": "vllm"
 },
 "bench": {
  "bullets_written": 3,
  "bullets_kept": 3,
  "digests_written": 9,
  "digests_kept": 9,
  "chips_written": 8,
  "chips_kept": 7,
  "chips_grounded": 0,
  "chips_published": 0,
  "dropped_by": {
   "new_noun": 1,
   "figure": 0,
   "length": 0,
   "cite": 0,
   "judge": 0,
   "quote": 0,
   "directive": 0,
   "redaction": 0,
   "profanity": 0
  },
  "new_noun_tokens_checked": [
   "1",
   "1.7×",
   "11.9",
   "122",
   "128.4",
   "134.4",
   "16-bit",
   "28%",
   "338",
   "4-bit",
   "45",
   "58.1",
   "70%",
   "77%",
   "92.2",
   "B",
   "Catan",
   "F32",
   "K",
   "K-quants",
   "M",
   "Q4",
   "Q4K",
   "Q4KM",
   "Q6K",
   "bf16",
   "q4",
   "q4km",
   "q8",
   "trained-4-bit"
  ],
  "new_noun_tokens_withheld": 1,
  "figure_definition": "v3",
  "drops": [
   {
    "kind": "chip",
    "reason": "new_noun",
    "item": null,
    "item_chars": 9,
    "withheld": "a string naming something the article does not",
    "detail": ""
   }
  ]
 },
 "generated_utc": "2026-09-09T01:27:57Z"
}
