{
 "schema_version": "1.0",
 "page_kind": "article",
 "pack_kind": "generate",
 "slug": "six-cards-the-same-questions",
 "title": "Six graphics cards, the same questions",
 "dek": "An RTX 3080, two of them in one desktop, an RTX 3090, two of those, an RTX 3090 Ti and an RTX PRO 6000 Blackwell, with a laptop part and an RTX 3080 Ti beside them — asked the same questions, in one table each: how fast each writes, how fast each draws, what the card itself pulled while doing it, how hot it got, and what it cost on the day it was priced. Not every card answers every question, and each table says which did not. Almost nothing here was measured for this page: every figure in its tables but two is already published on one of six earlier pages, and every one says which page and which section it came from. The two exceptions are prices for one card, read at a retailer and at a marketplace; the table says so, and names what else those reads found.",
 "published": "2026-10-01",
 "series": [
  "bench"
 ],
 "licence": "CC BY 4.0",
 "status": "pending-judge",
 "pack_note": "no judge is seated, so every chip is held",
 "notice": "items with published:false are held and are not this page's published words",
 "source": {
  "url": "https://research.strata2signal.com/six-cards-the-same-questions/",
  "md_url": "https://research.strata2signal.com/six-cards-the-same-questions/index.md",
  "md_sha": "72e082ab5a76104a48a18d28f04dedb72c8e2e825440c02ef96c3dd6159254c2",
  "html_sha": "d3b66c5445e7e74b84b5328649124e4675aa0a94cecf915dbfef91c4a5f1aa4b",
  "anchors_sha": "34cc0ed03b5be0a333f6482a94b00c2d49c4e16e246e06cfd8117ab3ce2ffd08"
 },
 "short": {
  "paragraph": "A power cap moves a card more than the card's own name does: one RTX 3090, on one day, wrote with a dense model at 33.3 tokens a second held at 250 W and 47.8 tokens a second at 300 W - a bigger jump than any gap between two different cards at the same cap in the dense model's table. Two used RTX 3080s beat one 3090 on the same dense model at a window both can hold (the memory reserved for the conversation) - 44.6 tokens a second against 33.1, for 250 W on each of two boards against one board's 250 W, a cap their page says costs the single 3090 about 30 per cent on this model and the pair 0.2 per cent - and then lose it entirely at the next window up, where the page prints does not fit. Drawing pictures on the eight-step sketch graph, the order is plainer: one 3080 takes 2.074 seconds a picture and one 3090 1.683 seconds, both at 250 W, and one 3090 Ti 1.369 seconds at its own 450 W limit. A used 3080 asked $375.00 on 17 September 2026; a renewed 3090 asked $1,879.99 the same day - one seller's card, its page says, not a market price; used 3090s sold for $1,094.44 on average in August 2026, by that marketplace's own record. Every number in this page's tables but two was published first somewhere else, with its conditions; the two are an RTX 3080 Ti's prices - a refurbished listing at $479.99 from a retailer, out of stock at both of its last two reads, and a used card's August 2026 average sale of $475.64 from a marketplace - both read on 1 October 2026 and both marked as such, with the other listings those reads turned up named under the table rather than in it. Where two cards were never measured comparably the cell is a dagger and a footnote rather than a number.",
  "from_dek": false,
  "counts": {
   "words": 15178,
   "minutes": 69,
   "tables": 17,
   "kit": false
  },
  "bullets": []
 },
 "sections": [
  {
   "id": "the-words-this-page-leans-on",
   "heading": "The words this page leans on",
   "level": 2,
   "span": [
    3416,
    5404
   ],
   "chunks": [
    [
     3416,
     5404
    ]
   ],
   "chars": 1988,
   "digest": "this section defines technical terms used throughout the article, including token, window, dense, and mixture models, as well as units like J / 1k tok and measurements like board power and power caps.",
   "digest_skipped": null
  },
  {
   "id": "the-cards-named-once",
   "heading": "The cards, named once",
   "level": 2,
   "span": [
    5404,
    8900
   ],
   "chunks": [
    [
     5404,
     8900
    ]
   ],
   "chars": 3496,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "how-to-read-the-tables",
   "heading": "How to read the tables",
   "level": 2,
   "span": [
    8900,
    13012
   ],
   "chunks": [
    [
     8900,
     13012
    ]
   ],
   "chars": 4112,
   "digest": "this section explains the table formatting, describing how figures, daggers, em dashes, and words are used to represent data, and clarifies how power caps, energy units, and second readings are presented.",
   "digest_skipped": null
  },
  {
   "id": "writing-how-fast-each-one-answers",
   "heading": "Writing: how fast each one answers",
   "level": 2,
   "span": [
    13012,
    22982
   ],
   "chunks": [
    [
     13012,
     22982
    ]
   ],
   "chars": 9970,
   "digest": "this section compares writing speeds for a mixture model, gemma4:26b, and a dense model, mistral-small3.2:24b, across different power caps, noting how the mixture model is faster for its size.",
   "digest_skipped": null
  },
  {
   "id": "when-the-model-is-bigger-than-the-card",
   "heading": "When the model is bigger than the card",
   "level": 2,
   "span": [
    22982,
    30276
   ],
   "chunks": [
    [
     22982,
     30276
    ]
   ],
   "chars": 7294,
   "digest": "this section discusses how context windows affect memory usage, explaining that if a model and its window do not fit in a card, the model spills into computer RAM, slowing performance.",
   "digest_skipped": null
  },
  {
   "id": "what-the-boards-actually-drew",
   "heading": "What the boards actually drew",
   "level": 2,
   "span": [
    30276,
    37067
   ],
   "chunks": [
    [
     30276,
     37067
    ]
   ],
   "chars": 6791,
   "digest": "this section clarifies that the figures represent card board power read from the driver rather than wall power, showing what each card actually drew under various power caps while writing.",
   "digest_skipped": null
  },
  {
   "id": "what-a-thousand-tokens-cost-in-energy",
   "heading": "What a thousand tokens cost in energy",
   "level": 2,
   "span": [
    37067,
    43870
   ],
   "chunks": [
    [
     37067,
     43870
    ]
   ],
   "chars": 6803,
   "digest": "this section presents the energy cost in joules per thousand tokens for both the mixture and dense models across different power caps, noting that lower energy usage is better.",
   "digest_skipped": null
  },
  {
   "id": "drawing-pictures",
   "heading": "Drawing pictures",
   "level": 2,
   "span": [
    43870,
    50381
   ],
   "chunks": [
    [
     43870,
     50381
    ]
   ],
   "chars": 6511,
   "digest": "this section provides data on image generation speeds using an eight-step graph, a thirty-two-step graph, and ten minutes of continuous drawing, measuring results in seconds per picture or total images finished.",
   "digest_skipped": null
  },
  {
   "id": "the-sketch-ladder-and-the-one-row-here-that-is-not-a-card",
   "heading": "The sketch ladder, and the one row here that is not a card",
   "level": 2,
   "span": [
    50381,
    54293
   ],
   "chunks": [
    [
     50381,
     54293
    ]
   ],
   "chars": 3912,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "heat-and-how-hard-the-fans-worked",
   "heading": "Heat, and how hard the fans worked",
   "level": 2,
   "span": [
    54293,
    63844
   ],
   "chunks": [
    [
     54293,
     63844
    ]
   ],
   "chars": 9551,
   "digest": "this section details core temperatures and fan speeds for various cards during writing and drawing tasks, noting that fan percentages are relative to each card's own maximum.",
   "digest_skipped": null
  },
  {
   "id": "one-box-two-cards",
   "heading": "One box, two cards",
   "level": 2,
   "span": [
    63844,
    68437
   ],
   "chunks": [
    [
     63844,
     68437
    ]
   ],
   "chars": 4593,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "what-each-one-costs-and-on-what-day",
   "heading": "What each one costs, and on what day",
   "level": 2,
   "span": [
    68437,
    79939
   ],
   "chunks": [
    [
     68437,
     79939
    ]
   ],
   "chars": 11502,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "what-this-comparison-cannot-say",
   "heading": "What this comparison cannot say",
   "level": 2,
   "span": [
    79939,
    82029
   ],
   "chunks": [
    [
     79939,
     82029
    ]
   ],
   "chars": 2090,
   "digest": "this section outlines the limitations of the comparison, such as the fact that the cards did not all sit in the same machine and that no memory-die temperatures are recorded.",
   "digest_skipped": null
  },
  {
   "id": "what-to-take-with-you",
   "heading": "What to take with you",
   "level": 2,
   "span": [
    82029,
    85358
   ],
   "chunks": [
    [
     82029,
     85358
    ]
   ],
   "chars": 3329,
   "digest": "this section summarizes key findings, such as how power caps affect dense models more than mixture models and how memory capacity determines if a model fits on a card.",
   "digest_skipped": null
  },
  {
   "id": "how-to-check-our-work",
   "heading": "How to check our work",
   "level": 2,
   "span": [
    85358,
    87584
   ],
   "chunks": [
    [
     85358,
     87584
    ]
   ],
   "chars": 2226,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "the-rest-of-the-seminar",
   "heading": "The rest of the seminar",
   "level": 2,
   "span": [
    87584,
    90152
   ],
   "chunks": [
    [
     87584,
     90152
    ]
   ],
   "chars": 2568,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "who-ran-this-and-thanks",
   "heading": "Who ran this, and thanks",
   "level": 2,
   "span": [
    90152,
    92892
   ],
   "chunks": [
    [
     90152,
     92892
    ]
   ],
   "chars": 2740,
   "digest": null,
   "digest_skipped": "credits"
  }
 ],
 "chips": [
  {
   "id": "c-44462f14",
   "q": "how are tokens and context windows defined?",
   "a": "A token is the chunk a model reads and writes, slightly less than a word. The context window is the amount of conversation a model holds in mind at once, measured in tokens. This window is reserved up front, meaning a longer window costs memory regardless of whether it is filled.",
   "cites": [
    "the-words-this-page-leans-on"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-3f901348",
   "q": "what determines the speed of a mixture model?",
   "a": "A mixture model is faster because it only wakes part of itself for each word. In the writing tests, the gemma4:26b mixture model achieved speeds such as 125.6 tok/s at a 250 W cap, whereas the dense mistral-small3.2:24b model reached 33.3 tok/s under the same conditions.",
   "cites": [
    "writing-how-fast-each-one-answers"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-9e4caf80",
   "q": "what happens when a model exceeds available memory?",
   "a": "When a model is larger than the card's memory, it spills into the computer's RAM. This causes the card to spend time waiting on the system memory. For example, a 10 GB card running a 24b model at a 65,536-token window results in a 'does not fit' status.",
   "cites": [
    "when-the-model-is-bigger-than-the-card"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-f6d66848",
   "q": "how is board power measured in these tables?",
   "a": "Board power is the actual amount a card draws, read directly from its own driver rather than from a wall socket. This distinguishes it from a power cap, which is a ceiling set by the workshop to limit how much a card may draw.",
   "cites": [
    "what-the-boards-actually-drew"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-a8054ffc",
   "q": "how is energy efficiency calculated for text generation?",
   "a": "Energy efficiency is measured in joules per thousand tokens (J / 1k tok). This represents the energy the board spent to write a thousand tokens. Lower values indicate better efficiency, such as 1,758.13 J / 1k tok for a mixture model at a 350 W cap.",
   "cites": [
    "what-a-thousand-tokens-cost-in-energy"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-f4149d7b",
   "q": "how does drawing speed vary with step counts?",
   "a": "The article compares an eight-step sketch graph to a thirty-two-step graph. An eight-step process is described as a sketch, while a thirty-two-step process is considered finished. For instance, an RTX 3090 at a 250 W cap takes 1.683 s for an eight-step picture and 6.297 s for a thirty-two-step picture.",
   "cites": [
    "drawing-pictures"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-c402ff31",
   "q": "how does temperature relate to workload?",
   "a": "Temperature measurements represent the maximum core load during a run. The article notes that a temperature reading without its workload context says little, as different runs may involve different durations or cooling conditions, such as a 3090 peaking at 67 °C while drawing at a 250 W cap.",
   "cites": [
    "heat-and-how-hard-the-fans-worked"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-9d23f259",
   "q": "what is the relationship between memory and speed?",
   "a": "Memory is the primary factor in model capacity and does not appear in speed figures. While two smaller cards might provide higher tokens-per-second at smaller windows, they fail when the model and context window exceed their combined capacity, whereas a single larger card continues to function.",
   "cites": [
    "what-to-take-with-you"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  }
 ],
 "related": [
  {
   "slug": "two-used-3080s-priced",
   "why": "shares ground with § Writing: how fast each one answers · § How fast does each one write?"
  },
  {
   "slug": "one-3090-against-one-3090-ti",
   "why": "shares ground with § What to take with you · § The two models want different caps"
  },
  {
   "slug": "three-at-the-table",
   "why": "shares ground with § Writing: how fast each one answers · § Measured on our own machines"
  }
 ],
 "thanks": "## Who ran this, and thanks {#who-ran-this-and-thanks}\n\nThe figures are six pages' work and not this one's. **[RTX 3090 vs RTX 3090 Ti](https://research.strata2signal.com/one-3090-against-one-3090-ti/)** supplies most of the ladders; **[Two 3080s against one 3090](https://research.strata2signal.com/two-used-3080s-priced/)** the pair, the single 3080 and the cheap end; **[A Short History of Mistral](https://research.strata2signal.com/a-short-history-of-mistral/)** the one-box reading and the spilled model; **[the rig cost page](https://research.strata2signal.com/home-inference-and-diffusion-rig-cost/)** its 3090 prices; **[What 150 Watts Buys](https://research.strata2signal.com/what-150-watts-buys/)** the 96 GB card's own temperature reading, which appears here named under a table rather than in one, because its cap is a range; and **[A Rig Your Friend Already Owns](https://research.strata2signal.com/a-rig-your-friend-already-owns/)** the sketch ladder, the one table whose rows include the laptop part. **[Two 3090s out, one 3090 Ti in](https://research.strata2signal.com/two-3090s-out-one-3090-ti-in/)** put twenty more records into the hardware roster this page reads from on 2026-09-19 and is thanked for the record rather than for a figure: none of them could fill a cell here, for the reason given above. The two prices this page took for itself were read at Newegg's product pages and in Jawa's published sale history, on 1 October 2026, by plain fetches of public pages.\n\n**EVGA** made the 24 GB boards measured here and one of the 10 GB ones, and published the specifications those pages checked themselves against; **Gigabyte** made the other 10 GB board, the one the single-card row's drawing and heat readings come from; **NVIDIA** published the launch figures, makes the Founders Edition named above, and makes the driver every temperature, clock and watt on this page came out of.\n\nThe writing was done on **ollama**, and under it **llama.cpp** and **NVIDIA's CUDA**; the pictures on **ComfyUI**. The models are Google's **Gemma 4**, **Mistral Small 3.2** from Mistral AI, Black Forest Labs' **FLUX.2 Klein 4B** and its decoder, and **Qwen** from Alibaba Cloud as the text encoder in those graphs. Each source page names the licence it found for the model it ran; this page adds none of its own, because it ran nothing.\n\nA small human team owns the cards, set the rules and the refusals before the runs, and signed the figures on each source page; a fleet of AI agents ran the harnesses, filled these tables from the record, and wrote the check that refuses this page when a cell and the record disagree. Thanks to the readers who kept asking the question this page exists to answer: *how does that one compare?*",
 "kit": null,
 "seat_class": {
  "writer": "gemma-class",
  "judge": null,
  "writer_runtime": "vllm"
 },
 "bench": {
  "bullets_written": 3,
  "bullets_kept": 0,
  "digests_written": 14,
  "digests_kept": 10,
  "chips_written": 8,
  "chips_kept": 8,
  "chips_grounded": 0,
  "chips_published": 0,
  "dropped_by": {
   "new_noun": 0,
   "figure": 0,
   "length": 0,
   "cite": 0,
   "judge": 0,
   "quote": 0,
   "directive": 0,
   "redaction": 3,
   "profanity": 0,
   "identifier": 0,
   "bare_figure": 0,
   "house_voice": 0,
   "hand-edit": 4
  },
  "new_noun_tokens_checked": [
   "1",
   "1.683",
   "10",
   "125.6",
   "1k",
   "24b",
   "250",
   "26b",
   "3080",
   "3090",
   "33.3",
   "350",
   "450",
   "5090",
   "536-token",
   "6.297",
   "6000",
   "65",
   "67",
   "758.13",
   "82",
   "C",
   "GB",
   "GPU",
   "J",
   "Laptop",
   "PRO",
   "RAM",
   "RTX",
   "Ti",
   "W",
   "gemma4",
   "mistral-small3.2"
  ],
  "new_noun_tokens_withheld": 3,
  "figure_definition": "v3",
  "figure_boundary": "guarded",
  "drops": [
   {
    "kind": "digest",
    "reason": "redaction",
    "item": null,
    "item_chars": 214,
    "withheld": "a string carrying a fenced literal",
    "detail": ""
   },
   {
    "kind": "digest",
    "reason": "redaction",
    "item": null,
    "item_chars": 162,
    "withheld": "a string carrying a fenced literal",
    "detail": ""
   },
   {
    "kind": "digest",
    "reason": "redaction",
    "item": null,
    "item_chars": 151,
    "withheld": "a string carrying a fenced literal",
    "detail": ""
   },
   {
    "kind": "bullet",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 27,
    "text_sha": "4071d080a5f5",
    "figure": "82",
    "withheld": "one bullet removed from this pack by hand",
    "removed_utc": "2026-10-01",
    "why": "lost its condition: 82 °C was the 3080 while drawing; the page says a temperature without its workload says little",
    "detail": "removed from this pack by hand on 2026-10-01 (UTC), after the writer ran and after the gates had counted: lost its condition: 82 °C was the 3080 while drawing; the page says a temperature without its workload says little. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "bullet",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 37,
    "text_sha": "5a4995690196",
    "figure": "450",
    "withheld": "one bullet removed from this pack by hand",
    "removed_utc": "2026-10-01",
    "why": "false: at its 450 W limit the 3090 Ti drew 283.01 W writing the mixture model (418.76 W on the dense one); 450 W is the cap, not a draw",
    "detail": "removed from this pack by hand on 2026-10-01 (UTC), after the writer ran and after the gates had counted: false: at its 450 W limit the 3090 Ti drew 283.01 W writing the mixture model (418.76 W on the dense one); 450 W is the cap, not a draw. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "bullet",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 31,
    "text_sha": "d7714e5997c0",
    "figure": "33.3",
    "withheld": "one bullet removed from this pack by hand",
    "removed_utc": "2026-10-01",
    "why": "lost its condition: 33.3 tok/s is the dense model at a 250 W cap and a 65,536-token window; the mixture model wrote at 125.6",
    "detail": "removed from this pack by hand on 2026-10-01 (UTC), after the writer ran and after the gates had counted: lost its condition: 33.3 tok/s is the dense model at a 250 W cap and a 65,536-token window; the mixture model wrote at 125.6. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "digest",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 179,
    "text_sha": "37934b0b7c8b",
    "section": "the-sketch-ladder-and-the-one-row-here-that-is-not-a-card",
    "withheld": "one section digest removed from this pack by hand",
    "removed_utc": "2026-10-01",
    "why": "mislabels the table: the sketch ladder holds a 3090, the 96 GB card and the laptop part, not the laptop part alone",
    "detail": "removed from this pack by hand on 2026-10-01 (UTC), after the writer ran and after the gates had counted: mislabels the table: the sketch ladder holds a 3090, the 96 GB card and the laptop part, not the laptop part alone. Not a gate drop, and recorded here so the ledger adds up."
   }
  ]
 },
 "generated_utc": "2026-10-01T12:56:11Z"
}
