{
 "schema_version": "1.0",
 "page_kind": "article",
 "pack_kind": "generate",
 "slug": "two-used-3080s-priced",
 "title": "Two 3080s against one 3090, and the cap decides",
 "dek": "Two GeForce RTX 3080 10 GB cards bought in 2021, in a desktop of the same vintage, against one GeForce RTX 3090 24 GB in the same box, on the same supply, through the same script. The single card is faster on three of the four things measured and uses about half the energy; the pair reads a long document in two thirds of the time and costs about two fifths as much.",
 "published": "2026-09-17",
 "series": [
  "bench"
 ],
 "licence": "CC BY 4.0",
 "status": "pending-judge",
 "pack_note": "no judge is seated, so every chip is held",
 "notice": "items with published:false are held and are not this page's published words",
 "source": {
  "url": "https://research.strata2signal.com/two-used-3080s-priced/",
  "md_url": "https://research.strata2signal.com/two-used-3080s-priced/index.md",
  "md_sha": "20b8d75965b2f4c03ba21def60cc926da53473eca3e00d78559f966093761f6f",
  "html_sha": "331cf02340adcce5263cf50abde2ee36fb17e81bcfd0a736ea09ad9dc73b3be1",
  "anchors_sha": "50bf4422ebf0f7806da772c3252bcd83dff2f9e2e494a3abaa6d10e63ccd05fd"
 },
 "short": {
  "paragraph": "Two used 10 GB cards are 20 GB of graphics memory between them, and 20 GB is enough to hold a 26-billion-parameter model whole with a 131,072-token context window, writing about 115 tokens a second - faster than you can read. One 3090 does the same job at about 136, holds more context on the other model, and uses roughly half the energy. Out of the box the pair gives you half its speed until you set one number, which no installer tells you about. The pair's one clear win is reading: it takes in a 98,000-token document in 28 seconds against 45. On prices read 17 September 2026, two 10 GB 3080s cost $750.00 against $1,879.99 for a renewed listing of the exact 3090 measured here. Around them, a working desktop is $1,394 to $1,579 - this is not a thousand-dollar machine, and the page shows its arithmetic.",
  "from_dek": false,
  "counts": {
   "words": 7371,
   "minutes": 34,
   "tables": 16,
   "kit": false
  },
  "bullets": [
   {
    "text": "the single 3090 at 300 W reaches a rate of 136.1",
    "figure": "136.1",
    "cite": "how-fast-does-each-one-write",
    "quote": "model two 3080s, 250 W two 3080s, 300 W one 3090, 250 W one 3090, 300 W mixture, at 131,072 114.8 115.4 125.6 136.1 dense, at 49,152 44.6 44.7 33.1 not run - its 300 W dense arm is a single point at 65,536, not a ladder dense, at 65,536 does not fit does not fit 33.3 47.8 The answer changes hands at the cap, and only on the dense model.",
    "span": [
     12680,
     12689
    ]
   },
   {
    "text": "the cheapest working desktop is priced at $1,393.81",
    "figure": "$1,393.81",
    "cite": "what-it-costs-priced-17-september-2026",
    "quote": "the build total as built - 48 GB and the HX1000i $1,578.80 48 GB and the RM1000e $1,483.80 32 GB and the HX1000i $1,488.81 32 GB and the RM1000e $1,393.81 board, processor and the two cards alone $1,080.83 Still to add to every figure in that table: a case, a drive, an operating system, and shipping - roughly $10 to $25 a used card, none of it included above.",
    "span": [
     35375,
     35387
    ]
   }
  ]
 },
 "sections": [
  {
   "id": "at-a-glance",
   "heading": "At a glance",
   "level": 2,
   "span": [
    1581,
    5219
   ],
   "chunks": [
    [
     1581,
     5219
    ]
   ],
   "chars": 3638,
   "digest": "this table compares four rigs across various metrics including tokens a second, window size, and energy. it shows that the single 3090 wins most categories, but the two 3080s at 250 W win the dense model by 34.9 %. the pair is 2.51× cheaper than the single card.",
   "digest_skipped": null
  },
  {
   "id": "the-question",
   "heading": "The question",
   "level": 2,
   "span": [
    5219,
    7695
   ],
   "chunks": [
    [
     5219,
     7695
    ]
   ],
   "chars": 2476,
   "digest": "the article asks if two older 10 GB cards are better than one newer card for running models in 2026. it assumes the reader knows what they are running and focuses on speed, capacity, power, and cost. it discloses that one card ran at PCIe x4.",
   "digest_skipped": null
  },
  {
   "id": "a-few-words-this-page-leans-on",
   "heading": "A few words this page leans on",
   "level": 2,
   "span": [
    7695,
    9817
   ],
   "chunks": [
    [
     7695,
     9817
    ]
   ],
   "chars": 2122,
   "digest": "this section defines technical terms used throughout the article, including token, parameters, context window, VRAM, decode, prefill, layer, dense, mixture of experts, the cap, and the server.",
   "digest_skipped": null
  },
  {
   "id": "the-four-rigs-named-once",
   "heading": "The four rigs, named once",
   "level": 2,
   "span": [
    9817,
    10566
   ],
   "chunks": [
    [
     9817,
     10566
    ]
   ],
   "chars": 749,
   "digest": "the four rigs used in the measurements are defined: two 3080s at 250 W, two 3080s at 300 W, one 3090 at 250 W, and one 3090 at 300 W. the cap is applied per card.",
   "digest_skipped": null
  },
  {
   "id": "the-instrument-and-how-the-control-was-run",
   "heading": "The instrument, and how the control was run",
   "level": 2,
   "span": [
    10566,
    11973
   ],
   "chunks": [
    [
     10566,
     11973
    ]
   ],
   "chars": 1407,
   "digest": "measurements used a frozen 2,101-byte prompt and reported the median of three scored runs. watts were card board power. the models were gemma4:26b and mistral-small3.2:24b. the single card was tested in the same box and slot as the pair.",
   "digest_skipped": null
  },
  {
   "id": "how-fast-does-each-one-write",
   "heading": "How fast does each one write?",
   "level": 2,
   "span": [
    11973,
    13821
   ],
   "chunks": [
    [
     11973,
     13821
    ]
   ],
   "chars": 1848,
   "digest": "the single 3090 wins the mixture and dense models at 300 W. at 250 W, the two 3080s win the dense model by 34.9 %. the single 3090 at 300 W reaches 47.8 tokens a second on a 65,536 window.",
   "digest_skipped": null
  },
  {
   "id": "how-much-can-each-one-hold",
   "heading": "How much can each one hold?",
   "level": 2,
   "span": [
    13821,
    16888
   ],
   "chunks": [
    [
     13821,
     16888
    ]
   ],
   "chars": 3067,
   "digest": "the 3090 holds larger windows whole. naming the layer count for the two 3080s doubled the rate. an unquantised cache on 24 GB holds a 131,072-token window at 133.6 tokens a second, but the pair requires 20,995 MiB for that setting.",
   "digest_skipped": null
  },
  {
   "id": "what-about-a-long-document",
   "heading": "What about a long document?",
   "level": 2,
   "span": [
    16888,
    17976
   ],
   "chunks": [
    [
     16888,
     17976
    ]
   ],
   "chars": 1088,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "what-if-more-than-one-person-is-asking",
   "heading": "What if more than one person is asking?",
   "level": 2,
   "span": [
    17976,
    19302
   ],
   "chunks": [
    [
     17976,
     19302
    ]
   ],
   "chars": 1326,
   "digest": "at 250 W and four conversations, the two 3080s lead by one per cent. at 300 W, the single 3090 leads by 7.8 %. running one 26-billion model split across both cards is 60.3 % faster than two 12-billion models.",
   "digest_skipped": null
  },
  {
   "id": "two-copies-or-one-bigger-model",
   "heading": "Two copies, or one bigger model?",
   "level": 2,
   "span": [
    19302,
    20411
   ],
   "chunks": [
    [
     19302,
     20411
    ]
   ],
   "chars": 1109,
   "digest": "running one 26-billion model split across both cards at 300 W yields 189.5 tokens a second. this is 60.3 % faster than running two copies of a 12-billion model, which yields 118.2 tokens a second.",
   "digest_skipped": null
  },
  {
   "id": "the-full-concurrency-ladder",
   "heading": "The full concurrency ladder",
   "level": 2,
   "span": [
    20411,
    21553
   ],
   "chunks": [
    [
     20411,
     21553
    ]
   ],
   "chars": 1142,
   "digest": "the single card leads at every level of concurrency except for four conversations at 250 W. at 300 W, the single card maintains a 7.8 % gap over the pair.",
   "digest_skipped": null
  },
  {
   "id": "the-slot-trap-if-you-serve-other-people",
   "heading": "The slot trap, if you serve other people",
   "level": 2,
   "span": [
    21553,
    22859
   ],
   "chunks": [
    [
     21553,
     22859
    ]
   ],
   "chars": 1306,
   "digest": "configuring a server for unnecessary concurrency can cause performance tails. on the pair, a 486 ms p95 can become 7,599 ms when using four parallel slots.",
   "digest_skipped": null
  },
  {
   "id": "what-the-cache-type-costs",
   "heading": "What the cache type costs",
   "level": 2,
   "span": [
    22859,
    24144
   ],
   "chunks": [
    [
     22859,
     24144
    ]
   ],
   "chars": 1285,
   "digest": "on the 3090, using an unquantised f16 cache allows 133.6 tokens a second. for the pair, the 8-bit cache is what fits, as the f16 cache results in out of memory at the full window.",
   "digest_skipped": null
  },
  {
   "id": "renders-one-3080-against-one-3090-seconds-per-image",
   "heading": "Renders: one 3080 against one 3090, seconds per image",
   "level": 2,
   "span": [
    24144,
    27918
   ],
   "chunks": [
    [
     24144,
     27918
    ]
   ],
   "chars": 3774,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "power-heat-and-the-fifty-watts",
   "heading": "Power, heat, and the fifty watts",
   "level": 2,
   "span": [
    27918,
    32740
   ],
   "chunks": [
    [
     27918,
     32740
    ]
   ],
   "chars": 4822,
   "digest": "on the pair, fifty watts buy nothing for speed. on the single 3090, fifty watts are worth 43.8 % on the dense model. the pair at 250 W on the mixture is the quietest, with 0 % fan.",
   "digest_skipped": null
  },
  {
   "id": "what-it-costs-priced-17-september-2026",
   "heading": "What it costs, priced 17 September 2026",
   "level": 2,
   "span": [
    32740,
    37303
   ],
   "chunks": [
    [
     32740,
     37303
    ]
   ],
   "chars": 4563,
   "digest": "the cheapest working desktop build is $1,393.81. the two 3080s are 2.51× cheaper than the single 3090. the build includes a motherboard, processor, memory, power supply, and two graphics cards.",
   "digest_skipped": null
  },
  {
   "id": "two-readers-two-answers",
   "heading": "Two readers, two answers",
   "level": 2,
   "span": [
    37303,
    38571
   ],
   "chunks": [
    [
     37303,
     38571
    ]
   ],
   "chars": 1268,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "still-in-the-bank",
   "heading": "Still in the bank",
   "level": 2,
   "span": [
    38571,
    40104
   ],
   "chunks": [
    [
     38571,
     40104
    ]
   ],
   "chars": 1533,
   "digest": "this section names measurements not included in the main text, such as a 350 W reading for a 3090 and a vision model's text path at 95.1 tokens a second.",
   "digest_skipped": null
  },
  {
   "id": "what-this-page-does-not-measure",
   "heading": "What this page does not measure",
   "level": 2,
   "span": [
    40104,
    41393
   ],
   "chunks": [
    [
     40104,
     41393
    ]
   ],
   "chars": 1289,
   "digest": "the article does not measure answer quality, memory-die temperature, wall power, or two 3080s at full slot width. it also does not measure a 350 W figure for either rig.",
   "digest_skipped": null
  },
  {
   "id": "how-to-check-our-work",
   "heading": "How to check our work",
   "level": 2,
   "span": [
    41393,
    42706
   ],
   "chunks": [
    [
     41393,
     42706
    ]
   ],
   "chars": 1313,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "who-ran-this-and-thanks",
   "heading": "Who ran this, and thanks",
   "level": 2,
   "span": [
    42706,
    44277
   ],
   "chunks": [
    [
     42706,
     44277
    ]
   ],
   "chars": 1571,
   "digest": null,
   "digest_skipped": "credits"
  }
 ],
 "chips": [
  {
   "id": "c-58652694",
   "q": "how do different power limits affect model performance?",
   "a": "The impact of power limits depends on the model type. For a single card, reducing the cap from 300 W to 250 W costs about 30 per cent on a dense model, whereas for a pair of cards, the same reduction costs under half a per cent of their speed.",
   "cites": [
    "how-fast-does-each-one-write"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-2c732e0f",
   "q": "how much context can different hardware configurations hold?",
   "a": "A single 24 GB card can hold a 131,072-token mixture model window entirely. In contrast, a pair of 10 GB cards requires the user to manually name the layer count to hold that same window, otherwise the server may spill parts of the model into slower system memory.",
   "cites": [
    "how-much-can-each-one-hold"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-1683f53e",
   "q": "how does serving multiple people affect token rates?",
   "a": "Scaling is not linear; four simultaneous conversations yield about 2.2 times the tokens of a single conversation. At a 250 W cap, four conversations result in roughly 268.0 tokens a second for the pair and 265.2 for the single card.",
   "cites": [
    "the-full-concurrency-ladder"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-b9b53334",
   "q": "how does running one large model compare to two smaller ones?",
   "a": "Running one 26-billion model split across two cards is 60.3 per cent faster than running two separate 12-billion models on each card. The split configuration also draws about a third less power.",
   "cites": [
    "two-copies-or-one-bigger-model"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-3f49ba0e",
   "q": "what happens when configuring high concurrency in a single slot?",
   "a": "Configuring a server for more parallel slots than necessary can cause significant delays. On a pair of cards, a worst-case response time can jump from 486 ms to 7,599 ms because the extra cache pushes the model off the cards.",
   "cites": [
    "the-slot-trap-if-you-serve-other-people"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-3dd3db69",
   "q": "how does cache compression affect token speeds?",
   "a": "Using an unquantised f16 cache allows for the fastest rates but may cause out of memory errors on 20 GB hardware. Compressing the cache to 8 bits allows the window to fit, costing the single card about 6 per cent of its rate.",
   "cites": [
    "what-the-cache-type-costs"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-9e9d8c64",
   "q": "how does power consumption change when a model is idle?",
   "a": "Keeping a model resident in memory while idle costs more power than an empty card. At a 250 W cap, the pair costs 85.7 W to hold a model, while the single card costs 53.1 W.",
   "cites": [
    "power-heat-and-the-fifty-watts"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  }
 ],
 "related": [
  {
   "slug": "home-inference-and-diffusion-rig-cost",
   "why": "shares ground with § What it costs, priced 17 September 2026 · § The bill, on one day"
  },
  {
   "slug": "the-two-hour-machine-priced",
   "why": "shares ground with § How fast does each one write? · § What the computer runs"
  },
  {
   "slug": "two-hours-on-battery",
   "why": "shares ground with § How fast does each one write? · § The one table to remember"
  }
 ],
 "thanks": "## Who ran this, and thanks {#who-ran-this-and-thanks}\n\nThe cards are a Gigabyte GeForce RTX 3080 10 GB, an EVGA GeForce RTX 3080 FTW3 Ultra 10 GB and an EVGA GeForce RTX 3090 XC3 Ultra 24 GB, and all three board partners get the credit for hardware still doing serious work five years after it was sold for games. **Gelid Solutions** made the thermal pads the first board was rebuilt with, in the several thicknesses a re-pad actually needs — memory, power stages and backplate are three different gaps.\n\nThe language work stands on **ollama** and, beneath it, **llama.cpp** and **NVIDIA's CUDA**; the pictures on **ComfyUI** and on **ComfyUI-MultiGPU**, pollockjj's maintained line of city96's loader nodes. The models are Google's **Gemma 4**, **Mistral Small 3.2** from Mistral AI, Black Forest Labs' **FLUX.2 Klein 4B** and its decoder, which drew every picture timed here, **Qwen** from Alibaba Cloud as the text encoder in those graphs, **Chroma1 HD** in the one oversized stack, and **nomic-embed-text** from Nomic AI, whose arm is reported as a failure rather than a figure. None of them owed us anything, and several ask for no credit, which is exactly why it is given.\n\nA small human team owns the cards, rebuilt one of them, set the gates before the runs and signed the numbers; a fleet of AI agents ran the harness, read the listings and did the arithmetic under that team's rulings. Thanks to the readers of the earlier pages in this series who asked the obvious next question — *what if I just buy two of the cheap ones?* — which is the whole of this one.",
 "kit": null,
 "seat_class": {
  "writer": "gemma-class",
  "judge": null,
  "writer_runtime": "vllm"
 },
 "bench": {
  "bullets_written": 3,
  "bullets_kept": 2,
  "digests_written": 19,
  "digests_kept": 16,
  "chips_written": 8,
  "chips_kept": 7,
  "chips_grounded": 0,
  "chips_published": 0,
  "dropped_by": {
   "new_noun": 3,
   "figure": 0,
   "length": 0,
   "cite": 0,
   "judge": 0,
   "quote": 0,
   "directive": 0,
   "redaction": 1,
   "profanity": 0,
   "identifier": 1
  },
  "new_noun_tokens_checked": [
   "0",
   "072-token",
   "1",
   "1.59×",
   "10",
   "101-byte",
   "118.2",
   "12-billion",
   "131",
   "133.6",
   "136.1",
   "18.0",
   "189.5",
   "2",
   "2.2",
   "2.51×",
   "20",
   "2026",
   "24",
   "24b",
   "250",
   "26-billion",
   "265.2",
   "268.0",
   "26b",
   "28.36",
   "30",
   "300",
   "3080s",
   "3090",
   "34.9",
   "350",
   "393.81",
   "43.8",
   "47.8",
   "486",
   "53.1",
   "536",
   "599",
   "6",
   "60.3",
   "65",
   "7",
   "7.8",
   "8",
   "8-bit",
   "85.7",
   "95.1",
   "98",
   "995",
   "GB",
   "MiB",
   "PCIe",
   "VRAM",
   "W",
   "f16",
   "gemma4",
   "mistral-small3.2",
   "p95",
   "x4"
  ],
  "new_noun_tokens_withheld": 3,
  "figure_definition": "v3",
  "drops": [
   {
    "kind": "digest",
    "reason": "new_noun",
    "item": null,
    "item_chars": 9,
    "withheld": "a string naming something the article does not",
    "detail": ""
   },
   {
    "kind": "digest",
    "reason": "new_noun",
    "item": null,
    "item_chars": 6,
    "withheld": "a string naming something the article does not",
    "detail": ""
   },
   {
    "kind": "digest",
    "reason": "redaction",
    "item": null,
    "item_chars": 193,
    "withheld": "a string carrying a fenced literal",
    "detail": ""
   },
   {
    "kind": "bullet",
    "reason": "identifier",
    "item": "the two 3080s at 300 W read a document in 28.36 s",
    "detail": "3080s is in the cited section only as code or in a table cell"
   },
   {
    "kind": "chip",
    "reason": "new_noun",
    "item": null,
    "item_chars": 9,
    "withheld": "a string naming something the article does not",
    "detail": ""
   }
  ]
 },
 "generated_utc": "2026-09-17T19:44:27Z"
}
