The bench — what this page measures and prices, on which day, and what it does not
Two 3080s against one 3090, and the cap decides
exhibit fifty The bench
Published 2026-09-17 (UTC)
updated 2026-09-17 (UTC)
A small (human) team and a fleet of AI agents.
Two GeForce RTX 3080 10 GB cards bought in 2021, in a desktop of the same vintage, against one GeForce RTX 3090 24 GB in the same box, on the same supply, through the same script. The single card is faster on three of the four things measured and uses about half the energy; the pair reads a long document in two thirds of the time and costs about two fifths as much. Which one wins changes with the power cap, and both were measured at two.
ask about this page → assistant.strata2signal.com · in beta, still being tested
the short version
Two used 10 GB cards are 20 GB of graphics memory between them, and 20 GB is enough to hold a 26-billion-parameter model whole with a 131,072-token context window, writing about 115 tokens a second — faster than you can read. One 3090 does the same job at about 136, holds more context on the other model, and uses roughly half the energy. Out of the box the pair gives you half its speed until you set one number, which no installer tells you about. The pair's one clear win is reading: it takes in a 98,000-token document in 28 seconds against 45. On prices read 17 September 2026, two 10 GB 3080s cost $750.00 against $1,879.99 for a renewed listing of the exact 3090 measured here. Around them, a working desktop is $1,394 to $1,579 — this is not a thousand-dollar machine, and the page shows its arithmetic.
- The single 3090 at 300 W reaches a rate of 136.1 (see How fast does each one write?)
- The cheapest working desktop is priced at $1,393.81 (see What it costs, priced 17 September 2026)
7,371 words, about 34 minutes to read.
The summary is this page’s own; the receipt lines were drafted by a model on this workshop’s network and every figure in them is in the article, checked before this page went out — what was dropped, and why, is in this page’s receipt file.
At a glance
All four rigs, because the answer changes with the cap — and the pair's cap is per card, so its "250 W" is 500 W between them. Rows marked ↓ are won by the smaller number. Every figure repeats further down with its window, its cap and its receipt file; this table is the shape of the answer, not the evidence for it. Four rows do not move with the cap — what fits, and what it costs — so one figure is printed across both of a rig's columns. A cell reading not run at this cap is an arm that was never measured there, not a zero and not a guess; every one of them is named again in "What this page does not measure". The render row is one card against one card, because a second card draws a second picture rather than the same one faster.
| two 3080s, 250 W each | two 3080s, 300 W each | one 3090, 250 W | one 3090, 300 W | winner | |
|---|---|---|---|---|---|
| out of the box, before you set anything — mixture at a 4,096 window | not run at this cap | 59.5 · 75.6 % on the cards | 129.4 · all of it on the cards | not run at this cap | one 3090, 2.17× — until the layer count is named, and then the pair doubles to 124.7. Each rig was measured at one cap only here, and they are not the same cap |
| mixture model, tokens a second, at 131,072 | 114.8 | 115.4 | 125.6 | 136.1 | one 3090 at both caps — by 9.4 % at 250 W, 18.0 % at 300 W |
| dense model, tokens a second, at the 49,152 window both rigs hold | 44.6 | 44.7 | 33.1 | not run at this window | two 3080s at 250 W, by 34.9 % — the page's turning point |
| dense model, tokens a second, each at its own largest window | 44.6 (49,152) | 44.7 (49,152) | 33.3 (65,536) | 47.8 (65,536) | one 3090 at 300 W, by 7.0 % — on a window a third larger |
| ↓ reading a 98,000-token document | not run at this cap | 28.36 s | 50.08 s | 45.12 s | two 3080s, 1.59× at 300 W; at 250 W only the single card's arm ran, at 50.08 s |
| ↓ one picture, 8 steps, 512 × 512, three at a time | 2.074 s (one card of the pair) | not run at 300 W | 1.683 s | 1.560 s | one 3090, by 1.23× — the narrowest margin on this page |
| largest window held whole, mixture model | 131,072 — only when the layer count is named | 131,072 — only when the layer count is named | 131,072 | 131,072 | a tie the pair reaches by hand: left alone, the server never got there at all |
| largest window held whole, dense model | 49,152 | 49,152 | 98,304 | 98,304 | one 3090, exactly twice the window |
| mean draw, dense model | 466 W (two cards) | 509 W (two cards) | 235 W | 290 W | one 3090 — about half the draw at either cap |
| ↓ energy, dense model, joules per 1,000 tokens | 10,442 | 11,385 | 7,076 | 6,059 | one 3090 — 32 % less at 250 W, 47 % less at 300 W |
| ↓ the two cards, priced 17 September 2026 | $750.00 (two, any make) | $750.00 | $1,879.99 (the exact board, renewed) | $1,879.99 | two 3080s — the cards are 2.51× cheaper |
| ↓ a working desktop around them, the same day's listings | $1,393.81 to $1,578.80 | $1,393.81 to $1,578.80 | no build is totalled around the single card here | the same — a price does not move with a cap | two 3080s — and even so, not a thousand-dollar machine |
Read the winner column and then the cap. The single card takes most of this table, and it takes it at 300 W. At 250 W — the setting this page recommends for the pair, where it is quieter and barely slower — the dense row turns over completely and the pair takes it by 34.9 per cent. That is the whole argument of this page and it has its own section below.
The question
Two graphics cards of the same model, bought new in 2021 for gaming, in a desktop of the same age. Each holds 10 GB, which is not enough for the models a person wants to run at home in 2026. Together they hold 20 GB, which is. So: is two of yesterday's cards a way to reach today's models, or should the money go to one better card?
What you would be running, since the page assumes you know and should not. The two models here are free to download and run entirely on a machine you own — nothing leaves the house, there is no account and there is no monthly bill. They answer questions, write, summarise and read documents you hand them. On hard reasoning they are behind the best paid services, and this page does not measure answer quality at all — only how fast this hardware runs them, what it holds, what it draws and what it costs.
This page is the middle rung of a three-tier series — below it a mini PC with no graphics card, above it one and two current-generation cards. It is written for two people who get different answers at the end: someone who already owns two older cards, and someone shopping with about a thousand dollars.
One disclosure travels with every number here. The two cards did not negotiate the same link: one runs at PCIe x16, the other at x4 of an x16-capable slot, because a consumer board's second full-length slot commonly drops to x4. A model split across two cards moves data between them on every token, so this sits in the path of the measurement rather than beside it. These figures are what a second card in a second slot should be expected to do — not a best case. Nothing here may be quoted for two 3080s both at full width, which was not measured.
Before you buy a second card, check three things on the desktop you own: a second full-length slot, two spare 8-pin power cables, and about 300 mm of clearance. One board here is a Gigabyte that shipped with thermal pads missing — no backplate contact, none on some memory modules — and was taken apart and re-padded by hand with Gelid Solutions pads before any of this ran. You do not need to do that to buy a used 3080. It is disclosed because it is one of the two differences between these two cards; the other is the stock EVGA FTW3 Ultra, which is the one in the x4 slot.
A few words this page leans on
- Token — the chunk a model reads and writes in, a little under a word. Around ten a second reads as fluent, so 115 is far faster than you can read.
- Parameters — how big a model is, counted in billions of adjustable numbers. A 26-billion model is bigger than a 12-billion one; bigger is usually better and always slower, and the whole game is fitting one into the memory you have.
- Context window — how much the model holds in mind at once, in tokens. 131,072 is a few hundred pages, and reserving that much costs memory whether you fill it or not.
- VRAM — the memory on the card itself: 10 GB each here, 20 GB for the pair, which the driver counts as 20,480 MiB. A model that fits runs on the cards alone; one that does not is split, the rest in the computer's own memory, and every word waits on the slower half.
- Decode, and prefill — decode is the model writing, one token at a time, and it waits on how fast one card can stream its own memory. Prefill is it reading what you handed it, all at once, and that is arithmetic two cards can share. That difference is why two cards help reading and not writing.
- Layer, and layer count — a model is a stack of layers, and each can sit on a card or in the computer's memory. The layer count is the number you hand the server to say put all of them on the cards.
- Dense, and mixture of experts — a dense model uses all of itself for every word; a mixture wakes only a few billion parameters per word. That is how the 26-billion mixture here outruns the 24-billion dense model by nearly three times.
- The cap — a limit you set on how many watts a card may draw; one command. Turning these two 3080s down from 300 W to 250 W costs under half a per cent of their speed. Turning a 3090 down from 300 W to 250 W costs it about 30 per cent on the dense model. That difference is most of this page's argument.
- The server — the free program that loads a model and answers requests. Everything here ran on ollama; the exact version is in the instrument section.
The four rigs, named once
Every table below uses these four labels, in this order, and nothing else.
| label | what it is |
|---|---|
| two 3080s, 250 W | two GeForce RTX 3080 10 GB, capped at 250 W each — 500 W between them |
| two 3080s, 300 W | the same pair at 300 W each — 600 W between them |
| one 3090, 250 W | one GeForce RTX 3090 24 GB, capped at 250 W |
| one 3090, 300 W | the same card at 300 W. Its own stock limit is 350 W, so neither rung is the setting it ships with |
The cap is set per card. A "250 W" row gives the single card a 250 W budget and the pair a 500 W one; the draw column appears in every power table because that is the comparison that tells the truth.
The instrument, and how the control was run
Every language figure comes from the harness the rest of this series uses: one frozen 2,101-byte prompt — 536 tokens as the mixture counts them, 1,023 as the dense model does — 256 tokens out, temperature and seed at zero. One warm-up run is discarded, three scored runs follow, the median is reported and all three are kept in the receipts. Watts are card board power read from the driver on both cards during every scored run, never the wall. The models are gemma4:26b (26 billion parameters, about 4 billion active per token) and mistral-small3.2:24b (24 billion dense), both at Q4_K_M, served by ollama 0.32.13 with an 8-bit cache. The pictures were drawn on ComfyUI 0.21.1 with each card's cap read off the hardware rather than asserted.
The control. The single card was put in the same box, the same slot and the same supply as the pair, and run through the same script — not quoted from another machine. The two sittings are both stamped 2026-09-17 UTC: the 250 W arms 03:15 to 03:52, the 300 W arms 14:55 to 15:04 the same day. The one cross-box check this workshop can make says the box matters little for the big arms and a great deal for small ones: the same card model measured here and on another machine reads +0.4 per cent on the mixture and −4.1 per cent on the dense model.
How fast does each one write?
Decode rate, tokens a second, median of three scored runs. Both models, both rigs, both caps, each at the largest window that rig holds whole. Same box, same script, 2026-09-17. Receipts: m5-split-250w.json, the 131,072 rung of m5-split-forced-ladder.json, m4-split-250w.json, the 49,152 rung of m4-split-forced-ladder.json, the 131,072 and 49,152 rungs of 3090-m5-one-auto-ladder.json and 3090-m4-one-auto-ladder.json, 3090-m5-one-300w.json, 3090-m4-one-300w.json.
| model | two 3080s, 250 W | two 3080s, 300 W | one 3090, 250 W | one 3090, 300 W |
|---|---|---|---|---|
| mixture, at 131,072 | 114.8 | 115.4 | 125.6 | 136.1 |
| dense, at 49,152 | 44.6 | 44.7 | 33.1 | not run — its 300 W dense arm is a single point at 65,536, not a ladder |
| dense, at 65,536 | does not fit | does not fit | 33.3 | 47.8 |
The answer changes hands at the cap, and only on the dense model. At 300 W the single card takes both: the mixture by 18.0 per cent and the dense model by 7.0. At 250 W it still takes the mixture by 9.4 per cent — and loses the dense model by 34.9 per cent, judged at the same 49,152 window on both sides. Fifty watts are worth 43.8 per cent to one 3090 on that model and 0.2 per cent to the pair, which is the entire mechanism: the pair is limited by how fast one card can read its own memory, and a single 3090 at 250 W is limited by the cap itself.
What the single card does at its own 350 W was not measured here. The rung was pre-registered on both benches; the trigger for running it was a cap fraction above 0.50, the fraction read 1.00, and an operator axed the rung anyway. This workshop has published a 350 W reading from another machine and it is named in "Still in the bank" rather than argued from here.
How much can each one hold?
Largest context window each rig holds entirely on the cards, and what it costs to get there. Same box, same script, 2026-09-17. The pair's dense model was never run as a ladder under the planner's own placement, so its cell is the residency the planner reached at the one window that arm used, from g2-doorman.json, and not a ceiling. Receipts: m5-split-forced-ladder.json, m5-split-auto-ladder.json, m5-split-auto-at4096.json, m4-split-forced-ladder.json, 3090-m5-one-auto-ladder.json, 3090-m4-one-auto-ladder.json, 3090-m4-one-forced-at98304.json.
| what was asked of the server | two 3080s (20,480 MiB) | one 3090 (24,576 MiB) |
|---|---|---|
| mixture, the server's own placement | never held whole at any window tested | 131,072 whole |
| mixture, the layer count named | 131,072 whole, 19,531 MiB used | 131,072 whole, 20,129 MiB used |
| dense, the server's own placement | never measured as a ladder; at a 49,152 window the planner kept 73.7 % on the cards | 65,536 whole |
| dense, the layer count named | 49,152 whole, 18,909 MiB used | 98,304 whole, 23,469 MiB used |
| mixture at a 4,096 window, out of the box | 59.5 tok/s, 75.6 % on the cards | 129.4 tok/s, 100 % on the cards |
This is the finding a buyer needs more than any speed on this page. Left to itself on two cards, the server put three quarters of the model on them, spilled the rest to the computer's memory, and left about 5 GB of the pair's 20 unused while doing it. Naming the layer count put the same model entirely on the cards and doubled the rate. A reader who installs the software, puts two 3080s in a desktop and asks for a 26-billion model gets the slow row, and nothing warns them. On 24 GB the problem does not arise: the server gets it right by itself, and the pair's headline figures are hand-set where the single card's are not.
What to actually do. The request that loads the model takes an option for how many layers go on the cards; set it to a number at least as large as the model has layers — this bench sent 99, and anything above the model's own count means all of them. Restart the server when you change it: changing it while the model is loaded makes the computer ask for a second copy before it has let go of the first, which is two 16 GiB copies in 20 GB. The load fails with an out-of-memory error.
And what four more gigabytes buy, as a number rather than as "headroom". With an unquantised cache the mixture holds its whole 131,072-token window on 24 GB and writes at 133.6 tokens a second at the lower cap — faster than anything the pair does anywhere. It needs 20,995 MiB, which is 515 more than two 10 GB cards have between them, so on the pair that setting refuses to load at that window at all and holds 65,536. The pair's 8-bit cache is not a preference; it is what fits, and it costs about 6 per cent of the rate (3090-m5-one-kvf16-at131072.json, m5-split-kvf16-at131072.json, m5-split-kvf16-at65536.json).
What about a long document?
One 131,072-token window with about 98,000 tokens really in it, read cold. Both rigs at a 300 W cap, same box, same script, 2026-09-17. Receipts: m5-split-filled.json, 3090-m5-one-300w-filled.json.
| two 3080s, 300 W | one 3090, 300 W | |
|---|---|---|
| reading 98,024 tokens | 28.36 s | 45.12 s |
| first token arrives | 30.04 s | 46.64 s |
| writing, once the window is full | 61.5 tok/s | 72.8 |
| energy at depth, joules per 1,000 tokens | 5,837 | 3,902 |
This is the pair's one clear win, and it is a real one: 1.59× on the read. Reading is arithmetic done all at once, so two cards bring twice the silicon to it; writing waits on one card's memory bandwidth, which two cards do not add up. A person who pastes whole documents and a person who chats are buying different machines. One caveat the bench states and this page carries: the filler is highly repetitive text built from the frozen prompt, and a real document of the same length may read faster or slower — that was not measured.
What if more than one person is asking?
The mixture at a 4,096 window, tokens a second summed across simultaneous conversations. Same box, same script, 2026-09-17. Receipts: concurrency-split-250w.json, concurrency-split-forced.json, 3090-concurrency-one.json, 3090-concurrency-one-300w.json.
| conversations at once | two 3080s, 250 W | two 3080s, 300 W | one 3090, 250 W | one 3090, 300 W |
|---|---|---|---|---|
| 1 | 124.4 | 124.0 | 126.7 | 134.8 |
| 2 | 188.9 | 189.5 | 195.5 | 210.4 |
| 4 | 268.0 | 267.6 | 265.2 | 288.4 |
At 250 W four conversations get the same total either way — 268.0 against 265.2, one per cent apart. At 300 W the single card takes it by 7.8 per cent. Each conversation falls to about 67 tokens a second at four, on both.
The two-card design question, answered, and the answer is not the one two cards invite. Two copies of a 12-billion model, one per card, serve two conversations at 118.2 tokens a second between them; one 26-billion model split across both serves the same two at 189.5 — the split wins by 60.3 per cent while drawing less power. A bigger copy does not help: a 14-billion model does not fit a 10 GB card once its window is reserved, at 92.1 per cent resident (g1-two-seats.json).
Two copies, or one bigger model?
Added 2026-09-17 (UTC). Two cards invite a question one card cannot ask: run one big model across both, or a smaller model on each so two people can be served at once? Measured 2026-09-17 01:17 to 01:22 UTC on the pair at a 300 W cap, both arrangements serving two callers. This is the arm the page previously named in passing and did not print.
| serving two callers on the pair, 300 W | total tokens a second | each caller | mean draw |
|---|---|---|---|
| one 26-billion model split across both cards | 189.5 | 94.8 | 356 W |
| two copies of a 12-billion model, one per card | 118.2 | 59.1 | 542 W |
The split wins by 60.3 per cent, and draws about a third less power doing it. A bigger copy is not available as an escape: a 14-billion model does not fit a 10 GB card once its window is reserved, at 92.1 per cent resident. So two cards are better used as one large machine than as two small ones, which is the opposite of what a second card looks like it offers (g1-two-seats.json, concurrency-split-forced.json).
The full concurrency ladder
Added 2026-09-17 (UTC). The page already gave four callers; this is every level for all four rigs, so the shape is visible rather than the endpoints. The mixture at a 4,096 window, tokens a second summed across simultaneous conversations, measured 2026-09-17 in the windows "How to check our work" dates.
| conversations at once | two 3080s, 250 W | two 3080s, 300 W | one 3090, 250 W | one 3090, 300 W |
|---|---|---|---|---|
| 1 | 124.4 | 124.0 | 126.7 | 134.8 |
| 2 | 188.9 | 189.5 | 195.5 | 210.4 |
| 4 | 268.0 | 267.6 | 265.2 | 288.4 |
| per conversation at 4 | 67.0 | 66.9 | 66.3 | 72.1 |
Neither rig scales anywhere near linearly: four conversations get about 2.2 times one conversation's tokens, not four. The single card leads at every level except one — four conversations at 250 W, where the pair is ahead by one per cent — and the gap it opens at 300 W, 7.8 per cent, is the same gap it has at one caller (concurrency-split-250w.json, concurrency-split-forced.json, 3090-concurrency-one.json, 3090-concurrency-one-300w.json).
The slot trap, if you serve other people
Added 2026-09-17 (UTC). A short gate-shaped call — eight tokens out, the kind a service makes to check something before doing work — on the dense model under the server's own placement. Measured 2026-09-17 on both rigs. This is the worst number either bench produced and it is a configuration mistake, not a hardware limit.
| the server was given | model on the cards | 1 caller, median / p95 | 4 callers, median / p95 |
|---|---|---|---|
| two 3080s, 1 parallel slot | 73.7 % | 450 ms / 486 ms | 782 ms / 1,098 ms |
| two 3080s, 4 parallel slots | 41.7 % | 709 ms / 712 ms | 1,076 ms / 7,599 ms |
| one 3090, 4 parallel slots | 56.2 % | 594 ms / 683 ms | 1,179 ms / 6,011 ms |
Configuring a server for concurrency it does not need is what produced both of those tails. Four slots reserve four windows' worth of cache, that cache pushes the model off the cards, and the tail follows it: on the pair, a 486 ms worst case becomes 7.6 seconds. The single card collapses the same way — 6.0 seconds — so this is not a thing two cards do and one does not. Set the slot count to the number of people who will actually ask at once (g2-doorman.json, g2-doorman-parallel4.json, 3090-g2-doorman.json).
What the cache type costs
Added 2026-09-17 (UTC). The page says the pair's 8-bit cache "is what fits" and gives one comparison; this is the whole ladder on both rigs, the mixture at its full 131,072-token window, measured 2026-09-17.
| cache type | two 3080s (20,480 MiB) | one 3090 (24,576 MiB) |
|---|---|---|
| f16, unquantised | out of memory — holds 65,536 instead, at 132.6 tok/s | 133.6 tok/s, 20,995 MiB |
| q8_0, the setting both rigs are measured at | 115.4 tok/s, 19,531 MiB | 125.6 tok/s, 20,129 MiB |
| q4_0 | 113.8 tok/s, 18,815 MiB | 124.7 tok/s, 19,415 MiB |
The fastest setting is the one that stores the conversation uncompressed, and on 20 GB it does not fit at the full window — it needs 20,995 MiB, which is 515 more than two 10 GB cards have between them. Compressing it to 8 bits costs the single card about 6 per cent of its rate and buys the pair the window outright. Going further, to 4 bits, buys nothing on either rig: a little slower, and the window was already there (m5-split-kvf16-at131072.json, m5-split-kvf16-at65536.json, m5-split-forced-ladder.json, m5-split-kvq4_0-at131072.json, 3090-m5-one-kvf16-at131072.json, 3090-m5-one-auto-ladder.json, 3090-m5-one-kvq4_0-at131072.json).
Renders: one 3080 against one 3090, seconds per image
The print lab's three certified graphs at batch 3, seconds per finished image, median. Same box, same slot, same card as every language table above, and the same three graphs by hash (79df5a8b…, 7dd85f01…, 94bfd633… on both sides). The two 3080 columns are the pair's own cards, one each, in the slots they sit in, at 250 W (r1-8st.json, r1-rest.json); the single-card columns are the same-box run of 17 September 2026, 17:11 and 17:24 UTC (r1-300w.json, r1-250w.json, in that run's own results directory). The ratio matches 250 W to 250 W, the cap both rigs' render arms ran at.
| graph, batch 3 | one 3080, x16 slot, 250 W | one 3080, x4 slot, 250 W | one 3090, 250 W | one 3090, 300 W | matched at 250 W |
|---|---|---|---|---|---|
| klein-4b-8st | 2.074 s | 2.639 s | 1.683 s | 1.560 s | 1.23× |
| klein-4b-8st-pixel4 | 2.067 s | 2.643 s | 1.721 s | 1.579 s | 1.20× |
| klein-4b-32st | 7.278 s | 7.821 s | 6.297 s | 5.827 s | 1.16× |
The single card wins every graph, and by less than anything else on this page. Twenty-three per cent on the short graphs and sixteen on the long one — against 18.0 per cent on the mixture and a doubling out of the box. For a card holding more than twice the memory, that is a small margin, and the reason is worth a paragraph.
It is a residency result, not a shader one. The three graphs load 16.13 GB of weights between them — a diffusion model, a text encoder and a decoder — against a card with 10 GB. So a 10 GB card pays to move models in and out on every render, and a 24 GB card does not. Three things in the traces say that is what is being measured rather than raw drawing speed. The gap is smallest on the 32-step graph, where the drawing takes longest and the movement is amortised over more work. The 3080 ran a higher clock and was still slower: 1,860 MHz median against 1,575 on the card that beat it. And on this pair, moving just the part that reads your words onto the second 3080 cut a picture from 3.462 s to 2.261 s — 1.53×, the same bytes and the same picture (r3a.json) — which is more than the whole generational gap above. What a 10 GB card lacks on this work is room, not silicon.
A second card does not make one picture arrive sooner — the renderer runs a picture's steps one after another on a single device. What two cards buy is two pictures at once, 24.96 images a minute against 16.37 for one of them alone, 1.52× — and that is still fewer than the 26.87 a minute one 3090 sustained over ten minutes on the same graph, at half the power. On drawing, two 3080s do not overtake one 3090 by any measure on this page. Asking for three pictures at a time does nearly erase the narrow slot, from 1.99 times the x16 board's seconds down to 1.27 — on two submissions each, which is thin and is not leant on.
Ten minutes of continuous drawing at 250 W, both same-box, same graph:
| one 3080, x16 slot | one 3090 | |
|---|---|---|
| images drawn | 158 | 269 |
| seconds an image, median | 3.737 | 2.156 |
| core temperature plateau | 77 °C at 5.6 min | 65 °C at 0.5 min |
| peak | 82 °C | 67 °C |
| fan, maximum | 92 % | 58 % |
| thermal slowdown reported | none | none |
That is not a cooler comparison and should not be read as one. The two cards did different amounts of work — 269 images against 158 — and the single card began the run 11 °C hotter, at 61 °C against 50. What the table does support is that neither card came near its own slowdown point, and that the faster card was also the calmer one while doing seventy per cent more work (each run's own r4.json).
Power, heat, and the fifty watts
Both models at each rig's headline window, both caps. Same box, same script, 2026-09-17; the cap is the only setting that differs between a rig's two rows, though the two sets were loaded forty minutes apart. Temperatures and fans are per card, the re-padded Gigabyte board first and the stock EVGA board second. Receipts: m5-split-250w.json, the 131,072 rung of m5-split-forced-ladder.json, m4-split-250w.json, the 49,152 rung of m4-split-forced-ladder.json, 3090-m5-one-auto-ladder.json, 3090-m5-one-300w.json, 3090-m4-one-auto-ladder.json, 3090-m4-one-300w.json — the single card's temperature and fan columns are the maxima across that arm's scored runs, from the same files. The single card's two 250 W arms ran hotter than its 300 W ones (63 °C against 47 and 57) because they are ladder arms of eight and six rungs against single points, so they ran far longer; that is a difference in how long the card was working, not in how hard. Its fan did not turn at all on either 300 W arm.
| rig and model | tok/s | mean draw | joules per 1,000 tokens | core temp max | fan max |
|---|---|---|---|---|---|
| two 3080s, 250 W, mixture | 114.8 | 290 W | 2,531 | 46 · 42 °C | 0 % · 0 % |
| two 3080s, 300 W, mixture | 115.4 | 346 W | 2,994 | 59 · 58 °C | 64 % · 0 % |
| two 3080s, 250 W, dense | 44.6 | 466 W | 10,442 | 58 · 52 °C | 62 % · 0 % |
| two 3080s, 300 W, dense | 44.7 | 509 W | 11,385 | 69 · 58 °C | 74 % · 0 % |
| one 3090, 250 W, mixture | 125.6 | 230 W | 1,832 | 63 °C | 42 % |
| one 3090, 300 W, mixture | 136.1 | 242 W | 1,772 | 47 °C | 0 % |
| one 3090, 250 W, dense | 33.3 | 235 W | 7,076 | 63 °C | 51 % |
| one 3090, 300 W, dense | 47.8 | 290 W | 6,059 | 57 °C | 0 % |
On the pair, fifty watts buy nothing. 0.5 per cent on the mixture and 0.2 on the dense model, for 18 and 9 per cent more energy per thousand tokens. On the single card the same fifty watts are worth 43.8 per cent on the dense model and 8.4 on the mixture. That asymmetry is why the dense row turns over between the two caps, and it is why this page recommends 250 W on the pair and would not recommend it on a 3090.
The quietest thing here is the pair at 250 W on the mixture: every scored run, both cards, neither fan turned at all, at 40 to 46 °C, writing 114.8 tokens a second. Two five-year-old gaming cards doing that in silence is the most surprising measurement on the page. Under the dense model the re-padded board reaches 69 °C and 74 per cent fan; ten minutes of continuous drawing, begun from a warm idle at 50 °C, took it to a 77 °C plateau and an 82 °C peak at 92 per cent fan with no thermal slowdown reported (r4.json). These cards' own specification puts slowdown at 95 °C.
Added 2026-09-17 (UTC): what each rig costs to leave switched on. A reader asked what the difference is for a machine that sits there all day with a model loaded and nobody asking it anything — the case for a home server, rather than the case for a bench. This block counts only that: card board power with the model resident and nothing generating, at a 250 W cap, measured in the window 2026-09-17 03:15 to 03:52 UTC for the single card and 01:43 to 01:50 for the pair. Nothing above it changes; those figures are the same arms read for a different question.
| at 250 W, nothing generating | two 3080s | one 3090 |
|---|---|---|
| cards empty | 44.0 W | 25.4 W |
| the model resident | 85.7 W | 53.1 W |
| what the model itself costs to hold | 41.7 W | 27.7 W |
The pair costs 32.6 W more than the single card to sit there holding the same model. Left on continuously that is 0.78 kWh a day and 286 kWh a year, which at 18.34 cents a kilowatt-hour — the United States residential average for June 2026 as the Energy Information Administration publishes it, the same rate this series uses elsewhere — is about $52 a year, or $4.37 a month. That is the honest scale of it: real, and smaller than one month of the price difference between the two rigs. It is card board power only, as everywhere on this page, so a wall meter would read more for both (idle-gemma4-250w.json on both benches).
On energy the pair is simply behind, at every cap and on both models. Stated as the single card's saving, it uses 40.8 per cent less per thousand tokens on the mixture at 300 W and 46.8 per cent less on the dense model; stated the other way round, the pair spends 69.0 and 87.9 per cent more for the same thousand tokens. Two old cards buy capability, not efficiency. The size of that in money was not metered and is not estimated here: over an hour of steady writing the difference is small change, and the single card's real advantages are heat, noise and one slot rather than a bill.
What it costs, priced 17 September 2026
This section prices parts from public listings read between 15:00 and 15:30 UTC on 17 September 2026, in United States dollars, before tax and before shipping except where a listing charged it. It prices a motherboard, a processor, memory, a power supply and two graphics cards — and nothing else: no drive, no case, no screen, no operating system. No earlier reading of this rig exists, so there is nothing to compare it against yet.
Both cards in this box were bought in 2021 and have no price on record, so the pair is priced from the day's asks for the model, and the single card from the day's ask for the exact board it is. Why those two are not the same kind of figure is set out under the table.
| line | the listing, as titled | price |
|---|---|---|
| motherboard | Newegg, "Refurbished MSI PRO Z490-A PRO LGA 1200 Intel Z490 ATX Intel Motherboard", in stock | $92.83 |
| processor | Newegg, "Intel Core i9-10850K Comet Lake 10-Core (CM8070104608302)", new, in stock | $238.00 |
| two used GeForce RTX 3080 10 GB, any make | Jawa, "Dell RTX 3080 (10GB) OEM", Used · Like New, 19 available, $375.00 each | $750.00 |
| memory, 48 GB as built | Newegg, "Rimlance 16GB (2X8GB) … DDR4 3200MHZ" $89.99 + "KingBank … DDR4 32GB (2 x 16GB) 3200MHz CL16" $172.99 | $262.98 |
| memory, 32 GB instead | the KingBank 2 × 16 GB kit alone | $172.99 |
| power supply as built | "Corsair HX1000i … 80 Plus Platinum", Newegg — the one part with a price on record, bought for $234.99 in September 2026 and reading the same on 17 September | $234.99 |
| power supply, lower tier | Newegg, "CORSAIR RMe Series RM1000e 1000 W ATX 3.1 Compatible Cybenetics Gold Full Modular Power Supply" | $139.99 |
The cards are priced by model, not by board, because that is what the bench measured and what the marketplace sells. Every figure on this page is a property of the silicon in a 10 GB RTX 3080; the two particular boards in this box differ from each other in cooler and in factory limit, and one of them differs in having been re-padded by hand, all of which is disclosed above. The two boards measured, as listed that day — not the price of the rig: the EVGA FTW3 Ultra asked $450.00 plus $12.00 shipping on Jawa, and the Gigabyte $620.00 from a marketplace seller shipping from overseas, the cheaper Gigabyte listings that day being out of stock.
| the build | total |
|---|---|
| as built — 48 GB and the HX1000i | $1,578.80 |
| 48 GB and the RM1000e | $1,483.80 |
| 32 GB and the HX1000i | $1,488.81 |
| 32 GB and the RM1000e | $1,393.81 |
| board, processor and the two cards alone | $1,080.83 |
Still to add to every figure in that table: a case, a drive, an operating system, and shipping — roughly $10 to $25 a used card, none of it included above. This is not a thousand-dollar machine. A thousand dollars buys the two cards, the board and the processor — $1,080.83 — and nothing else; a working desktop built entirely from these listings is about $1,394 at its cheapest. The tier below is a mini PC with no graphics card at $959.00, priced 10 September 2026; the tier above is one and two current-generation cards at $3,198.86 and $5,048.85, priced 13 September 2026 — that page carries a later addendum dated 2026-09-15, about training rather than prices, which moves no figure quoted here.
And the card these were measured against, priced the same way and the same day. The exact board: $1,879.99, "EVGA 24G-P5-3975-KR GeForce RTX 3090 XC3 Ultra Gaming, 24GB GDDR6X, iCX3 Cooling, ARGB LED, Metal Backplate (Renewed)", at Amazon's Renewed store, read 17 September 2026 in the same half hour, replacing an earlier figure this workshop published on 13 September that an operator ruled inaccurate and struck. It was read by hand at the store page, because that retailer serves no price to this workshop's fetcher. Against the two 10 GB 3080s this rig is priced with, it is 2.51 times as much. Supporting figure rather than a price: completed sales of 10 GB 3080s over the twelve months to 17 September 2026 averaged $358.00 each across 318 sales, and the day's asks sat just above that, from $375.00.
Two things this page could not read, said rather than estimated: no eBay completed sale and no r/hardwareswap ask — both refused every request made of them.
Two readers, two answers
If you already own two older 10 GB cards: keep them, put them both in, and set the layer count. You have a machine that holds a 26-billion model at a 131,072-token window and writes faster than you read, for the price of a slot and a power supply. And the two things that matter most on this page cost you nothing to do: name the layer count, and cap both cards at 250 W, where they are quieter, barely slower and cheaper to run.
If you are shopping with about a thousand dollars: a thousand dollars does not build this machine, and the page's own totals say so — the cheapest combination here is $1,393.81 before a case, a drive or an operating system. What a thousand dollars does buy is the two cards, the board and the processor, at $1,080.83 — so this rig makes sense if you already own a desktop, a case or a supply to put the cards into, and not if you are buying every part new. The one better card is $1,879.99 for a renewed listing, 88 per cent over that budget, for 18.0 per cent more on the mixture and 7.0 on the dense model. Buy the pair on price and on long documents; buy the one card for quiet, heat, one slot and 24 GB that never splits — and only if the budget reaches it.
Still in the bank
Measured, receipted, and not on this page: a 350 W reading for a 3090 taken on another machine, which this workshop has published and which would put a stock single card further ahead on the dense model than the 300 W rows show; a short gate-shaped call whose tail opens from a 486 ms p95 at one caller to 7,599 ms at four, purely because a four-slot server's cache pushes the model off the cards (g2-doorman.json, g2-doorman-parallel4.json); a vision model's text path at 95.1 tokens a second, which is not an image-reading figure and no image was sent (g3-one-x16-g3-vision.json); the resident cost of leaving a model loaded and idle, 43 W on the pair against 28 W on the single card; and an embedding arm that spread 18.5 per cent against a 15 per cent gate set before it ran, failed, and gets no number (embed-one-x16.json).
Added 2026-09-17 (UTC): two of those came out of the bank the same day. This paragraph was written when the page carried neither. The gate-shaped call now has its own section above, with the single card's rows beside the pair's, and the resident-idle figures are worked through in the power section as a cost per year. The sentence above is left as it was written rather than edited, so the list still reads as the list this page shipped with; what has moved is that two of its five items are now measured on the page instead of pointed at. The 350 W reading, the vision text path and the failed embedding arm are still only named here.
What this page does not measure
- Answer quality. Not once. Every figure here is speed, capacity, power or price.
- No memory-die temperature, at all. The driver does not expose the memory sensor on either 3080 board — asked three independent ways before anything ran, refused three times. The re-padded board's memory heat, which is the one thing a re-pad is for, is a question this bench cannot answer; the 77 °C plateau is a core reading.
- No wall power, and none is estimated. Every watt here is card board power from the driver. The only wall-side instrument in the room reported a figure that cannot be right, so nothing from it is printed and there is no monthly-dollars table.
- Nothing about two 3080s both at full slot width, a 12 GB RTX 3080, any card not measured here, or whether a used card a reader buys performs like these two.
- Nothing about two cards drawing one picture faster. The renders compare one card against one card, because a second card draws a second picture rather than the same one sooner; the two-at-once figure is the only pair render number here.
- No 350 W figure measured here, for either rig. The rung was pre-registered on both benches and axed by an operator's ruling before it ran.
How to check our work
Three benches stand behind this page, and all three keep their pre-registration, their dated amendments, their result files and their verdict in this workshop's own repository. All three are stamped 2026-09-17 UTC: the pair's language arms ran 00:33 to 01:51, the renders 01:53 to 02:41, the single card's 250 W arms 03:15 to 03:52, its 300 W arms 14:55 to 15:04, and its render arms 17:11 to 17:35, all the same day. On disk those result files carry a short box-class prefix, naming the box they ran on; this page drops that prefix and changes nothing else about the names. Two of the pair's own result files are also crossed against the card each one measured, which the bench found by watching which card's memory rose when a model loaded; the rows above are labelled from that measurement rather than from the file name.
The prices anyone can check: open the retailer or the marketplace, search the listing title exactly as printed, and read what it says on the day you do. If a figure does not reproduce for the half hour named above, say so at the contact desk, where a person reads every message. This page keeps its receipts the way every page here does.
Who ran this, and thanks
The cards are a Gigabyte GeForce RTX 3080 10 GB, an EVGA GeForce RTX 3080 FTW3 Ultra 10 GB and an EVGA GeForce RTX 3090 XC3 Ultra 24 GB, and all three board partners get the credit for hardware still doing serious work five years after it was sold for games. Gelid Solutions made the thermal pads the first board was rebuilt with, in the several thicknesses a re-pad actually needs — memory, power stages and backplate are three different gaps.
The language work stands on ollama and, beneath it, llama.cpp and NVIDIA's CUDA; the pictures on ComfyUI and on ComfyUI-MultiGPU, pollockjj's maintained line of city96's loader nodes. The models are Google's Gemma 4, Mistral Small 3.2 from Mistral AI, Black Forest Labs' FLUX.2 Klein 4B and its decoder, which drew every picture timed here, Qwen from Alibaba Cloud as the text encoder in those graphs, Chroma1 HD in the one oversized stack, and nomic-embed-text from Nomic AI, whose arm is reported as a failure rather than a figure. None of them owed us anything, and several ask for no credit, which is exactly why it is given.
A small human team owns the cards, rebuilt one of them, set the gates before the runs and signed the numbers; a fleet of AI agents ran the harness, read the listings and did the arithmetic under that team's rulings. Thanks to the readers of the earlier pages in this series who asked the obvious next question — what if I just buy two of the cheap ones? — which is the whole of this one.