# Six graphics cards, the same questions

*An RTX 3080, two of them in one desktop, an RTX 3090, two of those, an RTX 3090 Ti and an RTX PRO 6000 Blackwell, with a laptop part and an RTX 3080 Ti beside them — asked the same questions, in one table each: how fast each writes, how fast each draws, what the card itself pulled while doing it, how hot it got, and what it cost on the day it was priced. Not every card answers every question, and each table says which did not. Almost nothing here was measured for this page: every figure in its tables but two is already published on one of six earlier pages, and every one says which page and which section it came from. The two exceptions are prices for one card, read at a retailer and at a marketplace; the table says so, and names what else those reads found.*

*Published 2026-10-01 (UTC) · A small (human) team and a fleet of AI agents.*

**the short version:** A power cap moves a card more than the card's own name does: one RTX 3090, on one day, wrote with a dense model at 33.3 tokens a second held at 250 W and 47.8 tokens a second at 300 W — a bigger jump than any gap between two different cards at the same cap in the dense model's table. Two used RTX 3080s beat one 3090 on the same dense model at a window both can hold (the memory reserved for the conversation) — 44.6 tokens a second against 33.1, for 250 W on each of two boards against one board's 250 W, a cap their page says costs the single 3090 about 30 per cent on this model and the pair 0.2 per cent — and then lose it entirely at the next window up, where the page prints does not fit. Drawing pictures on the eight-step sketch graph, the order is plainer: one 3080 takes 2.074 seconds a picture and one 3090 1.683 seconds, both at 250 W, and one 3090 Ti 1.369 seconds at its own 450 W limit. A used 3080 asked $375.00 on 17 September 2026; a renewed 3090 asked $1,879.99 the same day — one seller's card, its page says, not a market price; used 3090s sold for $1,094.44 on average in August 2026, by that marketplace's own record. Every number in this page's tables but two was published first somewhere else, with its conditions; the two are an RTX 3080 Ti's prices — a refurbished listing at $479.99 from a retailer, out of stock at both of its last two reads, and a used card's August 2026 average sale of $475.64 from a marketplace — both read on 1 October 2026 and both marked as such, with the other listings those reads turned up named under the table rather than in it. Where two cards were never measured comparably the cell is a dagger and a footnote rather than a number.

15,178 words · about 69 minutes (at 220 words/min) · 17 tables · data kit: no

https://research.strata2signal.com/six-cards-the-same-questions/

---

*If you're new here: [strata→signal](https://strata2signal.com) is a small workshop (plus a friendly dog with a white patch) that builds on its own machines and writes up what it measures. This page measures almost nothing new: it collects what six earlier pages measured, out of the record they are kept in, and each of those figures links the page and the section it was published in. Two prices were read outside those pages for this page, at a retailer and at a marketplace, both for one card: they are printed in the price table and labelled as read rather than measured, and the other listings those reads turned up are named in the line beneath.*

## The words this page leans on {#the-words-this-page-leans-on}

- **Token** — the chunk a model reads and writes in, a little under a word. *Tokens a second* is how fast it writes; about ten a second reads as fluent.
- **Window** — the context window: how much of the conversation the model holds in mind at once, counted in tokens. It is reserved up front, so a longer one costs memory whether it is filled or not.
- **Writing, and drawing** — the two kinds of work here: a language model answering a prompt, and an image model making a picture.
- **Dense, and mixture** — a dense model uses all of itself for every word; a mixture (of experts) wakes only part of itself for each one, which is why the mixture model here is the faster of the two on every card.
- **Quantised, and Q4_K_M** — a model stored at fewer bits a weight than it was trained in, so it fits a smaller card; Q4_K_M is the format of every writing figure here whose page states one.
- **Cap, and board power** — a cap is a ceiling this workshop set on what a card may draw; board power is what the card actually drew, read from its own driver and never from a wall socket.
- **J / 1k tok** — joules a thousand tokens: the energy the board spent to write a thousand of them. A joule is one watt for one second, and 3,600 of them make a watt-hour.
- **Graph, and steps** — a graph is the recipe an image tool runs to make a picture; steps are how many passes it takes, so an eight-step picture is a sketch and a thirty-two-step one is finished.
- **First token** — the wait before the first word of an answer appears.
- **Ladder, rung, and arm** — a ladder is one card run at several caps in turn; each cap is a rung; an arm is one run of the harness, whatever its shape.
- **The shelf, and the record** — the shelf is this site's bench pages; the record, also called the hardware roster, is the list of every figure those pages published, each with its conditions, and it is what this page's tables are filled from.

## The cards, named once {#the-cards-named-once}

**RTX 3090 (24 GB)** — an EVGA GeForce RTX 3090 XC3 Ultra 24 GB, and the card with a figure in more of these tables than any other, sixteen of seventeen. This workshop has owned more than one of them, and its pages do not always say which ran: one calls its board a used card bought in 2021 on its third life, another prices a card bought renewed in May 2026, one of a pair, so a 3090 row here may not be one unit — except where the pages settle it: the two ladder pages cite the same result files for their 3090's 250 W and 300 W arms, so those cells are one board.

**RTX 3090 Ti (24 GB)** — an EVGA GeForce RTX 3090 Ti FTW3 Ultra 24 GB, new, first powered on for the bench that measured it.

**RTX PRO 6000 Blackwell (96 GB)** — an NVIDIA RTX PRO 6000 Blackwell Workstation Edition 96 GB, the workstation card this workshop serves its own models from.

**RTX 3080 (10 GB)** — one used card: in the drawing and heat tables, the Gigabyte GeForce RTX 3080 10 GB in the wide (x16) slot of a 2021 desktop, its thermal pads replaced by hand before any of this ran, as its page discloses; its split-model readings come from an earlier bench whose page does not name the board. And, as its own row, **two of them in that desktop** — that Gigabyte and an EVGA GeForce RTX 3080 FTW3 Ultra 10 GB.

**Two RTX 3090s in one desktop** — priced here and never benched as a pair for writing speed: the page that prices them queued that bench, and both cards left service on 2026-09-18 before it ran, which is why its row is nearly always empty.

**RTX 3080 Ti (12 GB)** — an NVIDIA GeForce RTX 3080 Ti Founders Edition 12 GB, refurbished, measured on two later pages of this shelf: [RTX 5080 vs RTX 3080 Ti](https://research.strata2signal.com/rtx-5080-vs-rtx-3080-ti/), which also prints what it cost, and [The same eighteen pictures, card by card](https://research.strata2signal.com/the-same-eighteen-pictures-card-by-card/). The hardware roster these tables are drawn from has not carried those readings yet, so its row is in every measurement table below with nothing in it until it does, and the sentence under each table says so.

**RTX 5090 Laptop GPU (24 GB)** — and this one is **not a graphics card**. It is an NVIDIA GeForce RTX 5090 Laptop GPU 24 GB, built into a gaming laptop at the factory: its graphics chip cannot be taken out or swapped, and it is not sold on its own. It is here because this shelf has measured it drawing pictures on the same work as two of the cards above, and because the answer is worth a reader's time. It fills no cell of any cap-keyed table on this page and it never will: its whole power envelope — 95 W by default, 175 W at the very most — sits below the 250 W column, which is the lowest column any of those tables has. Its row in each of them is a line of em dashes with a sentence under it saying why. The readings it does have are in **the sketch ladder**, further down, whose columns are readings rather than caps and which prints the cap each row ran at beside its figures. Its own page, [A laptop, asked the desktop's questions](https://research.strata2signal.com/a-laptop-asked-the-desktops-questions/), published 22 September 2026, puts it through this page's writing and drawing questions held to 95 W and then with NVIDIA's Dynamic Boost running, which held it between 140 and 150 W; every one of those readings sits below the lowest cap column here, so they stay on that page, linked rather than placed.

## How to read the tables {#how-to-read-the-tables}

Every table has the cards as rows. Most have **power caps** as columns — a cap is a ceiling this workshop set on what a card may draw, not the board's own limit, and on the dense model it is the loudest dial on this page. Four things can appear in a cell, and one kind of row:

- **a figure** — published, verbatim, on the page named under the table;
- **a dagger (†)** — this shelf *did* measure that card there, but under conditions that do not match the rest of the table, so it is not printed as a comparable number; the footnote says what differed and links it;
- **an em dash (—)** — no reading at all;
- **words** — *does not fit* where a page's own answer was that the model could not be held whole on the card or cards at that window, and *not found on* a named day, in the price table, where a listing was looked for and none was found;
- **a whole row of em dashes with a sentence under it** — a card this shelf *has* measured, whose readings no column of that particular table can hold. The RTX 5090 Laptop GPU is the one row of this kind: a 95-watt part has no place in a table whose lowest column is 250 W, so the row says so in words instead of pretending to a number. The RTX 3080 Ti's row looks the same, and its sentence says something different: that its readings are on two later pages and not yet in the record these tables are drawn from.

Under each table are its receipts — one line naming the page every figure came from, or, where a table draws on two, a letter beside each figure and the pages lettered to match — and, where they apply, four more kinds of line. **When** says the dates the records carry, and says so when they carry none — several of the drawing arms are dated nowhere in the section their figures were read from, and no date is lent to them here. **Off the ladder** names readings this shelf published for one of these cards that no cap column can hold, because their cap is written as a range, as a sentence, or not at all; they are named with their own cap wording rather than quietly left out. **Capped per card** quotes a two-card rig's own cap wording, so a 250 W column that means 250 W each and 500 W between them says so. **A second reading** appears where a second page published a different figure for the same cell under the conditions the table holds; both are printed and neither is adjusted toward the other. Five tables have no cap columns; four of them carry **the cap it ran at** as a column instead, and the fifth is the price table.

Every writing figure in the cap-keyed tables is the median of three scored runs on one frozen 2,101-byte prompt with 256 tokens out, after one discarded warm-up, as both ladder pages state; the window is room reserved, not filled; the spreads are on the source pages, and a gap smaller than its spread is not a finding. The decimals are each page's own — one page prints 47.8 where another prints 47.838 for the same run — and the extra digits are not extra precision. Three rows are nearly empty in every table, each for a reason its own sentence gives: two RTX 3090s, priced and never benched as a pair; the laptop part, which no cap column can hold; and the RTX 3080 Ti, whose readings are not in the record yet.

Two conditions are not constant anywhere on this page and are worth knowing before the first table. **These cards did not all sit in the same machine.** And **not one of the records filling a cell states how many people were asking at once** — two of the source pages measure that at length, in sections of their own, and three readings named under tables below do state it, two at one question at a time and one at sixteen callers at once; the readings in the cells do not.

A unit printed beside a bare figure — `tok/s`, `J / 1k tok` — is the record's label, not the source page's spelling, and it is added only where the page printed a bare number; where a page wrote the unit into the figure itself (`34.40 tokens a second`) that is what appears, so the same measure can wear two labels on this page. The digits are always the page's own.

## Writing: how fast each one answers {#writing-how-fast-each-one-answers}

Two models, kept apart because they behave differently. `gemma4:26b` is a **mixture** model: only part of it runs for each word, so it is fast for its size, though all of it must still fit in memory. `mistral-small3.2:24b` is **dense**: all of it runs for every word, and it is the one that finds a card's limits. Both were run **quantised** — stored at fewer bits a weight than they were trained in, in the format labelled `Q4_K_M`, which is how a 24-billion-parameter model fits a 24 GB card with room left for a long conversation.

<!-- s2s:compare:BEGIN gemma4-decode -->

**Writing speed, the mixture model, by power cap**

| The card | 250 W | 300 W | 350 W | 400 W | 420 W | 450 W |
|---|---|---|---|---|---|---|
| RTX 3090 (24 GB) | 125.6 tok/s (a) | 136.139 tok/s (b) | 136.618 tok/s (b) | — | — | — |
| RTX 3090 Ti (24 GB) | — | 145.908 tok/s (b) | 147.823 tok/s (b) | †1 | — | 147.560 tok/s (b) |
| RTX PRO 6000 Blackwell (96 GB) | — | — | — | — | †2 | — |
| RTX 3080 (10 GB), one card | — | — | — | — | — | — |
| Two RTX 3080s, one desktop | 114.8 tok/s (a) | 115.4 tok/s (a) | — | — | — | — |
| Two RTX 3090s, one desktop | — | — | — | — | — | — |
| RTX 5090 Laptop GPU (24 GB) | — | — | — | — | — | — |
| RTX 3080 Ti (12 GB) | — | — | — | — | — | — |

**Where each figure comes from.** (a) [Two 3080s against one 3090, and the cap decides § How fast does each one write?](https://research.strata2signal.com/two-used-3080s-priced/#how-fast-does-each-one-write) · (b) [RTX 3090 vs RTX 3090 Ti: an inference and diffusion bench § What a power cap takes, and what each card does at its own limit](https://research.strata2signal.com/one-3090-against-one-3090-ti/#what-a-power-cap-takes-and-what-each-card-does-at-its-own)

*Held constant: `gemma4:26b`, Q4_K_M, ollama 0.32.13 with an 8-bit cache, as both ladder pages state, a 131,072-token window, 256 tokens out.*

*When: 2026-09-17; 2026-09-17 (14:54 to 23:46 UTC); 2026-09-18, 00:22 to 00:54 UTC.*

*Capped per card, as its own page states it: **two RTX 3080s, one desktop** is 250 W each — 500 W between them in the 250 W column, 300 W each — 600 W between them in the 300 W column.*

*Off the ladder — this shelf publishes these too, and no column here can hold a cap given as a range, a sentence or nothing at all, so they are named instead of placed: **RTX 3090 (24 GB)** — 137 tok/s at a cap this page's columns cannot hold, stated as “measured before this workshop capped its cards”, one question at a time, on the model server of that day ([The cost to purchase and run a very capable home AI rig § What kind of performance to expect](https://research.strata2signal.com/home-inference-and-diffusion-rig-cost/#what-kind-of-performance-to-expect)).*

*RTX 5090 Laptop GPU (24 GB): a 95 W mobile part; its whole envelope sits below every cap column any table on this page holds, and its readings are in the sketch ladder on this page.*

*RTX 3080 Ti (12 GB): measured on two later pages of this shelf, linked above; the record these tables are filled from does not carry those readings yet, so its cells are empty rather than guessed.*

**†1** RTX 3090 Ti (24 GB), 400 W: measured, but the record behind this table carries no context window for it, so it cannot be placed in this row. Its page prints **147.599 tok/s** there — [RTX 3090 vs RTX 3090 Ti: an inference and diffusion bench § The two models want different caps](https://research.strata2signal.com/one-3090-against-one-3090-ti/#the-two-models-want-different-caps)
**†2** RTX PRO 6000 Blackwell (96 GB), 420 W: measured, but the record behind this table carries no quantisation, model-server version or context window for it, so it cannot be placed in this row. Its page prints **205.85 tokens a second** there — [A Short History of Mistral § What we measured on two cards in one box](https://research.strata2signal.com/a-short-history-of-mistral/#what-we-measured-on-two-cards-in-one-box)

*A note on two of the RTX PRO 6000 Blackwell (96 GB)'s figures, its 205.85 tokens a second on the mixture model and its 92.77 on the dense one: both runs were read on a model server that already had the model loaded when they began, so their own records show how fast the card wrote and not where the model sat — the load that would show that happened before either run, which this workshop's own audit of 2026-09-21 wrote down. The figures are printed here as speeds, and this page claims nothing about the fit.*

**A second reading — RTX 3090 (24 GB), 300 W.** A second page publishes **136.1 tok/s** for this cell under the same stated conditions ([Two 3080s against one 3090, and the cap decides § How fast does each one write?](https://research.strata2signal.com/two-used-3080s-priced/#how-fast-does-each-one-write)). This table prints the reading from the page that supplies the rest of the row; neither is adjusted toward the other.

<!-- s2s:compare:END gemma4-decode -->

<!-- s2s:compare:BEGIN mistral-decode -->

**Writing speed, the dense model, by power cap**

| The card | 250 W | 300 W | 350 W | 400 W | 420 W | 450 W |
|---|---|---|---|---|---|---|
| RTX 3090 (24 GB) | 33.3 tok/s (a) | 47.838 tok/s (b) | 46.880 tok/s (b) | — | — | — |
| RTX 3090 Ti (24 GB) | — | 42.969 tok/s (b) | 53.354 tok/s (b) | †1 | — | 57.396 tok/s (b) |
| RTX PRO 6000 Blackwell (96 GB) | — | — | — | — | †2 | — |
| RTX 3080 (10 GB), one card | — | — | — | — | — | — |
| Two RTX 3080s, one desktop | Does not fit (a) | Does not fit (a) | — | — | — | — |
| Two RTX 3090s, one desktop | — | — | — | — | — | — |
| RTX 5090 Laptop GPU (24 GB) | — | — | — | — | — | — |
| RTX 3080 Ti (12 GB) | — | — | — | — | — | — |

**Where each figure comes from.** (a) [Two 3080s against one 3090, and the cap decides § How fast does each one write?](https://research.strata2signal.com/two-used-3080s-priced/#how-fast-does-each-one-write) · (b) [RTX 3090 vs RTX 3090 Ti: an inference and diffusion bench § What a power cap takes, and what each card does at its own limit](https://research.strata2signal.com/one-3090-against-one-3090-ti/#what-a-power-cap-takes-and-what-each-card-does-at-its-own)

*Held constant: `mistral-small3.2:24b`, Q4_K_M, ollama 0.32.13 with an 8-bit cache, as both ladder pages state, a 65,536-token window, 256 tokens out.*

*When: 2026-09-17; 2026-09-17 (14:54 to 23:46 UTC); 2026-09-18, 00:22 to 00:54 UTC.*

*Capped per card, as its own page states it: **two RTX 3080s, one desktop** is 250 W each — 500 W between them in the 250 W column, 300 W each — 600 W between them in the 300 W column.*

*Off the ladder — this shelf publishes these too, and no column here can hold a cap given as a range, a sentence or nothing at all, so they are named instead of placed: **RTX 3090 (24 GB)** — 53 tok/s at a cap this page's columns cannot hold, stated as “measured before this workshop capped its cards”, one question at a time, on the model server of that day ([The cost to purchase and run a very capable home AI rig § What kind of performance to expect](https://research.strata2signal.com/home-inference-and-diffusion-rig-cost/#what-kind-of-performance-to-expect)).*

*RTX 5090 Laptop GPU (24 GB): a 95 W mobile part; its whole envelope sits below every cap column any table on this page holds, and its readings are in the sketch ladder on this page.*

*RTX 3080 Ti (12 GB): measured on two later pages of this shelf, linked above; the record these tables are filled from does not carry those readings yet, so its cells are empty rather than guessed.*

**†1** RTX 3090 Ti (24 GB), 400 W: measured, but the record behind this table carries no context window for it, so it cannot be placed in this row. Its page prints **56.936 tok/s** there — [RTX 3090 vs RTX 3090 Ti: an inference and diffusion bench § The two models want different caps](https://research.strata2signal.com/one-3090-against-one-3090-ti/#the-two-models-want-different-caps)
**†2** RTX PRO 6000 Blackwell (96 GB), 420 W: measured, but the record behind this table carries no model-server version or context window for it, so it cannot be placed in this row. Its page prints **92.77 tokens a second** there — [A Short History of Mistral § What we measured on two cards in one box](https://research.strata2signal.com/a-short-history-of-mistral/#what-we-measured-on-two-cards-in-one-box)

*A note on two of the RTX PRO 6000 Blackwell (96 GB)'s figures, its 205.85 tokens a second on the mixture model and its 92.77 on the dense one: both runs were read on a model server that already had the model loaded when they began, so their own records show how fast the card wrote and not where the model sat — the load that would show that happened before either run, which this workshop's own audit of 2026-09-21 wrote down. The figures are printed here as speeds, and this page claims nothing about the fit.*

**A second reading — RTX 3090 (24 GB), 300 W.** A second page publishes **47.8 tok/s** for this cell under the same stated conditions ([Two 3080s against one 3090, and the cap decides § How fast does each one write?](https://research.strata2signal.com/two-used-3080s-priced/#how-fast-does-each-one-write)). This table prints the reading from the page that supplies the rest of the row; neither is adjusted toward the other.

<!-- s2s:compare:END mistral-decode -->

Read along each row of the dense model's table and the cap is the loudest dial on this page; do the same in the mixture model's table and it barely moves. One cell reads backwards — a 3090's 46.880 tokens a second at 350 W under its 47.838 at 300 W — and its page says why: a ladder rung on a card eight degrees warmer against a single repeat on a cold one, not a cap effect. And a card's own limit is not always the best place to leave it: on the mixture model this shelf's 3090 Ti gains nothing past 350 W, with 147.823 tokens a second there and 147.560 at its own 450 W, a difference its page puts inside the bench's own run-to-run spread.

## When the model is bigger than the card {#when-the-model-is-bigger-than-the-card}

This is the part no tokens-a-second figure will tell you, and for anyone choosing between one big card and two small ones it decides the question. A model has to fit in the card's memory, *and so does the conversation* — the "context window", the number of words the model is asked to keep in mind at once. The window is reserved up front, so asking for a longer one is asking for memory the model itself then does not have. Both ladder pages held these windows with the conversation's memory kept at 8 bits — on the pair, its page says, the setting that fits, at about 6 per cent of the rate — and the pair reached its figures only with the layer count named by hand: left to itself, that page says, the server spilled part of the model to the computer's memory.

<!-- s2s:compare:BEGIN window-ladder -->

**The dense model, and how much window each rig can hold**

| The card | The cap it ran at | A 49,152-token window | A 65,536-token window |
|---|---|---|---|
| RTX 3090 (24 GB) | 250 W | 33.1 tok/s | 33.3 tok/s |
| RTX 3090 Ti (24 GB) | — | †1 | †2 |
| RTX PRO 6000 Blackwell (96 GB) | — | †3 | †4 |
| RTX 3080 (10 GB), one card | — | — | — |
| Two RTX 3080s, one desktop | 250 W each — 500 W between them | 44.6 tok/s | Does not fit |
| Two RTX 3090s, one desktop | — | — | — |
| RTX 5090 Laptop GPU (24 GB) | — | — | — |
| RTX 3080 Ti (12 GB) | — | — | — |

*Every figure in this table is published on [Two 3080s against one 3090, and the cap decides § How fast does each one write?](https://research.strata2signal.com/two-used-3080s-priced/#how-fast-does-each-one-write).*

*Held constant: `mistral-small3.2:24b`, Q4_K_M, ollama 0.32.13 with an 8-bit cache, 256 tokens out, one page, one day, and every board capped at 250 W — which for the two-card rig is 250 W each and 500 W between them, as its page states it. The columns are the size of the conversation the model is asked to hold, because that — not speed — is what a 10 GB board runs out of first. A cell reading “does not fit” is the page's own answer, not a missing measurement.*

*When: 2026-09-17.*

*RTX 5090 Laptop GPU (24 GB): a 95 W mobile part; its whole envelope sits below every cap column any table on this page holds, and its readings are in the sketch ladder on this page.*

*RTX 3080 Ti (12 GB): measured on two later pages of this shelf, linked above; the record these tables are filled from does not carry those readings yet, so its cells are empty rather than guessed.*

**†1** RTX 3090 Ti (24 GB), a 49,152-token window: measured, but its page ran it at a window of 65,536 tokens, not 49,152, and at 300 W, not 250 W, so it cannot be placed in this row. Its page prints **42.969 tok/s** there — [RTX 3090 vs RTX 3090 Ti: an inference and diffusion bench § What a power cap takes, and what each card does at its own limit](https://research.strata2signal.com/one-3090-against-one-3090-ti/#what-a-power-cap-takes-and-what-each-card-does-at-its-own)
**†2** RTX 3090 Ti (24 GB), a 65,536-token window: measured, but its page ran it at 300 W, not 250 W, so it cannot be placed in this row. Its page prints **42.969 tok/s** there — [RTX 3090 vs RTX 3090 Ti: an inference and diffusion bench § What a power cap takes, and what each card does at its own limit](https://research.strata2signal.com/one-3090-against-one-3090-ti/#what-a-power-cap-takes-and-what-each-card-does-at-its-own)
**†3** RTX PRO 6000 Blackwell (96 GB), a 49,152-token window: measured, but its page ran it at 420 W, not 250 W; the record behind this table carries no model-server version or context window for it, so it cannot be placed in this row. Its page prints **92.77 tokens a second** there — [A Short History of Mistral § What we measured on two cards in one box](https://research.strata2signal.com/a-short-history-of-mistral/#what-we-measured-on-two-cards-in-one-box)
**†4** RTX PRO 6000 Blackwell (96 GB), a 65,536-token window: measured, but its page ran it at 420 W, not 250 W; the record behind this table carries no model-server version or context window for it, so it cannot be placed in this row. Its page prints **92.77 tokens a second** there — [A Short History of Mistral § What we measured on two cards in one box](https://research.strata2signal.com/a-short-history-of-mistral/#what-we-measured-on-two-cards-in-one-box)

*A note on two of the RTX PRO 6000 Blackwell (96 GB)'s figures, its 205.85 tokens a second on the mixture model and its 92.77 on the dense one: both runs were read on a model server that already had the model loaded when they began, so their own records show how fast the card wrote and not where the model sat — the load that would show that happened before either run, which this workshop's own audit of 2026-09-21 wrote down. The figures are printed here as speeds, and this page claims nothing about the fit.*

<!-- s2s:compare:END window-ladder -->

Two used 10 GB cards, holding the model between them, are faster than one 24 GB card at the smaller window — and at the larger one they are not slower, they are absent. That is what a memory ceiling looks like in a table. The same page also records that only 92.1 per cent of a 14-billion-parameter model sits on one 10 GB board once its window is reserved; the rest spills to the computer's memory.

And when a model does not fit at all, it does not stop; it spills into the computer's own RAM, and the card spends its time waiting on it.

<!-- s2s:compare:BEGIN spilled -->

**What a card does when the model is bigger than it is**

| The card | The cap it ran at | `gemma4:26b`, the mixture model | `mistral-small3.2:24b`, the dense model |
|---|---|---|---|
| RTX 3090 (24 GB) | — | — | — |
| RTX 3090 Ti (24 GB) | — | — | — |
| RTX PRO 6000 Blackwell (96 GB) | — | — | — |
| RTX 3080 (10 GB), one card | Not stated | 34.40 tokens a second | 5.13 tok/s |
| Two RTX 3080s, one desktop | — | — | — |
| Two RTX 3090s, one desktop | — | — | — |
| RTX 5090 Laptop GPU (24 GB) | — | — | — |
| RTX 3080 Ti (12 GB) | — | — | — |

*Every figure in this table is published on [A Short History of Mistral § What we measured on two cards in one box](https://research.strata2signal.com/a-short-history-of-mistral/#what-we-measured-on-two-cards-in-one-box).*

*Held constant: One page, and no cap stated for either reading. Its own page files both under “two earlier benches on other boxes … controls, not like-for-like”, so they are not comparable with any other table here and are not compared with one. When a model will not fit in a card's memory it is split with the computer's own RAM — that page's words are “on a used consumer card too small to hold it, half in RAM”. Only one card on this page was measured this way.*

*When: 2026-09-10. Some of these arms carry no date in the section their figures were read from, and none is invented here.*

*RTX 5090 Laptop GPU (24 GB): a 95 W mobile part; its whole envelope sits below every cap column any table on this page holds, and its readings are in the sketch ladder on this page.*

*RTX 3080 Ti (12 GB): measured on two later pages of this shelf, linked above; the record these tables are filled from does not carry those readings yet, so its cells are empty rather than guessed.*

<!-- s2s:compare:END spilled -->

## What the boards actually drew {#what-the-boards-actually-drew}

A cap is a ceiling, not a reading. What a board draws under it is a separate measurement, and every figure below is card board power off the driver — the card's own driver software — not a wall socket. This shelf publishes whole-box wall draws elsewhere; they measure a different thing and are not in these tables. On the bench that ran the two 24 GB cards, the one instrument that could have read the whole box refused every request and that page records it as unreadable.

<!-- s2s:compare:BEGIN gemma4-draw -->

**What the board actually drew while writing the mixture model, by power cap**

| The card | 250 W | 300 W | 350 W | 450 W |
|---|---|---|---|---|
| RTX 3090 (24 GB) | 230 W (a) | 241.53 W (b) | 240.49 W (b) | — |
| RTX 3090 Ti (24 GB) | — | 240.58 W (b) | 282.64 W (b) | 283.01 W (b) |
| RTX PRO 6000 Blackwell (96 GB) | — | — | — | — |
| RTX 3080 (10 GB), one card | — | — | — | — |
| Two RTX 3080s, one desktop | 290 W (a) | 346 W (a) | — | — |
| Two RTX 3090s, one desktop | — | — | — | — |
| RTX 5090 Laptop GPU (24 GB) | — | — | — | — |
| RTX 3080 Ti (12 GB) | — | — | — | — |

**Where each figure comes from.** (a) [Two 3080s against one 3090, and the cap decides § Power, heat, and the fifty watts](https://research.strata2signal.com/two-used-3080s-priced/#power-heat-and-the-fifty-watts) · (b) [RTX 3090 vs RTX 3090 Ti: an inference and diffusion bench § What a power cap takes, and what each card does at its own limit](https://research.strata2signal.com/one-3090-against-one-3090-ti/#what-a-power-cap-takes-and-what-each-card-does-at-its-own)

*Held constant: `gemma4:26b`, Q4_K_M, ollama 0.32.13 with an 8-bit cache, 256 tokens out, each rig at its own largest window held whole, 131,072 tokens for this model on every rig here. Every figure is the mean the card's own driver reported over the run — card board power, never the wall.*

*When: 2026-09-17; 2026-09-17 (14:54 to 23:46 UTC); 2026-09-18, 00:22 to 00:54 UTC.*

*Capped per card, as its own page states it: **two RTX 3080s, one desktop** is 250 W each — 500 W between them in the 250 W column, 300 W each — 600 W between them in the 300 W column.*

*RTX 5090 Laptop GPU (24 GB): a 95 W mobile part; its whole envelope sits below every cap column any table on this page holds, and its readings are in the sketch ladder on this page.*

*RTX 3080 Ti (12 GB): measured on two later pages of this shelf, linked above; the record these tables are filled from does not carry those readings yet, so its cells are empty rather than guessed.*

**A second reading — RTX 3090 (24 GB), 300 W.** A second page publishes **242 W** for this cell under the same stated conditions ([Two 3080s against one 3090, and the cap decides § Power, heat, and the fifty watts](https://research.strata2signal.com/two-used-3080s-priced/#power-heat-and-the-fifty-watts)). This table prints the reading from the page that supplies the rest of the row; neither is adjusted toward the other.

<!-- s2s:compare:END gemma4-draw -->

<!-- s2s:compare:BEGIN mistral-draw -->

**What the board actually drew while writing the dense model, by power cap**

| The card | 250 W | 300 W | 350 W | 400 W | 450 W |
|---|---|---|---|---|---|
| RTX 3090 (24 GB) | 235 W (a) | 289.86 W (b) | 273.35 W (b) | — | — |
| RTX 3090 Ti (24 GB) | — | 296.64 W (b) | 338.05 W (b) | †1 | 418.76 W (b) |
| RTX PRO 6000 Blackwell (96 GB) | — | — | — | — | — |
| RTX 3080 (10 GB), one card | — | — | — | — | — |
| Two RTX 3080s, one desktop | 466 W (a) | 509 W (a) | — | — | — |
| Two RTX 3090s, one desktop | — | — | — | — | — |
| RTX 5090 Laptop GPU (24 GB) | — | — | — | — | — |
| RTX 3080 Ti (12 GB) | — | — | — | — | — |

**Where each figure comes from.** (a) [Two 3080s against one 3090, and the cap decides § Power, heat, and the fifty watts](https://research.strata2signal.com/two-used-3080s-priced/#power-heat-and-the-fifty-watts) · (b) [RTX 3090 vs RTX 3090 Ti: an inference and diffusion bench § What a power cap takes, and what each card does at its own limit](https://research.strata2signal.com/one-3090-against-one-3090-ti/#what-a-power-cap-takes-and-what-each-card-does-at-its-own)

*Held constant: `mistral-small3.2:24b`, Q4_K_M, ollama 0.32.13 with an 8-bit cache, 256 tokens out, each rig at its own largest window held whole: 49,152 tokens for the two 3080s, 65,536 for the single cards. Every figure is the mean the card's own driver reported over the run — card board power, never the wall.*

*When: 2026-09-17; 2026-09-17 (14:54 to 23:46 UTC); 2026-09-18, 00:22 to 00:54 UTC.*

*Capped per card, as its own page states it: **two RTX 3080s, one desktop** is 250 W each — 500 W between them in the 250 W column, 300 W each — 600 W between them in the 300 W column.*

*RTX 5090 Laptop GPU (24 GB): a 95 W mobile part; its whole envelope sits below every cap column any table on this page holds, and its readings are in the sketch ladder on this page.*

*RTX 3080 Ti (12 GB): measured on two later pages of this shelf, linked above; the record these tables are filled from does not carry those readings yet, so its cells are empty rather than guessed.*

**†1** RTX 3090 Ti (24 GB), 400 W: measured, but the record behind this table files its run as ladder rung, where this table holds 256 tokens out, so it cannot be placed in this row. Its page prints **382.16 W** there — [RTX 3090 vs RTX 3090 Ti: an inference and diffusion bench § What a power cap takes, and what each card does at its own limit](https://research.strata2signal.com/one-3090-against-one-3090-ti/#what-a-power-cap-takes-and-what-each-card-does-at-its-own)

**A second reading — RTX 3090 (24 GB), 300 W.** A second page publishes **290 W** for this cell under the same stated conditions ([Two 3080s against one 3090, and the cap decides § Power, heat, and the fifty watts](https://research.strata2signal.com/two-used-3080s-priced/#power-heat-and-the-fifty-watts)). This table prints the reading from the page that supplies the rest of the row; neither is adjusted toward the other.

<!-- s2s:compare:END mistral-draw -->

On this shelf's 3090 Ti the mixture model's draw stops climbing after 350 W — 282.64 W there, 283.01 W at the 450 W cap — while the dense model's keeps going, to 418.76 W; its page flags the mixture's step from 350 to 400 W and says it may not be read as a cap effect, because the fan read 49 per cent at one rung and 0 at the other, and board power contains the fans. One row reads backwards: a 3090's dense figure at 350 W, 273.35 W, is *below* its 300 W figure of 289.86 W. Its page explains why, and the explanation is a condition this table does not carry: the two rows are not the same arm — a single-window repeat on a cold card against a ladder rung on a card eight degrees warmer.

## What a thousand tokens cost in energy {#what-a-thousand-tokens-cost-in-energy}

The same runs, divided by the tokens they produced, as each page published them. Lower is better, and the cheapest cell is not always the fastest card.

<!-- s2s:compare:BEGIN gemma4-joules -->

**What a thousand tokens of the mixture model cost in energy, by power cap**

| The card | 250 W | 300 W | 350 W | 400 W | 450 W |
|---|---|---|---|---|---|
| RTX 3090 (24 GB) | 1,832 J / 1k tok (a) | 1,771.71 J / 1k tok (b) | 1,758.13 J / 1k tok (b) | — | — |
| RTX 3090 Ti (24 GB) | — | 1,648.85 J / 1k tok (c) | 1,909.96 J / 1k tok (c) | 1,723.81 J / 1k tok (c) | 1,915.13 J / 1k tok (c) |
| RTX PRO 6000 Blackwell (96 GB) | — | — | — | — | — |
| RTX 3080 (10 GB), one card | — | — | — | — | — |
| Two RTX 3080s, one desktop | 2,531 J / 1k tok (a) | 2,994 J / 1k tok (a) | — | — | — |
| Two RTX 3090s, one desktop | — | — | — | — | — |
| RTX 5090 Laptop GPU (24 GB) | — | — | — | — | — |
| RTX 3080 Ti (12 GB) | — | — | — | — | — |

**Where each figure comes from.** (a) [Two 3080s against one 3090, and the cap decides § Power, heat, and the fifty watts](https://research.strata2signal.com/two-used-3080s-priced/#power-heat-and-the-fifty-watts) · (b) [RTX 3090 vs RTX 3090 Ti: an inference and diffusion bench § What a power cap takes, and what each card does at its own limit](https://research.strata2signal.com/one-3090-against-one-3090-ti/#what-a-power-cap-takes-and-what-each-card-does-at-its-own) · (c) [RTX 3090 vs RTX 3090 Ti: an inference and diffusion bench § The two models want different caps](https://research.strata2signal.com/one-3090-against-one-3090-ti/#the-two-models-want-different-caps)

*Held constant: `gemma4:26b`, Q4_K_M, ollama 0.32.13 with an 8-bit cache, 256 tokens out, board power only — no wall reading — each rig at its own largest window held whole, which for this model its pages put at 131,072 tokens on every rig here.*

*Lower is better here: it is the energy the board spent to write a thousand tokens.*

*When: 2026-09-17; 2026-09-17 (14:54 to 23:46 UTC); 2026-09-18, 00:22 to 00:54 UTC. Some of these arms carry no date in the section their figures were read from, and none is invented here.*

*Capped per card, as its own page states it: **two RTX 3080s, one desktop** is 250 W each — 500 W between them in the 250 W column, 300 W each — 600 W between them in the 300 W column.*

*RTX 5090 Laptop GPU (24 GB): a 95 W mobile part; its whole envelope sits below every cap column any table on this page holds, and its readings are in the sketch ladder on this page.*

*RTX 3080 Ti (12 GB): measured on two later pages of this shelf, linked above; the record these tables are filled from does not carry those readings yet, so its cells are empty rather than guessed.*

*A note on two of the RTX 3090 Ti (24 GB)'s energy figures, 1,909.96 and 1,723.81 J a thousand tokens at 350 and 400 W: its page prints both with a flag and says the step between them may not be read as a cap effect — between those two caps the card drew 28.4 W less at identical core clocks and identical decode, because its fan read 49 per cent at one rung and 0 at the other, and board power contains the fans. The bench flags the step and corrects nothing, and neither does this page.*

**A second reading — RTX 3090 (24 GB), 300 W.** A second page publishes **1,772 J / 1k tok** for this cell under the same stated conditions ([Two 3080s against one 3090, and the cap decides § Power, heat, and the fifty watts](https://research.strata2signal.com/two-used-3080s-priced/#power-heat-and-the-fifty-watts)). This table prints the reading from the page that supplies the rest of the row; neither is adjusted toward the other.

<!-- s2s:compare:END gemma4-joules -->

<!-- s2s:compare:BEGIN mistral-joules -->

**What a thousand tokens of the dense model cost in energy, by power cap**

| The card | 250 W | 300 W | 350 W | 400 W | 450 W |
|---|---|---|---|---|---|
| RTX 3090 (24 GB) | 7,076 J / 1k tok (a) | 6,059.20 J / 1k tok (b) | 5,842.92 J / 1k tok (b) | — | — |
| RTX 3090 Ti (24 GB) | — | 6,902.38 J / 1k tok (c) | 6,343.81 J / 1k tok (c) | 6,713.05 J / 1k tok (c) | 7,295.49 J / 1k tok (c) |
| RTX PRO 6000 Blackwell (96 GB) | — | — | — | — | — |
| RTX 3080 (10 GB), one card | — | — | — | — | — |
| Two RTX 3080s, one desktop | 10,442 J / 1k tok (a) | 11,385 J / 1k tok (a) | — | — | — |
| Two RTX 3090s, one desktop | — | — | — | — | — |
| RTX 5090 Laptop GPU (24 GB) | — | — | — | — | — |
| RTX 3080 Ti (12 GB) | — | — | — | — | — |

**Where each figure comes from.** (a) [Two 3080s against one 3090, and the cap decides § Power, heat, and the fifty watts](https://research.strata2signal.com/two-used-3080s-priced/#power-heat-and-the-fifty-watts) · (b) [RTX 3090 vs RTX 3090 Ti: an inference and diffusion bench § What a power cap takes, and what each card does at its own limit](https://research.strata2signal.com/one-3090-against-one-3090-ti/#what-a-power-cap-takes-and-what-each-card-does-at-its-own) · (c) [RTX 3090 vs RTX 3090 Ti: an inference and diffusion bench § The two models want different caps](https://research.strata2signal.com/one-3090-against-one-3090-ti/#the-two-models-want-different-caps)

*Held constant: `mistral-small3.2:24b`, Q4_K_M, ollama 0.32.13 with an 8-bit cache, 256 tokens out, board power only — no wall reading — each rig at its own largest window held whole: 49,152 tokens for the two 3080s, 65,536 for the single cards, as their pages state.*

*Lower is better here: it is the energy the board spent to write a thousand tokens.*

*When: 2026-09-17; 2026-09-17 (14:54 to 23:46 UTC); 2026-09-18, 00:22 to 00:54 UTC. Some of these arms carry no date in the section their figures were read from, and none is invented here.*

*Capped per card, as its own page states it: **two RTX 3080s, one desktop** is 250 W each — 500 W between them in the 250 W column, 300 W each — 600 W between them in the 300 W column.*

*RTX 5090 Laptop GPU (24 GB): a 95 W mobile part; its whole envelope sits below every cap column any table on this page holds, and its readings are in the sketch ladder on this page.*

*RTX 3080 Ti (12 GB): measured on two later pages of this shelf, linked above; the record these tables are filled from does not carry those readings yet, so its cells are empty rather than guessed.*

**A second reading — RTX 3090 (24 GB), 300 W.** A second page publishes **6,059 J / 1k tok** for this cell under the same stated conditions ([Two 3080s against one 3090, and the cap decides § Power, heat, and the fifty watts](https://research.strata2signal.com/two-used-3080s-priced/#power-heat-and-the-fifty-watts)). This table prints the reading from the page that supplies the rest of the row; neither is adjusted toward the other.

<!-- s2s:compare:END mistral-joules -->

## Drawing pictures {#drawing-pictures}

Three tables. The first two time a named graph of one open-weight sketch model, FLUX.2 Klein 4B — its eight-step graph and its thirty-two-step one. The third is ten minutes of continuous drawing, and **neither source page names the graph that ran in it**, so this page does not either. The first two hold the work equal and let the seconds vary. The third does the opposite — every card draws for the same ten minutes and finishes a different number of pictures — so the two kinds of figure will not divide into each other, and neither page claims they should.

<!-- s2s:compare:BEGIN render-8st -->

**Seconds per finished picture, the eight-step sketch graph, by power cap**

| The card | 250 W | 300 W | 350 W | 400 W | 450 W |
|---|---|---|---|---|---|
| RTX 3090 (24 GB) | 1.683 s (a) | 1.560 s (a) | 1.509 s (a) | — | — |
| RTX 3090 Ti (24 GB) | — | 1.455 s (a) | 1.397 s (a) | 1.363 s (a) | 1.369 s (a) |
| RTX PRO 6000 Blackwell (96 GB) | — | — | — | — | — |
| RTX 3080 (10 GB), one card | 2.074 s (b) | — | — | — | — |
| Two RTX 3080s, one desktop | — | — | — | — | — |
| Two RTX 3090s, one desktop | — | — | — | — | — |
| RTX 5090 Laptop GPU (24 GB) | — | — | — | — | — |
| RTX 3080 Ti (12 GB) | — | — | — | — | — |

**Where each figure comes from.** (a) [RTX 3090 vs RTX 3090 Ti: an inference and diffusion bench § Renders](https://research.strata2signal.com/one-3090-against-one-3090-ti/#renders) · (b) [Two 3080s against one 3090, and the cap decides § Renders: one 3080 against one 3090, seconds per image](https://research.strata2signal.com/two-used-3080s-priced/#renders-one-3080-against-one-3090-seconds-per-image)

*Held constant: The `klein-4b-8st` graph on ComfyUI 0.21.1 — one picture, 8 steps, 512 × 512 — three pictures submitted at a time, the median of two timed submits (with two, their average). The single 3080's figure is its card in the wide slot; the same page prints its narrow-slot figure too, which is slower.*

*When: 2026-09-17. Some of these arms carry no date in the section their figures were read from, and none is invented here.*

*RTX 5090 Laptop GPU (24 GB): a 95 W mobile part; its whole envelope sits below every cap column any table on this page holds, and its readings are in the sketch ladder on this page.*

*RTX 3080 Ti (12 GB): measured on two later pages of this shelf, linked above; the record these tables are filled from does not carry those readings yet, so its cells are empty rather than guessed.*

<!-- s2s:compare:END render-8st -->

<!-- s2s:compare:BEGIN render-32st -->

**Seconds per finished picture, the thirty-two-step graph, by power cap**

| The card | 250 W | 300 W | 350 W | 400 W | 450 W |
|---|---|---|---|---|---|
| RTX 3090 (24 GB) | 6.297 s (a) | 5.827 s (a) | 5.670 s (a) | — | — |
| RTX 3090 Ti (24 GB) | — | 5.332 s (a) | 5.130 s (a) | 5.015 s (a) | 5.015 s (a) |
| RTX PRO 6000 Blackwell (96 GB) | — | — | — | — | — |
| RTX 3080 (10 GB), one card | 7.278 s (b) | — | — | — | — |
| Two RTX 3080s, one desktop | — | — | — | — | — |
| Two RTX 3090s, one desktop | — | — | — | — | — |
| RTX 5090 Laptop GPU (24 GB) | — | — | — | — | — |
| RTX 3080 Ti (12 GB) | — | — | — | — | — |

**Where each figure comes from.** (a) [RTX 3090 vs RTX 3090 Ti: an inference and diffusion bench § Renders](https://research.strata2signal.com/one-3090-against-one-3090-ti/#renders) · (b) [Two 3080s against one 3090, and the cap decides § Renders: one 3080 against one 3090, seconds per image](https://research.strata2signal.com/two-used-3080s-priced/#renders-one-3080-against-one-3090-seconds-per-image)

*Held constant: The `klein-4b-32st` graph on ComfyUI 0.21.1, three pictures submitted at a time, the median of two timed submits (with two, their average); neither page states this graph's picture size. The single 3080's figure is its card in the wide slot; the same page prints its narrow-slot figure too, which is slower.*

*When: 2026-09-17. Some of these arms carry no date in the section their figures were read from, and none is invented here.*

*RTX 5090 Laptop GPU (24 GB): a 95 W mobile part; its whole envelope sits below every cap column any table on this page holds, and its readings are in the sketch ladder on this page.*

*RTX 3080 Ti (12 GB): measured on two later pages of this shelf, linked above; the record these tables are filled from does not carry those readings yet, so its cells are empty rather than guessed.*

<!-- s2s:compare:END render-32st -->

<!-- s2s:compare:BEGIN render-ten-minutes -->

**Pictures finished in ten minutes of continuous drawing, by power cap**

| The card | 250 W | 300 W | 350 W | 400 W | 450 W |
|---|---|---|---|---|---|
| RTX 3090 (24 GB) | 269 images (a) | — | 311 images (a) | — | — |
| RTX 3090 Ti (24 GB) | — | 320 images (a) | 331 images (a) | 347 images (a) | 348 images (a) |
| RTX PRO 6000 Blackwell (96 GB) | — | — | — | — | — |
| RTX 3080 (10 GB), one card | 158 images (b) | — | — | — | — |
| Two RTX 3080s, one desktop | — | — | — | — | — |
| Two RTX 3090s, one desktop | — | — | — | — | — |
| RTX 5090 Laptop GPU (24 GB) | — | — | — | — | — |
| RTX 3080 Ti (12 GB) | — | — | — | — | — |

**Where each figure comes from.** (a) [RTX 3090 vs RTX 3090 Ti: an inference and diffusion bench § Renders](https://research.strata2signal.com/one-3090-against-one-3090-ti/#renders) · (b) [Two 3080s against one 3090, and the cap decides § Renders: one 3080 against one 3090, seconds per image](https://research.strata2signal.com/two-used-3080s-priced/#renders-one-3080-against-one-3090-seconds-per-image)

*Held constant: Ten minutes of continuous drawing on ComfyUI 0.21.1, counted by the render engine's own clock. These arms did unequal amounts of work by design — a faster card finishes more pictures in the same ten minutes — so the count is the measurement, not a rate held equal.*

*When: 2026-09-17. Some of these arms carry no date in the section their figures were read from, and none is invented here.*

*Which model drew these: the source pages do not name it for these arms, so this page does not either.*

*RTX 5090 Laptop GPU (24 GB): a 95 W mobile part; its whole envelope sits below every cap column any table on this page holds, and its readings are in the sketch ladder on this page.*

*RTX 3080 Ti (12 GB): measured on two later pages of this shelf, linked above; the record these tables are filled from does not carry those readings yet, so its cells are empty rather than guessed.*

<!-- s2s:compare:END render-ten-minutes -->

## The sketch ladder, and the one row here that is not a card {#the-sketch-ladder-and-the-one-row-here-that-is-not-a-card}

The three drawing tables above are keyed on power caps, and there is a part this shelf has measured that none of them can hold: the **RTX 5090 Laptop GPU**, built into a gaming laptop, allowed 95 W by default and 175 W at the very most. That whole envelope sits below the 250 W column, which is the lowest column any cap-keyed table on this page has. A row for it in those tables is a line of em dashes with a sentence under it, and that is the honest shape — a 95-watt figure printed in a 250-watt column would be a comparison nobody made.

What it does have is a table of its own, and the table's columns are **readings** rather than caps: one sketch size each, FLUX.2 klein-4B with the step count held at four, six renders per cell. Three of this page's rows have figures here, because one page put the same three sketch sizes to three parts in two computers: the 96 GB card's figures are an archived baseline from 23 August 2026, taken in the same desktop a 3090 later ran in, and that 3090's and the laptop part's runs carry no date on that page. The cap column says what each page states, and for these three rows every page states none — so the column reads *not stated* rather than borrowing a number from somewhere else.

Read the laptop's row before the workstation card's. At 512×512, the size the workshop's own sketch tool actually runs, a laptop part drew a picture in **0.71 s** against a desktop RTX 3090's **0.88 s** — and it is the 96 GB workstation card at **0.26 s** that puts both in proportion. The laptop's spread is the widest of the three at every size, and the top of each range is typically that cell's first render, its page says — the cache warming up.

<!-- s2s:compare:BEGIN laptop-readings -->

**The sketch ladder — how fast a 95-watt laptop part draws a picture, beside two desktop cards**

| The card | The cap it ran at | 512×512 *(the live default)* | 768×768 | 1024×1024 |
|---|---|---|---|---|
| RTX 3090 (24 GB) | Not stated | 0.88 s (0.88–1.19) | 1.65 s (1.64–1.96) | 2.86 s (2.84–3.17) |
| RTX 3090 Ti (24 GB) | — | — | — | — |
| RTX PRO 6000 Blackwell (96 GB) | Not stated | 0.26 s (0.26–0.40) | 0.48 s (0.47–0.61) | 0.82 s (0.80–0.96) |
| RTX 3080 (10 GB), one card | — | — | — | — |
| Two RTX 3080s, one desktop | — | — | — | — |
| Two RTX 3090s, one desktop | — | — | — | — |
| RTX 5090 Laptop GPU (24 GB) | Not stated | 0.71 s (0.69–1.27) | 1.37 s (1.36–2.02) | 2.40 s (2.35–2.98) |
| RTX 3080 Ti (12 GB) | — | — | — | — |

*Every figure in this table is published on [A Rig Your Friend Already Owns § The sketch ladder — the number the game lives on](https://research.strata2signal.com/a-rig-your-friend-already-owns/#the-sketch-ladder-the-number-the-game-lives-on).*

*Held constant: FLUX.2 klein-4B, steps held at 4, six renders per cell, at three sketch sizes — the median of each cell with its own spread, exactly as its page prints them. This is the only table on this page where the laptop part has figures, and it is a table of readings because no cap column here could hold one: that board's whole envelope, 95 W by default and 175 W at most, sits below the 250 W column that is the lowest any cap-keyed table above has. The cap column says what each page states, and for these three rows it states none. The record dates the 96 GB card's run in its page's own words, August 23rd — of 2026, the year that page was published — and the other two runs not at all.*

*When: August 23rd. Some of these arms carry no date in the section their figures were read from, and none is invented here.*

*RTX 3080 Ti (12 GB): measured on two later pages of this shelf, linked above; the record these tables are filled from does not carry those readings yet, so its cells are empty rather than guessed.*

<!-- s2s:compare:END laptop-readings -->

## Heat, and how hard the fans worked {#heat-and-how-hard-the-fans-worked}

Temperature is the measurement anyone can check at home with the tool these benches used, `nvidia-smi`, the driver's own, and it is the one most often quoted without saying what the card was doing. Every figure below says what the load was.

<!-- s2s:compare:BEGIN heat-writing -->

**How hot the core got while writing the dense model, by power cap**

| The card | 250 W | 300 W |
|---|---|---|
| RTX 3090 (24 GB) | 63 °C | 57 °C |
| RTX 3090 Ti (24 GB) | — | — |
| RTX PRO 6000 Blackwell (96 GB) | — | — |
| RTX 3080 (10 GB), one card | — | — |
| Two RTX 3080s, one desktop | 58 · 52 °C | 69 · 58 °C |
| Two RTX 3090s, one desktop | — | — |
| RTX 5090 Laptop GPU (24 GB) | — | — |
| RTX 3080 Ti (12 GB) | — | — |

*Every figure in this table is published on [Two 3080s against one 3090, and the cap decides § Power, heat, and the fifty watts](https://research.strata2signal.com/two-used-3080s-priced/#power-heat-and-the-fifty-watts).*

*Held constant: `mistral-small3.2:24b`, Q4_K_M, ollama 0.32.13 with an 8-bit cache, 256 tokens out, each rig at its own largest window held whole: 49,152 tokens for the two 3080s, 65,536 for the single card. Core temperature, the maximum over the run. No page on this shelf prints a memory-die temperature for any of these cards. A cell holding two figures is two cards: its page prints the re-padded Gigabyte board first and the stock EVGA board second, in that order and not by temperature.*

*When: 2026-09-17.*

*Capped per card, as its own page states it: **two RTX 3080s, one desktop** is 250 W each — 500 W between them in the 250 W column, 300 W each — 600 W between them in the 300 W column.*

*Off the ladder — this shelf publishes these too, and no column here can hold a cap given as a range, a sentence or nothing at all, so they are named instead of placed: **RTX PRO 6000 Blackwell (96 GB)** — 594.6 W mean across its first four-minute leg, 87 °C peak across the eight minutes at a cap this page's columns cannot hold, stated as “uncapped in that arm” ([A Short History of Mistral § What does it do on our own machines?](https://research.strata2signal.com/a-short-history-of-mistral/#what-does-it-do-on-our-own-machines)); **RTX PRO 6000 Blackwell (96 GB)** — 83 → 72 °C at a cap this page's columns cannot hold, stated as “600 → 450 W”, at a concurrency of 16, on Ollama 0.32.15 ([What 150 Watts Buys § The answer, in one paragraph](https://research.strata2signal.com/what-150-watts-buys/#the-answer-in-one-paragraph)).*

*RTX 5090 Laptop GPU (24 GB): a 95 W mobile part; its whole envelope sits below every cap column any table on this page holds, and its readings are in the sketch ladder on this page.*

*RTX 3080 Ti (12 GB): measured on two later pages of this shelf, linked above; the record these tables are filled from does not carry those readings yet, so its cells are empty rather than guessed.*

<!-- s2s:compare:END heat-writing -->

<!-- s2s:compare:BEGIN heat-drawing -->

**How hot the core got while drawing, by power cap**

| The card | 250 W | 300 W | 350 W | 400 W | 450 W |
|---|---|---|---|---|---|
| RTX 3090 (24 GB) | 67 °C (a) | — | 74 °C (a) | — | — |
| RTX 3090 Ti (24 GB) | — | 58 °C (a) | 60 °C (a) | 61 °C (a) | 62 °C (a) |
| RTX PRO 6000 Blackwell (96 GB) | — | — | — | — | — |
| RTX 3080 (10 GB), one card | 82 °C (b) | — | — | — | — |
| Two RTX 3080s, one desktop | — | — | — | — | — |
| Two RTX 3090s, one desktop | — | — | — | — | — |
| RTX 5090 Laptop GPU (24 GB) | — | — | — | — | — |
| RTX 3080 Ti (12 GB) | — | — | — | — | — |

**Where each figure comes from.** (a) [RTX 3090 vs RTX 3090 Ti: an inference and diffusion bench § Renders](https://research.strata2signal.com/one-3090-against-one-3090-ti/#renders) · (b) [Two 3080s against one 3090, and the cap decides § Renders: one 3080 against one 3090, seconds per image](https://research.strata2signal.com/two-used-3080s-priced/#renders-one-3080-against-one-3090-seconds-per-image)

*Held constant: Ten minutes of continuous drawing on ComfyUI 0.21.1, the same unnamed graph as the ten-minute drawing table. Core temperature, the peak over the ten minutes; the driver answers `[N/A]` for memory-die temperature on every one of these cards, so no page here prints one.*

*When: 2026-09-17. Some of these arms carry no date in the section their figures were read from, and none is invented here.*

*Which model drew these: the source pages do not name it for these arms, so this page does not either.*

*RTX 5090 Laptop GPU (24 GB): a 95 W mobile part; its whole envelope sits below every cap column any table on this page holds, and its readings are in the sketch ladder on this page.*

*RTX 3080 Ti (12 GB): measured on two later pages of this shelf, linked above; the record these tables are filled from does not carry those readings yet, so its cells are empty rather than guessed.*

<!-- s2s:compare:END heat-drawing -->

<!-- s2s:compare:BEGIN fan-drawing -->

**How hard the fans worked while drawing, by power cap**

| The card | 250 W | 300 W | 350 W | 400 W | 450 W |
|---|---|---|---|---|---|
| RTX 3090 (24 GB) | 58 % (a) | †1 | 67 % (a) | — | — |
| RTX 3090 Ti (24 GB) | — | 72 % (a) | 73 % (a) | 75 % (a) | 75 % (a) |
| RTX PRO 6000 Blackwell (96 GB) | — | — | — | — | — |
| RTX 3080 (10 GB), one card | 92 % (b) | — | — | — | — |
| Two RTX 3080s, one desktop | †2 | †3 | — | — | — |
| Two RTX 3090s, one desktop | — | — | — | — | — |
| RTX 5090 Laptop GPU (24 GB) | — | — | — | — | — |
| RTX 3080 Ti (12 GB) | — | — | — | — | — |

**Where each figure comes from.** (a) [RTX 3090 vs RTX 3090 Ti: an inference and diffusion bench § Renders](https://research.strata2signal.com/one-3090-against-one-3090-ti/#renders) · (b) [Two 3080s against one 3090, and the cap decides § Renders: one 3080 against one 3090, seconds per image](https://research.strata2signal.com/two-used-3080s-priced/#renders-one-3080-against-one-3090-seconds-per-image)

*Held constant: Ten minutes of continuous drawing on ComfyUI 0.21.1, the same unnamed graph again. The figure is the highest fan speed the driver reported, as a percentage of the board's own maximum — a percentage of different fans on different coolers, which is why it is not a ranking of anything.*

*When: 2026-09-17. Some of these arms carry no date in the section their figures were read from, and none is invented here.*

*Which model drew these: the source pages do not name it for these arms, so this page does not either.*

*RTX 5090 Laptop GPU (24 GB): a 95 W mobile part; its whole envelope sits below every cap column any table on this page holds, and its readings are in the sketch ladder on this page.*

*RTX 3080 Ti (12 GB): measured on two later pages of this shelf, linked above; the record these tables are filled from does not carry those readings yet, so its cells are empty rather than guessed.*

**†1** RTX 3090 (24 GB), 300 W: measured, but its page ran it on ollama, not ComfyUI, and on version 0.32.13, not 0.21.1; the record behind this table files its run as 256 tokens out, where this table holds ten minutes of continuous drawing, so it cannot be placed in this row. Its page prints **0 %** there — [Two 3080s against one 3090, and the cap decides § Power, heat, and the fifty watts](https://research.strata2signal.com/two-used-3080s-priced/#power-heat-and-the-fifty-watts)
**†2** Two RTX 3080s, one desktop, 250 W: measured, but its page ran it on ollama, not ComfyUI, and on version 0.32.13, not 0.21.1; the record behind this table files its run as 256 tokens out, where this table holds ten minutes of continuous drawing, so it cannot be placed in this row. Its page prints **0 % · 0 %** there — [Two 3080s against one 3090, and the cap decides § Power, heat, and the fifty watts](https://research.strata2signal.com/two-used-3080s-priced/#power-heat-and-the-fifty-watts)
**†3** Two RTX 3080s, one desktop, 300 W: measured, but its page ran it on ollama, not ComfyUI, and on version 0.32.13, not 0.21.1; the record behind this table files its run as 256 tokens out, where this table holds ten minutes of continuous drawing, so it cannot be placed in this row. Its page prints **64 % · 0 %** there — [Two 3080s against one 3090, and the cap decides § Power, heat, and the fifty watts](https://research.strata2signal.com/two-used-3080s-priced/#power-heat-and-the-fifty-watts)

<!-- s2s:compare:END fan-drawing -->

Three things in those tables are worth a reader's suspicion rather than their trust. A 3090's core ran **hotter** at the lower cap while writing the dense model than at the higher one — 63 °C at 250 W against 57 °C at 300 W — and its page gives the reason: the 250 W arms are ladder arms of eight and six rungs against single points, so the card was working far longer, which is a difference in duration and not in effort. A fan percentage is a percentage of a different fan on every row, so 92 % on one cooler and 75 % on another are not the same air, noise or headroom. And the drawing rows did unequal amounts of work by design; both source pages say plainly that theirs is not a cooler comparison, and neither is this.

No page here prints a memory-die temperature for any of these cards: one bench reports the driver answering `[N/A]` to the question, asked three ways; another reports the sensor not exposed on either board, asked three ways before anything ran; and for the 96 GB card no page here records asking. The honest cell is the one that is not there — and on a 3090, the memory temperature is the one owners watch most.

## One box, two cards {#one-box-two-cards}

One bench put a 3090 and the 96 GB workstation card in the same machine on the same night, at different caps. It is one of three pairings on this page whose cards shared a machine — the others are a 3090 and a 3090 Ti in the same slot of the same desktop on two nights, and two 3080s, one 3080 alone and a 3090 all in one box on one day.

<!-- s2s:compare:BEGIN same-box -->

**One box, two cards, two models — the 2026-09-16 reading**

| The card | The cap it ran at | `gemma4:26b`, decode | `mistral-small3.2:24b`, decode | First token on a prompt it has already seen, dense / mixture |
|---|---|---|---|---|
| RTX 3090 (24 GB) | 250 W | 126.12 tokens a second | 34.13 tokens a second | 98 ms / 167 ms |
| RTX 3090 Ti (24 GB) | — | †1 | †2 | — |
| RTX PRO 6000 Blackwell (96 GB) | 420 W | 205.85 tokens a second | 92.77 tokens a second | 28 ms / 162 ms |
| RTX 3080 (10 GB), one card | — | — | — | — |
| Two RTX 3080s, one desktop | — | †3 | †4 | — |
| Two RTX 3090s, one desktop | — | — | — | — |
| RTX 5090 Laptop GPU (24 GB) | — | — | — | — |
| RTX 3080 Ti (12 GB) | — | — | — | — |

*Every figure in this table is published on [A Short History of Mistral § What we measured on two cards in one box](https://research.strata2signal.com/a-short-history-of-mistral/#what-we-measured-on-two-cards-in-one-box).*

*Held constant: One box, one ollama build — the page states no version — 256 tokens out, the small hours of 2026-09-16 (UTC); its page states no window for these readings, and the quantisation, Q4_K_M, only for the dense model. Its columns are readings rather than caps, because only two cards were in that box. The cap each ran at is its own column, read off the same records as the figures beside it; the first-token column reads dense first, as its page prints it.*

*When: the small hours of 2026-09-16 (UTC).*

*RTX 5090 Laptop GPU (24 GB): a 95 W mobile part; its whole envelope sits below every cap column any table on this page holds, and its readings are in the sketch ladder on this page.*

*RTX 3080 Ti (12 GB): measured on two later pages of this shelf, linked above; the record these tables are filled from does not carry those readings yet, so its cells are empty rather than guessed.*

**†1** RTX 3090 Ti (24 GB), `gemma4:26b`, decode: measured, but its page ran it during 2026-09-17 (14:54 to 23:46 UTC), not the small hours of 2026-09-16 (UTC), so it cannot be placed in this row. Its page prints **145.908 tok/s** there — [RTX 3090 vs RTX 3090 Ti: an inference and diffusion bench § What a power cap takes, and what each card does at its own limit](https://research.strata2signal.com/one-3090-against-one-3090-ti/#what-a-power-cap-takes-and-what-each-card-does-at-its-own)
**†2** RTX 3090 Ti (24 GB), `mistral-small3.2:24b`, decode: measured, but its page ran it during 2026-09-17 (14:54 to 23:46 UTC), not the small hours of 2026-09-16 (UTC), so it cannot be placed in this row. Its page prints **42.969 tok/s** there — [RTX 3090 vs RTX 3090 Ti: an inference and diffusion bench § What a power cap takes, and what each card does at its own limit](https://research.strata2signal.com/one-3090-against-one-3090-ti/#what-a-power-cap-takes-and-what-each-card-does-at-its-own)
**†3** Two RTX 3080s, one desktop, `gemma4:26b`, decode: measured, but its page ran it during 2026-09-17, not the small hours of 2026-09-16 (UTC), so it cannot be placed in this row. Its page prints **114.8 tok/s** there — [Two 3080s against one 3090, and the cap decides § How fast does each one write?](https://research.strata2signal.com/two-used-3080s-priced/#how-fast-does-each-one-write)
**†4** Two RTX 3080s, one desktop, `mistral-small3.2:24b`, decode: measured, but its page ran it during 2026-09-17, not the small hours of 2026-09-16 (UTC), so it cannot be placed in this row. Its page prints **44.6 tok/s** there — [Two 3080s against one 3090, and the cap decides § How fast does each one write?](https://research.strata2signal.com/two-used-3080s-priced/#how-fast-does-each-one-write)

*A note on two of the RTX PRO 6000 Blackwell (96 GB)'s figures, its 205.85 tokens a second on the mixture model and its 92.77 on the dense one: both runs were read on a model server that already had the model loaded when they began, so their own records show how fast the card wrote and not where the model sat — the load that would show that happened before either run, which this workshop's own audit of 2026-09-21 wrote down. The figures are printed here as speeds, and this page claims nothing about the fit.*

<!-- s2s:compare:END same-box -->

## What each one costs, and on what day {#what-each-one-costs-and-on-what-day}

Prices last, because they are the figures that rot fastest. Three kinds of price are kept in three columns, because they are three different claims: what a **new** one is listed at, what a **used or renewed** one is listed at, and what they have actually **sold** for, as a marketplace averages its completed sales. A listing is an ask. A completed-sale average is a record of what people paid.

<!-- s2s:compare:BEGIN prices -->

**What each one costs, and on what day**

| The card | New, as listed | Used or refurbished, as listed | What they sold for |
|---|---|---|---|
| RTX 3090 (24 GB) | $1,829.99 | $1,879.99 | $1,094.44 |
| RTX 3090 Ti (24 GB) | $2,119.99 | — | — |
| RTX PRO 6000 Blackwell (96 GB) | No price printed | No price printed | No price printed |
| RTX 3080 (10 GB), one card | — | $375.00 each | $358.00 each across 318 sales |
| Two RTX 3080s, one desktop | — | $750.00 | — |
| Two RTX 3090s, one desktop | $3,659.98 | — | $2,188.88 |
| RTX 5090 Laptop GPU (24 GB) | — | — | — |
| RTX 3080 Ti (12 GB) | Not found on 1 October 2026 | $479.99 | $475.64 |

**What each reading is, and where it comes from.**

- **RTX 3090 (24 GB)** — new: the cheapest new card of any brand, 13 September 2026 · listed: one renewed listing at a retailer's renewed store, read by hand 17 September 2026 — its page calls it one refurbished card from one seller, not a market price · sold: the marketplace's own average completed sale for August 2026. [The cost to purchase and run a very capable home AI rig § The cards, in May and in September](https://research.strata2signal.com/home-inference-and-diffusion-rig-cost/#the-cards-in-may-and-in-september)
- **RTX 3090 Ti (24 GB)** — new: the card this workshop ordered, 15 September 2026. No used or completed-sale reading for this card on this shelf. [RTX 3090 vs RTX 3090 Ti: an inference and diffusion bench § What it cost](https://research.strata2signal.com/one-3090-against-one-3090-ti/#what-it-cost)
- **RTX PRO 6000 Blackwell (96 GB)** — no price for this card on any page here — it was not bought in a priced window and no page prices it.
- **RTX PRO 6000 Blackwell (96 GB), new** — **read outside this shelf, for this page** — no price is printed for this card, by a standing rule of this workshop: it holds none of record, and an invented one is worse than a blank. Bare boards were listed new on 1 October 2026, each read at its own product page — an NVIDIA RTX PRO 6000 Blackwell Max-Q Edition 96 GB, the 300 W board, sold and shipped by the retailer itself and in stock (item N82E16814132105), and an NVIDIA RTX PRO 6000 Blackwell Workstation Edition 96 GB, the 600 W board, from a third-party marketplace seller and in stock (item N82E16814132106) — alongside whole workstations built around the card (looked on 1 October 2026).
- **RTX PRO 6000 Blackwell (96 GB), used or refurbished** — **read outside this shelf, for this page** — no price is printed for this card, by the same rule. An open-box bare board was listed by the retailer itself on 1 October 2026, in stock and sold as final sale: item N82E16814132105R, the same NVIDIA RTX PRO 6000 Blackwell Max-Q Edition 96 GB, read at its own product page (looked on 1 October 2026).
- **RTX PRO 6000 Blackwell (96 GB), sold** — **read outside this shelf, for this page** — no completed-sale record for this card: a search of the marketplace that supplies the RTX 3080 Ti row's completed-sale figure found neither a price history nor a listing for it on 1 October 2026, and none would be printed in any case, by the same rule (looked on 1 October 2026) — read at a results page, which prints a condition badge and no seller block, so it is not a condition reading. Nothing is printed, by rule.
- **RTX 3080 (10 GB), one card** — listed: one marketplace listing's ask, 17 September 2026 · sold: that marketplace's own average over the twelve months to the same day, a figure that marketplace builds from used and new-condition sales alike. Its page prices no new card. [Two 3080s against one 3090, and the cap decides § What it costs, priced 17 September 2026](https://research.strata2signal.com/two-used-3080s-priced/#what-it-costs-priced-17-september-2026)
- **Two RTX 3080s, one desktop** — listed: the two cards at that day's ask, 17 September 2026 — the figure its own page prices this rig at. [Two 3080s against one 3090, and the cap decides § What it costs, priced 17 September 2026](https://research.strata2signal.com/two-used-3080s-priced/#what-it-costs-priced-17-september-2026)
- **Two RTX 3090s, one desktop** — this row is a card price doubled on its own page, not a price anybody was asked for a two-card machine · new: two of the cheapest new card of any brand, 13 September 2026 · sold: two at August 2026's average completed sale. [The cost to purchase and run a very capable home AI rig § The cards, in May and in September](https://research.strata2signal.com/home-inference-and-diffusion-rig-cost/#the-cards-in-may-and-in-september)
- **RTX 5090 Laptop GPU (24 GB)** — no card price: this part is built into a laptop and is not sold as a card; its own page prices the whole laptop.
- **RTX 3080 Ti (12 GB), new** — **read outside this shelf, for this page** — no new RTX 3080 Ti sold by the retailer itself on 1 October 2026: its own search, filtered to new cards it sells, found none. Every one of the 41 listings it labelled new was a third-party marketplace seller's ask, the cheapest $1,039.00 for an NVIDIA GeForce RTX 3080 Ti Founders Edition 12 GB (items 1FT-0004-006T6 and 1FT-0004-006W2, one seller, both in stock, each read at its own product page); a third listing from that seller at the same price, item 1FT-0004-007S0, is the one the first read, on 18 September, took as new before its product page was opened, and its own page still says Refurbished and spells the memory GDDR6 (looked on 1 October 2026).
- **RTX 3080 Ti (12 GB), used or refurbished** — **read outside this shelf, for this page** — refurbished, the retailer's own refurbished programme, sold and shipped by the retailer under NVIDIA's own refurbished part number 900-1G133-2518-RF2. Read at $479.99 on 18 September, 22 September and 1 October 2026, and out of stock at the last two. Of the 25 refurbished RTX 3080 Tis the retailer itself lists, the one in stock on 1 October 2026 was an EVGA GeForce RTX 3080 Ti FTW3 Ultra Hybrid 12 GB at $599.99 (item 1FT-001K-00F28T, read at its own product page); the ZOTAC GeForce RTX 3080 Ti AMP Holo 12 GB named here on 22 September at $509.99 (item N82E16814500515T) was out of stock too, read at Newegg 1 October 2026. Read at Newegg, 1 October 2026 — *NVIDIA GeForce RTX 3080 Ti Founders Edition Graphics Card 12GB GDDR6X - Titanium and Black* (item 1FT-0004-008K2T, the retailer's own stock, out of stock that minute, read at the listing's own product page).
- **RTX 3080 Ti (12 GB), sold** — **read outside this shelf, for this page** — a used RTX 3080 Ti's average sale price for August 2026 in the marketplace's own price history, the latest month it carried on 1 October 2026 — an average of that month's sales, not a listing, and the same kind of figure from the same marketplace and month as the RTX 3090 row's; its page gives no count for the month and does not say whether shipping is in it, read at Jawa 1 October 2026. Read at Jawa, 1 October 2026 — *NVIDIA GeForce RTX 3080 Ti Price Tracker*.

*Held constant: nothing. A listed price is a reading of one listing on one day, a sold price is a marketplace's own average of completed sales over the period its line names, and the day is in the row beneath. What each page says about tax and delivery is that page's own: the 3080 page prices before tax and before shipping except where a listing charged it, the page that prices a 3090 and two of them prices before tax and before shipping too, the page that prices a 3090 Ti says nothing about either, the listings read for this page are before tax and before delivery, and the completed-sale figure read for it is the marketplace's own, whose page does not say whether shipping is in it. No figure here is an estimate, an average of estimates, or a price this workshop expects to see again.*

<!-- s2s:compare:END prices -->

Two rows are not this shelf's own readings, and the table marks them: the RTX 3080 Ti, whose own readings and order prices are on a later page and not yet in the record these tables read, and the RTX PRO 6000 Blackwell, which no page here prices. What was read for them, and when: the 3080 Ti's three price cells and the 6000's three, outside this shelf, in three reads — the first on 18 September 2026 between 22:41 and 22:42 UTC, at a retailer's search pages; the second on 22 September 2026 between 03:38 and 03:44 UTC, at the listings' own product pages, bar one results page for the 6000; and the last, the first to reach all six cells, on 1 October 2026 between 05:33 and 05:48 UTC, at product pages again and at the marketplace whose published sale history fills the completed-sale cell. Against the 22 September read: the refurbished Founders Edition is still $479.99 and still out of stock; the retailer still sells no new 3080 Ti itself, so that cell says it found nothing and the line beneath names the marketplace asks it did find; and the completed-sale cell, empty until now, reads $475.64, the marketplace's own August 2026 average for a used card — the same kind of figure, from the same marketplace and the same month, as the RTX 3090 row's. The first read took a $1,039.00 marketplace listing as new, and the second did not: the retailer's search pages print a condition badge and no seller block, and that listing's own product page shows *Refurbished* in its spec table, a third-party seller, no manufacturer warranty, and GDDR6 in a title where this card's memory is GDDR6X. It is still listed that way, and the line beneath says so rather than dropping it. For the RTX PRO 6000 Blackwell all three cells read *No price printed*, and that is a rule rather than a failed search: this workshop holds no price of record for that card, an invented one is worse than a blank, and a listing read for it is not printed either, by the same rule. The first read recorded that the shop listed only whole computers built around it, which the second read found untrue — bare boards were listed too, and the third read opened their product pages — so the record keeps that sentence as withdrawn rather than quietly fixing it.

Two more things the column headers cannot say on their own. The *sold* column holds two kinds of average, and each line beneath says which: a month's — August 2026, for a 3090 and a 3080 Ti, used cards only — and a year's, the twelve months to 17 September 2026 for a 3080, a figure that marketplace builds from used and new-condition sales alike. And a 3090 Ti's figure under *new, as listed* is what this workshop paid for its card on the day it ordered it, not a listing read later; its line says so.

One price that used to be on this shelf is deliberately absent. A used-3090 figure of $908.00 — a trailing twelve-month average — was withdrawn by its own page on 2026-09-18 as a price no reader could pay in a market that had since climbed. Only figures still of record can reach a table here, so a withdrawn one cannot come back through one. The 3080's twelve-month figure stays because its page files it as a supporting figure rather than a price, beside the asks it read that day.

## What this comparison cannot say {#what-this-comparison-cannot-say}

- **These are not one bench, but more of it shared a box than you might expect.** Three pairings did: two 3080s, one of them alone and a 3090, all in the same box, the same slot and the same supply on one day; a 3090 and a 3090 Ti in the same slot of the same desktop on two nights; and a 3090 with the 96 GB card in one box on one night. Two of those pairings used the same 3090 in the same runs, and both pages cite the same result files; whether the 96 GB card's box held that board too is not something these pages settle — see the note beside it above.
- **Five different states of the model server, not one.** The cap ladders ran on ollama 0.32.13. The one-box reading says only “one ollama build” and names no version. The split-model readings name no server at all. The two figures named under the writing tables as off the ladder ran on “the model server of that day”, which no page identifies. And the 96 GB card's reading named under the heat table ran on ollama 0.32.15, with sixteen callers at once.
- **Every writing figure here whose server is named is ollama's.** Other model servers, and other ways of splitting one model across two cards, are not compared here.
- **Every watt here is card board power.** No wall figure in any table above, and none estimated. A room pays more than these numbers.
- **No memory-die temperature, on any card.**
- **A fan percentage is not comparable across coolers**, and the ten-minute rows did unequal work.
- **The RTX 3080 Ti has no reading in any table here.** It was measured on two later pages, and the record behind these tables has not carried those readings; here it has two prices read outside the shelf — a refurbished listing and a completed-sale average — and one price cell that says why it is empty.
- **This page names no winner.** Each source page ruled its own comparison in its own words; those rulings stay there and are not restated here as if this page had made them.
- **A dagger is not a small number.** It is the absence of a comparable one.

## What to take with you {#what-to-take-with-you}

- **A power cap does more to a dense model than the card does.** One RTX 3090, on one day, went from 33.3 to 47.8 tokens a second between a 250 W cap and a 300 W one — a bigger step than any card-to-card gap at the same cap in the same table. (The table's own 300 W cell prints 47.838 tokens a second from the deeper bench; the 3080s page prints the same run to one decimal, 47.8, and both pages cite the same result files for it.) Those two readings are not a straight line, and two of this shelf's pages mark two points on it: fifty watts, 250 to 300, were worth 43.8 per cent to this 3090 on the dense model, its page's own figure, while the rig page that runs every 3090 at 280 W measured that cap, on an earlier day and on the model server of that day, at 3 per cent of writing speed against the board's own 350 W. The deeper bench also says what the cap starves: at 250 W this 3090 streamed 55.4 per cent of the ceiling its memory clock implies, at 300 W 79.6 per cent.
- **A mixture model barely feels the cap.** That same 3090's mixture figures move only from 125.6 to 136.618 tokens a second across every cap it was held at — 250, 300 and 350 W — and this shelf's 3090 Ti's from 145.908 to 147.823 tokens a second and then back to 147.560 at its own limit.
- **Two smaller cards can beat one bigger one, right up until the model will not fit.** At a 49,152-token window, two RTX 3080s write the dense model at 44.6 tokens a second against one 3090's 33.1 — at 250 W a board, the cap their page says costs the single 3090 about 30 per cent on this model and the pair 0.2 per cent — and the pair reached its figures only with the layer count named by hand. Each board is capped at 250 W, so the pair is 500 W between them against one card's 250 W, and at each rig's own largest window, 49,152 tokens for the pair and 65,536 for the single card, it drew 466 W against 235 W. Ask for a 65,536-token window and the pair's page prints *does not fit*, while one 3090 keeps going at 33.3 tokens a second.
- **Memory is the thing you are buying, and it does not appear in any speed.** One 10 GB card cannot hold a 26-billion-parameter model at all. Split with the computer's own RAM, one wrote at 34.40 tokens a second — on a bench its own page files as a control rather than a like-for-like reading, which is why no table here divides it into anything.
- **Drawing pictures scales with the card in a way writing does not.** On the eight-step sketch graph, one 3080 at a 250 W cap takes 2.074 s a picture, one 3090 at the same cap 1.683 s, and a 3090 Ti at 450 W 1.369 s; at a shared 300 W, a 3090 Ti's 1.455 s against a 3090's 1.560 s.
- **A temperature without its workload says little, and these say nothing about coolers.** A 3080 peaked at 82 °C with its fan at 92 %; a 3090 Ti peaked at 62 °C with its fan at 75 %. Those two runs drew different numbers of pictures in different cases, and both source pages say in their own words that this is not a comparison of coolers.
- **Cheap cards are cheap, and that is the whole argument for them.** On 17 September 2026 a used 3080 asked $375.00 and a renewed 3090 asked $1,879.99; 10 GB 3080s had averaged $358.00 across 318 completed sales over the preceding year, by a marketplace figure that counts new-condition sales too.

## How to check our work {#how-to-check-our-work}

Three doors. Every table links the page and the section each figure came from, so a cell can be checked against its page in one click. The [hardware roster](https://research.strata2signal.com/hardware/) shows every record these tables are filled from, card by card, with its conditions and the page it was read from. And if a figure here does not reproduce, say so at the [contact desk](https://strata2signal.com/contact/).

Every figure these tables use lives in that roster as a record: the card, the figure exactly as its page printed it, and the conditions that page stated — model, quantisation, model server and version, context window, cap, tokens or steps, callers, and the window it was measured in. A small program selects from those records and writes the tables above; it computes nothing, averages nothing and ranks nothing. Each table declares what it holds constant, and a cell is filled only by a record that satisfies every one of those conditions — which is why the daggers exist, and why a check can refuse this page when a cell and the record stop agreeing.

One more page feeds that record and none of these tables. [Two 3090s out, one 3090 Ti in](https://research.strata2signal.com/two-3090s-out-one-3090-ti-in/) put twenty records into the roster on 2026-09-19 — a note on the day two 3090s came out of service and one 3090 Ti went in — and not one of them fills a cell here or raises a dagger. The reason is the unit, not the card: that note publishes seconds a batch of three plates and milliseconds a sleeve — the pictures its two services actually draw, by their own names — while every drawing table here is keyed on seconds a finished picture, pictures finished in ten minutes, a core temperature or a fan percentage. A figure in a different unit is a different figure, which is the rule these tables are built on. What that note re-prints — a 3090 Ti's 60 °C peak and a 3090's 74 °C at the shared 350 W cap, and their fans at 73 % against 67 % — was already here, credited to the bench that measured it: the roster carries a reading once, from the page that took it, and what the note measured that is new lives on the roster's own card pages.

## The rest of the seminar {#the-rest-of-the-seminar}

- [RTX 3090 vs RTX 3090 Ti: an inference and diffusion bench](https://research.strata2signal.com/one-3090-against-one-3090-ti/) — the two cards in the same slot, four caps, the deepest source here.
- [Two 3090s out, one 3090 Ti in](https://research.strata2signal.com/two-3090s-out-one-3090-ti-in/) — the move that followed that bench: what one 3090 Ti did for two live services on their first evening on it, and why none of its figures could reach a table here.
- [Two 3080s against one 3090, and the cap decides](https://research.strata2signal.com/two-used-3080s-priced/) — the pair, the single 3080, and the fifty watts.
- [A Short History of Mistral](https://research.strata2signal.com/a-short-history-of-mistral/) — where the one-box reading and the spilled-model figures come from.
- [The cost to purchase and run a very capable home AI rig](https://research.strata2signal.com/home-inference-and-diffusion-rig-cost/) — where its 3090 prices come from, and the correction that withdrew one of them.
- [What 150 Watts Buys](https://research.strata2signal.com/what-150-watts-buys/) — the 96 GB card's own cap ladder. No table here could join it to the rest, because its caps are written as ranges; one of its readings is named under a table all the same.
- [A Rig Your Friend Already Owns](https://research.strata2signal.com/a-rig-your-friend-already-owns/) — the sketch ladder: a 3090, the 96 GB card and the laptop part at three sketch sizes, with no cap stated; the one table here whose rows include the laptop part.
- [A laptop, asked the desktop's questions](https://research.strata2signal.com/a-laptop-asked-the-desktops-questions/) — the laptop part's own bench: this page's writing and drawing questions held to 95 W and under Dynamic Boost, every reading below the lowest cap column here.
- [RTX 5080 vs RTX 3080 Ti](https://research.strata2signal.com/rtx-5080-vs-rtx-3080-ti/) — the 3080 Ti Founders Edition measured, with what it cost, beside a 16 GB card this page has no row for.
- [The same eighteen pictures, card by card](https://research.strata2signal.com/the-same-eighteen-pictures-card-by-card/) — the 3080 Ti, the 96 GB card and the laptop part among six boards drawing the same eighteen kinds of picture on one software stack, each desktop card at 300 W and at its stock limit (a 3090 Ti at its everyday 350 W instead), and the laptop part under Dynamic Boost.
- [The hardware roster](https://research.strata2signal.com/hardware/) — every figure above, with every condition, card by card.

## Who ran this, and thanks {#who-ran-this-and-thanks}

The figures are six pages' work and not this one's. **[RTX 3090 vs RTX 3090 Ti](https://research.strata2signal.com/one-3090-against-one-3090-ti/)** supplies most of the ladders; **[Two 3080s against one 3090](https://research.strata2signal.com/two-used-3080s-priced/)** the pair, the single 3080 and the cheap end; **[A Short History of Mistral](https://research.strata2signal.com/a-short-history-of-mistral/)** the one-box reading and the spilled model; **[the rig cost page](https://research.strata2signal.com/home-inference-and-diffusion-rig-cost/)** its 3090 prices; **[What 150 Watts Buys](https://research.strata2signal.com/what-150-watts-buys/)** the 96 GB card's own temperature reading, which appears here named under a table rather than in one, because its cap is a range; and **[A Rig Your Friend Already Owns](https://research.strata2signal.com/a-rig-your-friend-already-owns/)** the sketch ladder, the one table whose rows include the laptop part. **[Two 3090s out, one 3090 Ti in](https://research.strata2signal.com/two-3090s-out-one-3090-ti-in/)** put twenty more records into the hardware roster this page reads from on 2026-09-19 and is thanked for the record rather than for a figure: none of them could fill a cell here, for the reason given above. The two prices this page took for itself were read at Newegg's product pages and in Jawa's published sale history, on 1 October 2026, by plain fetches of public pages.

**EVGA** made the 24 GB boards measured here and one of the 10 GB ones, and published the specifications those pages checked themselves against; **Gigabyte** made the other 10 GB board, the one the single-card row's drawing and heat readings come from; **NVIDIA** published the launch figures, makes the Founders Edition named above, and makes the driver every temperature, clock and watt on this page came out of.

The writing was done on **ollama**, and under it **llama.cpp** and **NVIDIA's CUDA**; the pictures on **ComfyUI**. The models are Google's **Gemma 4**, **Mistral Small 3.2** from Mistral AI, Black Forest Labs' **FLUX.2 Klein 4B** and its decoder, and **Qwen** from Alibaba Cloud as the text encoder in those graphs. Each source page names the licence it found for the model it ran; this page adds none of its own, because it ran nothing.

A small human team owns the cards, set the rules and the refusals before the runs, and signed the figures on each source page; a fleet of AI agents ran the harnesses, filled these tables from the record, and wrote the check that refuses this page when a cell and the record disagree. Thanks to the readers who kept asking the question this page exists to answer: *how does that one compare?*

<!-- derived 2026-10-01 (UTC) by tools/derive_md.py from the pour source.
     source html sha256: 4b4549f6ac986c474d6d51a7f5cacd30888225fa752efeafad7202e327245120
     derivation sha256:  81bc5adddde65e5a3263fe0a35f4e474995bf4b42caa9c695d72b277389b7747
     the {#id} on each heading is the anchor that heading carries on the page. -->
