# RTX 3090 vs RTX 3090 Ti

*One EVGA GeForce RTX 3090 XC3 Ultra 24 GB and one EVGA GeForce RTX 3090 Ti FTW3 Ultra 24 GB, in the same x16 slot of the same desktop, on the same supply and the same script, 17–18 September 2026. Numbers only: what each writes, what each holds, what a picture costs, how hot each ran, and what a power cap takes away. Run as they come, this Ti takes both models; held at 300 watts it loses the dense one, which is not a fact about the boards — 300 watts is 86 per cent of what one card was built to draw and 67 per cent of what the other was.*

*Published 2026-09-18 (UTC) · A small (human) team and a fleet of AI agents.*

**the short version:** This 3090 Ti holds its memory at 10,251 MHz and a 3090 at 9,501 — 7.9 per cent more bandwidth, present at every setting, but whether the work can reach it is decided by a number nobody sets: held at 300 W a Ti pulls 66.3 per cent of the ceiling that clock implies, at its own 450 W limit 88.5. Run each board at the limit it ships with and this Ti is 8.0 per cent faster on a mixture model, 22.4 on a dense one and 9 to 12 per picture, for 17.7 and 53.2 per cent more board power. It cost $2,119.99 new; a renewed listing of the exact 3090 read $1,879.99.

6,236 words · about 28 minutes (at 220 words/min) · 10 tables · data kit: no

https://research.strata2signal.com/one-3090-against-one-3090-ti/

---

## The two cards, named once {#the-two-cards-named-once}

The two boards are an **EVGA GeForce RTX 3090 XC3 Ultra Gaming 24 GB** and an **EVGA GeForce RTX 3090 Ti FTW3 Ultra Gaming 24 GB** — the same maker, the same 24 GB, the same x16 slot of the same desktop on two nights — so **a 3090** and **a 3090 Ti** everywhere below are those two boards and no others. **Both are stock**: factory thermal pads, factory cooler, factory BIOS, nothing re-padded, re-shrouded or flashed. That is deliberate, and deliberately unlike [the pair of 3080s this workshop priced](https://research.strata2signal.com/two-used-3080s-priced/), where one of the two cards had been re-padded by hand. **The only difference between these two units is age and use.**

**This Ti was new.** Ordered 15 September 2026, unboxed on the 17th, first powered on for this bench — and given **no burn-in**: the first hours of load in its life are the four cap blocks on this page, its language arms 19:07 to 23:46 UTC and its renders 19:41 to 00:01. Every Ti figure here is a new unit's first day, with whatever a reader takes that to mean.

**This 3090 is a used board bought in 2021, on its third life**, and it still carries its original thermal pads: it mined Ethereum, it flew Flight Simulator, and now it serves language models, draws pictures and trains. Ethereum stopped paying for that kind of mining in September 2022, so the first of the three is over and dated. The ten-minute temperature rows further down are therefore a 2021 board on factory pads against a board that was days old.

**And the Ti is out of production.** The GeForce RTX 3090 Ti went on sale on **29 March 2022 at $1,999** — [NVIDIA's own announcement that day](https://www.nvidia.com/en-us/geforce/news/geforce-rtx-3090-ti-out-now/) calls it “24GB of the fastest 21Gbps GDDR6X memory” — the last and fastest of the GeForce RTX 30 line. On **16 September 2022**, less than six months later, EVGA told two interviewers it was ending its partnership with NVIDIA and would build no further graphics cards, selling through the RTX 30 stock it had. **It published no last-built date**, so how long this board kept being made is not on the record and this page does not say; what is on the record is that nobody has built one since. The unit measured here is new, boxed stock of a 2022 card, bought for **$2,119.99** on 15 September 2026 — $120.99 above its own launch price, four and a half years later.

| label | the board | its own default limit, read from the driver | each cap as a fraction of that |
|---|---|---|---|
| **a 3090** | one EVGA GeForce RTX 3090 XC3 Ultra 24 GB, 24,576 MiB | **350 W** (the driver's `max_limit` is 366 W) | 250 W = **71 %** · 300 W = **86 %** · 350 W = **100 %** |
| **a 3090 Ti** | one EVGA GeForce RTX 3090 Ti FTW3 Ultra 24 GB, 24,564 MiB | **450 W** (`max_limit` 480 W) | 300 W = **67 %** · 350 W = **78 %** · 400 W = **89 %** · 450 W = **100 %** |

Both limits are read off the card in every result file, never from a specification.

## At a glance {#at-a-glance}

*Both cards, one slot, two nights. **The run-as-it-comes rows are for a reader who runs a card the way it ships** — 350 W against 450 W; the rows that say “both at 300 W” are the controlled comparison, and the asymmetry inside that cap is the page: **300 W is 86 per cent of a 3090's own 350 W limit and 67 per cent of a Ti's 450 W**. Rows marked ↓ are won by the smaller number; every figure repeats below with its cap and its receipt file.*

| | a 3090 | a 3090 Ti | winner |
|---|---:|---:|---|
| **run as it comes — mixture model, tokens a second at 131,072** | 136.618 *(350 W)* | **147.560** *(450 W)* | **a 3090 Ti, by 8.0 %** |
| **run as it comes — dense model, tokens a second at 65,536** | 46.880 *(350 W)* | **57.396** *(450 W)* | **a 3090 Ti, by 22.4 %** |
| mixture model, tokens a second at 131,072 — both at 300 W | 136.139 | **145.908** | a 3090 Ti, by 7.2 % |
| dense model, tokens a second at 65,536 — both at 300 W | **47.838** | 42.969 | a 3090, by 11.3 % — **and the two rows above reverse it**; one cap step is the whole difference |
| ↓ one picture, 8 steps, 512 × 512, three at a time — each run as it comes | 1.509 s *(350 W)* | **1.369 s** *(450 W)* | **a 3090 Ti, by 9.3 %** — 0.907× the seconds |
| ↓ the same picture — both at 300 W | 1.560 s | **1.455 s** | a 3090 Ti, by 6.7 % — 0.933× the seconds |
| four conversations at once, tokens a second — each run as it comes | 294.636 *(350 W)* | **337.620** *(450 W)* | **a 3090 Ti, by 14.6 %** — its widest margin here |
| memory clock held under load | 9,501 MHz | **10,251 MHz** | a 3090 Ti, by 7.9 % — and neither moved at any cap |
| largest window held whole, mixture model | 131,072 | 131,072 | a tie, at every cap either card ran |
| largest window held whole, dense model, layer count named | 98,304 | 98,304 | a tie — and the window rows do not move with a cap |
| ↓ energy, mixture model, joules per 1,000 tokens — each run as it comes | **1,758.13** *(350 W)* | 1,915.13 *(450 W)* | a 3090, 8.2 % less — but at the shared 300 W the Ti is 6.9 % less |
| ↓ energy, dense model, joules per 1,000 tokens — each run as it comes | **5,842.92** *(350 W)* | 7,295.49 *(450 W)* | a 3090, 19.9 % less — and 7.9 % less at the shared 350 W |
| ↓ core temperature over ten minutes of continuous drawing, each run as it comes | 74 °C *(350 W, 311 images)* | **62 °C** *(450 W, 348 images)* | **a 3090 Ti, twelve degrees cooler** — while drawing 31 W more and making 37 more images, so not equal work |
| ↓ the card itself, priced | no price on record; a renewed listing of this exact board read **$1,879.99** on 17 September 2026 | $2,119.99, new, ordered 15 September 2026 | *not a throughput row* — **a 3090 costs $240.00 less** on those two readings, and one is a renewed listing against one new card |
| ↓ what the $240 buys, per $100 of price — dense model, each run as it comes | 2.49 tok/s *(350 W)* | **2.71 tok/s** *(450 W)* | **a 3090 Ti, by 8.8 %** — the only per-dollar row it wins; a 3090 takes the rest in “What the $240 buys” below |

## The question {#the-question}

Two cards a generation apart in name and eighteen months apart in age, both holding 24 GB, both in the same slot within a day and a half. A 3090 Ti's argument over a 3090 is **memory bandwidth**: 21 Gbps per pin against 19.5, over the same 384-bit bus — 1,008 GB/s against 936 on paper. **Both boards held their memory one bin under their own reported maximum throughout**: 9,501 MHz against a 9,751 MHz reading, 10,251 against a reported 10,501. So every ceiling below (912.1 and 984.1 GB/s) is a ceiling at the clock these two units actually held, about 2.5 per cent under those paper figures. Writing a token on a quantised model waits on how fast a card streams its own weights, so if that argument is real it has to appear here and nowhere else.

A **dense** model uses all of itself for every word and a **mixture** wakes only a few billion, which at 300 W is why the mixture writes 2.8 times a 3090's dense rate and 3.4 times this Ti's. The **cap-active flag** below is the driver's report of whether the cap held the clock down; mean draw ÷ cap, beside it, says how hard the card actually pushed. **Answer quality is not measured here at all.**

## The instrument, and the control {#the-instrument-and-the-control}

One harness produced every language figure: a frozen 2,101-byte prompt checked by sha256, 256 tokens out, temperature and seed at zero, one discarded warm-up then three scored runs, the median reported, with **card board power** read from the driver at 2 Hz and never the wall. The models are `gemma4:26b` (26 billion parameters, ~4 billion active per token) and `mistral-small3.2:24b` (24 billion dense), the same blobs by digest on both benches, Q4_K_M, on ollama 0.32.13 with an 8-bit cache; the pictures came from ComfyUI 0.21.1 on three certified graphs identical by hash on both sides (`79df5a8b…`, `7dd85f01…`, `94bfd633…`). **The control is the slot** — one card out, the other in, one box, one case. An image service sat on the card throughout, unavailable to every arm: 262 MiB of the Ti and 256 MiB of the older board when the render bench read it, 304 and 298 MiB when the language bench did.

## What a power cap takes, and what each card does at its own limit {#what-a-power-cap-takes-and-what-each-card-does-at-its-own}

*The two caps both cards were measured at, then each card at its own limit. The 300 W columns and all three of the Ti's higher columns ran 2026-09-17 (14:54 to 23:46 UTC); a 3090's own 350 W column ran 2026-09-18, 00:22 to 00:54 UTC, after the card went back in the slot. Receipts: `3090-m5-one-300w.json`, `3090-m4-one-300w.json`, `3090-m5-one-350w-auto-ladder.json`, `3090-m4-one-350w-auto-ladder.json`, and `3090ti-m5-one-<cap>w-auto-ladder.json` / `3090ti-m4-one-<cap>w-auto-ladder.json` at 300, 350 and 450 W.*

| | a 3090, 300 W *(86 % of its limit)* | a 3090 Ti, 300 W *(67 %)* | a 3090, 350 W *(its own limit)* | a 3090 Ti, 350 W *(78 %)* | a 3090 Ti, 450 W *(its own limit)* |
|---|---:|---:|---:|---:|---:|
| mixture, tok/s at 131,072 | 136.139 | **145.908** | 136.618 | **147.823** | 147.560 |
| dense, tok/s at 65,536 | **47.838** | 42.969 | 46.880 | **53.354** | **57.396** |
| mean draw, mixture | 241.53 W | 240.58 W | 240.49 W | 282.64 W | 283.01 W |
| mean draw, dense | 289.86 W | 296.64 W | 273.35 W | 338.05 W | 418.76 W |
| draw as a share of the cap, dense | 96.6 % | 98.9 % | **78.1 %** | 96.6 % | 93.1 % |
| the cap flag, dense | active in 100 % of samples | active in 100 % | active in 100 % | active in 100 % | active in 100 % |
| the cap flag, mixture | active in 100 % | active in 83.3 % | active in 83.3 % | **0 %** | **0 %** |
| core clock median, dense | 1,275 MHz | **960 MHz** | 1,215 MHz | 1,305 MHz | 1,965 MHz |
| ↓ energy, dense, J/1,000 tokens | 6,059.20 | 6,902.38 | **5,842.92** | 6,343.81 | 7,295.49 |
| ↓ energy, mixture, J/1,000 tokens | 1,771.71 | **1,648.85** | **1,758.13** | 1,909.96 | 1,915.13 |

**Read the 350 W columns first: that is a 3090 at the limit it ships with.** There the Ti wins both models, by **13.8 per cent** on the dense one and **8.2** on the mixture; at each card's own limit, by **22.4** and **8.0**. Only at 300 W, below both limits, does the dense row turn over — and the two energy rows do not agree with each other. **On the dense model a 3090 spends less per thousand tokens at both shared caps** — 12.2 per cent less at 300 W and 7.9 at 350, where it is also the slower board. **On the mixture the answer changes with the cap**: at the shared 350 W a 3090 spends 7.9 per cent less, and at the shared 300 W the Ti spends **6.9 per cent less**. Cheaper to run is a property of the pairing, not of the board. **"Cap active" is not "pinned at the limit":** the driver called it active for 100 per cent of a 3090's 350 W dense arm while the card drew **78.1 per cent** of what it was allowed, which is why every table prints the flag and the draw together.

*The dense model's whole ladder on both cards, from each card's `m4-one-<cap>w-auto-ladder.json` files and `3090-m4-one-300w.json`. The **ceiling** is derived from the memory clock observed at that rung and the bus width — a ceiling at an observed clock, never an achieved rate; the **streamed** figure is what the decode really pulled. Both are dense-only, because a mixture reads a subset of its weights per token and neither bench has a receipt for which. **A 3090's two bandwidth columns are not written by that bench's own per-arm harness, which predates the instrument** — they are computed from what its files do carry (the median decode, a memory clock whose floor and max are both 9,501 MHz, the model's 15,177,384,862-byte store, and a **384-bit bus read live from NVML**, receipted in `3090-measured-constants-350w.json`) by the formula that file records, and the same six cells appear in that bench's own 350 W addendum. **The Ti's 384 bits is declared rather than probed.** **The `arm shape` column is load-bearing**: a step into or out of a single-window repeat is not a cap effect alone. A 3090's 250 W arms predate the cap-in-the-name convention and are the un-suffixed files — `3090-m4-one-auto-ladder.json` here, `r4.json` in the ten-minute table and `3090-concurrency-one.json` in the concurrency one.*

| card and cap | % of its own limit | arm shape | mean draw | % of the cap drawn | cap flag | memory clock | ceiling at that clock | streamed | **% of the ceiling** | dense tok/s |
|---|---:|---|---:|---:|---:|---:|---:|---:|---:|---:|
| a 3090, 250 W | 71 % | ladder rung | 235.38 W | 94.2 % | active 100 % | 9,501 MHz | 912.1 GB/s | 504.9 GB/s | **55.4 %** | 33.265 |
| a 3090, 300 W | 86 % | **single-window repeat** | 289.86 W | 96.6 % | active 100 % | 9,501 MHz | 912.1 GB/s | 726.1 GB/s | **79.6 %** | 47.838 |
| a 3090, 350 W | 100 % | ladder rung | 273.35 W | 78.1 % | active 100 % | 9,501 MHz | 912.1 GB/s | 711.5 GB/s | **78.0 %** | 46.880 |
| a 3090 Ti, 300 W | 67 % | ladder rung | 296.64 W | 98.9 % | active 100 % | 10,251 MHz | 984.1 GB/s | 652.2 GB/s | **66.3 %** | 42.969 |
| a 3090 Ti, 350 W | 78 % | ladder rung | 338.05 W | 96.6 % | active 100 % | 10,251 MHz | 984.1 GB/s | 809.8 GB/s | **82.3 %** | 53.354 |
| a 3090 Ti, 400 W | 89 % | ladder rung | 382.16 W | 95.5 % | active 100 % | 10,251 MHz | 984.1 GB/s | 864.1 GB/s | **87.8 %** | 56.936 |
| a 3090 Ti, 450 W | 100 % | ladder rung | 418.76 W | 93.1 % | active 100 % | 10,251 MHz | 984.1 GB/s | 871.1 GB/s | **88.5 %** | 57.396 |

**Within a card the ceiling column is one number, repeated** — neither memory clock moved, on any arm. What moved was the reach: a Ti goes from **66.3 per cent of its ceiling at 300 W to 88.5 at 450**, under core clocks of 960 and 1,965 MHz. **A power cap does not slow the memory down; it slows the core down, and a core with no power cannot drain a fast bus.** **A 3090's own last rung buys nothing, and the step that looks like a loss is not a cap effect.** Its dense rung at 350 W reads **2.0 per cent slower** than at 300 W on spreads that do not overlap (46.783 to 47.105 against 47.835 to 47.898) — a separated regression, not noise — but the two rows are not the same arm: a single-window repeat on a cold card against a ladder rung on a card eight degrees warmer, drawing **16.5 W less** at 78.1 per cent of its cap. Its mixture gains **0.35 per cent**. **A 3090 has no reason to sit below its 350 W default and nothing to gain from it either.**

**And the same fifty watts are a starvation to one workload and an irrelevance to another, on one card.** At 350 W a 3090's dense model drew **273.35 W — 78.1 per cent of its cap** — and gained nothing over 300. At that identical 350 W, ten minutes of continuous drawing drew a median of **325.8 W with the cap flag active in 100 per cent of 1,153 samples** (`r4-350w.json`, one picture at a time), and on the render ladder the pictures got faster at every step up to it: 1.683 → 1.560 → **1.509 seconds** at batch three (`r1-250w.json`, `r1-300w.json`, `r1-350w.json`). Same board, same setting, two jobs — one wants every watt, the other has stopped asking. **A cap is not a property of a wattage. It is a property of a card at a workload** — those two files, `3090-m4-one-350w-auto-ladder.json` and `r4-350w.json`, are the whole of that claim.

## How a cap is set {#how-a-cap-is-set}

Every cap on this page was set with the driver's own tool, one card at a time, and read back in the same call that set it:

```
nvidia-smi --query-gpu=power.default_limit,power.max_limit --format=csv   # what this board allows
nvidia-smi -pm 1                                                          # persistence mode
nvidia-smi -i <index> -pl <watts>                                         # the cap
```

**Persistence mode matters**: without it the limit goes back to the card's default at the next reboot, so the durable form is a boot-time unit that sets it again. A value below the board's own minimum is refused rather than clamped, and the two ends of the allowed range are what the first command prints — 350 W and 366 W for one of these boards, 450 W and 480 W for the other. **This is a power limit, not an undervolt**: it caps what the card may draw and lets the driver find the clock, and nothing on this page touched a voltage curve.

## The two models want different caps {#the-two-models-want-different-caps}

*A 3090 Ti, both models at their headline windows, all four caps. The step is against the rung below it in watts. Receipts: `3090ti-m5-one-<cap>w-auto-ladder.json` and `3090ti-m4-one-<cap>w-auto-ladder.json`, four caps each.*

| cap | % of its own limit | mixture tok/s | step | mixture J/1,000 | dense tok/s | step | dense J/1,000 |
|---|---:|---:|---:|---:|---:|---:|---:|
| 300 W | 67 % | 145.908 | — | **1,648.85** | 42.969 | — | 6,902.38 |
| 350 W | 78 % | **147.823** | +1.31 % | 1,909.96 ⚑ | 53.354 | **+24.17 %** | **6,343.81** |
| 400 W | 89 % | 147.599 | −0.15 % | 1,723.81 ⚑ | 56.936 | +6.71 % | 6,713.05 |
| 450 W | 100 % | 147.560 | −0.03 % | 1,915.13 | **57.396** | +0.81 % | 7,295.49 |

**The mixture is finished at 350 W; the dense model is not** — everything above 350 W on the mixture sits inside the bench's own run-to-run spread, while the dense model is still buying 6.71 per cent at 400, and each model's cheapest rung sits one below its fastest. **A rule written before the data** — *the lowest rung within 1.0 per cent of every higher one, refused outright if the within-rung spread is wider than that* — names **350 W** as the mixture's knee and **names no knee at all for the dense model**, whose 300 W spread is 1.925 per cent.

**⚑ Two watt figures above may not be read as cap effects.** Between 350 and 400 W the mixture drew **28.4 W less** at identical core clocks and identical decode, because the fan read 49 per cent at one rung and 0 at the other, and board power contains the fans. The bench flags the step and corrects nothing.

## Renders {#renders}

*Three certified graphs at batch 3, seconds per finished image, median of two timed submits. **Two pairs of columns are like for like — 300 W against 300 W and 350 against 350**; the rest are each card at its own settings, labelled so. Receipts: `r1-<cap>w.json` on each card's own render bench.*

| graph, batch 3 | a 3090 Ti, 300 W *(67 %)* | 350 W *(78 %)* | 400 W *(89 %)* | 450 W *(100 %)* | a 3090, 250 W *(71 %)* | 300 W *(86 %)* | 350 W *(100 %)* |
|---|---:|---:|---:|---:|---:|---:|---:|
| `klein-4b-8st` | 1.455 s | 1.397 s | **1.363 s** | 1.369 s | 1.683 s | 1.560 s | 1.509 s |
| `klein-4b-8st-pixel4` | 1.436 s | 1.379 s | 1.351 s | **1.350 s** | 1.721 s | 1.579 s | 1.528 s |
| `klein-4b-32st` | 5.332 s | 5.130 s | 5.015 s | 5.015 s | 6.297 s | 5.827 s | 5.670 s |

**The Ti takes every graph at both shared caps, and by more at the higher one** — 0.933×, 0.909× and 0.915× of the seconds at 300 W, 0.926×, 0.902× and 0.905× at 350, and **0.907×, 0.884× and 0.884× run as they come**: 9 to 12 per cent off every picture. The core clocks differ too — 2,048 against 1,875 MHz median on the 8-step graph run as they come, 1,920 against 1,785 at the shared 300 W — so this is not single-variable.

**The Ti's ladder saturates at 400 W**: its cap flag over the ten-minute arm (`r4-<cap>w.json`) collapses from 92.9 per cent of samples at 300 W and 87.9 at 350 to **16.0 at 400 and 14.5 at 450** — and the render ladder agrees, 0.94 of samples at 300 and 350 W against 0.00 to 0.06 at 400 and 450 — and the last 50 W bought 0.006 s **slower** on the 8-step graph and one more image in ten minutes. A 3090, by contrast, is still power-limited at its own limit — cap flag 100 per cent on all three graphs at 350 W, drawing 336 to 338 W of the 350 allowed.

*Ten minutes of continuous drawing, one card and one cap a row. **Not a cooler comparison** — the rows did different amounts of work at different caps, and the image counts are printed so a reader can see that. Receipts: `r4-<cap>w.json`.*

| card and cap | % of its own limit | images | s/image | plateau | peak | fan max | median draw | cap flag |
|---|---:|---:|---:|---|---:|---:|---:|---:|
| a 3090 Ti, 300 W | 67 % | 320 | 1.811 | 56 °C at 0.7 min | 58 °C | 72 % | 292.0 W | 92.9 % |
| a 3090 Ti, 350 W | 78 % | 331 | 1.715 | 55 °C at 0.4 min | 60 °C | 73 % | 316.2 W | 87.9 % |
| a 3090 Ti, 400 W | 89 % | 347 | 1.676 | 59 °C at 0.3 min | 61 °C | 75 % | 355.4 W | 16.0 % |
| a 3090 Ti, 450 W | 100 % | **348** | **1.672** | 57 °C at 0.3 min | **62 °C** | 75 % | 357.1 W | 14.5 % |
| a 3090, 250 W | 71 % | 269 | 2.156 | 65 °C at 0.5 min | 67 °C | 58 % | 244.3 W | 99.9 % |
| a 3090, 350 W | 100 % | 311 | 1.886 | **71 °C at 0.9 min** | **74 °C** | 67 % | 325.8 W | **100 %** |

**The hottest this Ti ran was 62 °C**, its board power holding a median of 357 W. **A 3090 at its own limit ran twelve degrees hotter — 74 °C against 62 — while drawing 31 W less and making 37 fewer pictures**: the one pairing here where each board sits at the setting it ships with, and still not equal work at the settings they ship with. The 83 °C stop armed at every rung and fired at none.

## What fits, and what the cache costs {#what-fits-and-what-the-cache-costs}

*Largest window each card holds on itself, and what the cache settings cost at the full mixture window. **Every cache row is at 350 W on both cards**; the window rows do not move with a cap at all. Receipts: each card's `m5-one-350w-auto-ladder`, `m4-one-350w-auto-ladder`, `m4-one-350w-forced-at98304`, `m4-one-350w-forced-at131072`, `m5-one-350w-kvf16-at131072` and `m5-one-350w-kvq4_0-at131072` files.*

| | a 3090 (24,576 MiB) | a 3090 Ti (24,564 MiB) |
|---|---|---|
| mixture, the server's own placement | **131,072** whole, 20,129 MiB | **131,072** whole, 20,142 MiB |
| dense, the server's own placement | 65,536 whole, 20,589 MiB | 65,536 whole, 20,602 MiB |
| dense, the layer count named | **98,304** whole, 23,469 MiB | **98,304** whole, 23,482 MiB |
| dense at 131,072, layer count named | refuses to load — the allocator's own out-of-memory error, in the file | refuses to load — the same error, at all four caps |
| mixture at 131,072, `f16` cache, both at 350 W | 146.308 tok/s, 20,995 MiB | **159.410 tok/s**, 21,008 MiB |
| mixture at 131,072, `q8_0` — what both were measured at | 136.618, 20,129 MiB | 147.823, 20,142 MiB |
| mixture at 131,072, `q4_0` | 135.249, 19,415 MiB | 146.780, 19,428 MiB |

**Two cards, one answer: 24 GB is 24 GB** — every window identical, every allocation within thirteen mebibytes of its twin, and on both the planner leaves a third of the dense model's window unclaimed until the layer count is named. **The live suggestion is the cache**: at the same 350 W the unquantised setting is fastest on both boards and fits the full window on both, **+7.8 per cent on the Ti** and **+7.1 on a 3090**, while four bits buys nothing.

## Three more arms {#three-more-arms}

*Several people asking at once — the mixture at a 4,096-token window, tokens a second summed across simultaneous conversations (`concurrency-one-<cap>w.json` on both benches).*

| conversations at once | a 3090, 250 W *(71 %)* | a 3090, 300 W *(86 %)* | a 3090, 350 W *(100 %)* | a 3090 Ti, 300 W *(67 %)* | 350 W *(78 %)* | 400 W *(89 %)* | 450 W *(100 %)* |
|---|---:|---:|---:|---:|---:|---:|---:|
| 1 | 126.650 | 134.754 | 136.274 | 145.817 | 147.201 | 146.996 | **147.364** |
| 2 | 195.522 | 210.419 | 214.039 | 228.601 | 231.908 | **232.647** | 231.240 |
| 4 | 265.164 | 288.438 | 294.636 | 327.124 | 337.044 | **337.842** | 337.620 |

The Ti leads by 8.2, 8.6 and 13.4 per cent at the shared 300 W and **by 8.1, 8.0 and 14.6 run as they come** — four conversations at once being its widest margin here. **A short gate-shaped call is the one arm where the boards are indistinguishable**: 594.3 ms at the median against 591.7 (`3090-g2-doorman-350w.json`, `3090ti-g2-doorman-350w.json`). **A long document read cold** — 98,000 tokens in a 131,072-token window, both cards at the shared 350 W — is taken in at **2,549.966 tokens a second** by the Ti against **2,270.547** by a 3090, and written through at **80.056** against **73.327** (`3090ti-m5-one-350w-filled.json`, `3090-m5-one-350w-filled.json`). Run as they come, the Ti's 450 W file reads 2,628.756 and 79.740.

**What each costs to leave switched on, and why the figure has two answers.** *From `3090ti-idle-gemma4-300w.json` and `3090-idle-gemma4-300w.json`, 57 samples over thirty seconds each.* With the mixture resident and nothing generating, a Ti reads a **median of 18.74 W and a mean of 43.04**; a 3090, **22.26 and 52.76**. The gap is a load spike inside the window — a 3090's own 350 W file proves it, its empty and loaded windows sharing the identical **22.39 W** median while their means differ by 26 W. Left on continuously the medians come to **about $30 a year against $36**, the means to **$69 against $85**, at 18.34 cents a kilowatt-hour (the United States residential average for June 2026, as the Energy Information Administration publishes it). **The median is the idle figure and the mean an upper bound containing a load**, so **the faster card is the cheaper one to leave on, by $6 a year on the medians and $16 on the means** — board power, so a wall meter would read more for both.

## The card itself, as an object {#the-card-itself-as-an-object}

**All 24 GB is on the front.** A 3090 Ti carries twelve 2 GB memory chips, every one face-up beside the processor; a 3090 carries twenty-four 1 GB chips, twelve on the front and twelve on the *back* — which is the heat a person feels through a 3090's backplate, and the reason that backplate is a cooling part. On this Ti the backplate is much thicker, and the memory it would have had to cool is not behind it.

**It is a 450 W cooler, and it is the larger board by a measured margin — this version of it, at least**; other 3090 Ti boards are built differently and none was handled here. **Measured on 17 September 2026 by an operator with a kitchen scale and a tape measure**, not a caliper or a bench balance: the XC3 Ultra **1,203 g** and **about 2 inches** thick, the FTW3 Ultra Ti **2,157 g** and **about 2.75 inches** — **954 g heavier, 1.79 times the weight, and about three quarters of an inch thicker**. The weights are scale readings; both thicknesses are tape readings and carry their “about”. **On the consumer board in this machine it covers both PCIe slots**, so two of these did not go unmeasured: it is not possible, and the render bench deleted its two-card arms rather than leave them looking unrun. Where this workshop's other pages say "the pair", they mean two 10 GB cards, never two of these. Nothing else about either board was put on an instrument — no dimension beyond that one thickness, and no backplate measurement — and it ran cool to the touch at every cap, said here as the ten-minute plateaus above and not as a feeling.

## A short history {#a-short-history}

The GeForce RTX 3090 Ti went on sale on **29 March 2022 at $1,999**, in NVIDIA's own words with "24GB of the fastest 21Gbps GDDR6X memory" — the last and fastest of the GeForce RTX 30 line's flagships, six months before the generation above it replaced it, and its 2 GB chips are what let the whole 24 GB sit on one face of the board. **Then its maker left.** On **16 September 2022**, less than six months after this card shipped, EVGA told Gamers Nexus and JayzTwoCents in closed-door interviews that it was ending its partnership with NVIDIA and leaving the graphics-card business altogether, and would sell through the RTX 30 stock it had; the reporting of those interviews put GPU sales at close to 80 per cent of its gross revenue, which is the interviewers' figure and not a number EVGA published. **No last-built date was ever given**, so this page does not claim one: what the record carries is that EVGA has not built a graphics card since, and that stock of a 2022 card was still boxed and unsold in 2026 — a finite pile, and nobody making another.

## What it cost {#what-it-cost}

**This Ti has a price on record**: bought **new** for **$2,119.99**, ordered **15 September 2026**, listed as *"EVGA GeForce RTX 3090 Ti FTW3 Ultra Gaming, 24G-P5-4985-KR, 24GB GDDR6X, iCX3, ARGB LED, Backplate, Free eLeash"* — $120.99 above its own March 2022 launch price, four and a half years later. **This 3090 has no price on record**; it was already in the workshop, and the nearest dated figure is a listing of the exact board, *"EVGA 24G-P5-3975-KR GeForce RTX 3090 XC3 Ultra Gaming, 24GB GDDR6X, iCX3 Cooling, ARGB LED, Metal Backplate (Renewed)"*, at **$1,879.99**, read by hand at the store page on **17 September 2026**.

**The two prices above, re-read.** On 18 September 2026 between 00:25 and 00:35 UTC **both retailers served no price to this workshop's fetcher**, which is why the 17 September reading was taken by hand. One purchase and one dated hand reading, then — neither a claim about what either card costs today, both checkable the same way: open the retailer, search the listing title exactly as printed, read what it says on the day you do.

## What the $240 buys, on those two readings {#what-the-240-buys-on-those-two-readings}

*Not a recommendation — the arithmetic, and the reader's call, on the two prices printed above and nothing else: **$1,879.99** and **$2,119.99**, a difference of **$240.00**, 12.8 per cent. Each row names its cap; the picture rate is the 8-step graph at batch 3 turned into images a minute.*

| per $100 of the price as read | a 3090 *($1,879.99, a renewed listing)* | a 3090 Ti *($2,119.99, new)* |
|---|---:|---:|
| mixture model, tokens a second, both at 300 W | **7.24** | 6.88 |
| dense model, tokens a second, both at 300 W | **2.54** | 2.03 |
| mixture model, each run as it comes *(350 W / 450 W)* | **7.27** | 6.96 |
| dense model, each run as it comes *(350 W / 450 W)* | 2.49 | **2.71** |
| pictures a minute, 8 steps, batch 3, both at 300 W | **2.05** | 1.95 |
| pictures a minute, each run as it comes *(350 W / 450 W)* | **2.11** | 2.07 |

**Per dollar the cheaper card wins every row but one**, losing the one that matters most to a dense-model buyer: **2.71** tokens a second per $100 against 2.49. **What the $240 buys that per-dollar arithmetic cannot show** is the road above 350 W: a Ti goes from 42.969 tokens a second on the dense model to 53.354 and then 57.396, **22.4 per cent** past anything a 3090 reaches at its own limit; four conversations open a **14.6 per cent** gap; every picture comes **9 to 12 per cent** sooner; ten minutes of drawing runs **twelve degrees cooler** while making 37 more images. Against that it draws **17.7 per cent** more board power on the mixture model and **53.2** more on the dense one, it is dearer to run per thousand tokens at every pairing here, and it covers both PCIe slots.

## The setting this page would choose, and the dissent {#the-setting-this-page-would-choose-and-the-dissent}

A rule written before any of this ran asked for **the lower of 350 and 400 W unless the higher wins at least 10 per cent on the 32-step graph and plateaus under 70 °C**. Its first condition fails: **400 W wins 2.24 per cent** there (5.1296 → 5.0150 s per image), so the second never decides anything. **The setting is 350 W**, and three lines land on it — that rule, the mixture's knee, and the cap flag going inactive at the same rung. **The dissent: the dense model is not done at 350 W**, gaining **6.71 per cent** at 400, so a reader whose machine mostly runs one should read 400 W as their ceiling. **450 W is a measurement here, not a recommendation.** What this workshop then did with its own two boards, stated as what was done and not as advice: the Ti went into the render box at **350 W**, by the rule printed above, and a 3090 went back to language work at **300 W**.

## What this page does not measure {#what-this-page-does-not-measure}

- **Answer quality.** Not once. Every figure here is speed, capacity, power, heat or price.
- **No memory-die temperature**, on either card at any cap: the driver answers `[N/A]`, asked three ways, on every arm, so every temperature here is a **core** temperature. **And no wall power, none estimated** — every watt is card board power, a lower bound on what a room pays, and the only whole-box instrument in the room refused every request at every rung and is recorded as **unreadable**, which is not the same as quiet.
- **A 3090 Ti at 250 W was not run**, on an operator's ruling made before the bench started: *"so far below the default max, it will likely limit the clock speeds too much."* None is interpolated from the rungs around it, nor is 480 W, its `max_limit` rather than its shipped default. **And no 400 or 450 W rung for a 3090** — above its own limit.
- **Loudness.** Fan speed is printed on almost every table and **no sound was metered**: a fan percentage is not a noise figure. **Weight and one thickness ARE measured** — by a kitchen scale and a tape measure on 17 September 2026, printed above with the instrument named — and nothing else about either board's dimensions is: no width, no length, no backplate thickness, and the remaining size sentences are an operator's plain description.
- **No cross-box row**, **nothing about two of these cards** (impossible on this board, not merely un-run), and **no streamed-bandwidth figure for the mixture model on either card**, which is a printed refusal rather than an invented number.

## How to check our work {#how-to-check-our-work}

Four benches stand behind this page, with their pre-registrations, dated amendments, result files and verdicts in this workshop's repository. All of it ran between **17 September 2026 03:15 UTC and 18 September 01:09 UTC**: a 3090's 250 and 300 W language arms 03:15 to 15:04 and its renders 17:09 to 17:35; a 3090 Ti's language arms 19:07 to 23:46 and its renders 19:41 to 00:01; then a 3090 back in the slot for its own 350 W default, language 00:22 to 00:54 and renders 00:56 to 01:09 on the 18th. The Ti's bank is **48 result files over 228 scored runs, zero arms refused**. If a figure here does not reproduce, say so at the [contact desk](https://strata2signal.com/contact/). This page keeps its receipts the way [every page here does](https://research.strata2signal.com/how-we-work/).

## Who ran this, and thanks {#who-ran-this-and-thanks}

The cards are an **EVGA GeForce RTX 3090 XC3 Ultra 24 GB** and an **EVGA GeForce RTX 3090 Ti FTW3 Ultra 24 GB**, and **EVGA** gets the credit twice: for boards still doing serious work four and five years after they were sold for games, and for publishing the figures this page checked itself against. **NVIDIA** published the launch specifications quoted in the history, and its driver is the instrument every reading came out of.

The language work stands on **ollama** and, beneath it, **llama.cpp** and **NVIDIA's CUDA**; the pictures on **ComfyUI**. The models are Google's **Gemma 4**, **Mistral Small 3.2** from Mistral AI, Black Forest Labs' **FLUX.2 Klein 4B** and its decoder, which drew every picture timed here, and **Qwen** from Alibaba Cloud as the text encoder in those graphs. None of them owed us anything, and several ask for no credit, which is exactly why it is given.

A small human team owns the cards, moved one out of the slot and the other in, set the rules and the refusals before the runs and signed the numbers; a fleet of AI agents ran the harness and did the arithmetic under that team's rulings. Thanks to the readers of the page before this one, who asked the question this one answers: *what happens if you put the faster card in the same slot?*

<!-- derived 2026-09-18 (UTC) by tools/derive_md.py from the pour source.
     source html sha256: e0eae5eab4a20454cc8612ea77f6e7f20fcb782b58506035ae6b19e746f163c8
     derivation sha256:  29c53f47786eddea85fcc2c128d2655c88d286b44d3974a342bfbc6db871d781
     the {#id} on each heading is the anchor that heading carries on the page. -->
