The notes — which card left and which came in, why the cap is 350 W, which figures are bench and which are live, and what this page cannot say

One 3090 out, one 3090 Ti in, and the cap that decided it

exhibit sixty-one The notes
Published 2026-09-25 (UTC)
A small (human) team and a fleet of AI agents.

On 2026-09-25 (UTC) an EVGA GeForce RTX 3090 XC3 Ultra Gaming left the workstation, the computer behind the workshop's public services, and an EVGA GeForce RTX 3090 Ti FTW3 Ultra Gaming took its place at 350 W, not the old card's 300, because at 300 W our earlier bench had it behind a 3090. On the same computer and test, nine days apart in a different slot, it writes text about 9.5 per cent faster; its first live replies, eight long ones, agree. Whether it runs cooler here is not settled: its first hour read warmer, on a fresh boot.

ask about this page → assistant.strata2signal.com · in beta, still being tested

hardware on this page: RTX PRO 6000 Blackwell (96 GB) · RTX 3090 (24 GB) · RTX 3090 Ti (24 GB) → research.strata2signal.com/hardware/ · the roster is in beta, still being tested

the short version

The workstation's small language model, mistral-small3.2:24b, moved from a 3090 at 300 W to a 3090 Ti at 350 W, and the cap is the story: held to 300 W, two-thirds of what it ships with, this 3090 Ti was the slower card on that model; at 350 W it led. In service, 350 W against 300, it is about 9.5 per cent faster (8.1 to 11.2 across the runs) on about the same energy per thousand tokens, for 7 to 13 per cent more board power. Beyond that one dense model it is the more efficient card: it idles lower (18.74 W against 22.26 on the render computer, a model loaded), draws less and finishes sooner when drawing pictures (316.2 W against 325.8, 9.1 per cent less time per image), and spends 10.1 per cent less energy on the mixture model. "Runs much cooler" held while drawing pictures on the render computer, the workshop's picture machine, against a used 2021 3090, and has not shown here.

3,186 words, about 14 minutes to read.

The summary is this page’s own; what was dropped, and why, is in this page’s receipt file.

strata→signal is a small workshop that runs its own machines and writes up what it measures.

The workstation sitting on its battery backup: a pink-lit 3090 Ti beside the bracket of the NVIDIA RTX PRO 6000 Blackwell Workstation Edition, the parts numbered 1 to 5 as in the key below.

The workstation sitting on its battery backup, taken 2026-09-25: every part was bought new; this 3090 Ti, out of its box on 2026-09-17, served two other computers first.

  1. The CPU cooler, an ARCTIC Freezer 4U-M, on an AMD EPYC 7532.
  2. The server board, an ASRock Rack board, with its memory.
  3. This 3090 Ti, the EVGA GeForce RTX 3090 Ti FTW3 Ultra Gaming, lit pink with OpenRGB (violet to the camera).
  4. The NVIDIA RTX PRO 6000 Blackwell Workstation Edition, its bracket.
  5. The battery backup the machine sits on, a CyberPower CP2000PFCRM2U.

An operator's crop, straightened 1.3 degrees, resized to 1,600 pixels, metadata stripped; two stickers on the RTX PRO 6000's bracket painted out; the key's numbers drawn on; nothing else changed.

A photograph of the same computer with the old 3090 still in it is coming, taken at a quiet hour: this machine serves the workshop's public apps, and we didn't photograph it before the swap.

Why the swap

The RTX PRO 6000 carries the big seats (models kept loaded for the public services); beside it, until 2026-09-25, a 3090 carried the small one, mistral-small3.2:24b (its Q4_K_M build), a dense model (every weight read for every token written), at a 300 W cap below the card's 350 W default, set at the knee of its own ladder, where the next 50 W bought only 3.83 per cent. An operator proposed the swap at 09:22 UTC on two expectations: that a 3090 Ti "runs much cooler" than a 3090 and runs dense models "so much faster" (the workshop's decision log paraphrases the second as over 10 per cent). Both matched what the render computer had measured eight days earlier before moving its pictures onto this Ti.

What the earlier bench said

The swap was first planned at 300 W; the render computer's files said that would be a downgrade. In one slot, hours apart on 2026-09-17, in a 65,536-token window, this 3090 Ti ran the dense model at 42.969 tokens a second at 300 W on one cap-ladder step. The render computer's 3090, another EVGA GeForce RTX 3090 XC3 Ultra Gaming (the same model, a different board), ran 47.838 in a single-window repeat on a cold card: a loss of 10.2 per cent (11.3 the other way, as that page prints it). At 350 W this Ti ran 53.354, on 6,343.81 joules per thousand tokens against 6,902.38 at 300: 8.1 per cent less energy for 50 W more, because it finishes each answer sooner. At 300 W its chip reached only 66.3 per cent of the bandwidth its memory clock allows, at 350 W 82.3, the clock unchanged: the cap starved the chip, not the memory. And 300 W is two-thirds of a 3090 Ti's 450 W default, against 86 per cent of a 3090's 350. So this Ti went in at 350 W, where those files, from another computer, said an operator's "over 10 per cent" would hold; the same computer, below, says about 9.5 per cent. Why not this Ti's 450 W default, no record says; the ruling weighed 300 against 350, and the table's 450 W row shows what the last 100 W buys.

The same computer, the same script

Both cards ran the workshop's cap ladder on the workstation: one program, one frozen 1,023-token prompt in a 32,768-token window, 256 tokens out, three scored runs after a discarded warm-up, the same Q4_K_M model file, flash attention on, the cache at q8_0, one request at a time, ollama 0.32.15. What differed: the day and hour (2026-09-16 06:27 UTC against 09-25 12:08), the slot, the RTX PRO 6000 also serving the seat during this Ti's runs, the device-selection settings, and a safety guard polling the battery backup and the card through this Ti's runs only (any cost it had fell on this Ti); no room temperature was recorded either day. This Ti's runs spread 0.1 to 0.8 per cent at each cap; the old card's 300 W baseline spreads 2.1 per cent (48.626 · 48.181 · 47.597).

Speed at each cap

Cap and modelThe old 3090, tokens/sThis 3090 Ti, tokens/sChange
300 W, mistral-small3.2:24b48.18142.222−12.4 %
350 W, mistral-small3.2:24b50.02452.749+5.4 %
450 W, mistral-small3.2:24b—56.982—
350 W, gemma4:26b136.492146.691+7.5 %

Tokens written a second (decode), the median of three scored runs; the old 3090 on 2026-09-16, this Ti on 2026-09-25; the old card's maximum is 366 W. In service, this Ti's 350 W cell against the old card's 300 W: 9.48 per cent faster at the medians, 8.1 to 11.2 across the nine run pairs. Energy per thousand tokens overlaps run to run (5,510 to 6,160 joules against 5,346 to 5,718), so no saving is printed. Board power, the card's own draw, is 7 to 13 per cent more: 6.86 by the medians (292.08 W against 273.33), 12.6 by the means. The 9.48 is two gains multiplied: 3.83 from raising the old card to 350 W, then 5.45 from the change of card at 350.

Speed at each cap Tokens written a second (decode) by the seat's dense model, mistral-small3.2:24b, the median of three scored runs, on the workstation: one frozen 1,023-token prompt, a 32,768-token window, 256 tokens out. The old 3090 ran on 2026-09-16 and this 3090 Ti on 2026-09-25, nine days apart and in a different slot.

The old 3090 (2026-09-16) This 3090 Ti (2026-09-25)

300 W 48.181 tok/s 42.222 tok/s The old 3090 served at this cap. 350 W 50.024 tok/s 52.749 tok/s This 3090 Ti serves at this cap. 450 W — 56.982 tok/s The old card's maximum is 366 W.

The workstation's two cap-ladder result files for mistral-small3.2:24b, one per card, as the table under 'Speed at each cap' prints them.

At the two defaults, 450 W against 350, this Ti leads by 13.9 per cent on 20.3 per cent more energy. So "over 10 per cent" holds at the defaults and lands within a point of it at the caps the cards run at, which three runs nine days apart cannot separate. On gemma4:26b, a mixture-of-experts model, this Ti at 350 W was 7.5 per cent faster on 10.1 per cent less energy, the ranges not overlapping.

The mixture-of-experts control at 350 W Tokens written a second (decode) by gemma4:26b, a mixture-of-experts model, the median of three scored runs on the workstation's cap ladder: the old 3090 on 2026-09-16, this 3090 Ti on 2026-09-25, nine days apart and in a different slot. Drawn on a scale of its own, apart from the dense model's.

The old 3090 (2026-09-16) This 3090 Ti (2026-09-25)

350 W 136.492 tok/s 146.691 tok/s

The workstation's two cap-ladder result files for gemma4:26b, one per card, as the table under 'Speed at each cap' prints them.

The earlier bench's 22.4 per cent, default against default, rests on the render computer's 3090 reading 46.880 at 350 W, below its own 300 W reading, where the workstation's 3090 read 50.024; the Ti side agrees across computers and windows (56.982 here, 57.396 there): the gap is between two 3090 boards in two computers, not a correction.

Drawing pictures, on the render computer

Cap and measureAnother 3090This 3090 TiChange
300 W, seconds per image—1.811—
350 W, seconds per image1.8861.715−9.1 %
350 W, images drawn in ten minutes311331+6.4 %
350 W, peak core temperature, °C7460−14 °C
450 W, seconds per image—1.672—

The render computer, the same slot, this Ti on 2026-09-17 and the other 3090 on 09-18: ten minutes of continuous drawing with FLUX.2 klein 4B at 8 steps, 512 × 512, one image at a time; this Ti's peaks at 300 and 450 W were 58 and 62 °C. That 3090 is another card of the same model, not the one that left the workstation. This Ti drew the workshop's pictures there before it came here, and may draw them in the workstation one day; today it serves text.

Drawing pictures at 350 W, on the render computer FLUX.2 klein 4B at 8 steps, 512 × 512, one image at a time, ten minutes of continuous drawing each. The render computer, one slot, a 350 W cap on both: this 3090 Ti on 2026-09-17, that computer's own 3090 on 09-18, another card of the same model and not the one that left the workstation.

The render computer's 3090 This 3090 Ti

Seconds per image, median (lower is faster) 350 W 1.886 s 1.715 s Images drawn in the ten minutes 350 W 311 images 331 images

The render computer's ten-minute render files at 350 W (r4-350w.json, one per card), as the table under 'Drawing pictures, on the render computer' and RTX 3090 vs RTX 3090 Ti print them.

What these cards serve

The NVIDIA RTX PRO 6000 (96 GB) keeps gemma4:26b, the mixture model in the speed table, loaded as the writer, with a second Gemma 4 build (gemma-4-26b-a4b in FP8, under vLLM) and a small reranker; this 3090 Ti (24 GB, at 350 W) keeps mistral-small3.2:24b, the dense model, for the short calls; each has Nomic's nomic-embed-text embedder beside it. By service:

  • RuleSage, the board-game rules helper, uses both cards in turn: this 3090 Ti's mistral-small3.2:24b screens every question before anything is looked up; the RTX PRO 6000's reranker, the ms-marco-MiniLM-L6-v2 cross-encoder, re-orders the rulebook passages the search found so the best-matching rule text reaches the writer first (Four homes for one reranker is that model's own page); then the RTX PRO 6000's gemma4:26b reads those passages and writes the answer.
  • Amble asks the RTX PRO 6000's gemma4:26b for its answers.
  • The Beat Lab has a doorman and a genie: this 3090 Ti's mistral-small3.2:24b screens each wish, then the RTX PRO 6000's gemma4:26b turns the wish into a beat.
  • The long table seats its guests on the RTX PRO 6000's Gemma 4 build under vLLM, and its chairs on this 3090 Ti's mistral-small3.2:24b, called directly, so its rounds are the live readings below.
  • The assistant that answers questions about this site's pages (the "ask about this page" line on every article) shares the long table's Gemma 4 build on the RTX PRO 6000.
  • The print lab draws its plates on the render computer, on neither card here.

The split: screens, doorman and chairs on the dense model; every writer on Gemma 4, which the speed table shows writing far faster. Measured 2026-09-25: when this 3090 Ti's seat went dark for a 90-second drill, RuleSage's screen was answered in 5.1 s by a standby on the render computer's processor, and came back when the seat returned; the Beat Lab's doorman has no such standby yet.

The seat, live

The seat's own log, the same model each side, public services' calls only; the unit is one dinner round at the long table. Two old-card baselines, kept apart: its 71 long public replies at 300 W from 2026-09-19 14:28 to 09-25 09:11 UTC, and its last four rounds alone (2026-09-22 to 09-25), 21 of those. This Ti's side is one round, an operator's own test, seven minutes after a boot. Writing speed (decode) on replies of 64 tokens or more (this Ti's two short classifier answers are left out): 49.01 tokens a second over the 71 against 53.11 over eight, +8.4 per cent; the four rounds, round by round, 49.10 · 49.15 · 49.23 · 49.43 against 53.11, +7.4 to +8.2. The card itself confirms the old card's 300 W only from 2026-09-22 15:53 UTC; before that, the workstation's cap setting is the record. Reading speed (prefill) on the uncached part of each prompt (segments of 128 to 1,024 tokens): 1,598.7 tokens a second over 80 prompts against 1,895.4 over 10, +18.6 per cent. Prompt dates: the old card 2026-09-19 13:52 to 09-25 09:11 UTC; this Ti 09-25 09:49 to 09:58. This Ti had the harder round: writing slows as the transcript grows, and its round carried 3,891 uncached prompt tokens against 997 to 2,862 in the old card's rounds. Eight replies are a direction, not a rate.

The swap, minute by minute (UTC)

On 2026-09-25 the old card's last request came at 09:24:14. Power-off after 09:36:05, boot at 09:49:15: the workstation was down between eleven and thirteen minutes. At 09:53:26 this Ti read back a 350 W cap and a 450 W default; the RTX PRO 6000 held its 420 W cap. From 12:08:55 to 12:14:48 the seat ran on the RTX PRO 6000 while this Ti ran the ladder; it was offline only for two restarts, nine and twelve seconds.

Cooler?

"Runs much cooler" holds for one thing measured and is not shown for another. Ten minutes of continuous drawing at 350 W on the render computer, the table above: this Ti peaked at 60 °C (2026-09-17 20:42 UTC), that computer's 3090 at 74 (09-18 00:58), fourteen degrees cooler on a lower median draw, 316.2 W against 325.8. The caveats: the earlier page labels that table "not a cooler comparison", and its own headline gap, each card at its default, is twelve degrees; that 3090 was bought used in 2021, on its original thermal pads; the coolers differ, an XC3 against an FTW3. On the dense model at 300 W and 65,536 tokens in that slot this Ti read 54 to 57 °C and that 3090 53 to 57, this Ti's fans up to 56 per cent and that 3090's off: equal in degrees, not in cooling.

Peak and idle temperatures The card's core, in degrees Celsius, and a different 3090 in each panel. The render computer, one slot: ten minutes of drawing at 350 W, this 3090 Ti on 2026-09-17 making 331 images at a median 316.2 W, that computer's own 3090 on 09-18 making 311 at 325.8 W; the coolers differ, an FTW3 against an XC3, and that 3090 was bought used in 2021, on its original thermal pads; the earlier page calls its render table not a cooler comparison. Its dense rows are each run's peak at 300 W in the 65,536-token window, hours apart on 2026-09-17. The workstation, live, in a different slot: idle medians, the old 3090 from 09:49 to 10:10 UTC on 2026-09-23 and 24 (115 readings), this 3090 Ti from 10:00 to 10:10 UTC on 2026-09-25 on a fresh boot (27 readings); round peaks, the old 3090 over four long-table rounds, this 3090 Ti over one.

A 3090 (a different board in each panel) This 3090 Ti

The render computer: its own 3090, this Ti Render, 350 W 74 °C 60 °C Ten-minute peak; 3090 09-18, Ti 09-17. Dense, 300 W 53–57 °C 54–57 °C Run peaks at 65,536; Ti fans to 56 %, 3090 off. 20 °C 80 °C The workstation, live: the old 3090, this Ti Idle 32 °C 42 °C Median; the Ti's first hour, still settling. Round peak 43–50 °C 52 °C 3090 four rounds, to 299.51 W; Ti one, 346.79 W. 20 °C 80 °C

Render and dense rows: the render computer's bench files (the render peaks are printed on RTX 3090 vs RTX 3090 Ti). Workstation rows: the workshop's 15-second card collector.

On the workstation nothing is settled. During the ladder this Ti's peaks read about ten degrees below the old card's, but so do its starts, on different days in different slots with no room-temperature reading. Live, in its first hour on a fresh boot, it idled at 42 °C (10:00 to 10:10 UTC, still settling) against the old card's 32 over the same clock hours on 2026-09-23 and 09-24. Under one round at the long table it peaked at 52 °C, a rise of 13, at 346.79 W with its fans up to 71 per cent, against the old card's 43 to 50 over four rounds, rises of 12 to 16, fan off, at most 299.51 W: the same rise, on about 47 W more.

At idle the render computer's same-slot bench does give a like-for-like reading, at a 300 W cap on both: this Ti drew 18.74 W with a model loaded against that 3090's 22.26, and 18.77 against 22.57 with the card empty.

Idle board power, on the render computer Watts, the card's own draw, the median of 57 samples over thirty seconds per window. The render computer, one slot, a 300 W cap on both, on 2026-09-17, hours apart (15:02 and 19:37 UTC): that computer's own 3090, not the one that left the workstation. The resident model is gemma4:26b, not the seat's. The text above prints all four; RTX 3090 vs RTX 3090 Ti prints the resident pair.

The render computer's 3090 This 3090 Ti

Model resident 22.26 W 18.74 W gemma4:26b loaded, not generating. Empty 22.57 W 18.77 W No model loaded.

The render computer's two idle files, 3090-idle-gemma4-300w.json and 3090ti-idle-gemma4-300w.json, as RTX 3090 vs RTX 3090 Ti names them; that page prints the resident pair.

What this page cannot say

  • Cooler on the workstation? One fresh boot, one round; a sustained load here would settle it.
  • Memory heat. No board here reports its memory's temperature; every degree on this page is the chip's core.
  • Idle draw on the workstation. Too soon to read, an hour after a boot; the render computer's same-slot reading is the idle chart above.
  • The slot. Different slots, both PCIe generation 4, x16; no card was measured in both.
  • A same-day comparison. The old card's ladder is nine days older.
  • The live gain. Eight replies, one round; a dated block follows when a week of traffic exists.
  • Boards in general. One board of each model per computer.

How to check our work

  • Every render-computer figure here is on RTX 3090 vs RTX 3090 Ti, except the dense-model readings at 300 W and the render start times, from that bench's result files.
  • The workstation's rows come from four cap-ladder result files, two per card; the live rows from the seat's log and the workshop's 15-second card collector.
  • Percentages are computed from the figures as printed, medians unless named; the tables round to one decimal.
  • The files follow as a data kit once private details are removed; the cards are on the hardware roster.

Who ran this, and thanks

EVGA built the three boards and NVIDIA the chips, the driver and CUDA. The seat stands on ollama and ggml from llama.cpp; the model is Mistral AI's Mistral Small 3.2 with Nomic's nomic-embed-text; Google's Gemma 4 ran the mixture-of-experts control and writes the answers above, on ollama and on vLLM; the reranker is the Sentence-Transformers project's ms-marco-MiniLM-L6-v2; Black Forest Labs' FLUX.2 klein 4B drew the render rows on ComfyUI and PyTorch. AMD, ASRock Rack, ARCTIC and CyberPower made the parts in the picture, OpenRGB lit the card, and systemd's journal held every live speed figure. None of them owed us anything. An operator proposed the swap, moved the cards and ran the first round; agents ran the ladder and did the arithmetic; an operator signed the numbers.

elsewhere in the workshop

a strata→signal property · hello@strata2signal.com · say hello