The bench — what this page measures, on what, and what it cannot say

The same eighteen pictures, card by card

exhibit sixty-three The bench
Published 2026-09-26 (UTC)
A small (human) team and a fleet of AI agents.

Held to the same 300 W, at 2048×1280 on FLUX.2 [klein] 4B, each card in its own computer, an NVIDIA RTX PRO 6000 Blackwell Workstation Edition 96 GB does each sampling step 1.6 times as fast as a ZOTAC GeForce RTX 5080 Solid CORE OC 16 GB, down from 2.8 times at their stock 600 W and 360 W. Over ten minutes at 1024×1024 on the same model at 8 steps, an NVIDIA GeForce RTX 5090 Laptop GPU 24 GB, on the 150 W its Dynamic Boost held through that run, spent the same board energy per picture as the workstation card at 300 W, 663 against 667 J, each in its own computer. On an NVIDIA GeForce RTX 3080 Ti Founders Edition 12 GB, FLUX.2 [klein] 4B in bf16 sat whole on the card in every cell; Z-Image-Turbo in bf16 and FLUX.1 [dev] in fp8 paged their weights in every cell that ran, down to 68.7 per cent resident. In one computer, on one ComfyUI checkout, both held to 300 W on a 512×512 recipe, a 5080 did a 3090 Ti's work in 0.69 of the time for 0.64 of the board energy; at their stock limits, against the second 3090 Ti in that computer, 0.74 of the time for 0.53 of the energy.

ask about this page → assistant.strata2signal.com · in beta, still being tested

hardware on this page: RTX PRO 6000 Blackwell (96 GB) · RTX 3090 (24 GB) · RTX 5090 Laptop GPU (24 GB) · RTX 3090 Ti (24 GB) · RTX 3080 Ti (12 GB) · RTX 5080 (16 GB) → research.strata2signal.com/hardware/ · the roster is in beta, still being tested

the short version

Four cards ran the same eighteen kinds of picture on one ComfyUI stack in four computers: every cross-card figure is "same software, different computer". One FLUX.2 [klein] 4B picture at 1024×1024 and 8 steps, at stock: 1.51 s on the workstation card, 3.59 on the 5080, 4.50 on the laptop, 5.73 on the 3080 Ti. On the 12 GB card klein 4B in bf16 sat whole in every cell; Z-Image-Turbo in bf16 and FLUX.1 [dev] in fp8 paged, down to 68.7 per cent resident. Ten minutes of that klein 4B picture, board energy only: 252 at 667 J from the workstation card at 300 W, 164 at 977.6 to 991.6 J from the 5080 at 300 W or stock, 132 at 663 J from the laptop at 150 W, 101 at 1,908.4 J from the 3080 Ti at its stock 350 W beside an idle board. In one computer at 300 W a 5080 did a 3090 Ti's work in 0.69 of the time for 0.64 of the energy.

6,511 words, about 30 minutes to read.

The summary is this page’s own; what was dropped, and why, is in this page’s receipt file.

strata→signal is a small workshop that runs its own machines and writes up what it measures. This is its first page about pictures only, and the first where every grid card ran the identical Python stack: the same ComfyUI (the free program that runs image models on your own computer) and the same 103 packages installed fresh on each computer; the graphics driver differed by computer, and the bench recorded no operating system. Two of the three models are Black Forest Labs'; a short history of that company is the companion to this page.

Who ran where

Four cards, four computers, one grid; on the second test, the render bench, three cards took turns in one of those computers (Three cards in one computer). Drivers: 595.84 in the workstation and bench computers, 595.71.05 in the i9 computer, 595.91.07 on the laptop.

CardComputer and processorSystem memoryLink under loadSettings run
NVIDIA RTX PRO 6000 Blackwell Workstation Edition 96 GBWorkstation computer: AMD EPYC 7532, 64 threads126,353 MBPCIe x16, generation 4600 W (stock) and 300 W
ZOTAC GeForce RTX 5080 Solid CORE OC 16 GBi9 computer: Intel Core i9-10850K, 20 threads47,041 MBPCIe x16, generation 3 (the processor's limit)360 W (stock) and 300 W
NVIDIA GeForce RTX 5090 Laptop GPU 24 GBLaptop: Intel Core Ultra 9 290HX Plus, 24 threads62,699 MBPCIe x8, generation 5138 to 150 W, Dynamic Boost
NVIDIA GeForce RTX 3080 Ti Founders Edition 12 GBBench computer, the next day: AMD Ryzen 7 3700X, 16 threads30,990 MBPCIe x16, generation 4350 W (stock) and 300 W

Below they are the workstation card, the 5080, the laptop and the 3080 Ti; in the tables, RTX PRO 6000, RTX 5080, RTX 5090 Laptop and RTX 3080 Ti FE. System memory is what ComfyUI's startup line printed on each computer; its MB, like every MB the runtime prints, are MiB, and paged weights live there. The 5080 was in two computers that day (the kit's Who ran where has the hours). The 3080 Ti ran on 2026-09-25 in the bench computer's x16 slot with an EVGA GeForce RTX 3080 Ti XC3 Ultra 12 GB held idle in the chipset x4 slot directly below it, seated there for a later test of the two 3080 Tis (finding 4).

More cards are coming to the grid, each under its own dated section when it lands (How to check our work):

  • an EVGA GeForce RTX 3090 XC3 Ultra 24 GB, the grid's first desktop card with 24 GB;
  • the EVGA GeForce RTX 3090 Ti FTW3 Ultra 24 GB of the second test below, whose grid pass was stopped early when, under the combined load, the battery backup of the computer it ran in reported a low runtime estimate and that computer's power safety released its other work;
  • the ZOTAC GeForce RTX 3090 Ti AMP Extreme Holo 24 GB, also of the second test.

Until then the grid has no desktop card with 24 GB.

Words this page leans on

  • A sampling step. An image model draws by cleaning up noise a step at a time. The grid times 8 and 32 steps as two points on a line, not as quality settings; at another step count your seconds should lie near the line through them (two points were timed; each model's own workflow count is in How it was measured).
  • Seconds per step, the headline. The 32-step time minus the 8-step time, over the 24 extra steps, on ComfyUI's own clock. Once-per-picture work (prompt, decode, save) cancels; work the computer does every step, the debug log included (What this page cannot say), does not.
  • A cell. One model at one size and one step count, eighteen per card: the median of three timed pictures (their spread is in the kit's cell-tables note).
  • The ninety-second rule. A picture past 90 s is stopped, its cell reads "over 90 s", and every costlier cell of that model on that card is skipped; ruled before the first picture to keep each pass short, and not a verdict on the picture. A first load's limit is 600 s.
  • Board power, and joules per picture. What the card itself draws, read from its driver twice a second and summed over the picture's own run (3,600 J is a watt-hour); the processor, memory, fans, power-supply loss and, on the laptop, the processor's side of Dynamic Boost are outside it, and nothing is measured at the wall. A picture under about two seconds has too few readings and prints a dash.
  • Stock, cap and posture. Stock is the limit a desktop card ships with (600, 360 and 350 W here); a cap, here 300 W, is a lower limit a person set and read back in every sample; 300 W is the cap the workshop runs its cards at day to day, most of a card's speed for less heat and power, on a card whose stock limit is near it (the 600 W card's step takes 1.69 to 1.73 times as long at 300 W, finding 2). The laptop's Dynamic Boost budget (default 95 W, up to the board's own 175 W maximum; 138 to 150 W here, set by the driver against the processor's draw) is a posture, not a cap.
  • "Same software, different computer." The label on every grid figure that compares cards: every card ran in a different computer, so a ratio between cards includes the computer; per-step figures are the fairest comparison, and no whole-picture second is set in a ratio across computers. A picture with no model in it took 0.021 s on the workstation computer and 0.007 s on the i9 computer, so a fast card's smallest cells partly measure the computer.

The grid, and the one thing it holds fixed

Eighteen cells per card: three models at three sizes (512×512, 1024×1024, 2048×1280) and two step counts (8 and 32), batch 1. FLUX.2 [klein] 4B in bf16 runs from the workshop's everyday 8-step workflow (euler, Flux2Scheduler, cfg 1.0; the print lab, the workshop's picture service, draws with it) with only the size, the step count and the batch changed; Z-Image-Turbo in bf16 runs from its official ComfyUI template on the same text encoder; FLUX.1 [dev], one fp8 all-in-one checkpoint, runs from the official template at its default guidance of 3.5. All three are diffusion transformers, timed here and nothing more; a UNet model such as SDXL was not measured. The prompts are two real print-lab prompts with their style words removed, and the seeds follow a formula, so no two pictures in one run are the same request; the formula is the same on every card, so every card was sent the same requests in the same order.

The one thing held fixed is the ComfyUI stack: commit 3216c62e (v0.34.0) with the same 103 packages from one hashed lock on every card, without --fast, so FLUX.1 [dev]'s weights are stored in fp8 and its math is done in bf16 on every card. The stack in full, the five differences from a desktop ComfyUI and the run order are in How it was measured.

Seven things the grid found

Every figure here is from grid.json or sustained.json, one pair per card and setting, which follow in the kit (How to check our work); every figure is in its tables file now. Cross-card figures carry the label same software, different computer.

1 · Twelve gigabytes: which models fit whole

The runtime prints every model's resident megabytes after every node, and the harness reads those lines for the sampling model over the sampler's own span (re-read by residency.py v2, which follows with the raw files; the verdicts the files were first written with are kept beside it). Every instance loaded through ComfyUI's dynamic VRAM path (each log reads "DynamicVRAM support detected and enabled" and every model "prepared for dynamic VRAM loading"), which streams what does not fit from system memory each step; the 12 GB verdicts are this version's. A cell whose transformer sat whole on the card through that span is whole; one that did not is paged, its per-step figure has weight traffic inside it and carries the label, and a pair measured two ways gets no figure. On the three cards with 16 GB or more, every transformer read whole through every step of every cell: klein's 7,392 MB, Z-Image-Turbo's 11,744 of 11,738 staged, FLUX.1 [dev]'s 11,360 of 11,350 (the runtime rounds up). On the 12 GB card, at stock:

Model on the 3080 TiTransformer staged, MBResident inside the sampler's span, MBVerdict
FLUX.2 [klein] 4B, bf167,3927,392Whole, 6 of 6 cells (18 timed pictures)
Z-Image-Turbo, bf1611,7388,064 to 11,104Paged, 6 of 6 (18)
FLUX.1 [dev], fp811,3508,640 to 11,008Paged, 5 of 5 (15); the sixth cell ran over 90 s
Lowest share of the transformer on the card, over the timed pictures at that size512×5121024×10242048×1280
Z-Image-Turbo92.4 %83.4 %68.7 %
FLUX.1 [dev]93.0 %87.7 %76.1 % (the 8-step cell; the 32-step cell ran over 90 s)

The bigger the picture, the more of the transformer the runtime evicts for the step's working memory; the rest of the graph left first (two of the three text encoders and all three decoders fell to nothing on the card during the steps; the megabytes, and the 300 W pass's, are in the kit's residency note), and at 300 W the same two models paged the same way. Paged weights cross the computer's link and memory every step (here PCIe x16 at generation 4, a Ryzen 7 3700X and 30,990 MB of system memory), so a paged figure is more the computer's than any other here, and this page sets none in a ratio with a whole one. No cell failed, ran out of memory or needed a tiled decode; the one cell the 12 GB card lost, it lost to the clock (finding 7), and that cell's 300 W twin to the rule.

The 16 GB card is not paged, and not free of moves (its text encoder and part of its transformer cross its generation-3 link around each encode and decode; the kit's residency note has the megabytes), which sits in its fixed cost and ten-minute count, not its per-step figure. Peak memory is a budget on three of the four cards, not a footprint: ComfyUI fills a card to what it has, so the 12 GB card peaked within 388 to 650 MiB of its 12,288 on every cell, the 16 GB card within 465 to 851 of its 16,303, and the laptop's 22,409 MiB on Z-Image-Turbo at 2048×1280 and 32 steps sits under the 96 GB card's 27,306 on the same cell. The 96 GB card's row (the kit's peak-memory table) is what ComfyUI holds when nothing forces it out, not what a model needs: klein peaked at 22,978 MiB there and still sat whole on the 12 GB card. Another Z-Image-Turbo does fit 12 GB: on the RTX 5080 vs RTX 3080 Ti page its INT8 build (6.2 GB, decimal), timed warm with the prompt's encoding cached, sat whole on this same Founders Edition; the grid ran the bf16 file with the encoder re-run for every picture.

2 · Seconds per sampling step, and what 300 W does to the gap

First the unit most people look at: the median seconds per picture at 1024×1024 and stock (lower is faster; the 3080 Ti's Z-Image-Turbo and FLUX.1 [dev] cells are paged, finding 1). 8 steps is Z-Image-Turbo's template count and klein's recipe count, so those rows are a default workflow's time; FLUX.1 [dev]'s template runs 20, between its two rows.

Card, setting and steps, 1024×1024FLUX.2 [klein] 4BZ-Image-TurboFLUX.1 [dev]
RTX PRO 6000, 600 W, 8 steps1.5061.9932.758
RTX PRO 6000, 600 W, 32 steps4.3686.7189.601
RTX 5080, 360 W, 8 steps3.5935.6507.033
RTX 5080, 360 W, 32 steps12.02318.57025.689
RTX 5090 Laptop, 138–150 W, 8 steps4.5016.4369.777
RTX 5090 Laptop, 138–150 W, 32 steps16.01124.24037.231
RTX 3080 Ti FE, 350 W, 8 steps5.7258.604, paged10.857, paged
RTX 3080 Ti FE, 350 W, 32 steps19.00430.289, paged39.809, paged

Every size, and every cell's joules, are in the kit's tables. The page's own instrument is seconds per sampling step, here at 2048×1280, the one size every setting ran:

Card and setting, seconds per step at 2048×1280FLUX.2 [klein] 4BZ-Image-TurboFLUX.1 [dev]
RTX PRO 6000, 600 W0.35890.59850.7875
RTX PRO 6000, 300 W0.62151.02891.3312
RTX 5080, 360 W0.99331.63542.2085
RTX 5080, 300 W1.00901.65102.2965
RTX 5090 Laptop, 138–150 W1.40912.3272No figure
RTX 3080 Ti FE, 350 W1.58122.6002, pagedNo figure
RTX 3080 Ti FE, 300 W1.64422.6899, pagedNo figure

it/s is one divided by the seconds (the 3080 Ti's klein column, for instance, reads 0.632 and 0.608 it/s). ComfyUI's console bar times only the sampler's last interval and differs from this figure by up to 7 per cent on 32-step renders and 32 on 8-step ones across these cells at both settings (6 and 29 at stock alone), so compare your bar against a 32-step render. No figure: the laptop's and the 3080 Ti's 32-step FLUX.1 [dev] cells ran over 90 s, and the 3080 Ti's 300 W twin was skipped by the rule. Read, under the label, on the cells that are whole everywhere:

  • At stock the workstation card does a klein step in 0.36 of the 5080's time, a lead of 2.77 times; 2.73 on Z-Image-Turbo, 2.80 on FLUX.1 [dev]. The 3080 Ti takes 1.59 times as long per klein step as the 5080; the laptop takes 1.42 times as long as the 5080 on both models with a figure. The lead depends on size (2.57 times at 512×512, 2.94 at 1024×1024 on klein; the kit's per-step tables): the subtraction cancels once-per-picture work, not the computer's share of each step.
  • With the desktop cards held to 300 W the workstation card's lead over the 5080 is 1.62 times (0.6215 against 1.0090); 1.60 on Z-Image-Turbo, 1.73 on FLUX.1 [dev]. The cap makes the workstation card's step 1.69 to 1.73 times as long; the 5080's 1.6 per cent longer on klein, 1.0 on Z-Image-Turbo and 4.0 on FLUX.1 [dev]; the 3080 Ti's 4.0 per cent longer on klein and 3.4 on its paged Z-Image-Turbo. 300 W is half the workstation card's stock limit, five-sixths of the 5080's and six-sevenths of the 3080 Ti's; that asymmetry is the finding.
  • The stock figures are a hot card's. The cells' core peaks ran up to 91 °C on the workstation card (each model's instance starting on a card still warm), 73 on the 5080, 83 on the 3080 Ti with a second board in the air it breathes, and 72 on the laptop under boost; at 300 W the three desktop cards peaked at 70, 68 and 77. No thermal slowdown was raised on the workstation card; on the 3080 Ti the driver's thermal-slowdown flag was set on 2.4 to 5 per cent of the samples of four timed renders in three stock cells. The ranges by setting, and the clocks, are in the kit's temperatures note.

3 · Joules per picture, and per step

Joules per picture is the trapezoid over the card's power readings inside the picture's own window; every cell's figure is in the kit's tables, and the mean draw of every 2048×1280 cell in its mean-draw table. The wall pays more, unmeasured. Across computers the fair figure is joules per step: the 32-step picture's joules minus the 8-step picture's, over 24.

Card and setting, J per step at 2048×1280FLUX.2 [klein] 4BZ-Image-TurboFLUX.1 [dev]
RTX PRO 6000, 600 W210.5365.4467.7
RTX PRO 6000, 300 W190.4305.0398.5
RTX 5080, 360 W312.6520.8772.8
RTX 5080, 300 W298.0491.4690.7
RTX 5090 Laptop, 138–150 W209.5347.0No figure
RTX 3080 Ti FE, 350 W552.4909.0, pagedNo figure
RTX 3080 Ti FE, 300 W491.8803.7, pagedNo figure

Read, under the label: the 5080 spends more energy per step than the workstation card and the laptop on every whole cell, 1.57 to 1.73 times the workstation card's with both at 300 W and 1.43 to 1.65 times at stock. The 3080 Ti spends the most: per klein step 1.77 times the 5080's at stock and 1.65 at 300 W, 2.62 and 2.58 times the workstation card's, 2.64 times the laptop's. The laptop spends 1.10 to 1.14 times the workstation card's at 300 W, 0.5 and 5.0 per cent below it at stock, and about two-thirds of the 5080's (0.67 against the 5080 at stock, 0.70 to 0.71 against it at 300 W).

4 · Ten minutes of drawing at 1024×1024

FLUX.2 [klein] 4B at 8 steps, back to back, one prompt always queued, on ComfyUI's own clock from the first counted picture's start to the last one's end; from each run's own file, the count, the median seconds, the mean draw, the joules and the peak temperature. The 3080 Ti's row is its stock run; it has no 300 W run on this cell. The last column, kilowatt-hours per thousand pictures, is the J per picture over 3,600 (3,600 J is a watt-hour).

Card and settingPicturesSeconds per picture, medianMean board WJ per picturekWh per thousand pictures
RTX PRO 6000, 300 W2522.295279.1666.70.185
RTX 5080, 360 W1643.608270.1991.60.275
RTX 5080, 300 W1643.611265.8977.60.272
RTX 5090 Laptop, boost, 150 W throughout1324.482145.4663.00.184
RTX 3080 Ti FE, 350 W1015.814321.21,908.40.530

Core peaks were 70, 67, 67, 71 and 80 °C, and the card was busy for 0.966 to 0.983 of each window. Read as counts, under the label. The laptop (limit 150 W, 145.4 W mean) and the workstation card (limit 300 W, 279.1 W mean) spent the same per picture, 663 against 667 J, within sensor noise; on the fairer figure across computers, joules per step at 2048×1280 (finding 3), the laptop spends 1.10 to 1.14 times the workstation card's. These counts carry more of the computer than the per-step figures do: once-per-picture work, some of it the computer's, is 0.55 s of the workstation card's 1.51 s picture at stock (its 300 W row has no grid cell at this size), 0.78 s of the 5080's 3.59 s, 0.66 s of the laptop's 4.50 s and 1.30 s of the 3080 Ti's 5.73 s.

Two conditions on the 3080 Ti's row. Its core sat at 79 to 80 °C for most of the window (68 °C at the open, 66 at its lowest, 80 °C peak, the 83 °C stop never met), the driver's power-cap flag was set in 98 per cent of the samples (mean draw 321.2 W) and its thermal-slowdown flag on 2.7 per cent of the window's 1,081 samples, so its count is a warm card's at its limit; and a second board sat idle in the air it breathes (27.3 W, 1 MiB, 32 °C at most over the run), a seating registered before the run as not directly comparable with a card measured alone. Its seconds, watts and memory are read by UUID and are its own.

Why there is no stock row for the workstation card. Its stock run started 22 seconds after the grid's last cell (which had peaked at 90 °C), opened at 76 °C and met the ten-minute run's 83 °C stop at 31.1 s, 20 pictures inside a 31.5 s counted window, so no figure is printed from it; the stop is the bench's rule, below the card's own slowdown point, the absence is the queue order's as much as the card's, and at 300 W the same card ran the full ten minutes and peaked at 70 °C.

5 · What 300 W costs, cell by cell

Within one card, stock against 300 W on the same 2048×1280 cells, same computer, software and prompts: no label. From the seconds and joules in each card's two result files; the cell-by-cell change table, the seconds, the joules and the mean draws are in the kit's tables.

CardStock limitTime at 300 WBoard energy at 300 W
Workstation card600 W+54 to +68 %−9 to −22 %
5080360 W−0.1 to +3.8 %−3 to −11 %
3080 Ti350 W+1.3 to +3.3 %−10 to −12 %

On the workstation card, 300 W is half its stock limit, and it shows. One caveat: these stock cells ran at 70 to 91 °C and the 300 W cells at 56 to 70 °C, so the saving is the cap and the cooler silicon together, and this grid cannot split them. A cap acts on an average: single readings reached 530 W on the workstation card at 300 W and 186 W on the laptop in their ten-minute runs. On the 5080, 300 W is 60 W under its 360 W stock limit, and at stock it drew 289 to 344 W at the mean on these cells, above 300 W on four of the six. On the 3080 Ti, 300 W is 50 W under its 350 W stock limit, and it was using the watts: 331 to 345 W at the mean on the five of these cells it finished at stock (its Z-Image-Turbo and FLUX.1 [dev] cells paged at both settings), with its stock cells at 79 to 83 °C and its 300 W cells at 68 to 77, the workstation card's caveat again. The render bench's own cap ladders on other cards, the two 3090 Tis' among them, are in the kit's tables (What 300 watts buys). A cap saves something only on a card that was using the watts.

6 · A 2048×1280 step costs about three 1024×1024 steps

Before the first picture, the bench registered that the per-step cost would rise 2.5 to 4.5 times from 1024×1024 to 2048×1280 on every model: these models cut a picture into patches called tokens, 2048×1280 has 2.5 times as many, and attention, which compares every token with every other, grows faster than the token count. Measured at stock, all nine whole pairs with a figure fall between 2.76 and 3.14 times (the kit's size table; the 3080 Ti's paged Z-Image-Turbo pair is printed there under its label and read against no band), near the linear end of the registered band: a 2048×1280 step costs about three 1024×1024 steps, not two and a half, on all nine whole pairs with a figure.

7 · The ninety-second rule fired twice

The costliest cell of the grid is FLUX.1 [dev] at 2048×1280 and 32 steps. It finished on the workstation card (26.194 s at 600 W, 43.646 at 300 W) and on the 5080 (72.198 at 360 W, 74.933 at 300 W). The laptop's warm-up was stopped at 92.9 s at step 28 of 32, its progress bar reading 3.23 s a step, so the 32 steps finish in about 105 s (the 8-step median, 27.711 s, plus 24 × 3.23); the 3080 Ti's at 93.45 s at step 27 of 32, its bar at 3.31 s a step, about 108 s the same way (28.867 plus 24 × 3.31), and its 300 W twin was skipped by the rule. Since at 2048×1280 every warm-up on every card sat within 3 per cent of its cell's timed median, the timed renders would have run over too. Every cell that ran finished under the limit: 36 at stock and 12 at 300 W on the workstation card and the 5080, 17 of 18 on the laptop, 17 of 18 and 5 of 5 on the 3080 Ti.

Three cards in one computer, one checkout, one recipe

This section needs no label, with the 5080's caveat below. On 2026-09-24, between 07:19Z and 22:40Z, a ZOTAC GeForce RTX 5080 Solid CORE OC 16 GB, an EVGA GeForce RTX 3090 Ti FTW3 Ultra 24 GB and a ZOTAC GeForce RTX 3090 Ti AMP Extreme Holo 24 GB (bought refurbished) took turns in the same x16 slot of the bench computer, at PCIe generation 4, on the same ComfyUI checkout: the print lab's three saved ComfyUI workflows (its recipes) in batches of three, then its everyday 8-step recipe (graph 79df5a8b, 512×512 pictures with klein 4B) one picture at a time for ten minutes, a new prompt each picture, at 300 W and at stock. From the render bench's own files (the count, the median seconds, the mean draw; joules as mean draw × window ÷ pictures):

Card and cap, ten minutes on the 8-step recipePicturesSeconds per pictureMean board WJ per picture
ZOTAC RTX 5080, 300 W4561.217256.1338.3
ZOTAC RTX 5080, 360 W (stock)4661.212265.9343.8
EVGA RTX 3090 Ti, 300 W3221.771282.3528.0
ZOTAC RTX 3090 Ti, 300 W3051.873268.2528.9
ZOTAC RTX 3090 Ti, 450 W (stock)3411.645369.9652.9

The conditions (caps read back, ECC, cores, fans) are in the kit's render-bench note; the one that needs saying here: the build and the driver (595.84) were read at 17:55Z, and the 5080's files, written hours earlier (07:19 to 09:59Z, its 360 W rows first), record the torch but neither the commit nor the driver, with the card swapped out in between. Only these rows may be set in a ratio with each other; the older render-bench rows, on other builds in other computers, are in the kit's tables and never share a column with the grid's joules.

Per picture, at 300 W: the 5080 did the work in 0.69 of the EVGA board's time for 0.64 of its energy; against the ZOTAC board, 0.65 and 0.64. At stock (the 5080 at 360 W, the ZOTAC board at 450 W; the EVGA board has no stock ten-minute run in this computer), 0.74 of the time for 0.53 of the energy. Per sampling step, at 300 W, 512×512, one picture's share of a batch of three (32-step recipe median minus 8-step median, over 24, from the kit's recipes table): 0.1050 s on the 5080, 0.1540 on the EVGA 3090 Ti, 0.1606 on the ZOTAC; the two 3090 Tis take 1.47 and 1.53 times as long, and the per-batch column, where the text encoder ran fresh on every card, gives 1.47 and 1.52. The two 3090 Tis have their own page, next on this site; the EVGA board's grid column is not in this version (Who ran where). What the computer and the build moved on one of them, the EVGA 3090 Ti a week earlier in the i9 computer on the old build, is measured beside the old rows in the kit's tables.

fp8, briefly: on the RTX 5080 vs RTX 3080 Ti page an fp8 build of klein 4B (its text encoder in fp4), warm with the prompt's encoding cached, drew 1.64 to 1.79 times as fast as the bf16 file on the 5080 and 1.05 to 1.06 times on the 3080 Ti Founders Edition, which has no fp8 math; the workstation card and the laptop are Blackwell too, with the same generation of fp8-capable tensor cores, and neither ran that file. The full story is the fp8 door on that page.

The full tables

Every table this page does not print is in the data kit, the tables file published beside this page in its data/ folder, 29 tables with their headings and conditions: every cell's seconds and joules at every size at stock and at 2048×1280 at 300 W, the mean draw of every 2048×1280 cell, the per-step figures at every size, peak memory, the recipes, every ten-minute row this site has measured, what 300 watts buys on the render bench, the change and size tables of findings 5 and 6, the cold loads, the predictions, the dated windows, and the notes this page points at (the cells' spread, residency, temperatures, the neighbours, the UPS logs and the render bench's conditions).

What this page cannot say

  • A ratio between two grid cards as anything but "same software, different computer". No card has run the grid in two computers. The nearest measurement of the computer's share is the 5080 in two computers the same day: 1.237 s on the grid's klein 512×512 8-step cell in the i9 computer, 1.212 s a picture on the same recipe at the same 360 W over ten minutes in the bench computer; two per cent, one reading on one cell from two instruments and two Python installs, not a bound and not a correction.
  • What a 12 GB card does with the encoder kept off the card, a smaller encoder or a quantised transformer, or what paging costs. Not measured here; the INT8 build is in finding 1, and no model ran both whole and paged on one card.
  • Batches above one, torch.compile, Sage attention, --fast and fp8 math, other samplers, clock locks and undervolts. Held fixed or absent.
  • The cost of the debug log. Every instance ran --verbose DEBUG so the runtime would print its residency, and it prints every sampling step, not once a picture (about 30 lines a step on FLUX.2 [klein] 4B and 60 on FLUX.1 [dev]), so its cost sits inside the per-step figure as well as the fixed cost, on the host, heaviest where a step is shortest (0.040 s, the workstation card's klein step at 512×512); it was not measured. ComfyUI took the same attention path on every card (every log reads "Using pytorch attention"); which PyTorch kernel that path picked on each chip was not recorded.
  • Power at the wall socket. The three desktop UPSes were a safety stop only, none a wall meter: the workstation and bench computers' read on mains throughout; the i9 computer's log read a frozen 9 per cent through the 5080's 300 W pass while the card drew 272 to 297 W at the mean, and it will ship in the kit marked as such (the kit's UPS note has the sample counts).
  • The room temperature, or the laptop's fan. No room temperature was read; the laptop's driver cannot report its fan, so it was set to maximum by hand; the bench computer's three USB fans are turned up by hand before every run, and the run records that they were.
  • The cold loads, as a headline. Recorded, and in the kit's tables, but the page cache was not controlled: the bench re-hashes every model file before it runs, and what that leaves in memory differs by computer.
  • Any picture, judged. Every picture is kept with its sha256 and none is judged; no picture from any model is published.

How it was measured

The grid. One harness, built in its own folder on each computer, its model files linked read-only where their sha256 matched the pin, else downloaded and checked. Timing is ComfyUI's own execution_start to execution_success. Power, temperature, clocks, fan and the enforced limit were read twice a second on every card in the computer; the headline power is power.draw.instant.

The stack, in full. Every card ran ComfyUI at commit 3216c62e (v0.34.0), CPython 3.14.4 built by uv and torch 2.13.0+cu130, with the same 103 packages from one hashed lock. ComfyUI started with --verbose DEBUG --cache-classic --preview-method none, locked to its card by UUID, in its normal memory mode (no low-memory flag) and without --fast. That flag and the file, a plain fp8 cast with no scale keys, decide what the FLUX.1 [dev] cells measure: the weights are stored in fp8 and the math is done in bf16 on every card, so the fp8 hardware in the Blackwell cards was never used. Five differences from a desktop ComfyUI: every render re-ran its text encoder (the part that reads the prompt), as a new prompt does, which re-rolling seeds on one prompt skips (0.1 to 0.7 s a picture here, on the driver's websocket clock); previews were off; every model ran at cfg 1.0, one model call a step, where cfg above 1 makes two and roughly doubles the per-step cost; every instance ran the debug log, about 30 to 60 lines a step (What this page cannot say), so a normal ComfyUI should match or beat these per-step figures by an amount not measured; and the laptop ran with its fan at maximum and its usual processor services running, so on its own fan curve, or with a busier processor, it may get less; neither was measured.

What ran on each card, in order. Each model starts in a fresh ComfyUI, its first picture (load included) kept apart; then its six cells, small to large, each three timed pictures on prompts A, B, A after a warm-up on B, the text encoder run every time. A check after every picture stops the run if anything else used the card, the limit moved in any sample or the UPS left mains. The stock grid is followed by the ten-minute run (klein 4B, 1024×1024, 8 steps, stopping at 83 °C), then the desktop cards' 2048×1280 cells at 300 W; every cap change was typed by a person. Not ranked here: which card is "better", any single picture, anything in money.

ModelSamplerSchedulerGuidance (cfg)Steps in the workflow run
FLUX.2 [klein] 4B, bf16, Apache-2.0eulerFlux2Scheduler1.08
Z-Image-Turbo, bf16, Apache-2.0res_multistepsimple, shift 31.08
FLUX.1 [dev], fp8 checkpoint (Comfy-Org's all-in-one file), FLUX.1 [dev] Non-Commercial Licenseeulersimple1.0, guidance 3.520

klein's 8 is the workshop's recipe count; the other two are their official ComfyUI templates' counts.

The neighbours. The workstation card's neighbour, the EVGA GeForce RTX 3090 XC3 Ultra 24 GB, kept 18,072 MiB of live language services and idled at 26.5 to 27.0 W; its one 17 s request during the 300 W ten-minute run moved nothing measurable in that run's median. The i9 computer's own services were paused for the pass. In the bench computer a second board is allowed only when named, held idle and watched: the EVGA GeForce RTX 3080 Ti XC3 Ultra 12 GB below the 3080 Ti held at 1 MiB through all 7 loads, at 27.4 and 27.2 W at the median over the two passes, in the air the Founders Edition cooler draws (finding 4). The rule and every reading are in the kit's neighbours note; the 2 Hz trace follows with the raw files.

The dated windows (UTC). The workstation card, the 5080 and the laptop ran the grid on 2026-09-24 (16:18 to 16:52Z, 17:47 to 18:47Z and 19:20 to 20:04Z), the 3080 Ti on 2026-09-25 (13:31 to 14:29Z); the bench computer's render bench ran on 2026-09-24 between 07:19 and 22:40Z; the older render-bench rows ran in the i9 computer between 2026-09-17 01:55Z and 2026-09-18 01:09Z and on the laptop on 2026-09-21. Every window to the second, every cap's read-back in every sample and the idle draws before each pass are in the kit's windows tables.

How to check our work

The pre-registration was frozen before the first picture, revised once before that freeze after two review panels, and amended in words for the 12 GB card before its pass. Its predictions: the three this page can score held; for the 12 GB card, the two testable predictions registered in words before its pass held and the three rules registered with them were followed (the kit's predictions table); the four registered for the EVGA GeForce RTX 3090 Ti FTW3 Ultra's grid pass, which is not in this version (Who ran where), are scored as registered when its column is added, and any that cannot be is printed as unscored. The kit's first drop, published with this page, is the tables file alone. The raw files the tables come from (grid.json with the runtime's residency lines, sustained.json, the cell graphs, the 2 Hz traces, the UPS logs, the websocket events, ComfyUI's logs, the lock file) and reduce.py, which re-derives every table from the raw render records (ComfyUI's stamps, the 2 Hz power samples and the runtime's log lines) with a --check that refuses on a single moved digit, follow once cleared of the workshop's private paths and names, with a manifest of every file's size and sha256, and this section will say when; the render bench's and the older tests' files follow with them. That check passed on every result file in our own run, and you can repeat it when they ship; the tables file passed the harness's gate before this page was published. The pictures, the prompts' full text and the harness code are withheld, and named as withheld. Anything measured again or corrected later is added under its own dated section, and the byline says when.

More on this site

Who ran this, and thanks

The cards are named in full in Who ran where and Three cards in one computer; the laptop is an MSI Raider 16 Max HX. NVIDIA for the silicon, the driver every reading came from and the Founders Edition board; ZOTAC, EVGA and MSI for the boards, and Gigabyte for one in the kit's tables. The pictures stand on ComfyUI and the comfy-kitchen and comfy-aimdo libraries beneath it, on PyTorch and NVIDIA's CUDA, on Astral's uv and on Network UPS Tools. The models are Black Forest Labs' FLUX.2 [klein] 4B and FLUX.1 [dev], Alibaba's Tongyi Lab's Z-Image-Turbo, Alibaba Cloud's Qwen 3 4B, and Google's T5-XXL and OpenAI's CLIP ViT-L/14, the text encoders inside the FLUX.1 [dev] checkpoint; Hugging Face hosted every pinned file. None of them owed us anything.

A small human team owns the cards, ruled the grid and its refusals before a picture was drawn, pasted every cap and signed the numbers; a fleet of AI agents built the harness, ran it and did the arithmetic under that team's rulings. Thanks to the people who ask, on every forum this workshop reads, what a picture costs on the card they have.

Measured on 2026-09-24 and 2026-09-25 by one pinned stack on four computers and one instrument in the fourth, with the label on every cross-computer figure.

elsewhere in the workshop

a strata→signal property · hello@strata2signal.com · say hello