# The same eighteen pictures, card by card

*Held to the same 300 W, at 2048×1280 on FLUX.2 \[klein\] 4B, each card in its own computer, an NVIDIA RTX PRO 6000 Blackwell Workstation Edition 96 GB does each sampling step 1.6 times as fast as a ZOTAC GeForce RTX 5080 Solid CORE OC 16 GB, down from 2.8 times at their stock 600 W and 360 W. Over ten minutes at 1024×1024 on the same model at 8 steps, an NVIDIA GeForce RTX 5090 Laptop GPU 24 GB, on the 150 W its Dynamic Boost held through that run, spent the same board energy per picture as the workstation card at 300 W, 663 against 667 J, each in its own computer. On an NVIDIA GeForce RTX 3080 Ti Founders Edition 12 GB, FLUX.2 \[klein\] 4B in bf16 sat whole on the card in every cell; Z-Image-Turbo in bf16 and FLUX.1 \[dev\] in fp8 paged their weights in every cell that ran, down to 68.7 per cent resident. In one computer, on one ComfyUI checkout, both held to 300 W on a 512×512 recipe, a 5080 did a 3090 Ti's work in 0.69 of the time for 0.64 of the board energy; at their stock limits, against the second 3090 Ti in that computer, 0.74 of the time for 0.53 of the energy.*

*Published 2026-09-26 (UTC) · A small (human) team and a fleet of AI agents.*

**the short version:** Four cards ran the same eighteen kinds of picture on one ComfyUI stack in four computers: every cross-card figure is "same software, different computer". One FLUX.2 [klein] 4B picture at 1024×1024 and 8 steps, at stock: 1.51 s on the workstation card, 3.59 on the 5080, 4.50 on the laptop, 5.73 on the 3080 Ti. On the 12 GB card klein 4B in bf16 sat whole in every cell; Z-Image-Turbo in bf16 and FLUX.1 [dev] in fp8 paged, down to 68.7 per cent resident. Ten minutes of that klein 4B picture, board energy only: 252 at 667 J from the workstation card at 300 W, 164 at 977.6 to 991.6 J from the 5080 at 300 W or stock, 132 at 663 J from the laptop at 150 W, 101 at 1,908.4 J from the 3080 Ti at its stock 350 W beside an idle board. In one computer at 300 W a 5080 did a 3090 Ti's work in 0.69 of the time for 0.64 of the energy.

6,511 words · about 30 minutes (at 220 words/min) · 10 tables · data kit: yes

https://research.strata2signal.com/the-same-eighteen-pictures-card-by-card/

---

*[strata→signal](https://strata2signal.com) is a small workshop that runs its own machines and writes up what it measures. This is its first page about pictures only, and the first where every grid card ran the identical Python stack: the same ComfyUI (the free program that runs image models on your own computer) and the same 103 packages installed fresh on each computer; the graphics driver differed by computer, and the bench recorded no operating system. Two of the three models are Black Forest Labs'; [a short history of that company](https://research.strata2signal.com/who-makes-flux/) is the companion to this page.*

## Who ran where {#who-ran-where}

Four cards, four computers, one grid; on the second test, the render bench, three cards took turns in one of those computers (Three cards in one computer). Drivers: 595.84 in the workstation and bench computers, 595.71.05 in the i9 computer, 595.91.07 on the laptop.

| Card | Computer and processor | System memory | Link under load | Settings run |
|---|---|---:|---|---|
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition 96 GB | Workstation computer: AMD EPYC 7532, 64 threads | 126,353 MB | PCIe x16, generation 4 | 600 W (stock) and 300 W |
| ZOTAC GeForce RTX 5080 Solid CORE OC 16 GB | i9 computer: Intel Core i9-10850K, 20 threads | 47,041 MB | PCIe x16, generation 3 (the processor's limit) | 360 W (stock) and 300 W |
| NVIDIA GeForce RTX 5090 Laptop GPU 24 GB | Laptop: Intel Core Ultra 9 290HX Plus, 24 threads | 62,699 MB | PCIe x8, generation 5 | 138 to 150 W, Dynamic Boost |
| NVIDIA GeForce RTX 3080 Ti Founders Edition 12 GB | Bench computer, the next day: AMD Ryzen 7 3700X, 16 threads | 30,990 MB | PCIe x16, generation 4 | 350 W (stock) and 300 W |

Below they are the workstation card, the 5080, the laptop and the 3080 Ti; in the tables, RTX PRO 6000, RTX 5080, RTX 5090 Laptop and RTX 3080 Ti FE. System memory is what ComfyUI's startup line printed on each computer; its MB, like every MB the runtime prints, are MiB, and paged weights live there. The 5080 was in two computers that day (the kit's Who ran where has the hours). The 3080 Ti ran on 2026-09-25 in the bench computer's x16 slot with an EVGA GeForce RTX 3080 Ti XC3 Ultra 12 GB held idle in the chipset x4 slot directly below it, seated there for a later test of the two 3080 Tis (finding 4).

More cards are coming to the grid, each under its own dated section when it lands (How to check our work):

- an EVGA GeForce RTX 3090 XC3 Ultra 24 GB, the grid's first desktop card with 24 GB;
- the EVGA GeForce RTX 3090 Ti FTW3 Ultra 24 GB of the second test below, whose grid pass was stopped early when, under the combined load, the battery backup of the computer it ran in reported a low runtime estimate and that computer's power safety released its other work;
- the ZOTAC GeForce RTX 3090 Ti AMP Extreme Holo 24 GB, also of the second test.

Until then the grid has no desktop card with 24 GB.

## Words this page leans on {#words-this-page-leans-on}

- **A sampling step.** An image model draws by cleaning up noise a step at a time. The grid times 8 and 32 steps as two points on a line, not as quality settings; at another step count your seconds should lie near the line through them (two points were timed; each model's own workflow count is in How it was measured).
- **Seconds per step, the headline.** The 32-step time minus the 8-step time, over the 24 extra steps, on ComfyUI's own clock. Once-per-picture work (prompt, decode, save) cancels; work the computer does every step, the debug log included (What this page cannot say), does not.
- **A cell.** One model at one size and one step count, eighteen per card: the median of three timed pictures (their spread is in the kit's cell-tables note).
- **The ninety-second rule.** A picture past 90 s is stopped, its cell reads "over 90 s", and every costlier cell of that model on that card is skipped; ruled before the first picture to keep each pass short, and not a verdict on the picture. A first load's limit is 600 s.
- **Board power, and joules per picture.** What the card itself draws, read from its driver twice a second and summed over the picture's own run (3,600 J is a watt-hour); the processor, memory, fans, power-supply loss and, on the laptop, the processor's side of Dynamic Boost are outside it, and nothing is measured at the wall. A picture under about two seconds has too few readings and prints a dash.
- **Stock, cap and posture.** Stock is the limit a desktop card ships with (600, 360 and 350 W here); a cap, here 300 W, is a lower limit a person set and read back in every sample; 300 W is the cap the workshop runs its cards at day to day, most of a card's speed for less heat and power, on a card whose stock limit is near it (the 600 W card's step takes 1.69 to 1.73 times as long at 300 W, finding 2). The laptop's Dynamic Boost budget (default 95 W, up to the board's own 175 W maximum; 138 to 150 W here, set by the driver against the processor's draw) is a posture, not a cap.
- **"Same software, different computer."** The label on every grid figure that compares cards: every card ran in a different computer, so a ratio between cards includes the computer; per-step figures are the fairest comparison, and no whole-picture second is set in a ratio across computers. A picture with no model in it took 0.021 s on the workstation computer and 0.007 s on the i9 computer, so a fast card's smallest cells partly measure the computer.

## The grid, and the one thing it holds fixed {#the-grid-and-the-one-thing-it-holds-fixed}

Eighteen cells per card: three models at three sizes (512×512, 1024×1024, 2048×1280) and two step counts (8 and 32), batch 1. **FLUX.2 [klein] 4B** in bf16 runs from the workshop's everyday 8-step workflow (`euler`, `Flux2Scheduler`, cfg 1.0; the print lab, the workshop's picture service, draws with it) with only the size, the step count and the batch changed; **Z-Image-Turbo** in bf16 runs from its official ComfyUI template on the same text encoder; **FLUX.1 [dev]**, one fp8 all-in-one checkpoint, runs from the official template at its default guidance of 3.5. All three are diffusion transformers, timed here and nothing more; a UNet model such as SDXL was not measured. The prompts are two real print-lab prompts with their style words removed, and the seeds follow a formula, so no two pictures in one run are the same request; the formula is the same on every card, so every card was sent the same requests in the same order.

**The one thing held fixed is the ComfyUI stack:** commit `3216c62e` (v0.34.0) with the same 103 packages from one hashed lock on every card, **without `--fast`**, so FLUX.1 [dev]'s weights are stored in fp8 and its math is done in bf16 on every card. The stack in full, the five differences from a desktop ComfyUI and the run order are in How it was measured.

## Seven things the grid found {#seven-things-the-grid-found}

Every figure here is from `grid.json` or `sustained.json`, one pair per card and setting, which follow in the kit (How to check our work); every figure is in its tables file now. Cross-card figures carry the label **same software, different computer**.

### 1 · Twelve gigabytes: which models fit whole {#twelve-gigabytes-which-models-fit-whole}

The runtime prints every model's resident megabytes after every node, and the harness reads those lines for the sampling model over the sampler's own span (re-read by `residency.py` v2, which follows with the raw files; the verdicts the files were first written with are kept beside it). Every instance loaded through ComfyUI's dynamic VRAM path (each log reads "DynamicVRAM support detected and enabled" and every model "prepared for dynamic VRAM loading"), which streams what does not fit from system memory each step; the 12 GB verdicts are this version's. A cell whose transformer sat whole on the card through that span is **whole**; one that did not is **paged**, its per-step figure has weight traffic inside it and carries the label, and a pair measured two ways gets no figure. On the three cards with 16 GB or more, every transformer read whole through every step of every cell: klein's 7,392 MB, Z-Image-Turbo's 11,744 of 11,738 staged, FLUX.1 [dev]'s 11,360 of 11,350 (the runtime rounds up). On the 12 GB card, at stock:

| Model on the 3080 Ti | Transformer staged, MB | Resident inside the sampler's span, MB | Verdict |
|---|---:|---|---|
| FLUX.2 [klein] 4B, bf16 | 7,392 | 7,392 | Whole, 6 of 6 cells (18 timed pictures) |
| Z-Image-Turbo, bf16 | 11,738 | 8,064 to 11,104 | Paged, 6 of 6 (18) |
| FLUX.1 [dev], fp8 | 11,350 | 8,640 to 11,008 | Paged, 5 of 5 (15); the sixth cell ran over 90 s |

| Lowest share of the transformer on the card, over the timed pictures at that size | 512×512 | 1024×1024 | 2048×1280 |
|---|---:|---:|---:|
| Z-Image-Turbo | 92.4 % | 83.4 % | 68.7 % |
| FLUX.1 [dev] | 93.0 % | 87.7 % | 76.1 % (the 8-step cell; the 32-step cell ran over 90 s) |

The bigger the picture, the more of the transformer the runtime evicts for the step's working memory; the rest of the graph left first (two of the three text encoders and all three decoders fell to nothing on the card during the steps; the megabytes, and the 300 W pass's, are in the kit's residency note), and at 300 W the same two models paged the same way. Paged weights cross the computer's link and memory every step (here PCIe x16 at generation 4, a Ryzen 7 3700X and 30,990 MB of system memory), so a paged figure is more the computer's than any other here, and this page sets none in a ratio with a whole one. No cell failed, ran out of memory or needed a tiled decode; the one cell the 12 GB card lost, it lost to the clock (finding 7), and that cell's 300 W twin to the rule.

**The 16 GB card is not paged, and not free of moves** (its text encoder and part of its transformer cross its generation-3 link around each encode and decode; the kit's residency note has the megabytes), which sits in its fixed cost and ten-minute count, not its per-step figure. **Peak memory is a budget on three of the four cards, not a footprint:** ComfyUI fills a card to what it has, so the 12 GB card peaked within 388 to 650 MiB of its 12,288 on every cell, the 16 GB card within 465 to 851 of its 16,303, and the laptop's 22,409 MiB on Z-Image-Turbo at 2048×1280 and 32 steps sits under the 96 GB card's 27,306 on the same cell. The 96 GB card's row (the kit's peak-memory table) is what ComfyUI holds when nothing forces it out, not what a model needs: klein peaked at 22,978 MiB there and still sat whole on the 12 GB card. **Another Z-Image-Turbo does fit 12 GB:** on the [RTX 5080 vs RTX 3080 Ti](/rtx-5080-vs-rtx-3080-ti/) page its INT8 build (6.2 GB, decimal), timed warm with the prompt's encoding cached, sat whole on this same Founders Edition; the grid ran the bf16 file with the encoder re-run for every picture.

### 2 · Seconds per sampling step, and what 300 W does to the gap {#seconds-per-sampling-step-and-what-300-w-does-to-the-gap}

First the unit most people look at: the median seconds per picture at 1024×1024 and stock (lower is faster; the 3080 Ti's Z-Image-Turbo and FLUX.1 [dev] cells are paged, finding 1). 8 steps is Z-Image-Turbo's template count and klein's recipe count, so those rows are a default workflow's time; FLUX.1 [dev]'s template runs 20, between its two rows.

| Card, setting and steps, 1024×1024 | FLUX.2 [klein] 4B | Z-Image-Turbo | FLUX.1 [dev] |
|---|---:|---:|---:|
| RTX PRO 6000, 600 W, 8 steps | 1.506 | 1.993 | 2.758 |
| RTX PRO 6000, 600 W, 32 steps | 4.368 | 6.718 | 9.601 |
| RTX 5080, 360 W, 8 steps | 3.593 | 5.650 | 7.033 |
| RTX 5080, 360 W, 32 steps | 12.023 | 18.570 | 25.689 |
| RTX 5090 Laptop, 138–150 W, 8 steps | 4.501 | 6.436 | 9.777 |
| RTX 5090 Laptop, 138–150 W, 32 steps | 16.011 | 24.240 | 37.231 |
| RTX 3080 Ti FE, 350 W, 8 steps | 5.725 | 8.604, paged | 10.857, paged |
| RTX 3080 Ti FE, 350 W, 32 steps | 19.004 | 30.289, paged | 39.809, paged |

Every size, and every cell's joules, are in the kit's tables. The page's own instrument is seconds per sampling step, here at 2048×1280, the one size every setting ran:

| Card and setting, seconds per step at 2048×1280 | FLUX.2 [klein] 4B | Z-Image-Turbo | FLUX.1 [dev] |
|---|---:|---:|---:|
| RTX PRO 6000, 600 W | 0.3589 | 0.5985 | 0.7875 |
| RTX PRO 6000, 300 W | 0.6215 | 1.0289 | 1.3312 |
| RTX 5080, 360 W | 0.9933 | 1.6354 | 2.2085 |
| RTX 5080, 300 W | 1.0090 | 1.6510 | 2.2965 |
| RTX 5090 Laptop, 138–150 W | 1.4091 | 2.3272 | No figure |
| RTX 3080 Ti FE, 350 W | 1.5812 | 2.6002, paged | No figure |
| RTX 3080 Ti FE, 300 W | 1.6442 | 2.6899, paged | No figure |

it/s is one divided by the seconds (the 3080 Ti's klein column, for instance, reads 0.632 and 0.608 it/s). ComfyUI's console bar times only the sampler's last interval and differs from this figure by up to 7 per cent on 32-step renders and 32 on 8-step ones across these cells at both settings (6 and 29 at stock alone), so compare your bar against a 32-step render. No figure: the laptop's and the 3080 Ti's 32-step FLUX.1 [dev] cells ran over 90 s, and the 3080 Ti's 300 W twin was skipped by the rule. Read, under the label, on the cells that are whole everywhere:

- **At stock** the workstation card does a klein step in 0.36 of the 5080's time, a lead of 2.77 times; 2.73 on Z-Image-Turbo, 2.80 on FLUX.1 [dev]. The 3080 Ti takes 1.59 times as long per klein step as the 5080; the laptop takes 1.42 times as long as the 5080 on both models with a figure. The lead depends on size (2.57 times at 512×512, 2.94 at 1024×1024 on klein; the kit's per-step tables): the subtraction cancels once-per-picture work, not the computer's share of each step.
- **With the desktop cards held to 300 W** the workstation card's lead over the 5080 is 1.62 times (0.6215 against 1.0090); 1.60 on Z-Image-Turbo, 1.73 on FLUX.1 [dev]. The cap makes the workstation card's step 1.69 to 1.73 times as long; the 5080's 1.6 per cent longer on klein, 1.0 on Z-Image-Turbo and 4.0 on FLUX.1 [dev]; the 3080 Ti's 4.0 per cent longer on klein and 3.4 on its paged Z-Image-Turbo. 300 W is half the workstation card's stock limit, five-sixths of the 5080's and six-sevenths of the 3080 Ti's; that asymmetry is the finding.
- **The stock figures are a hot card's.** The cells' core peaks ran up to 91 °C on the workstation card (each model's instance starting on a card still warm), 73 on the 5080, 83 on the 3080 Ti with a second board in the air it breathes, and 72 on the laptop under boost; at 300 W the three desktop cards peaked at 70, 68 and 77. No thermal slowdown was raised on the workstation card; on the 3080 Ti the driver's thermal-slowdown flag was set on 2.4 to 5 per cent of the samples of four timed renders in three stock cells. The ranges by setting, and the clocks, are in the kit's temperatures note.

### 3 · Joules per picture, and per step {#joules-per-picture-and-per-step}

Joules per picture is the trapezoid over the card's power readings inside the picture's own window; every cell's figure is in the kit's tables, and the mean draw of every 2048×1280 cell in its mean-draw table. The wall pays more, unmeasured. Across computers the fair figure is **joules per step**: the 32-step picture's joules minus the 8-step picture's, over 24.

| Card and setting, J per step at 2048×1280 | FLUX.2 [klein] 4B | Z-Image-Turbo | FLUX.1 [dev] |
|---|---:|---:|---:|
| RTX PRO 6000, 600 W | 210.5 | 365.4 | 467.7 |
| RTX PRO 6000, 300 W | 190.4 | 305.0 | 398.5 |
| RTX 5080, 360 W | 312.6 | 520.8 | 772.8 |
| RTX 5080, 300 W | 298.0 | 491.4 | 690.7 |
| RTX 5090 Laptop, 138–150 W | 209.5 | 347.0 | No figure |
| RTX 3080 Ti FE, 350 W | 552.4 | 909.0, paged | No figure |
| RTX 3080 Ti FE, 300 W | 491.8 | 803.7, paged | No figure |

Read, under the label: **the 5080 spends more energy per step than the workstation card and the laptop on every whole cell**, 1.57 to 1.73 times the workstation card's with both at 300 W and 1.43 to 1.65 times at stock. **The 3080 Ti spends the most:** per klein step 1.77 times the 5080's at stock and 1.65 at 300 W, 2.62 and 2.58 times the workstation card's, 2.64 times the laptop's. **The laptop spends 1.10 to 1.14 times the workstation card's at 300 W**, 0.5 and 5.0 per cent below it at stock, and about two-thirds of the 5080's (0.67 against the 5080 at stock, 0.70 to 0.71 against it at 300 W).

### 4 · Ten minutes of drawing at 1024×1024 {#ten-minutes-of-drawing-at-1024-1024}

FLUX.2 [klein] 4B at 8 steps, back to back, one prompt always queued, on ComfyUI's own clock from the first counted picture's start to the last one's end; from each run's own file, the count, the median seconds, the mean draw, the joules and the peak temperature. The 3080 Ti's row is its stock run; it has no 300 W run on this cell. The last column, kilowatt-hours per thousand pictures, is the J per picture over 3,600 (3,600 J is a watt-hour).

| Card and setting | Pictures | Seconds per picture, median | Mean board W | J per picture | kWh per thousand pictures |
|---|---:|---:|---:|---:|---:|
| RTX PRO 6000, 300 W | 252 | 2.295 | 279.1 | 666.7 | 0.185 |
| RTX 5080, 360 W | 164 | 3.608 | 270.1 | 991.6 | 0.275 |
| RTX 5080, 300 W | 164 | 3.611 | 265.8 | 977.6 | 0.272 |
| RTX 5090 Laptop, boost, 150 W throughout | 132 | 4.482 | 145.4 | 663.0 | 0.184 |
| RTX 3080 Ti FE, 350 W | 101 | 5.814 | 321.2 | 1,908.4 | 0.530 |

Core peaks were 70, 67, 67, 71 and 80 °C, and the card was busy for 0.966 to 0.983 of each window. Read as counts, under the label. The laptop (limit 150 W, 145.4 W mean) and the workstation card (limit 300 W, 279.1 W mean) spent the same per picture, 663 against 667 J, within sensor noise; on the fairer figure across computers, joules per step at 2048×1280 (finding 3), the laptop spends 1.10 to 1.14 times the workstation card's. These counts carry more of the computer than the per-step figures do: once-per-picture work, some of it the computer's, is 0.55 s of the workstation card's 1.51 s picture at stock (its 300 W row has no grid cell at this size), 0.78 s of the 5080's 3.59 s, 0.66 s of the laptop's 4.50 s and 1.30 s of the 3080 Ti's 5.73 s.

**Two conditions on the 3080 Ti's row.** Its core sat at 79 to 80 °C for most of the window (68 °C at the open, 66 at its lowest, 80 °C peak, the 83 °C stop never met), the driver's power-cap flag was set in 98 per cent of the samples (mean draw 321.2 W) and its thermal-slowdown flag on 2.7 per cent of the window's 1,081 samples, so its count is a warm card's at its limit; and a second board sat idle in the air it breathes (27.3 W, 1 MiB, 32 °C at most over the run), a seating registered before the run as **not directly comparable** with a card measured alone. Its seconds, watts and memory are read by UUID and are its own.

**Why there is no stock row for the workstation card.** Its stock run started 22 seconds after the grid's last cell (which had peaked at 90 °C), opened at 76 °C and met the ten-minute run's 83 °C stop at 31.1 s, 20 pictures inside a 31.5 s counted window, so no figure is printed from it; the stop is the bench's rule, below the card's own slowdown point, the absence is the queue order's as much as the card's, and at 300 W the same card ran the full ten minutes and peaked at 70 °C.

### 5 · What 300 W costs, cell by cell {#what-300-w-costs-cell-by-cell}

Within one card, stock against 300 W on the same 2048×1280 cells, same computer, software and prompts: no label. From the seconds and joules in each card's two result files; the cell-by-cell change table, the seconds, the joules and the mean draws are in the kit's tables.

| Card | Stock limit | Time at 300 W | Board energy at 300 W |
|---|---:|---:|---:|
| Workstation card | 600 W | +54 to +68 % | −9 to −22 % |
| 5080 | 360 W | −0.1 to +3.8 % | −3 to −11 % |
| 3080 Ti | 350 W | +1.3 to +3.3 % | −10 to −12 % |

**On the workstation card, 300 W is half its stock limit, and it shows.** One caveat: these stock cells ran at 70 to 91 °C and the 300 W cells at 56 to 70 °C, so the saving is the cap and the cooler silicon together, and this grid cannot split them. A cap acts on an average: single readings reached 530 W on the workstation card at 300 W and 186 W on the laptop in their ten-minute runs. **On the 5080, 300 W is 60 W under its 360 W stock limit,** and at stock it drew 289 to 344 W at the mean on these cells, above 300 W on four of the six. **On the 3080 Ti, 300 W is 50 W under its 350 W stock limit,** and it was using the watts: 331 to 345 W at the mean on the five of these cells it finished at stock (its Z-Image-Turbo and FLUX.1 [dev] cells paged at both settings), with its stock cells at 79 to 83 °C and its 300 W cells at 68 to 77, the workstation card's caveat again. The render bench's own cap ladders on other cards, the two 3090 Tis' among them, are in the kit's tables (What 300 watts buys). A cap saves something only on a card that was using the watts.

### 6 · A 2048×1280 step costs about three 1024×1024 steps {#a-2048-1280-step-costs-about-three-1024-1024-steps}

Before the first picture, the bench registered that the per-step cost would rise 2.5 to 4.5 times from 1024×1024 to 2048×1280 on every model: these models cut a picture into patches called tokens, 2048×1280 has 2.5 times as many, and attention, which compares every token with every other, grows faster than the token count. Measured at stock, all nine whole pairs with a figure fall between 2.76 and 3.14 times (the kit's size table; the 3080 Ti's paged Z-Image-Turbo pair is printed there under its label and read against no band), near the linear end of the registered band: a 2048×1280 step costs about three 1024×1024 steps, not two and a half, on all nine whole pairs with a figure.

### 7 · The ninety-second rule fired twice {#the-ninety-second-rule-fired-twice}

The costliest cell of the grid is FLUX.1 [dev] at 2048×1280 and 32 steps. It finished on the workstation card (26.194 s at 600 W, 43.646 at 300 W) and on the 5080 (72.198 at 360 W, 74.933 at 300 W). The laptop's warm-up was stopped at 92.9 s at step 28 of 32, its progress bar reading 3.23 s a step, so the 32 steps finish in about 105 s (the 8-step median, 27.711 s, plus 24 × 3.23); the 3080 Ti's at 93.45 s at step 27 of 32, its bar at 3.31 s a step, about 108 s the same way (28.867 plus 24 × 3.31), and its 300 W twin was skipped by the rule. Since at 2048×1280 every warm-up on every card sat within 3 per cent of its cell's timed median, the timed renders would have run over too. Every cell that ran finished under the limit: 36 at stock and 12 at 300 W on the workstation card and the 5080, 17 of 18 on the laptop, 17 of 18 and 5 of 5 on the 3080 Ti.

## Three cards in one computer, one checkout, one recipe {#three-cards-in-one-computer-one-checkout-one-recipe}

This section needs no label, with the 5080's caveat below. On 2026-09-24, between 07:19Z and 22:40Z, a ZOTAC GeForce RTX 5080 Solid CORE OC 16 GB, an EVGA GeForce RTX 3090 Ti FTW3 Ultra 24 GB and a ZOTAC GeForce RTX 3090 Ti AMP Extreme Holo 24 GB (bought refurbished) took turns in the same x16 slot of the bench computer, at PCIe generation 4, on the same ComfyUI checkout: the print lab's three saved ComfyUI workflows (its recipes) in batches of three, then its everyday 8-step recipe (graph `79df5a8b`, 512×512 pictures with klein 4B) one picture at a time for ten minutes, a new prompt each picture, at 300 W and at stock. From the render bench's own files (the count, the median seconds, the mean draw; joules as mean draw × window ÷ pictures):

| Card and cap, ten minutes on the 8-step recipe | Pictures | Seconds per picture | Mean board W | J per picture |
|---|---:|---:|---:|---:|
| ZOTAC RTX 5080, 300 W | 456 | 1.217 | 256.1 | 338.3 |
| ZOTAC RTX 5080, 360 W (stock) | 466 | 1.212 | 265.9 | 343.8 |
| EVGA RTX 3090 Ti, 300 W | 322 | 1.771 | 282.3 | 528.0 |
| ZOTAC RTX 3090 Ti, 300 W | 305 | 1.873 | 268.2 | 528.9 |
| ZOTAC RTX 3090 Ti, 450 W (stock) | 341 | 1.645 | 369.9 | 652.9 |

The conditions (caps read back, ECC, cores, fans) are in the kit's render-bench note; the one that needs saying here: the build and the driver (595.84) were read at 17:55Z, and the 5080's files, written hours earlier (07:19 to 09:59Z, its 360 W rows first), record the torch but neither the commit nor the driver, with the card swapped out in between. Only these rows may be set in a ratio with each other; the older render-bench rows, on other builds in other computers, are in the kit's tables and never share a column with the grid's joules.

**Per picture, at 300 W:** the 5080 did the work in 0.69 of the EVGA board's time for 0.64 of its energy; against the ZOTAC board, 0.65 and 0.64. At stock (the 5080 at 360 W, the ZOTAC board at 450 W; the EVGA board has no stock ten-minute run in this computer), 0.74 of the time for 0.53 of the energy. **Per sampling step, at 300 W, 512×512, one picture's share of a batch of three** (32-step recipe median minus 8-step median, over 24, from the kit's recipes table): 0.1050 s on the 5080, 0.1540 on the EVGA 3090 Ti, 0.1606 on the ZOTAC; the two 3090 Tis take 1.47 and 1.53 times as long, and the per-batch column, where the text encoder ran fresh on every card, gives 1.47 and 1.52. The two 3090 Tis have their own page, next on this site; the EVGA board's grid column is not in this version (Who ran where). What the computer and the build moved on one of them, the EVGA 3090 Ti a week earlier in the i9 computer on the old build, is measured beside the old rows in the kit's tables.

**fp8, briefly:** on the RTX 5080 vs RTX 3080 Ti page an fp8 build of klein 4B (its text encoder in fp4), warm with the prompt's encoding cached, drew 1.64 to 1.79 times as fast as the bf16 file on the 5080 and 1.05 to 1.06 times on the 3080 Ti Founders Edition, which has no fp8 math; the workstation card and the laptop are Blackwell too, with the same generation of fp8-capable tensor cores, and neither ran that file. The full story is [the fp8 door](/rtx-5080-vs-rtx-3080-ti/#pictures-and-the-fp8-door) on that page.

## The full tables {#the-tables}

Every table this page does not print is in [the data kit](data/), the tables file published beside this page in its `data/` folder, 29 tables with their headings and conditions: every cell's seconds and joules at every size at stock and at 2048×1280 at 300 W, the mean draw of every 2048×1280 cell, the per-step figures at every size, peak memory, the recipes, every ten-minute row this site has measured, what 300 watts buys on the render bench, the change and size tables of findings 5 and 6, the cold loads, the predictions, the dated windows, and the notes this page points at (the cells' spread, residency, temperatures, the neighbours, the UPS logs and the render bench's conditions).

## What this page cannot say {#what-this-page-cannot-say}

- **A ratio between two grid cards as anything but "same software, different computer".** No card has run the grid in two computers. The nearest measurement of the computer's share is the 5080 in two computers the same day: 1.237 s on the grid's klein 512×512 8-step cell in the i9 computer, 1.212 s a picture on the same recipe at the same 360 W over ten minutes in the bench computer; two per cent, one reading on one cell from two instruments and two Python installs, not a bound and not a correction.
- **What a 12 GB card does with the encoder kept off the card, a smaller encoder or a quantised transformer, or what paging costs.** Not measured here; the INT8 build is in finding 1, and no model ran both whole and paged on one card.
- **Batches above one, `torch.compile`, Sage attention, `--fast` and fp8 math, other samplers, clock locks and undervolts.** Held fixed or absent.
- **The cost of the debug log.** Every instance ran `--verbose DEBUG` so the runtime would print its residency, and it prints every sampling step, not once a picture (about 30 lines a step on FLUX.2 [klein] 4B and 60 on FLUX.1 [dev]), so its cost sits inside the per-step figure as well as the fixed cost, on the host, heaviest where a step is shortest (0.040 s, the workstation card's klein step at 512×512); it was not measured. ComfyUI took the same attention path on every card (every log reads "Using pytorch attention"); which PyTorch kernel that path picked on each chip was not recorded.
- **Power at the wall socket.** The three desktop UPSes were a safety stop only, none a wall meter: the workstation and bench computers' read on mains throughout; the i9 computer's log read a frozen 9 per cent through the 5080's 300 W pass while the card drew 272 to 297 W at the mean, and it will ship in the kit marked as such (the kit's UPS note has the sample counts).
- **The room temperature, or the laptop's fan.** No room temperature was read; the laptop's driver cannot report its fan, so it was set to maximum by hand; the bench computer's three USB fans are turned up by hand before every run, and the run records that they were.
- **The cold loads, as a headline.** Recorded, and in the kit's tables, but the page cache was not controlled: the bench re-hashes every model file before it runs, and what that leaves in memory differs by computer.
- **Any picture, judged.** Every picture is kept with its sha256 and none is judged; no picture from any model is published.

## How it was measured {#how-it-was-measured}

**The grid.** One harness, built in its own folder on each computer, its model files linked read-only where their sha256 matched the pin, else downloaded and checked. Timing is ComfyUI's own `execution_start` to `execution_success`. Power, temperature, clocks, fan and the enforced limit were read twice a second on every card in the computer; the headline power is `power.draw.instant`.

**The stack, in full.** Every card ran ComfyUI at commit `3216c62e` (v0.34.0), CPython 3.14.4 built by uv and torch 2.13.0+cu130, with the same 103 packages from one hashed lock. ComfyUI started with `--verbose DEBUG --cache-classic --preview-method none`, locked to its card by UUID, in its normal memory mode (no low-memory flag) and **without `--fast`**. That flag and the file, a plain fp8 cast with no scale keys, decide what the FLUX.1 [dev] cells measure: the weights are stored in fp8 and the math is done in bf16 on every card, so the fp8 hardware in the Blackwell cards was never used. Five differences from a desktop ComfyUI: every render re-ran its text encoder (the part that reads the prompt), as a new prompt does, which re-rolling seeds on one prompt skips (0.1 to 0.7 s a picture here, on the driver's websocket clock); previews were off; every model ran at cfg 1.0, one model call a step, where cfg above 1 makes two and roughly doubles the per-step cost; every instance ran the debug log, about 30 to 60 lines a step (What this page cannot say), so a normal ComfyUI should match or beat these per-step figures by an amount not measured; and the laptop ran with its fan at maximum and its usual processor services running, so on its own fan curve, or with a busier processor, it may get less; neither was measured.

**What ran on each card, in order.** Each model starts in a fresh ComfyUI, its first picture (load included) kept apart; then its six cells, small to large, each three timed pictures on prompts A, B, A after a warm-up on B, the text encoder run every time. A check after every picture stops the run if anything else used the card, the limit moved in any sample or the UPS left mains. The stock grid is followed by the ten-minute run (klein 4B, 1024×1024, 8 steps, stopping at 83 °C), then the desktop cards' 2048×1280 cells at 300 W; every cap change was typed by a person. Not ranked here: which card is "better", any single picture, anything in money.

| Model | Sampler | Scheduler | Guidance (cfg) | Steps in the workflow run |
|---|---|---|---:|---:|
| FLUX.2 [klein] 4B, bf16, Apache-2.0 | `euler` | `Flux2Scheduler` | 1.0 | 8 |
| Z-Image-Turbo, bf16, Apache-2.0 | `res_multistep` | `simple`, shift 3 | 1.0 | 8 |
| FLUX.1 [dev], fp8 checkpoint (Comfy-Org's all-in-one file), FLUX.1 [dev] Non-Commercial License | `euler` | `simple` | 1.0, guidance 3.5 | 20 |

klein's 8 is the workshop's recipe count; the other two are their official ComfyUI templates' counts.

**The neighbours.** The workstation card's neighbour, the EVGA GeForce RTX 3090 XC3 Ultra 24 GB, kept 18,072 MiB of live language services and idled at 26.5 to 27.0 W; its one 17 s request during the 300 W ten-minute run moved nothing measurable in that run's median. The i9 computer's own services were paused for the pass. In the bench computer a second board is allowed only when named, held idle and watched: the EVGA GeForce RTX 3080 Ti XC3 Ultra 12 GB below the 3080 Ti held at 1 MiB through all 7 loads, at 27.4 and 27.2 W at the median over the two passes, in the air the Founders Edition cooler draws (finding 4). The rule and every reading are in the kit's neighbours note; the 2 Hz trace follows with the raw files.

**The dated windows (UTC).** The workstation card, the 5080 and the laptop ran the grid on 2026-09-24 (16:18 to 16:52Z, 17:47 to 18:47Z and 19:20 to 20:04Z), the 3080 Ti on 2026-09-25 (13:31 to 14:29Z); the bench computer's render bench ran on 2026-09-24 between 07:19 and 22:40Z; the older render-bench rows ran in the i9 computer between 2026-09-17 01:55Z and 2026-09-18 01:09Z and on the laptop on 2026-09-21. Every window to the second, every cap's read-back in every sample and the idle draws before each pass are in the kit's windows tables.

## How to check our work {#how-to-check-our-work}

The pre-registration was frozen before the first picture, revised once before that freeze after two review panels, and amended in words for the 12 GB card before its pass. **Its predictions:** the three this page can score held; for the 12 GB card, the two testable predictions registered in words before its pass held and the three rules registered with them were followed (the kit's predictions table); the four registered for the EVGA GeForce RTX 3090 Ti FTW3 Ultra's grid pass, which is not in this version (Who ran where), are scored as registered when its column is added, and any that cannot be is printed as unscored. The kit's first drop, published with this page, is the tables file alone. The raw files the tables come from (`grid.json` with the runtime's residency lines, `sustained.json`, the cell graphs, the 2 Hz traces, the UPS logs, the websocket events, ComfyUI's logs, the lock file) and `reduce.py`, which re-derives every table from the raw render records (ComfyUI's stamps, the 2 Hz power samples and the runtime's log lines) with a `--check` that refuses on a single moved digit, follow once cleared of the workshop's private paths and names, with a manifest of every file's size and sha256, and this section will say when; the render bench's and the older tests' files follow with them. That check passed on every result file in our own run, and you can repeat it when they ship; the tables file passed the harness's gate before this page was published. The pictures, the prompts' full text and the harness code are withheld, and named as withheld. Anything measured again or corrected later is added under its own dated section, and the byline says when.

## More on this site {#more-on-this-site}

- [A short history of Black Forest Labs](https://research.strata2signal.com/who-makes-flux/) — the makers of FLUX; the companion to this page.
- [RTX 5080 vs RTX 3080 Ti](/rtx-5080-vs-rtx-3080-ti/) — the fp8 door in full.
- [A laptop, asked the desktop's questions](/a-laptop-asked-the-desktops-questions/) — three postures.
- [RTX 3090 vs RTX 3090 Ti](/one-3090-against-one-3090-ti/) — the i9 computer's rows.
- [Two 3080s against one 3090, and the cap decides](/two-used-3080s-priced/) — the pair.
- [What 150 watts buys](/what-150-watts-buys/) — the workstation card's earlier cap ladder.
- [How the print lab works](/how-the-print-lab-works/) — the recipe's home.
- [The hardware roster](/hardware/) — every published figure, card by card.

## Who ran this, and thanks {#who-ran-this-and-thanks}

The cards are named in full in Who ran where and Three cards in one computer; the laptop is an **MSI Raider 16 Max HX**. **NVIDIA** for the silicon, the driver every reading came from and the Founders Edition board; **ZOTAC**, **EVGA** and **MSI** for the boards, and **Gigabyte** for one in the kit's tables. The pictures stand on **ComfyUI** and the **comfy-kitchen** and **comfy-aimdo** libraries beneath it, on **PyTorch** and **NVIDIA's CUDA**, on **Astral's uv** and on **Network UPS Tools**. The models are Black Forest Labs' **FLUX.2 [klein] 4B** and **FLUX.1 [dev]**, Alibaba's Tongyi Lab's **Z-Image-Turbo**, Alibaba Cloud's **Qwen 3 4B**, and Google's **T5-XXL** and OpenAI's **CLIP ViT-L/14**, the text encoders inside the FLUX.1 [dev] checkpoint; **Hugging Face** hosted every pinned file. None of them owed us anything.

A small human team owns the cards, ruled the grid and its refusals before a picture was drawn, pasted every cap and signed the numbers; a fleet of AI agents built the harness, ran it and did the arithmetic under that team's rulings. Thanks to the people who ask, on every forum this workshop reads, what a picture costs on the card they have.

*Measured on 2026-09-24 and 2026-09-25 by one pinned stack on four computers and one instrument in the fourth, with the label on every cross-computer figure.*

<!-- derived 2026-09-26 (UTC) by tools/derive_md.py from the pour source.
     source html sha256: 6dab24b0549bbc0faa35f38fadb18990190ddb9fc58455a93bfad9b93de42b77
     derivation sha256:  dbdd1c7fe13093e6b8cb9df48e2b3bf21ef45c0bf35035ae503e9fc4fa0c02a2
     the {#id} on each heading is the anchor that heading carries on the page. -->
