The bench — what this page measures, on what, and what it cannot say

A laptop, asked the desktop's questions

exhibit fifty-seven The bench
Published 2026-09-22 (UTC)
A small (human) team and a fleet of AI agents.

One NVIDIA GeForce RTX 5090 Laptop GPU, 24 GB, soldered into a gaming laptop, put through the same language and render bench this shelf ran on a 3090, a 3090 Ti and a pair of 3080s — and measured three times: held to its 95-watt default; with NVIDIA's Dynamic Boost left to float the budget, which it held between 140 and 150 W; and in this machine's own boot state, boost and a clock lock together, which changed nothing under load and four watts at idle. The finding is a ratio, not a deficit: a part allowed 95 watts, where the desktop cards beside it are allowed 300 to 450, does most of the work on the model this workshop runs, loses a third of it on a dense one, and gets the third back the moment it is allowed 150.

ask about this page assistant.strata2signal.com · in beta, still being tested

hardware on this page: RTX 3090 (24 GB) · RTX 5090 Laptop GPU (24 GB) · RTX 3080 (10 GB) · RTX 3090 Ti (24 GB) · Rig B — the gaming laptop research.strata2signal.com/hardware/ · the roster is in beta, still being tested

the short version

Yes, a laptop GPU is good for this. Held to 95 W, this part wrote at 118.4 tokens a second on the mixture-of-experts model this workshop actually runs, where a 3090 at 350 W wrote at 136.6 — 87 per cent of the speed on a third of the power the 3090 drew — and held the same 131,072-token window. On the dense 24-billion-parameter model it lost outright: 28.8 tokens a second against 46.9. In ten minutes of continuous drawing it made 294 images to the 3090's 311, peaking at 58 °C. With Dynamic Boost running, which is how a laptop like this arrives, the board's own limit read 140 to 150 W across the pass (n=374 at five-second intervals; 150 W in 335 of them) and never approached its 175 W maximum; on that budget the part wrote at 155.3 tokens a second on the mixture model — past the 3090 and past a 3090 Ti at its own 450 W limit (147.6) — brought the dense model to 44.9 against the 3090's 46.9, and drew 378 images in ten minutes to the 3090 Ti's 348 at 450 W, at between a third and two-fifths of the energy an image. A third pass put the boot-time clock lock this laptop normally runs with back on, and nothing moved under load; the lock costs about four watts at idle. The desktop comparator in the headlines is a 3090 at its own limit; a 3090 Ti is in every table, and it is ahead at both of its rungs on the dense model.

7,606 words, about 35 minutes to read.

The summary is this page’s own; the receipt lines were drafted by a model on this workshop’s network and every figure in them is in the article, checked before this page went out — what was dropped, and why, is in this page’s receipt file.

strata→signal is a small workshop that runs its own machines and writes up what it measures. This page is the fifth board through one instrument: the same language bank and the same render recipes that produced the 3090 and 3090 Ti and two-3080 pages on this shelf, with the same version-frozen runtime and prompt for the language rows, so the rows can sit beside each other.

Six words this page leans on

  • Mixture of experts (MoE) — a model that keeps all of its weights on the card but consults only a slice of them for each token. gemma4:26b, the model this workshop's products run, is one, and this page calls it the mixture model throughout. That slice is why it stays fast on modest power.
  • Dense — a model that reads every one of its weights for every token. mistral-small3.2:24b is one, this page's dense model, and it is where a power budget shows.
  • The power budget, and Dynamic Boost — a laptop GPU has a power budget its maker sets, and the same GPU name ships at different budgets in different laptops. This board's default is 95 W and its maximum is 175 W, both read off the board; a buyer looking at another machine with this part should read that machine's own budget before expecting these numbers. NVIDIA's Dynamic Boost is a background program that moves budget between the processor and the graphics chip as each one asks for it; with it running the board's limit floats, and with it stopped the board holds its default. Nothing can set a limit on this board.
  • Context window — how many tokens the model holds at once; 131,072 tokens is roughly a long novel. The cache that holds them, the KV cache, lives on the card, so a bigger window is a memory question before it is a speed one. The ladders below open a window of a given size and run a short prompt in it; the one arm that fills the window is named as such. The KV cache can be stored at three precisions, and this page tries all three. The planner is the part of the runtime that decides how much of a model goes on the card and how much stays in the computer's own memory; left alone it is cautious, and it can be told to put every layer on the card.
  • First token — the wait before the answer starts, as a program asking the runtime sees it. Most of it, on every board here, is the runtime's own per-request setup rather than the pass over the prompt; the page separates the two where it matters.
  • Board power, and an arm — every watt on this page is the card's own draw, read from its driver at 2 Hz, and not the laptop's or the wall's. An arm is one timed job of the bench — a model at one window size, three scored runs — and a pass is every arm run once under one posture.

If you own one of these, in plain words

Yes. For the mixture model this workshop runs its own products on, a laptop 5090 on its stock 95-watt budget writes at 118.4 tokens a second against a desktop 3090's 136.6 at 350 W, and holds the same 131,072-token window — on these two models there is nothing a 3090 fits that this board does not. The one real loss is the dense model: 28.8 tokens a second against 46.9. Letting Dynamic Boost run, which is the state a laptop like this arrives in, raised the board's draw by about 43 W at the headline window on the mixture model, closed that gap to 4 per cent and put the mixture model ahead of both desktop cards. Four people asking at once got 284 tokens a second between them over the wall clock under boost, and 231 at 95 W.

Which row is yours: a laptop you buy runs with Dynamic Boost on and no clock lock, so in every table below the dynamic boost row is the one to read. This shelf has not measured a desktop 5090; the desktop cards here are a 3090, a 3090 Ti at two of its rungs, and a pair of 3080s, and the 3090 Ti at its own 450 W limit is still 22 per cent ahead on the dense model. Two things to carry away with the numbers: every watt here is the card's own and not the wall's, and the fan was held at maximum by hand for the passes the headlines come from — what this board does on its own fan curve, on a lap, on battery, or how loud it is doing any of this, this page does not say.

The board, and the two things it will not do

The part is an NVIDIA GeForce RTX 5090 Laptop GPU with 24 GB of memory, soldered to the mainboard of an MSI Raider 16 Max HX gaming laptop (B2WJ-002US) with 64 GB of system memory, which ran on mains power throughout. The same GPU name ships in other laptops at other budgets; these figures are this machine's. Under load its link reads PCIe x8 at generation 5, and nothing on this page turns on that. Its memory clock held 14,001 MHz at the median through every rung of both context ladders and both render arms — the two small 95 W arms, the embedding batches and the one-caller gate call, sat at 9,001 — and that figure matters below: it is the ceiling the mixture model turned out to be waiting on.

It will not take a power cap. The bench was planned as a two-rung ladder, 175 W then 95 W, the way every desktop board on this shelf was measured. The first command an operator ran answered:

Changing power management limit is not supported for GPU: 00000000:01:00.0

With Dynamic Boost stopped, the board read a current limit of 95.00 W, equal to its default, with a maximum of 175.00 W. So the ladder was withdrawn before a single arm ran, and postures replaced it. The withdrawn rung was deleted rather than renamed: a result file named for 175 W would claim a setting that was never made.

PostureMechanismHow its power is printed below
Fixed, 95 WDynamic Boost stopped; the board enforces its own default, 95 W of a possible 175, and nothing moves itThe number 95 W, read back from the card at the open and the close of every stage and at 2 Hz inside every scored run; a change would have voided the stage, and none happened
Dynamic boostDynamic Boost running and no limit set, since none can be; the board floats its budget between its 95 W default and its 175 W maximum against the processor's draw. The state a laptop like this arrives inA range with its sample count, out of the pass's own trace, never a single wattage — and where a window read flat, the flat reading with its count
This machine's boot stateDynamic Boost running and a clock lock an operator added to this laptop, 1,200 to 2,550 MHz, read back from the card as in force at every gate. Not a state a buyer's machine is in unless they set itThe same rule as boost; in this pass the limit read 150 W in every one of its samples

It will not report its fan. The driver answers [N/A] for fan speed on this board, so no result file carries a fan column and none borrows one from another board's row. An operator set the laptop's fan to maximum by hand, with the machine's own max-fan key, before the first pass and did not change it — and, asked afterwards whether it was still at maximum for the third pass, was not sure. The card's temperature is in every file; the fan's speed is in none. Read every temperature on this page as a temperature under that condition — and every figure that might have throttled on an ordinary fan curve as one that did not get the chance.

One more condition. This laptop normally starts up with its GPU's clock pinned to a floor of 1,200 MHz and a ceiling of 2,550 — a setting an operator added so the picture service answers its first request quickly. The first two passes ran with that lock released, because the desktop boards these rows sit beside run no lock at all, and a board that cannot clock below 1,200 MHz is not honouring a 95 W budget the way an unlocked one does. A third pass ran the machine as it boots, lock and all, and its figures sit in a short section below and in a row of the tables where the lock could have mattered.

Five things the 95 W pass found

  • A quarter of the power allowed buys most of the speed, on the mixture model. gemma4:26b at a 131,072-token window: 118.424 tokens a second here, allowed 95 W and drawing 79.64 W, against 136.618 at 240.49 W on a 3090 held at 350 W. That is 148.7 tokens a second per 100 W against 56.8 — 2.6 times the work per watt, at 87 per cent of the throughput; the joule figures say the same from the other side, 672.50 J per 1,000 tokens against 1,758.13. One caution rides every watt figure on this model, and it is printed beside the comparison table: a mixture-model run here lasts about two seconds, and a two-second run sampled at 2 Hz gives a coarse mean.
  • The dense model is where the mobile part actually loses. mistral-small3.2:24b at 65,536: 28.839 tokens a second here against 46.880 on the 3090 — 61.5 per cent of it. A dense 24-billion-parameter model reads its whole weight set for every token, and a 95 W budget cannot feed that the way it feeds a mixture model's active slice. The efficiency lead survives but shrinks: 31.9 tokens a second per 100 W against 17.2. These runs last twenty seconds and their watt figures are tight.
  • The context ceiling is identical to the 3090's, on both models — a memory fact a power limit cannot move. gemma4:26b held every rung to 131,072; the ladder ran out of rungs before the card ran out of memory. mistral-small3.2:24b held 65,536 whole and spilled at 98,304 with 92.2 per cent of the model on the card. Both are the 3090's own headlines, rung for rung, and the two later passes repeated them exactly. Filling that 131,072-token window rather than opening it is a different job, and the gap widens there at 95 W: 56.174 tokens a second against the 3090's 73.327.
  • The planner left the card on the table, and forcing the layer count took the rung back. At 98,304 the dense model's automatic placement did not fit — 92.2 per cent on the card, 7.8 per cent in system memory — and that is where the headline ladder stops. The same rung with every layer forced onto the card fits whole, 21.52 GiB resident, at 28.851 tokens a second, inside the spread of the 65,536 rung. One rung higher, 131,072 forced, the runtime answers cudaMalloc failed: out of memory, and that refusal is written into its own file as the rung's outcome. Under boost the forced rung fits again, at 44.946, and the next one refuses again.
  • A 16-bit KV cache is faster than the 8-bit default on this board, and on the desktop card too. At 131,072: 121.379 tokens a second at f16 against 118.424 at q8_0, the default, and 116.125 at q4_0, the smallest — and f16 drew the least, 76.53 W and 622.11 J per 1,000 tokens. A 3090 at 350 W shows the same order in its own files, 146.308 against 136.618 and 135.249, so this is the runtime's doing, not the budget's. q4_0 did buy a much faster prefill — the pass over the prompt before the first word — at either posture, so the two are a trade, and the bench's tables carry both.

The render leg at 95 W: in ten minutes of continuous drawing this board made 294 images where a 3090 at 350 W made 311 — 94.5 per cent of the throughput on 29 per cent of the board power, 0.0529 watt-hours an image against 0.1763, at a peak of 58 °C against 74. The board's own power ceiling held its clocks down for 93 to 100 per cent of every timed window at 95 W.

What the extra watts bought

The second pass started Dynamic Boost and changed nothing else. Inside every scored run the board's limit was sampled at 2 Hz beside its power; over the fifty-seven runs of the two ladders and the scored arms it read 150 W in thirty-four of them from the first sample to the last and touched 140 W in twenty-three, and a slower witness over the whole pass — the runs and the gaps between them, every five seconds — read 150 W in 335 of 374 samples and 140 W in the other 39. Through both render arms the limit read 150 W in every one of their 1,223 samples at 2 Hz. It never floated near its 175 W maximum, and it never fell back to 95. So "dynamic boost" on this bench means a budget of 140 to 150 W, and mostly 150.

  • The mixture model's speed stopped following the power. gemma4:26b decoded at 155.2 to 156.0 tokens a second at every window from 4,096 to 131,072 — about 31 per cent faster than at 95 W at every rung (30.8 to 31.9 per cent), past a 3090 at 350 W (136.6), and past a 3090 Ti whether it is held to 350 W (147.8) or allowed its own 450 W limit (147.6). The memory clock sat at 14,001 MHz through both ladders at both postures, and the graphics clock went from a median of 1,523 MHz across the 95 W runs to 2,175 under boost. What this page cannot say is how much of the budget the board actually spent doing it: a mixture-model run is about two seconds long, eight samples at 2 Hz, and the three runs at 131,072 under boost read means of 98.6, 146.1 and 122.6 W for decode speeds within 0.3 per cent of each other. The cap-active flag was raised at seven of the eight rungs, so the board reached its budget in every run; how much of each run it spent there, a two-second trace cannot say.
  • The dense model is still power-bound, and every extra watt the second posture allowed went into speed. mistral-small3.2:24b decoded at 44.8 to 44.9 tokens a second at every window, 56 to 68 per cent faster than at 95 W, at 131 to 141 W mean board power — twenty-second runs, tight means — with the card's own power-cap flag raised for the whole of every scored rung (one run of eighteen read 0.900). That closes the gap to a 3090 at 350 W from 61.5 per cent of its speed to 95.8 per cent, at 3,019 joules per thousand tokens against 5,843 — and, unlike the mixture model, it got cheaper per token than at 95 W (3,132). Against a 3090 Ti allowed its own 450 W it is a different comparison: 57.4 tokens a second there, 78 per cent of it here, at 7,295 joules per thousand tokens against 3,019. A dense model on this board runs exactly as fast as the wattage it is allowed, and it was still gaining speed when the budget stopped at 150.
  • The renders went from slightly slower than the desktop cards to faster than both, at the caps their pages share. On the print lab's 8-step recipe — the print lab is this workshop's picture service, and a recipe is what its software calls a graph — 1.246 s an image here against a 3090's 1.509 s at 350 W and a 3090 Ti's 1.397 s at 350 W, from 1.640 at 95 W. In ten minutes of continuous drawing, 378 images here; 311 for the 3090 and 331 for the 3090 Ti at 350 W; 348 for the 3090 Ti at its own 450 W limit; 294 at 95 W. Energy per image went up from 0.0529 watt-hours at 95 W to 0.0633, which is 36 per cent of the 3090's 0.1763 and 39 per cent of the 3090 Ti's 0.1640. The board drew 145.9 W median through the ten minutes and its core plateaued at 68 °C, peaking at 71 — thirteen degrees hotter at the peak than at 95 W, and three under a 3090's 74 °C at 350 W, which drew 311 images in its ten minutes against this board's 378. The filled-window arm tells the same story from the language side: 72.159 tokens a second under boost against the 3090's 73.327.

The first token did not pay for the extra watts and did not gain from them either: across the mixture model's eight windows it moved between 18 ms faster and 29 ms slower with no direction. It is shorter here than on the desktop cards at both postures — 372 ms against the 3090's 553 at the 131,072-token window, 211 against 294 on the dense model — but most of that wait, on every board here, is the runtime's own per-request setup and not the card: the pass over the 536-token prompt itself took 75 ms here and 81 ms on the 3090, and the other 292 and 459 ms were setup in the runtime, which this bench did not isolate. The one arm that barely benefited was the embedding arm: 269.0 texts a second against 252.6 at 95 W, for 83.8 joules per thousand texts against 65.7. A small model that was never power-bound gained 6.5 per cent of speed for 28 per cent more energy.

The way it boots, measured

The third pass ran the machine exactly as it starts up: Dynamic Boost running and the boot-time clock lock in force, read back from the card as in force at every gate — the driver snaps the requested 1,200 MHz floor to 1,192 at rest — on an empty card two hours and twenty minutes after the boost pass closed, with nothing else changed on the box except a fan setting nobody could confirm. Under load the lock changed nothing — every arm came back within half a per cent of the boost pass's figure, the worst 0.43 per cent, four of seven outside the boost rung's own three-run spread and all of them tiny — and the lock's 2,550 MHz ceiling never bit: the highest graphics-clock reading of the boost pass was 2,392 MHz. Three things moved, and all three are the floor's doing. At idle the card cannot clock down below 1,200 MHz, so it drew 11.57 W empty where the unlocked board drew 7.14, and 14.42 W with the mixture model resident where the unlocked board drew 11.34 — 4.4 and 3.1 W, all day, for a lock that buys the picture service its first-load latency. Its idle link reads PCIe x8 at generation 2 where the unlocked board drops to generation 1. And the board's limit never dipped: 150 W in every one of the pass's 385 language samples and 1,224 render samples, where the unlocked boost pass had read 140 W in a tenth of its samples, all in the idle gaps between runs.

WhatDynamic boost, lock releasedThis machine's boot state
gemma4:26b at 131,072, tokens/s155.251155.288
mistral-small3.2:24b at 65,536, tokens/s44.89744.702
Four callers, over the wall clock, tokens/s283.934284.359
Ten minutes of drawing, images378382
The card empty, idle W7.1411.57
gemma4:26b held resident, idle W11.3414.42
The board's limit over the pass140 to 150 W; 150 in nine samples of ten150 W in every sample

The one figure that moved materially is the embedding arm, 250.7 texts a second against 268.988, back to where it stood at 95 W: a small model that finishes a batch in a quarter of a second is the arm most exposed to whatever else the processor was doing, and three runs are not enough to say more. Every other figure of the pass is in the bench's own tables.

The tables

Every figure below is the median of three scored runs unless its cell says otherwise. Board watts are printed one way throughout: the median, over the three runs, of each run's mean draw at 2 Hz. On the mixture model those runs are about two seconds long, eight to ten samples, and the three means spread widely for the same speed — the comparison table prints all three beside the median so the coarseness is visible; on the dense model the runs are twenty seconds long and the three means agree within a few per cent. Rows from the desktop boards were measured in another computer, at their own caps, by the same version-frozen runtime and prompt, on that computer's own driver, and the cap prints beside each of them. Those rows come from the sibling benches' own result files; their pages published the same readings, some at a coarser precision, and two columns here are computed from the files rather than quoted from the pages — the first-token times, which those pages do not print, and tokens a second per 100 W. Where a desktop row carries no spread its file carries none. The two-3080 pair's language rows are its 300 W run and its render rows its 250 W one, as that page and its files record. Watt-hours an image is computed the same way for every board — energy over the timed window divided by the images finished in it, from each run's own power trace — over ten minutes for the ten-minute table and over the timed batches for the recipes table, so the two tables' columns are not the same quantity. A flat wattage cell is a finding, not a range: under a floating posture the limit was never seen to move inside that pass. A dash carries its reason.

This board beside the desktop rows — gemma4:26b at a 131,072-token window

ConfigurationPower settingTokens/sFirst token, msBoard W (the three runs' means)J per 1,000 tokensTokens/s per 100 W
This board, fixed95 W118.424372.4979.64 (76.95 · 79.64 · 94.22)672.50148.7
This board, dynamic boost140–150 W (n=374; 150 W in 335)155.251372.92122.57 (98.58 · 146.10 · 122.57)787.76126.7
This board, its boot state150 W, flat (n=385)155.288371.01116.12 (116.12 · 115.76 · 146.42)747.97133.7
A GeForce RTX 3090 24 GB350 W, its own limit136.618552.75240.49 (275.82 · 240.49 · 222.39)1,758.1356.8
A GeForce RTX 3090 Ti 24 GB350 W, the shared-cap rung147.823541.49282.64 (284.40 · 254.06 · 282.64)1,909.9652.3
A GeForce RTX 3090 Ti 24 GB450 W, its own limit147.560539.43283.01 (229.38 · 302.00 · 283.01)1,915.1352.1
Two GeForce RTX 3080 10 GB, the model split across them, every layer forced onto the cards300 W each, as the file records it115.366571.40346, both boards, as the pair's own page prints it2,994, as that page prints it33.3

The three runs' means are the spread the median hides, and on this model it is wide on every board: a two-second run gives the 2 Hz trace three to seven busy samples. The joule and per-100-W cells inherit it. The dense table below does not have this problem.

This board beside the desktop rows — mistral-small3.2:24b at a 65,536-token window

ConfigurationPower settingTokens/sFirst token, msBoard W (mean)J per 1,000 tokensTokens/s per 100 W
This board, fixed95 W28.839210.9490.313,131.5631.9
This board, dynamic boost140–150 W (n=374; 150 W in 335)44.897203.72135.763,019.4033.1
This board, its boot state150 W, flat (n=385)44.702191.65129.902,904.3534.4
A GeForce RTX 3090 24 GB350 W, its own limit46.880294.27273.355,842.9217.2
A GeForce RTX 3090 Ti 24 GB350 W, the shared-cap rung53.354272.09338.056,343.8115.8
A GeForce RTX 3090 Ti 24 GB450 W, its own limit57.396275.63418.767,295.4913.7
Two GeForce RTX 3080 10 GB, the model split across them, every layer forced — at 49,152, the pair's largest whole window300 W each, as the file records it44.702318.59509.01, both boards11,385.408.8

The largest window each configuration held whole

ModelThis board, at every postureA 3090 at 350 WA 3090 Ti at 350 W and at 450 WTwo 3080s at 300 W each
gemma4:26b131,072 — the ladder ran out of rungs first131,072131,072131,072, layers forced
mistral-small3.2:24b65,536 — spilled at 98,30465,536 — spilled at 98,30465,536 — spilled at 98,30449,152, layers forced

Ten minutes of continuous drawing — the 8-step recipe, one image at a time

Board and posturePower settingImages in ten minutesSeconds per image (spread)Wh per imagePlateau (°C, and when reached)Core max °CBoard W median/maxFan max %
This board, fixed95 W294 (600.9 s)1.958 (1.918–2.019, n=294)0.052954 °C at 3.3 min5893.8 / 100.6Unreadable; set to maximum by hand
This board, dynamic boost150 W, flat — every one of 1,141 samples at 2 Hz378 (600.6 s)1.521 (1.476–1.574, n=378)0.063368 °C at 5.5 min71145.9 / 156.9Unreadable; set to maximum by hand
This board, its boot state150 W, flat — every one of 1,140 samples at 2 Hz382 (600.5 s)1.507 (1.475–1.590, n=382)0.062968 °C at 4.6 min70146.4 / 155.2Unreadable; not confirmed for this pass
A GeForce RTX 3090 24 GB350 W, its own limit311 (601.6 s)1.886 (1.851–1.975, n=311)0.176371 °C at 0.9 min74325.8 / 344.867, on its own curve
A GeForce RTX 3090 Ti 24 GB350 W, the shared-cap rung331 (601.4 s)1.715 (1.702–1.819, n=331)0.164055 °C at 0.4 min60316.2 / 353.973, on its own curve
A GeForce RTX 3090 Ti 24 GB450 W, its own limit348 (602.1 s)1.672 (1.660–1.711, n=348)0.174257 °C at 0.3 min62357.1 / 390.475, on its own curve
A GeForce RTX 3080 10 GB, one of a pair250 W158 (602.6 s)3.737 (3.512–3.853, n=158)0.237177 °C at 5.6 min82220.0 / 252.692, on its own curve

The rows did not do the same amount of work. The image counts are in the table, and a board that draws more images in ten minutes makes more heat for that reason and not because it runs hotter; read every temperature against the image count beside it, every plateau against the power setting beside that, and this board's rows against a fan an operator held at maximum by hand where the desktop boards ran their own curves.

The print lab's own recipes — FLUX.2 klein-4B, batches of three

The three certified recipes the print lab ships: eight steps, eight steps with a pixel pass, and thirty-two steps. For every board, two timed batches of three images each after a warm-up submit that is excluded; the spread is across the two batches. This board's render rows ran the same ComfyUI release as the desktop boards on a newer torch and driver, which a Blackwell board needs; the ratios are box against board, not card against card.

Board and posturePower setting8 steps · s per image (spread, n=2)8 steps · Wh per image8 steps + pixel pass · s per image (spread, n=2)8 steps + pixel pass · Wh per image32 steps · s per image (spread, n=2)32 steps · Wh per imageBoard W median/max, 8 stepsCore max °C, 8 steps
This board, fixed95 W1.640 (1.583–1.698)0.04311.627 (1.579–1.675)0.04296.086 (6.028–6.145)0.160794.3 / 96.044
This board, dynamic boost150 W, flat — 82 samples at 2 Hz across the three recipes' timed batches1.246 (1.202–1.290)0.05001.245 (1.194–1.296)0.05134.649 (4.584–4.715)0.1934148.1 / 151.350
A GeForce RTX 3090 24 GB350 W, its own limit1.509 (1.457–1.562)0.13691.528 (1.477–1.580)0.14275.670 (5.598–5.741)0.5292337.3 / 341.5— the comparison carries none
A GeForce RTX 3090 Ti 24 GB350 W, the shared-cap rung1.397 (1.351–1.443)0.13091.379 (1.332–1.426)0.13075.130 (5.082–5.177)0.4956347.5 / 349.9
A GeForce RTX 3090 Ti 24 GB450 W, its own limit1.369 (1.330–1.409)0.13911.350 (1.308–1.392)0.14185.015 (4.972–5.059)0.5381378.5 / 382.7
A GeForce RTX 3080 10 GB, one of a pair250 W2.074 (1.825–2.323)0.13982.067 (1.830–2.305)0.13977.278 (7.025–7.531)0.5014248.6 / 249.8

The scored arms at 95 W

ArmTokens/s (spread of three)First token, msBoard W (mean)J per 1,000 tokensCore max °C
gemma4:26b, KV cache q8_0 (the default), 131,072 — the ladder's own rung, for reference118.424 (117.802–119.239)372.4979.64672.5042
gemma4:26b, KV cache f16, 131,072121.379 (120.822–123.017)435.9676.53622.1148
gemma4:26b, KV cache q4_0, 131,072116.125 (115.846–116.716)326.2786.87744.2956
gemma4:26b, the window filled to 131,07256.174 (56.039–56.374)744.9288.381,573.3264
A 3090 at 350 W, the window filled to 131,072, from its own page's files — for reference73.327 (73.042–73.392)1,113.07293.383,997.4562
mistral-small3.2:24b, every layer forced onto the card, 98,30428.851 (28.721–28.993)214.3793.133,227.9945
mistral-small3.2:24b, every layer forced onto the card, 131,072Refused — cudaMalloc failed: out of memory
minicpm-v4.5, the vision model the workshop's products use to read pictures, measured here on its text path only86.808 (85.795–87.364)149.8679.27913.1641

The scored arms under dynamic boost

The f16 arm's limit read 150 W in every one of the 24 samples of its own 2 Hz trace; the other arms touched 140 W.

ArmTokens/s (spread of three)First token, msBoard W (mean)J per 1,000 tokensCore max °C
gemma4:26b, KV cache q8_0 (the default), 131,072 — the ladder's own rung, for reference155.251 (155.128–155.593)372.92122.57787.7644
gemma4:26b, KV cache f16, 131,072165.656 (165.545–166.175)439.78116.01700.3144
gemma4:26b, KV cache q4_0, 131,072155.712 (155.641–155.908)336.23148.71955.0344
gemma4:26b, the window filled to 131,07272.159 (72.024–72.171)721.95123.991,718.0148
mistral-small3.2:24b, every layer forced onto the card, 98,30444.946 (44.917–44.963)208.97140.483,124.3548
mistral-small3.2:24b, every layer forced onto the card, 131,072Refused — cudaMalloc failed: out of memory
minicpm-v4.5, the vision model, text path only121.733 (121.722–121.762)132.94144.741,188.7243

Under boost the two quantised caches changed places by 0.461 tokens a second, three-tenths of a per cent — the two arms' three-run spreads do not overlap, 155.641–155.908 against 155.128–155.593, so it is a small real difference rather than noise — while f16's lead widened to ten tokens a second.

Several callers at once — gemma4:26b, one, two and four streams

The summed figure adds each stream's own rate; the wall-clock figure divides every token produced by the wall time the slowest stream took, which is what four people asking at once actually get between them. The per-stream figure is what each of them sees. The boost rows were read with the board's limit at 140 to 150 W over the arm (n=33 at five-second intervals; 150 W in 30 of them).

PostureStreamsSummed tokens/sOver the wall clock, tokens/sPer streamFirst token, msJ per 1,000 tokens
Fixed, 95 W1118.368102.226118.368345.52857.75
Fixed, 95 W2199.413165.32199.950525.93504.27
Fixed, 95 W4275.732231.23969.255702.42396.86
Dynamic boost1155.231124.544155.231401.70829.49
Dynamic boost2256.955202.821128.550528.06626.27
Dynamic boost4351.997283.93487.994694.10405.97

The small arms — a gate-shaped call, embeddings, and what it costs to leave a model loaded

A gate-shaped call is the short question a product asks before it answers a visitor — is this in scope, which page — and the embedding arm is the model that turns text into the vectors a search runs on. The boost cells were read with the limit flat at 150 W over their short stages (n=11 and n=5 at five-second intervals) and the boost idle cells with the limit at 140 to 150 W (n=19); every boot-state cell with the limit flat at 150 W.

ArmFixed 95 WDynamic boostThis machine's boot state
A short gate-shaped call on mistral-small3.2:24b, one caller, with the model split 56.2 % on the card and 43.8 % in system memory at every posture — wall latency, median of ten461.64 ms (p95 493.52; first token 314.52 ms; 27.81 W mean)438.98 ms (p95 468.26; first token 298.34 ms; 32.11 W mean)430.85 ms (p95 448.20; first token 293.91 ms; 32.37 W mean)
The same, four callers at once — median of eight928.54 ms (p95 4,671.24; first token 465.75 ms; 66.59 W mean)948.56 ms (p95 4,567.46; first token 569.39 ms; 73.56 W mean)885.73 ms (p95 4,452.80; first token 470.94 ms; 70.93 W mean)
nomic-embed-text, batches of 64 texts, 768 dimensions252.570 texts a second · 253.40 ms a batch · 65.67 J per 1,000 texts268.988 texts a second · 237.93 ms a batch · 83.79 J per 1,000 texts250.655 texts a second · 255.33 ms a batch · 90.19 J per 1,000 texts
The card empty, idle — mean over 30 s7.18 W (57 samples)7.14 W (57 samples)11.57 W (58 samples)
gemma4:26b held resident at 131,072, idle — mean over 30 s9.03 W, so residency costs 1.85 W11.34 W, so residency costs 4.20 W14.42 W, so residency costs 2.85 W

Every rung of both context ladders at all three postures — eight windows for the mixture model, seven for the dense one, with the first-token, watt, joule and temperature figures of each — is in the bench's own tables and not repeated here; the speeds do not move with the window on either model, and the comparison tables above carry the rung the desktop pages published.

What this page cannot say

  • Wall watts. There is no metered supply on this laptop, and its battery exposes no energy counter to a program running without administrator rights. The desktop pages carry wall watts from a metered supply; those columns are not compared here.
  • Noise, battery, and heat where your hands are. The fan ran at maximum for every pass and nothing here measures how loud that is; every pass ran on mains power and nothing here says what this board does unplugged, or at what budget; every temperature is the core's, and nothing was read at the keyboard, the palm rest or the underside. On a laptop those are the first three questions, and this page answers none of them.
  • An ordinary fan curve. The fan was set to maximum by hand before the first pass and confirmed there for the first two; for the third, an operator was not sure. No instrument on this box can read it, and nothing here says what the same board does with the fan on its own curve.
  • A control seat. Every earlier board on this shelf came out of the same slot as the one it was compared with, so the swap was the control. This part is soldered to its mainboard: there is no seat, nothing to swap, and no desktop board shares either of this page's wattages. The desktop rows are another machine's measurements at their own caps, printed with the cap in the cell — not a control.
  • How much of the budget the mixture model actually used. Its runs are two seconds long, and a 2 Hz trace over two seconds cannot say more than the comparison table's three means do. The dense model's runs are ten times longer, and its watt figures are the tight ones.
  • The training arm. The 24 GB desktop pages carry a row for training a small adapter on the workshop's own images. The prepared tensors for it live on a machine that has been unreachable since 15:03 UTC on the 21st, before this bench ran, so the arm is absent rather than failed, and no figure on this page depends on it.
  • A memory-bandwidth ceiling. This board's memory is GDDR7, and the harness's peak-bytes-per-second formula is a GDDR6 and GDDR6X fact checked against three published bandwidths, so the ceiling is withheld rather than computed under a formula that does not describe this memory. The memory clock above is a reading, and prints.
  • Where the extra watts would stop paying. The boost pass only ever saw a 140 to 150 W budget. Whether the dense model keeps climbing toward the board's 175 W maximum is a question this bench could not ask, because nothing on this board lets anyone set the budget.
  • A desktop 5090. This shelf has not measured one. Nothing here says how the laptop part compares with the card that shares its name.
  • A price, or a row on the comparison page. This part is soldered into a laptop and is not sold as a card. The whole machine was bought new from Micro Center for $3,099.99 before tax, ordered 2026-07-03; a buyer is weighing that whole machine against a card plus the rest of a computer, and the desktop half of that sum is on this shelf's price pages. So the price tables on this shelf's comparison page carry no row for it — and nor do its other tables: they key their columns on 250 to 600 W, and this board's whole envelope, a 95 W default and a 175 W maximum, sits below the lowest of them, so no reading from this bench can fill a cell keyed on a cap it was never measured at. This page is where its rows live, with the wattage in the cell instead of in the column head.

How it was measured

The language bank ran against a separate copy of ollama at the version the desktop legs ran, 0.32.13, on its own port, with keep-alive at zero, a q8_0 KV cache and one parallel slot — four for the concurrency arm — on this box's own driver, 595.91.07, where the desktop box ran 595.71.05. The render arms ran against a separate ComfyUI 0.21.1 checkout that refuses to start on any other version, drawing the print lab's three certified recipes with their bytes verified before the first submit — on torch 2.12.1 and that driver, where the desktop boards ran torch 2.11.0 on theirs, because a Blackwell board cannot run the older build; that is the difference the render ratios carry, and the reason they are read as a box-and-board comparison rather than a like-for-like one. Every rung is three scored runs of one frozen prompt; the first-token and decode figures are the median of the three. Board power, clocks, temperature and the enforced power limit were sampled at 2 Hz through every scored run, and the limit was read again at the open and the close of every stage: under the fixed posture a change would have voided the stage, and none happened; under the floating postures the readings are the ranges the wattage cells print.

This is not a quiet box, and the bench did not make it one. The laptop serves four small web services and four database containers, and a system copy of ollama holds an embedding model on the same card whenever something has asked for one. A contention gate refused to start any scored run that found the card more than 5 per cent busy over the ten seconds before it and tried again, on the rule that six refusals in a row would void an arm: it found the card busy on 23 of its 77 checks across the three passes and waited, the busiest read 99 per cent, and no run needed more than three tries, so nothing was voided. What the gate reads is the card; what the processor was doing during a run is not gated, and under the two floating postures the processor's draw is what moves the board's budget. A memory guard that fires every 30 s on this machine was paused for the duration of each pass and restarted on the way out.

The 95 W pass ran from 18:43 to 19:34 UTC on 2026-09-21 with the card otherwise empty. The boost pass that the tables read ran from 21:01 to 21:50 UTC the same day, on a card this pass emptied and kept empty: its first attempt, from 20:24, had to be set aside when the render leg's own gate refused to start on a card another process was holding — the system's embedding model had been loaded onto it at 19:51 by an unrelated call and held there — so the model was unloaded before the re-run (it reloads on its next call, and nothing was lost), the hourly job that could have reloaded it was paused for the window and started again afterwards, and the card was checked empty at the re-run's open and again at its render leg's gate. That first attempt's language leg is kept with the receipts as a replication; it agrees with the re-run within nine-tenths of a token a second at every window of the mixture model. Between the first two passes an operator started Dynamic Boost and nothing else. The third pass, boost and the clock lock, ran from 00:10 to 00:56 UTC on 2026-09-22, again on a card checked empty at its open and at its render leg's gate, with the lock put back by an operator's one command and read back from the card as in force at every gate. The desktop rows are those benches' own result files, read the way the tables' preamble says.

How to check our work

The pre-registration — written before the first arm, amended once when the board refused a cap and once more to register the third posture, each amendment dated with its reason quoted — all three passes' logs, an operator's conditions file, the superseded first attempt, and every result file named in the tables exist as this bench's record, and they follow as this page's data kit once they have been cleared of the workshop's private paths and names; this section will say when. Each language result carries the instance's environment, the runtime's own version answer, the card's identity and the prompt's hash; each render result carries the recipe's hash and the version the checkout answered. Anything measured again or corrected later is added under its own section, dated at both ends, with the comparison to what stood before, and the byline says when.

The rest of the seminar

Who ran this, and thanks

The board is an NVIDIA GeForce RTX 5090 Laptop GPU 24 GB in an MSI Raider 16 Max HX, and NVIDIA gets the credit for the board, for the driver every watt, clock and temperature on this page was read from, and for Dynamic Boost, the mechanism the second pass measured; MSI for the chassis that carried it and the max-fan key an operator leaned on. The language work stands on ollama and, beneath it, llama.cpp and NVIDIA's CUDA; the pictures on ComfyUI. The models are Google's Gemma 4, Mistral Small 3.2 from Mistral AI, MiniCPM-V 4.5 from OpenBMB, nomic-embed-text from Nomic, and Black Forest Labs' FLUX.2 Klein 4B with its decoder, which drew every picture timed here, and Qwen from Alibaba Cloud as the text encoder in those recipes. None of them owed us anything, and every one of them ran on a laptop.

A small human team owns the laptop, set the rules and the refusals before the runs, held the fan and signed the numbers; a fleet of AI agents ran the harness and did the arithmetic under that team's rulings. Thanks to the people asking, on every forum this workshop reads, whether a laptop GPU is any good for this — that question is the whole of this page.

Measured on 2026-09-21 and 2026-09-22 by one instrument on one laptop, with the fan held by hand and the wattage beside every figure.

Planned maintenance tonight, 04:00 to 04:45 UTC: the parts of this workshop that think — the long table's chairs, the assistant, and the answers behind RuleSage, amble and the Beat Lab's genie — are offline for a hardware test. Every page stays up. Come back after. · D-20260921-206

elsewhere in the workshop

a strata→signal property · hello@strata2signal.com · say hello