# RESULTS — a 24 GB laptop GPU, allowed 95 watts, asked the 24 GB desktop ladder's own questions

*What this is, for a reader who scrolled straight here: the language bank the estate's 24 GB
desktop cards ran — a 3090, a 3090 Ti, and a pair of 3080s before them — run again on the
**NVIDIA GeForce RTX 5090 Laptop GPU 24 GB** soldered into a gaming laptop, on the pinned
ollama the published comparison tables demand, against the same frozen prompt by digest and the
same model blobs. The measured pass ran 2026-09-21 between **18:43:29Z and 19:19:53Z**. Every
table below is generated from the JSON receipts in `results/` by `summarize_5090laptop.py`;
nothing here is typed in by hand and no figure is rounded past what its file carries. The rules
that produced the numbers are in `PREREG.md`, written before the first scored run.*

*The render half of this bench — arms R1 and R4, ComfyUI 0.21.1, the print lab's certified
graphs — is a separate document:
[`../render-5090-laptop-2026-09-21/RESULTS.md`](../render-5090-laptop-2026-09-21/RESULTS.md).*

---

## The one sentence

**On the model this shelf actually runs, a laptop part held to 95 watts wrote at 118.4 tokens a
second where a 3090 at 350 watts wrote at 136.6 — 87 % of the speed on a third of the board
power, and 672.5 joules per thousand tokens against 1,758.1.** It held the same 131,072-token
window the 3090 held, on the same planner setting, with the whole model resident. The dense
model is where the gap opens: 28.8 tokens a second against 46.9.

Both halves of that are readings of a board drawing **95 W**, and the cap is printed beside
every figure in this document because it is the whole point of it.

## What changed the shape of this bench, before any number

**There is no cap ladder here, and there was supposed to be one.** The operator's read-back at
about 17:20Z on the day answered:

```text
Changing power management limit is not supported for GPU: 00000000:01:00.0
```

**This board refuses to have a power limit SET, at any wattage.** So the planned 175 W → 95 W
ladder was withdrawn — `leg-laptop-175w.sh` is deleted and a fence keeps it deleted — and what
replaced it is three POSTURES, told apart by whether a daemon is floating the board's limit and by
whether this box's own boot-time GPU clock lock is in force:

| posture | what it is | its cap cell | state |
|---|---|---|---|
| **fixed, 95 W** | `nvidia-powerd` stopped, clocks released. The board enforces its own default limit and nothing moves it. | the number **95 W**, read back from the card at every stage boundary | **measured — this document** |
| **dynamic-boost** | `nvidia-powerd` running, no cap set, clocks **released**. The limit floats between 95 W and 175 W against the CPU's draw. | a **RANGE** with its `n`, from that pass's own witness — never a single wattage | **measured — this document** |
| **shipped** | `nvidia-powerd` running **and** this box's boot-time 1,200–2,550 MHz clock lock IN FORCE. **This is how the machine actually boots**, so it is what the laptop does on an ordinary day. | the same shape — and this pass's limit never moved, so it reads **`150 W, flat`** with its `n` and the finding stated | **measured — this document** |

**All three postures are in the tables below, and each pass's rows were ADDED without re-cutting a
table, a heading or a sentence.** That was the one structural decision worth stating when this
document was written with an empty dynamic-boost column, it held when that column filled, and it
held again on 2026-09-22 when the third pass landed: the tables it went into were the tables that
had been printing `— not yet measured` in its cells since the hour it was registered. A posture
with no files still prints that placeholder, never a blank and never a borrowed figure.

**175 W is reachable on this board only by letting Dynamic Boost float the limit there**, which
is the second posture. A file named `…-175w…` would claim a setting nobody made, so there is no
such file and no such row.

**The dynamic-boost pass, and the attempt before it — written 2026-09-21.** What this paragraph is
about, for a reader who scrolled straight to it: every `dynamic-boost` figure in this document
comes from ONE pass over both legs of this bench, run **2026-09-21 between 21:01:51Z and
21:50:30Z** with the card otherwise empty (the receipt is
[`RUNG-DYNBOOST-RECEIPT.txt`](RUNG-DYNBOOST-RECEIPT.txt), and both legs exited 0).
**It is the second attempt.** The first started at **2026-09-21 20:24Z** and its render leg
refused to start at **20:56Z**, because another process held the card — the system `ollama`'s
resident embedding model, 610 MiB, loaded by the archive's hourly embed job at about 19:51Z — and
this bench's render instrument requires an empty card, which is what the 95 W pass had (15 MiB on
its idle receipt against that attempt's 633). So the tenant was unloaded, the hourly timer was
paused for the window and put back, and the whole pass was re-run into fresh files. **That
attempt's language leg is kept rather than deleted**, under
`results/superseded-tenant-resident-20260921T2024Z/` with a README of its own explaining why, and
it reads as a replication rather than a casualty: against the re-run its widest disagreement over
gemma4's eight windows is **0.897 tokens a second** — 0.58 % of the reading, at the 32,768 rung —
and over mistral's six it is **0.142**. The operator-side half of the same story, the paused timer
included, is in [`OPERATOR-CONDITIONS.md`](OPERATOR-CONDITIONS.md).

**The shipped pass — what this laptop does on an ordinary day — written 2026-09-22.** What this
paragraph is about, for a reader who scrolled straight to it: a THIRD pass over both legs of this
bench, on the same soldered NVIDIA GeForce RTX 5090 Laptop GPU 24 GB, in the posture this machine
actually **boots** in — NVIDIA Dynamic Boost running **and** this box's own `ai-perf.service` clock
lock (`nvidia-smi -lgc 1200,2550`) IN FORCE. The first two passes released that lock on purpose, so
their figures could sit beside benchbox's unlocked desktop rows; this one keeps it, because that is
what the owner's machine does. It ran **2026-09-22 between 00:10:49Z and 00:56:33Z**, both legs,
rc=0, on a card verified empty at the open and again at the render leg's gate (the receipt is
[`RUNG-SHIPPED-RECEIPT.txt`](RUNG-SHIPPED-RECEIPT.txt)). an operator ruled it at about 23:10Z the night
before, against the registered question `Q-BOOST-1`; `PREREG.md` Amendment 4 carries the
registration and what these rows may be compared with.

**The comparison, in one sentence: under load the shipped posture replicates the unlocked
dynamic-boost pass at every arm, and the lock shows up only at idle.** `gemma4:26b` decoded
**155.103–155.987 tok/s** across its eight windows against the boost pass's **155.236–156.025**;
`mistral-small3.2:24b` **44.702–44.799** against **44.846–44.935**; four concurrent streams summed
**354.050 tok/s** against **351.997** — two of those three land a whisker below the boost pass and
one a whisker above, which is run-to-run noise and not an effect. What the lock does change is the
board at rest: **11.57 W on an empty card and 14.42 W with `gemma4:26b` resident, against 7.14 W
and 11.34 W unlocked**, because the floor holds `clocks.sm` at **1,192 MHz** where the unlocked card
drops to **180 MHz**. And the enforced limit **never moved from 150 W** — all **385** witness
samples of this pass, every stage — so its cap cell is `150 W, flat` with its `n` and the finding
stated, never a span. With the clock floor up the card never idles low enough for the platform to
shift budget away, so the 140 W dips the unlocked boost pass showed in its idle gaps do not occur.
The lock's verdict is **ON-BY-CARD**, read from the card at every gate and not from any unit's
state; see the next section.

## Two conditions ride every table, and no instrument on this box can read either

1. **The laptop's fan was at MAXIMUM, set by the operator's hand.** `fan.speed` reads `[N/A]` on
   this board — the driver reports no fan for it, and every `fan.speed` reduction in every result
   file carries `n: 0`. The operator held the machine's own max-fan key (Fn+Up) from before this
   pass's start at 18:43:09Z until it finished. That is recorded in
   [`OPERATOR-CONDITIONS.md`](OPERATOR-CONDITIONS.md) because there is nowhere else
   it can live. **It is a condition of the measurement, never a figure**, and the fan column is a
   declared non-figure rather than a gap somebody fills later.
2. **The boot-time clock lock was OFF for two of the three passes and ON for the third, and the
   CARD says which.** This box boots with `ai-perf.service` running `nvidia-smi -lgc 1200,2550`,
   and the desktop cards that produced every comparison row below ran unlocked. So the lock was
   **released** for the fixed 95 W rung and for `dynamic-boost` — that is what makes them
   comparable at all — and **kept** for `shipped`, whose whole subject is the machine as it boots.
   The receipt is **not** the unit's state in either direction: `ai-perf.service` is `Type=oneshot`
   and reads `active` for the whole uptime whatever `-rgc` did since, so the unit is not evidence.
   Both verdicts are the CARD's, they are in the receipts' own file names, and they are printed in
   the posture table below beside the verdict each posture REQUIRES:
   - **OFF** is proved by a sampled SM clock **below the lock's floor**: read at **180 MHz against
     a 1,200 MHz floor, over twelve reads**, for both unlocked passes
     (`results/gpu-5090-laptop-24g-clock-lock-OFF-BY-CARD-95w.txt`, and the `-dynboost` file
     beside it). A locked card cannot clock below its floor.
   - **ON** is proved by a card **at rest** that cannot fall: `clocks.sm` pinned at **1,192 MHz**
     over twelve reads at **0 % utilisation**
     (`results/gpu-5090-laptop-24g-clock-lock-ON-BY-CARD-shipped.txt`). ⚠ **1,192, not 1,200** —
     the driver snaps the requested floor to the nearest supported SM clock, eight MHz low on this
     board, which is why the floor carries a 25 MHz tolerance and why a probe that tested
     `clocks.sm < floor` would have called this very state OFF. Unlocked and at rest this board
     reads 180–232 MHz; merely holding a model resident puts it at **1,590 MHz with the lock off**,
     which is why the ON verdict requires rest and is taken after the models are unloaded.

## What a reader needs to know before the first number

**The card.** One NVIDIA GeForce RTX 5090 Laptop GPU, **24,463 MiB** of VRAM, **soldered** to the
machine's board at `00000000:01:00.0`. There is no seat, no partner, and no card that can replace
it for a control. Its link is **PCIe x8, generation 5 under load** — a reading of this machine,
not a choice, and not the x16 the desktop legs ran on.

**Watts are CARD BOARD POWER** from `nvidia-smi` at 2 Hz, never the wall. There is no readable
UPS on this box, no NUT, no battery power sensor, and RAPL is root-only. **No whole-box figure
appears anywhere in this document**, the 80 % UPS stop could not arm, and the refusal is recorded
as a receipt that is read and quoted at the foot of this page rather than merely produced. That
is the same absence the desktop bench box published its whole 24 GB ladder under, which is what makes the two
comparable in this respect.

**Two version pins decide whether a figure from this bench can ever enter a published comparison
table**, and they are proved from the result files rather than asserted: **ollama 0.32.13** (this
leg) and **ComfyUI 0.21.1** (the render leg). Both legs ran on pinned copies staged into the
bench directory — never on this box's live installs, which are one patch and several versions
ahead.

**The layer count was left to ollama's own planner (`auto`) for every ladder rung**, which is what
a buyer gets out of the box. Where that choice is itself the finding — the dense model at 98,304
— the bench forces `num_gpu 99` beside it and prints both.

**No price appears in this document and none may be derived from it.**

**The two-3080 rig's language rows are its 300 W run, and that is a correction made 2026-09-21.**
Both files those rows read record the cap themselves — their
`instance_env.power_limits_at_start_w` reads `0, 300.00 W;1, 300.00 W;` — and the pair bench's fold
lists them as its **300 W** rows beside separate 250 W ones. They were labelled `250 W each` in
this document until today; the 3090 and 3090 Ti legs that quote the same file have always said
300 W, so the mislabel was this document's alone and no other page needs correcting for it. The
render leg's pair rows are a different measurement and **do** record 250 W a board, so they stand.

## The five findings, before the tables

1. **A fifth of the power buys most of the speed, on the MoE.** `gemma4:26b` at a 131,072-token
   window: **118.424 tok/s at 79.64 W median board power** here, against **136.618 at 240.49 W**
   on a 3090 held at 350 W. That is **148.7 tokens a second per 100 W against 56.8** — the laptop
   is **2.6× as efficient per token**, at **0.867×** the throughput. The joule figures say the
   same thing from the other side: **672.50 J per 1,000 tokens against 1,758.13**.
2. **The dense model is where the mobile part actually loses.** `mistral-small3.2:24b` at 65,536:
   **28.839 tok/s** here against **46.880** on the 3090 — **0.615×**. A dense 24-billion-parameter
   model reads its whole weight set for every token, and a 95 W envelope cannot feed that the way
   it can feed an MoE's active slice. The efficiency advantage survives but shrinks: **31.9 tok/s
   per 100 W against 17.2**.
3. **The context ceiling is IDENTICAL to the 3090's, on both models, and that is a memory fact a
   power limit cannot move.** `gemma4:26b` held every rung tested, to 131,072 — the ladder ran out
   of rungs before the card ran out of memory. `mistral-small3.2:24b` held 65,536 whole and spilled
   at 98,304 with 92.2 % on the card. Both are the 3090's own headlines, rung for rung.
4. **The planner left the card on the table, and forcing the layer count took the rung back.**
   At 98,304 the dense model's `auto` rung did NOT fit — 92.2 % VRAM / 7.8 % RAM, which is where
   the ladder stopped. The same rung with `options.num_gpu 99` **fits whole**, 21.52 GiB
   fully resident, at **28.851 tok/s** — inside the spread of the 65,536 rung. The next rung up,
   131,072 forced, answers **`cudaMalloc failed: out of memory`**, and that refusal is written into
   its own file as the rung's outcome. So the headline ladder is honest about what `auto` does and
   the probe is honest about what the card can actually hold.
5. **A KV cache at `f16` is FASTER than the `q8_0` default on this board, not slower.** At 131,072:
   **121.379 tok/s at `f16`** against **118.424 at `q8_0`** and **116.125 at `q4_0`** — and `f16`
   also drew the least (76.53 W median, 622.11 J/1,000 tokens). On a 95 W envelope the cheaper
   cache is not the faster one; the quantised cache costs arithmetic the board has less room to pay
   for. `q4_0` did buy a much faster prefill (**10,755 tok/s against 4,407**), so the two are a
   trade and the table prints both.

## What this bench did NOT measure, named here and again at the foot of the page

- **No comparison between the SHIPPED pass and any benchbox row**, and this is the one absence the
  third pass created rather than closed. The shipped posture IS measured now — it ran
  2026-09-22 00:10:49–00:56:33Z and its rows are in every table below — but benchbox, which produced
  every desktop row on this page, runs no clock lock at all. `PREREG.md` §11.2 registered that in
  advance and Amendment 4 does not weaken it: a board that cannot clock below 1,200 MHz is not
  honouring a power posture the way an unlocked one does, so **no cell from the shipped pass is
  offered for a cross-box comparison**. Its rows sit beside THIS machine's other two postures, and
  the posture table prints the lock verdict each row was measured under so the line is visible
  rather than remembered.
- **No wall watts.** Board power only; the receipt is quoted at the foot of the page.
- **No arm T, the trainer.** Its preprocessed tensor set lives on benchbox, unreachable since
  2026-09-21T15:03Z. It is a **second train** and legs 1 and 2 never depended on it. Absent, not
  failed.
- **No fan figure, no memory-die temperature, no peak-bandwidth ceiling** (GDDR7 against a GDDR6
  formula — the memory *clock* is a reading and prints).
- **No control seat.** The card is soldered. The fourteen CPU-only rows this same laptop has
  published were measured by another bench on another day with the card parked, and folding them
  in as this leg's control would be the substitution this bench refuses everywhere else.

## ⚠ Three fields in the result files describe a different machine, and no table here reads them

Stated rather than left to be found by someone grepping the JSON. `thermal_layout_note`,
`pcie_note` and the render leg's `question` field were inherited from the 3080 Ti leg this
harness was forked from and **were not re-cut for this box while it ran**: they describe "ONE
GeForce RTX 3090 Ti 24 GB in a consumer desktop" on "a x16 slot … generation 3". This board is
soldered, its link is x8 gen 5, and there is no desktop.

**They were left alone on purpose.** Editing the harness between the passes would mean one posture
ran a different instrument from another, and instrument identity across the three is worth more than
three tidy strings. **No table in this document or in the render leg's document quotes any of
them**, and the reducer never reads them.

**Corrected in the harness on 2026-09-21, for runs after ALL THREE of these.** For a reader who
scrolled straight to this section: `thermal_layout_note` and `pcie_note` are free-text strings the
harness (`bench_gpu.py`) stamps into every scored result file — the first says what physical
arrangement a temperature is a reading OF, the second says what the link reading means — and
through every one of this leg's three measured passes (fixed 95 W,
**2026-09-21T18:43:29Z – 19:19:53Z**; dynamic-boost, **2026-09-21T21:02:13Z – 21:34:25Z**; shipped,
**2026-09-22T00:10:49Z – 00:56:33Z**, all UTC) they carried the desktop wording quoted above. They
were re-cut on **2026-09-21 UTC, after the second pass closed at 21:34:25Z**, and from there on
they describe this machine: one NVIDIA GeForce RTX 5090 Laptop GPU 24 GB soldered into a gaming
laptop, a single board with no slot, its link a reading traced under load — **x8 gen 5 under load,
x8 gen 1 at rest** — rather than a slot width assumed. Set against the inherited wording, exactly
one thing changed: which machine the two strings describe. No figure moves, no table changes, no
result file is touched, and no bandwidth is claimed from the link in either version. `bench_gpu.py`
read sha256 `595258bdf0a6aa51e658e43f874439949b1175e76bf40a1bb5854283a3d37d79` before the re-cut
and `67e22dbd5aa35906a970ec902b73244efadffcc3b2efee6601dc7604798d9b6c` after it — both digests
recorded, which is what this section promised.

**⚠ AND THE THIRD PASS RAN AFTER THAT RE-CUT AND STILL CARRIES THE OLD WORDING, which is a
decision and not a slip — added 2026-09-22.** The re-cut lives in the branch; the bench runs a
STAGED copy of the harness, and the third pass's staging deliberately did not carry `bench_gpu.py`
across. So the box's copy read `595258bd…` for all three passes and the repo's reads `67e22dbd…`,
and the difference between them is two prose strings and nothing else. The alternative — staging
the re-cut before the third pass — would have made `shipped` the only posture measured by a
different `bench_gpu.py` than the two it is compared against, which is the one comparison this
whole third pass exists to make. Instrument identity across the three postures is worth more than
three tidy strings, and this is the second time that sentence has decided something here.

**Every result file of all three passes still carries the inherited wording**, and that is the
record rather than an oversight: a result file is a receipt of what the instrument stamped at the
time, and it is never edited afterwards. No table in this document or in the render leg's document
reads either field; the render leg's `question` field is not part of this change; and the fence
`test_the_inherited_notes_that_describe_another_machine_are_never_quoted`
(`tests/test_summarize.py`) now holds all three ends of it — the inherited wording stays out of
the tables, the harness's own notes describe this machine (read off the module, never typed into
the fence), and the result files still carry what they were measured with.

---

<!-- GENERATED:BEGIN — everything to the END marker is written by ./report.sh -->

# The tables

### The board, and the three postures every table below is keyed on

| field | reading |
|---|---|
| card | **an NVIDIA GeForce RTX 5090 Laptop GPU 24 GB** — soldered (mobile; no socket — link width is traced under load, never assumed) at 00000000:01:00.0 - vendor unread - GPU-edff232c |
| uuid | `GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff` |
| bus id | `00000000:01:00.0` |
| board power envelope | **95 W default · 175 W maximum**, both READ off the card; `power.limit` answers `[N/A]` on this board and always will — nothing has ever set one and nothing can |
| ollama | **0.32.13** (client), API **0.32.13**, pin `0.32.13` — `results/gpu-5090-laptop-24g-m5-one-95w-auto-ladder.json` › `instance_env` |
| models store | `/usr/share/ollama/.ollama/models`, writable by this user: **False** |
| the frozen prompt | `90eedd0c53f9554ae3837674504fcb7090013d9a432a653352183c0f25a7ce5c` |
| fan | **a declared non-figure.** `fan.speed` reads `[N/A]` on this board — the driver reports no fan for it — so no result file can carry a fan column and none is filled from another board's row. The operator held the laptop's own maximum-fan key by hand for the whole of this pass (`OPERATOR-CONDITIONS.md`, 2026-09-21T17:30Z and 19:10Z); that is a condition of the measurement and not a reading of it. |


| posture | state | cap cell | what the cap cell is made of | clock lock, read FROM THE CARD |
|---|---|---|---|---|
| **fixed · 95 W** | measured | 95 W (`enforced.power.limit` 95.00 W; `power.limit` reads [N/A]; the board's own default 95.00 W, max 175.00 W) | the number 95 W, read back from the card at every stage boundary | **OFF-BY-CARD** (**as this posture REQUIRES**) — evidence clocks.sm was observed at 180 MHz, BELOW the lock's floor of 1200 MHz, over 12 reads (min 180, max 180, clocks.max.sm 3090). A locked card cannot clock below its floor, so the lock is not in force. (`results/gpu-5090-laptop-24g-clock-lock-OFF-BY-CARD-95w.txt`) |
| **dynamic-boost · 95–175 W** | measured | 140–150 W (median 150, n=374) | a RANGE with its n, from this pass's own 0.2 Hz witness — never a single wattage | **OFF-BY-CARD** (**as this posture REQUIRES**) — evidence clocks.sm was observed at 180 MHz, BELOW the lock's floor of 1200 MHz, over 12 reads (min 180, max 180, clocks.max.sm 3090). A locked card cannot clock below its floor, so the lock is not in force. (`results/gpu-5090-laptop-24g-clock-lock-OFF-BY-CARD-dynboost.txt`) |
| **shipped · boost + clock lock** | measured | 150 W, flat (n=385) — ⚠ THE LIMIT DID NOT MOVE, which is a FINDING and not a range: under this posture `nvidia-powerd` is supposed to be floating it | a RANGE with its n, from this pass's own 0.2 Hz witness — never a single wattage; measured with this box's boot-time 1,200–2,550 MHz clock lock IN FORCE, which is this posture's declared condition and is read FROM THE CARD | **ON-BY-CARD** (**as this posture REQUIRES**) — evidence the card was AT REST for the whole sample (utilization.gpu read 0 % on all 12 reads) and clocks.sm never left the lock's floor: min 1192 MHz, max 1192 MHz, against a declared floor of 1200 MHz and the driver's 25 MHz quantization tolerance (the driver snaps 1200 to 1192 on this board, measured 2026-09-21T23:13Z). Unlocked and at rest this board reads 232 MHz or less (232 MHz at 17:30Z, 180 MHz at 18:11Z), so a resting card pinned at the floor is the lock holding it there. clocks.max.sm reads 3090 MHz, which is this board's own unlocked maximum with OR without the lock and is therefore not part of this verdict. (`results/gpu-5090-laptop-24g-clock-lock-ON-BY-CARD-shipped.txt`) |

> The clock-lock verdict is read from `clocks.sm`, never from `ai-perf.service`: that unit is `Type=oneshot` and reads `active` for the whole uptime whatever `-rgc` did since, so the unit is not evidence. A sampled SM clock below the lock's 1,200 MHz floor — by more than the driver's quantization tolerance — is proof the lock is not in force; a card AT REST whose SM clock is pinned at that floor is proof it is.
> **The lock is a confound for two of these postures and the DECLARED CONDITION of the third.** `fixed · 95 W` and `dynamic-boost` are measured with the clocks released, which is what lets them sit beside benchbox's unlocked desktop rows; `shipped` is measured with this box's own boot-time lock IN FORCE, because that is what the machine does for its owner on an ordinary day. Each pass REFUSES on the wrong side of that line rather than footnoting it, and the required verdict is printed in the cell beside the one the card gave.
> ⚠ **1,192 MHz, not 1,200.** `ai-perf.service` asks for a 1,200 MHz floor and the driver snaps it to the nearest supported SM clock on this board, eight MHz low (twelve consecutive reads, 2026-09-21T23:13Z, utilisation 0 %). Unlocked and at rest this board reads 180–232 MHz instead, and merely holding a model resident puts it at 1,590 MHz with the lock off — which is why the ON verdict requires a card at rest and the OFF verdict allows the floor a tolerance.

#### The two version pins the published comparison tables demand

| posture | ollama client version | the API's own `/api/version` | verdict | read from |
|---|---|---|---|---|
| fixed · 95 W | `0.32.13` | `0.32.13` | **MATCHES the pin** (`0.32.13`) | `/workshop/bench-laptop-5090-2026-09-21/ollama-0.32.13/bin/ollama` · `results/gpu-5090-laptop-24g-m5-one-95w-auto-ladder.json` › `instance_env` |
| dynamic-boost · 95–175 W | `0.32.13` | `0.32.13` | **MATCHES the pin** (`0.32.13`) | `/workshop/bench-laptop-5090-2026-09-21/ollama-0.32.13/bin/ollama` · `results/gpu-5090-laptop-24g-m5-one-dynboost-auto-ladder.json` › `instance_env` |
| shipped · boost + clock lock | `0.32.13` | `0.32.13` | **MATCHES the pin** (`0.32.13`) | `/workshop/bench-laptop-5090-2026-09-21/ollama-0.32.13/bin/ollama` · `results/gpu-5090-laptop-24g-m5-one-shipped-auto-ladder.json` › `instance_env` |

> ⚠ **`ollama --version` prints the SERVER's version first** — it contacts whatever is listening, and on this box that is a live seat on a different version. The line that matters is `client version`, and the figures above are the bench instance's own `/api/version` beside the binary that served it.
> The render leg's pin is **ComfyUI 0.21.1**, proved in `../render-5090-laptop-2026-09-21/RESULTS.md` from its own inventory arm. Both legs ran on the pinned copies staged into the bench directory, not on this box's live installs.

## The arms

#### Context ladder — `gemma4:26b` on `one`

| num_ctx | fits whole? | size | size_vram | in VRAM | spilled to RAM | card MiB used | KV over base | runner RSS GiB | fixed · 95 W · decode tok/s (all runs) | fixed · 95 W · decode median | fixed · 95 W · TTFT ms median | fixed · 95 W · median W (the card) | fixed · 95 W · J / 1,000 tokens | fixed · 95 W · mem clock median MHz | fixed · 95 W · core temp max °C | fixed · 95 W · sw_power_cap active fraction | dynamic-boost · 95–175 W · decode tok/s (all runs) | dynamic-boost · 95–175 W · decode median | dynamic-boost · 95–175 W · TTFT ms median | dynamic-boost · 95–175 W · median W (the card) | dynamic-boost · 95–175 W · J / 1,000 tokens | dynamic-boost · 95–175 W · mem clock median MHz | dynamic-boost · 95–175 W · core temp max °C | dynamic-boost · 95–175 W · sw_power_cap active fraction | shipped · boost + clock lock · decode tok/s (all runs) | shipped · boost + clock lock · decode median | shipped · boost + clock lock · TTFT ms median | shipped · boost + clock lock · median W (the card) | shipped · boost + clock lock · J / 1,000 tokens | shipped · boost + clock lock · mem clock median MHz | shipped · boost + clock lock · core temp max °C | shipped · boost + clock lock · sw_power_cap active fraction |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 4,096 | **yes** | 16.02 GiB | 16.02 GiB | 100.0% VRAM | 0.0% RAM | 17,933 | 0.00 GiB | 1.72 | 119.197 · 119.215 · 118.691 | **119.197** | 398.83 | 78.85 | 664.33 | 14,001 | 39.0 | 1.000 | 155.937 · 156.157 · 156.025 | **156.025** | 388.03 | 106.51 | 682.65 | 14,001 | 42.0 | 0.667 | 155.588 · 155.833 · 156.042 | **155.833** | 380.67 | 122.33 | 785.01 | 14,001 | 42.0 | 1.000 |
| 8,192 | **yes** | 16.18 GiB | 16.18 GiB | 100.0% VRAM | 0.0% RAM | 18,141 | 0.16 GiB | 1.41 | 118.484 · 118.860 · 119.272 | **118.860** | 403.96 | 79.87 | 669.64 | 14,001 | 39.0 | 0.750 | 155.648 · 155.996 · 155.378 | **155.648** | 387.20 | 146.66 | 943.36 | 14,001 | 43.0 | 1.000 | 155.987 · 155.611 · 156.193 | **155.987** | 367.51 | 104.48 | 669.43 | 14,001 | 43.0 | 0.833 |
| 16,384 | **yes** | 16.06 GiB | 16.06 GiB | 100.0% VRAM | 0.0% RAM | 18,095 | 0.04 GiB | 1.40 | 118.869 · 118.560 · 119.469 | **118.869** | 383.91 | 93.53 | 788.89 | 14,001 | 40.0 | 1.000 | 155.416 · 155.835 · 155.800 | **155.800** | 365.56 | 145.92 | 938.90 | 14,001 | 44.0 | 1.000 | 155.937 · 155.539 · 155.733 | **155.733** | 377.85 | 103.31 | 664.21 | 14,001 | 44.0 | 0.667 |
| 32,768 | **yes** | 16.13 GiB | 16.13 GiB | 100.0% VRAM | 0.0% RAM | 18,345 | 0.11 GiB | 1.42 | 117.403 · 119.483 · 118.791 | **118.791** | 365.44 | 78.84 | 669.41 | 14,001 | 41.0 | 0.875 | 155.734 · 155.701 · 155.891 | **155.734** | 392.78 | 137.15 | 880.85 | 14,001 | 43.0 | 1.000 | 155.103 · 154.992 · 155.489 | **155.103** | 374.45 | 102.86 | 663.17 | 14,001 | 44.0 | 0.667 |
| 49,152 | **yes** | 16.21 GiB | 16.21 GiB | 100.0% VRAM | 0.0% RAM | 18,595 | 0.19 GiB | 1.43 | 119.016 · 118.272 · 119.493 | **119.016** | 367.37 | 80.08 | 672.85 | 14,001 | 40.0 | 1.000 | 155.657 · 155.741 · 155.307 | **155.657** | 374.58 | 102.87 | 662.36 | 14,001 | 44.0 | 1.000 | 155.296 · 155.104 · 155.418 | **155.296** | 355.00 | 116.62 | 750.36 | 14,001 | 43.0 | 1.000 |
| 65,536 | **yes** | 16.29 GiB | 16.29 GiB | 100.0% VRAM | 0.0% RAM | 18,845 | 0.27 GiB | 1.45 | 118.717 · 119.163 · 118.366 | **118.717** | 401.65 | 77.62 | 655.77 | 14,001 | 40.0 | 1.000 | 155.236 · 155.045 · 155.536 | **155.236** | 430.94 | 146.41 | 941.32 | 14,001 | 44.0 | 1.000 | 155.458 · 155.269 · 155.493 | **155.458** | 370.26 | 104.46 | 672.77 | 14,001 | 44.0 | 0.833 |
| 98,304 | **yes** | 16.45 GiB | 16.45 GiB | 100.0% VRAM | 0.0% RAM | 19,345 | 0.43 GiB | 1.48 | 119.032 · 118.129 · 117.938 | **118.129** | 376.85 | 79.93 | 676.63 | 14,001 | 40.0 | 1.000 | 155.643 · 155.779 · 155.838 | **155.779** | 380.47 | 128.14 | 822.58 | 14,001 | 44.0 | 1.000 | 155.529 · 155.368 · 155.232 | **155.368** | 380.40 | 102.31 | 657.82 | 14,001 | 43.0 | 0.667 |
| 131,072 | **yes** | 16.60 GiB | 16.60 GiB | 100.0% VRAM | 0.0% RAM | 19,845 | 0.58 GiB | 1.52 | 119.239 · 118.424 · 117.802 | **118.424** | 372.49 | 79.64 | 672.50 | 14,001 | 42.0 | 0.750 | 155.128 · 155.251 · 155.593 | **155.251** | 372.92 | 122.57 | 787.76 | 14,001 | 44.0 | 1.000 | 155.246 · 155.288 · 155.430 | **155.288** | 371.01 | 116.12 | 747.97 | 14,001 | 43.0 | 1.000 |

> **fixed · 95 W** — `results/gpu-5090-laptop-24g-m5-one-95w-auto-ladder.json` · layer policy: layers left to ollama's own planner (`auto`) · quantization **Q4_K_M** · store 16.75 GiB · the model's own `context_length` ceiling 262,144
> ⚠ **every rung tested fits, up to and including 131,072 -- the ladder ran out of rungs before the card ran out of memory, so the figure reported is the largest window TESTED, not a measured ceiling**
> **Headline:** the biggest context this arm holds for gemma4:26b is 131,072 tokens
> **dynamic-boost · 95–175 W** — `results/gpu-5090-laptop-24g-m5-one-dynboost-auto-ladder.json` · layer policy: layers left to ollama's own planner (`auto`) · quantization **Q4_K_M** · store 16.75 GiB · the model's own `context_length` ceiling 262,144
> ⚠ **every rung tested fits, up to and including 131,072 -- the ladder ran out of rungs before the card ran out of memory, so the figure reported is the largest window TESTED, not a measured ceiling**
> **Headline:** the biggest context this arm holds for gemma4:26b is 131,072 tokens
> **shipped · boost + clock lock** — `results/gpu-5090-laptop-24g-m5-one-shipped-auto-ladder.json` · layer policy: layers left to ollama's own planner (`auto`) · quantization **Q4_K_M** · store 16.75 GiB · the model's own `context_length` ceiling 262,144
> ⚠ **every rung tested fits, up to and including 131,072 -- the ladder ran out of rungs before the card ran out of memory, so the figure reported is the largest window TESTED, not a measured ceiling**
> **Headline:** the biggest context this arm holds for gemma4:26b is 131,072 tokens

> The shared left-hand columns are properties of the MODEL and the WINDOW, not of the power posture: a power limit does not change how many bytes a model is. They are the fixed posture's readings, and the reducer compares the other posture's against them row by row and prints a FINDING above if any differs.

#### Context ladder — `mistral-small3.2:24b` on `one`

| num_ctx | fits whole? | size | size_vram | in VRAM | spilled to RAM | card MiB used | KV over base | runner RSS GiB | fixed · 95 W · decode tok/s (all runs) | fixed · 95 W · decode median | fixed · 95 W · TTFT ms median | fixed · 95 W · median W (the card) | fixed · 95 W · J / 1,000 tokens | fixed · 95 W · mem clock median MHz | fixed · 95 W · core temp max °C | fixed · 95 W · sw_power_cap active fraction | dynamic-boost · 95–175 W · decode tok/s (all runs) | dynamic-boost · 95–175 W · decode median | dynamic-boost · 95–175 W · TTFT ms median | dynamic-boost · 95–175 W · median W (the card) | dynamic-boost · 95–175 W · J / 1,000 tokens | dynamic-boost · 95–175 W · mem clock median MHz | dynamic-boost · 95–175 W · core temp max °C | dynamic-boost · 95–175 W · sw_power_cap active fraction | shipped · boost + clock lock · decode tok/s (all runs) | shipped · boost + clock lock · decode median | shipped · boost + clock lock · TTFT ms median | shipped · boost + clock lock · median W (the card) | shipped · boost + clock lock · J / 1,000 tokens | shipped · boost + clock lock · mem clock median MHz | shipped · boost + clock lock · core temp max °C | shipped · boost + clock lock · sw_power_cap active fraction |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 4,096 | **yes** | 13.60 GiB | 13.60 GiB | 100.0% VRAM | 0.0% RAM | 15,073 | 0.00 GiB | 0.99 | 26.774 · 27.137 · 26.392 | **26.774** | 214.12 | 93.28 | 3,437.32 | 14,001 | 61.0 | 1.000 | 44.930 · 44.942 · 44.935 | **44.935** | 198.02 | 141.22 | 3,143.13 | 14,001 | 47.0 | 1.000 | 44.716 · 44.727 · 44.784 | **44.727** | 212.19 | 135.14 | 3,017.58 | 14,001 | 46.0 | 1.000 |
| 8,192 | **yes** | 14.16 GiB | 14.16 GiB | 100.0% VRAM | 0.0% RAM | 15,661 | 0.56 GiB | 0.86 | 26.864 · 26.941 · 26.620 | **26.864** | 221.30 | 91.60 | 3,400.00 | 14,001 | 60.0 | 1.000 | 44.935 · 44.930 · 44.935 | **44.935** | 190.77 | 131.02 | 2,915.75 | 14,001 | 47.0 | 1.000 | 44.813 · 44.756 · 44.681 | **44.756** | 199.77 | 134.96 | 3,011.60 | 14,001 | 46.0 | 1.000 |
| 16,384 | **yes** | 14.84 GiB | 14.84 GiB | 100.0% VRAM | 0.0% RAM | 16,357 | 1.24 GiB | 0.87 | 26.732 · 26.449 · 27.449 | **26.732** | 200.57 | 92.41 | 3,493.91 | 14,001 | 61.0 | 1.000 | 44.938 · 44.877 · 44.874 | **44.877** | 183.66 | 141.04 | 3,143.06 | 14,001 | 47.0 | 1.000 | 44.799 · 44.718 · 44.800 | **44.799** | 213.49 | 131.20 | 2,928.66 | 14,001 | 47.0 | 1.000 |
| 32,768 | **yes** | 15.95 GiB | 15.95 GiB | 100.0% VRAM | 0.0% RAM | 17,481 | 2.35 GiB | 0.85 | 27.788 · 27.840 · 28.373 | **27.840** | 216.00 | 93.60 | 3,352.11 | 14,001 | 52.0 | 1.000 | 44.838 · 44.888 · 44.880 | **44.880** | 185.63 | 141.11 | 3,144.14 | 14,001 | 46.0 | 1.000 | 44.828 · 44.795 · 44.761 | **44.795** | 193.60 | 135.20 | 3,020.51 | 14,001 | 47.0 | 1.000 |
| 49,152 | **yes** | 17.30 GiB | 17.30 GiB | 100.0% VRAM | 0.0% RAM | 18,865 | 3.70 GiB | 0.87 | 28.285 · 28.815 · 28.851 | **28.815** | 197.14 | 89.92 | 3,120.59 | 14,001 | 46.0 | 1.000 | 44.930 · 44.846 · 44.799 | **44.846** | 205.62 | 140.72 | 3,141.14 | 14,001 | 46.0 | 1.000 | 44.808 · 44.765 · 44.751 | **44.765** | 199.37 | 130.81 | 2,919.35 | 14,001 | 47.0 | 1.000 |
| 65,536 | **yes** | 18.71 GiB | 18.71 GiB | 100.0% VRAM | 0.0% RAM | 20,305 | 5.11 GiB | 0.89 | 28.672 · 28.874 · 28.839 | **28.839** | 210.94 | 90.31 | 3,131.56 | 14,001 | 45.0 | 1.000 | 44.963 · 44.888 · 44.897 | **44.897** | 203.72 | 135.76 | 3,019.40 | 14,001 | 47.0 | 1.000 | 44.702 · 44.680 · 44.726 | **44.702** | 191.65 | 129.90 | 2,904.35 | 14,001 | 46.0 | 1.000 |
| 98,304 | no | 23.00 GiB | 21.20 GiB | 92.2% VRAM | 7.8% RAM | 22,855 | 9.40 GiB | 14.18 | N/A | **N/A** | N/A | N/A | N/A | N/A | N/A | N/A | N/A | **N/A** | N/A | N/A | N/A | N/A | N/A | N/A | N/A | **N/A** | N/A | N/A | N/A | N/A | N/A | N/A |

> **fixed · 95 W** — `results/gpu-5090-laptop-24g-m4-one-95w-auto-ladder.json` · layer policy: layers left to ollama's own planner (`auto`) · quantization **Q4_K_M** · store 14.14 GiB · the model's own `context_length` ceiling 131,072
> **The ladder stopped at num_ctx 98,304** — the first rung that did not fit whole on the card: `does not fit: only part of the model is on the card (92.2% VRAM / 7.8% RAM)`, with 22,855 MiB used on the card
> Rungs never reached (the ladder stopped first, so nothing is reported for them): 131,072
> **Headline:** the biggest context this arm holds for mistral-small3.2:24b is 65,536 tokens
> **dynamic-boost · 95–175 W** — `results/gpu-5090-laptop-24g-m4-one-dynboost-auto-ladder.json` · layer policy: layers left to ollama's own planner (`auto`) · quantization **Q4_K_M** · store 14.14 GiB · the model's own `context_length` ceiling 131,072
> **The ladder stopped at num_ctx 98,304** — the first rung that did not fit whole on the card: `does not fit: only part of the model is on the card (92.2% VRAM / 7.8% RAM)`, with 22,855 MiB used on the card
> Rungs never reached (the ladder stopped first, so nothing is reported for them): 131,072
> **Headline:** the biggest context this arm holds for mistral-small3.2:24b is 65,536 tokens
> **shipped · boost + clock lock** — `results/gpu-5090-laptop-24g-m4-one-shipped-auto-ladder.json` · layer policy: layers left to ollama's own planner (`auto`) · quantization **Q4_K_M** · store 14.14 GiB · the model's own `context_length` ceiling 131,072
> **The ladder stopped at num_ctx 98,304** — the first rung that did not fit whole on the card: `does not fit: only part of the model is on the card (92.2% VRAM / 7.8% RAM)`, with 22,855 MiB used on the card
> Rungs never reached (the ladder stopped first, so nothing is reported for them): 131,072
> **Headline:** the biggest context this arm holds for mistral-small3.2:24b is 65,536 tokens
> ⚠ at num_ctx 8,192 the `size` reading under dynamic-boost · 95–175 W (14963703807) differs from the fixed posture's (15202789621) — the shared left-hand columns print the fixed posture's, and this difference is a FINDING
> ⚠ at num_ctx 8,192 the `size_vram` reading under dynamic-boost · 95–175 W (14963703807) differs from the fixed posture's (15202789621) — the shared left-hand columns print the fixed posture's, and this difference is a FINDING
> ⚠ at num_ctx 16,384 the `size` reading under dynamic-boost · 95–175 W (15685124095) differs from the fixed posture's (15932598517) — the shared left-hand columns print the fixed posture's, and this difference is a FINDING
> ⚠ at num_ctx 16,384 the `size_vram` reading under dynamic-boost · 95–175 W (15685124095) differs from the fixed posture's (15932598517) — the shared left-hand columns print the fixed posture's, and this difference is a FINDING

> The shared left-hand columns are properties of the MODEL and the WINDOW, not of the power posture: a power limit does not change how many bytes a model is. They are the fixed posture's readings, and the reducer compares the other posture's against them row by row and prints a FINDING above if any differs.

#### The scored arms

| arm | fixed · 95 W · decode median tok/s | fixed · 95 W · decode spread | fixed · 95 W · TTFT ms median | fixed · 95 W · prefill tok/s median | fixed · 95 W · median W (the card) | fixed · 95 W · J / 1,000 tokens | fixed · 95 W · core temp max °C | fixed · 95 W · fits whole? | fixed · 95 W · source file | dynamic-boost · 95–175 W · decode median tok/s | dynamic-boost · 95–175 W · decode spread | dynamic-boost · 95–175 W · TTFT ms median | dynamic-boost · 95–175 W · prefill tok/s median | dynamic-boost · 95–175 W · median W (the card) | dynamic-boost · 95–175 W · J / 1,000 tokens | dynamic-boost · 95–175 W · core temp max °C | dynamic-boost · 95–175 W · fits whole? | dynamic-boost · 95–175 W · source file | shipped · boost + clock lock · decode median tok/s | shipped · boost + clock lock · decode spread | shipped · boost + clock lock · TTFT ms median | shipped · boost + clock lock · prefill tok/s median | shipped · boost + clock lock · median W (the card) | shipped · boost + clock lock · J / 1,000 tokens | shipped · boost + clock lock · core temp max °C | shipped · boost + clock lock · fits whole? | shipped · boost + clock lock · source file |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| **A2** KV `f16` at 131,072 — `gemma4:26b` | **121.379** | n=3; 120.822–123.017 | 435.96 | 4,407.2 | 76.53 | 622.11 | 48.0 | **yes** | `results/gpu-5090-laptop-24g-m5-one-95w-kvf16-at131072.json` | **165.656** | n=3; 165.545–166.175 | 439.78 | 4,837.0 | 116.01 | 700.31 | 44.0 | **yes** | `results/gpu-5090-laptop-24g-m5-one-dynboost-kvf16-at131072.json` | **165.991** | n=3; 165.967–166.369 | 452.09 | 4,685.2 | 104.05 | 626.93 | 43.0 | **yes** | `results/gpu-5090-laptop-24g-m5-one-shipped-kvf16-at131072.json` |
| **A2** KV `q4_0` at 131,072 — `gemma4:26b` | **116.125** | n=3; 115.846–116.716 | 326.27 | 10,755.1 | 86.87 | 744.29 | 56.0 | **yes** | `results/gpu-5090-laptop-24g-m5-one-95w-kvq4_0-at131072.json` | **155.712** | n=3; 155.641–155.908 | 336.23 | 13,981.6 | 148.71 | 955.03 | 44.0 | **yes** | `results/gpu-5090-laptop-24g-m5-one-dynboost-kvq4_0-at131072.json` | **155.889** | n=3; 155.534–156.131 | 367.21 | 17,713.2 | 124.90 | 799.97 | 44.0 | **yes** | `results/gpu-5090-laptop-24g-m5-one-shipped-kvq4_0-at131072.json` |
| **A3** the FILLED window at 131,072 — `gemma4:26b` | **56.174** | n=3; 56.039–56.374 | 744.92 | 1,421,822.4 | 88.38 | 1,573.32 | 64.0 | fits: fully resident on the card | `results/gpu-5090-laptop-24g-m5-one-95w-filled.json` | **72.159** | n=3; 72.024–72.171 | 721.95 | 1,306,089.8 | 123.99 | 1,718.01 | 48.0 | fits: fully resident on the card | `results/gpu-5090-laptop-24g-m5-one-dynboost-filled.json` | **72.364** | n=3; 72.221–72.390 | 678.53 | 1,323,518.7 | 108.96 | 1,505.71 | 47.0 | fits: fully resident on the card | `results/gpu-5090-laptop-24g-m5-one-shipped-filled.json` |
| **D-probe** `num_gpu 99` forced at 98,304 — `mistral-small3.2:24b` | **28.851** | n=3; 28.721–28.993 | 214.37 | 22,118.0 | 93.13 | 3,227.99 | 45.0 | **yes** | `results/gpu-5090-laptop-24g-m4-one-95w-forced-at98304.json` | **44.946** | n=3; 44.917–44.963 | 208.97 | 50,384.2 | 140.48 | 3,124.35 | 48.0 | **yes** | `results/gpu-5090-laptop-24g-m4-one-dynboost-forced-at98304.json` | **44.801** | n=3; 44.767–44.813 | 208.19 | 35,654.5 | 130.17 | 2,904.72 | 47.0 | **yes** | `results/gpu-5090-laptop-24g-m4-one-shipped-forced-at98304.json` |
| **D-probe** `num_gpu 99` forced at 131,072 — `mistral-small3.2:24b` | **REFUSED** — warmup: HTTP 500: {"error":"llama-server startup failed after projector CPU offload retry: llama-server process has terminated: exit status 1: cudaMalloc failed: out of memory\nall | — | — | — | — | — | — | — | `results/gpu-5090-laptop-24g-m4-one-95w-forced-at131072.json` | **REFUSED** — warmup: HTTP 500: {"error":"llama-server startup failed after projector CPU offload retry: llama-server process has terminated: exit status 1: cudaMalloc failed: out of memory\nall | — | — | — | — | — | — | — | `results/gpu-5090-laptop-24g-m4-one-dynboost-forced-at131072.json` | **REFUSED** — warmup: HTTP 500: {"error":"llama-server startup failed after projector CPU offload retry: llama-server process has terminated: exit status 1: cudaMalloc failed: out of memory\nall | — | — | — | — | — | — | — | `results/gpu-5090-laptop-24g-m4-one-shipped-forced-at131072.json` |
| **G3** vision model, TEXT path only — `minicpm-v4.5` | **86.808** | n=3; 85.795–87.364 | 149.86 | 43,500.2 | 79.27 | 913.16 | 41.0 | **yes** | `results/gpu-5090-laptop-24g-g3-one-95w-g3-vision.json` | **121.733** | n=3; 121.722–121.762 | 132.94 | 30,117.7 | 144.74 | 1,188.72 | 43.0 | **yes** | `results/gpu-5090-laptop-24g-g3-one-dynboost-g3-vision.json` | **121.581** | n=3; 121.145–121.650 | 140.80 | 59,370.7 | 98.15 | 807.28 | 43.0 | **yes** | `results/gpu-5090-laptop-24g-g3-one-shipped-g3-vision.json` |

> Every decode figure is the SERVER's own `eval_count / eval_duration`; the wall figure is carried in the same files and is never substituted for it.
> A **REFUSED** cell is the arm's own outcome written into its file, not a gap: an out-of-memory answer to a forced layer count is a result about the card.

#### Concurrency — `gemma4:26b`, 1 / 2 / 4 streams

| streams | fixed · 95 W · aggregate tok/s (summed) | fixed · 95 W · per-stream decode tok/s | fixed · 95 W · per-stream TTFT ms median | fixed · 95 W · median W (the card) | fixed · 95 W · J / 1,000 tokens | fixed · 95 W · core temp max °C | fixed · 95 W · voided? | fixed · 95 W · source file | dynamic-boost · 95–175 W · aggregate tok/s (summed) | dynamic-boost · 95–175 W · per-stream decode tok/s | dynamic-boost · 95–175 W · per-stream TTFT ms median | dynamic-boost · 95–175 W · median W (the card) | dynamic-boost · 95–175 W · J / 1,000 tokens | dynamic-boost · 95–175 W · core temp max °C | dynamic-boost · 95–175 W · voided? | dynamic-boost · 95–175 W · source file | shipped · boost + clock lock · aggregate tok/s (summed) | shipped · boost + clock lock · per-stream decode tok/s | shipped · boost + clock lock · per-stream TTFT ms median | shipped · boost + clock lock · median W (the card) | shipped · boost + clock lock · J / 1,000 tokens | shipped · boost + clock lock · core temp max °C | shipped · boost + clock lock · voided? | shipped · boost + clock lock · source file |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | **118.368** | 118.368 | 345.52 | 86.32 | 857.75 | 42.0 | no | `results/gpu-5090-laptop-24g-concurrency-one-95w.json` | **155.231** | 155.231 | 401.70 | 105.33 | 829.49 | 46.0 | no | `results/gpu-5090-laptop-24g-concurrency-one-dynboost.json` | **155.078** | 155.078 | 350.44 | 105.39 | 825.16 | 45.0 | no | `results/gpu-5090-laptop-24g-concurrency-one-shipped.json` |
| 2 | **199.413** | 99.950 | 525.93 | 84.38 | 504.27 | 43.0 | no | `results/gpu-5090-laptop-24g-concurrency-one-95w.json` | **256.955** | 128.550 | 528.06 | 127.02 | 626.27 | 47.0 | no | `results/gpu-5090-laptop-24g-concurrency-one-dynboost.json` | **258.544** | 129.420 | 534.68 | 104.13 | 519.47 | 45.0 | no | `results/gpu-5090-laptop-24g-concurrency-one-shipped.json` |
| 4 | **275.732** | 69.255 | 702.42 | 80.75 | 396.86 | 43.0 | no | `results/gpu-5090-laptop-24g-concurrency-one-95w.json` | **351.997** | 87.994 | 694.10 | 115.27 | 405.97 | 47.0 | no | `results/gpu-5090-laptop-24g-concurrency-one-dynboost.json` | **354.050** | 88.506 | 704.25 | 121.92 | 433.72 | 46.0 | no | `results/gpu-5090-laptop-24g-concurrency-one-shipped.json` |

> aggregate_sum_tok_s is the sum of the streams' own decode rates -- what the card produced while producing. aggregate_wall_tok_s is every stream's tokens over the wall time of the whole level -- what a caller experiences, queueing and ragged finish included. Both are printed; neither is the 'real' one on its own.
> **Headline:** one-95w takes 4 parallel streams at 275.7 tok/s aggregate (sum of streams) / 231.2 tok/s over wall

#### The doorman — a short gate-shaped call on `mistral-small3.2:24b`

| streams | fixed · 95 W · calls scored | fixed · 95 W · errors | fixed · 95 W · wall latency ms median | fixed · 95 W · p95 | fixed · 95 W · max | fixed · 95 W · TTFT ms median | fixed · 95 W · mean W (the card) | fixed · 95 W · source file | dynamic-boost · 95–175 W · calls scored | dynamic-boost · 95–175 W · errors | dynamic-boost · 95–175 W · wall latency ms median | dynamic-boost · 95–175 W · p95 | dynamic-boost · 95–175 W · max | dynamic-boost · 95–175 W · TTFT ms median | dynamic-boost · 95–175 W · mean W (the card) | dynamic-boost · 95–175 W · source file | shipped · boost + clock lock · calls scored | shipped · boost + clock lock · errors | shipped · boost + clock lock · wall latency ms median | shipped · boost + clock lock · p95 | shipped · boost + clock lock · max | shipped · boost + clock lock · TTFT ms median | shipped · boost + clock lock · mean W (the card) | shipped · boost + clock lock · source file |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 10 | 0 | **461.64** | 493.52 | 493.52 | 314.52 | 27.81 | `results/gpu-5090-laptop-24g-g2-doorman-95w.json` | 10 | 0 | **438.98** | 468.26 | 468.26 | 298.34 | 32.11 | `results/gpu-5090-laptop-24g-g2-doorman-dynboost.json` | 10 | 0 | **430.85** | 448.20 | 448.20 | 293.91 | 32.37 | `results/gpu-5090-laptop-24g-g2-doorman-shipped.json` |
| 4 | 8 | 0 | **928.54** | 4,671.24 | 4,671.24 | 465.75 | 66.59 | `results/gpu-5090-laptop-24g-g2-doorman-95w.json` | 8 | 0 | **948.56** | 4,567.46 | 4,567.46 | 569.39 | 73.56 | `results/gpu-5090-laptop-24g-g2-doorman-dynboost.json` | 8 | 0 | **885.73** | 4,452.80 | 4,452.80 | 470.94 | 70.93 | `results/gpu-5090-laptop-24g-g2-doorman-shipped.json` |

> **Headline:** a doorman call on mistral-small3.2:24b costs 461.64 ms at the median and 493.52 ms at p95, one caller at a time

#### Embeddings — the same deterministic batch

| model | fixed · 95 W · batches scored | fixed · 95 W · batch size | fixed · 95 W · dimensions | fixed · 95 W · texts/s median | fixed · 95 W · ms per batch median | fixed · 95 W · J / 1,000 texts median | fixed · 95 W · source file | dynamic-boost · 95–175 W · batches scored | dynamic-boost · 95–175 W · batch size | dynamic-boost · 95–175 W · dimensions | dynamic-boost · 95–175 W · texts/s median | dynamic-boost · 95–175 W · ms per batch median | dynamic-boost · 95–175 W · J / 1,000 texts median | dynamic-boost · 95–175 W · source file | shipped · boost + clock lock · batches scored | shipped · boost + clock lock · batch size | shipped · boost + clock lock · dimensions | shipped · boost + clock lock · texts/s median | shipped · boost + clock lock · ms per batch median | shipped · boost + clock lock · J / 1,000 texts median | shipped · boost + clock lock · source file |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| `nomic-embed-text:latest` | 5 | 64 | 768 | **252.570** | 253.40 | 65.67 | `results/gpu-5090-laptop-24g-embed-one-95w.json` | 5 | 64 | 768 | **268.988** | 237.93 | 83.79 | `results/gpu-5090-laptop-24g-embed-one-dynboost.json` | 5 | 64 | 768 | **250.655** | 255.33 | 90.19 | `results/gpu-5090-laptop-24g-embed-one-shipped.json` |

> **Headline:** one-95w: 252.6 texts/s on a batch of 64 (253.4 ms per batch)

#### The idle receipt — what the card draws with the model RESIDENT and not generating

| model held resident | fixed · 95 W · card empty W (mean) | fixed · 95 W · model resident W (mean) | fixed · 95 W · the cost of residency W | fixed · 95 W · window s | fixed · 95 W · samples | fixed · 95 W · source file | dynamic-boost · 95–175 W · card empty W (mean) | dynamic-boost · 95–175 W · model resident W (mean) | dynamic-boost · 95–175 W · the cost of residency W | dynamic-boost · 95–175 W · window s | dynamic-boost · 95–175 W · samples | dynamic-boost · 95–175 W · source file | shipped · boost + clock lock · card empty W (mean) | shipped · boost + clock lock · model resident W (mean) | shipped · boost + clock lock · the cost of residency W | shipped · boost + clock lock · window s | shipped · boost + clock lock · samples | shipped · boost + clock lock · source file |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| `gemma4:26b` at num_ctx 131,072 | 7.18 | **9.03** | 1.85 | 30.0 | 57 | `results/gpu-5090-laptop-24g-idle-gemma4-95w.json` | 7.14 | **11.34** | 4.20 | 30.0 | 57 | `results/gpu-5090-laptop-24g-idle-gemma4-dynboost.json` | 11.57 | **14.42** | 2.85 | 30.0 | 58 | `results/gpu-5090-laptop-24g-idle-gemma4-shipped.json` |

> **Headline:** with gemma4:26b resident and not generating, the card draws 9.03 W (board power), against 7.18 W with the cards empty -- a resident cost of 1.85 W

#### Which instance ran on which card — MEASURED

| posture | instance | port | probe model | card MiB before | card MiB after | growth MiB | the card that grew | verdict | source file |
|---|---|---|---|---|---|---|---|---|---|
| fixed · 95 W | one | 11470 | `qwen3-vl:2b` | 15 | 4,841 | **4,826** | GPU-edff232c… | MATCHES | `results/gpu-5090-laptop-24g-card-mapping-95w.json` |
| dynamic-boost · 95–175 W | one | 11470 | `qwen3-vl:2b` | 15 | 4,841 | **4,826** | GPU-edff232c… | MATCHES | `results/gpu-5090-laptop-24g-card-mapping-dynboost.json` |
| shipped · boost + clock lock | one | 11470 | `qwen3-vl:2b` | 15 | 4,841 | **4,826** | GPU-edff232c… | MATCHES | `results/gpu-5090-laptop-24g-card-mapping-shipped.json` |

> A declared `CUDA_VISIBLE_DEVICES` mapping is not a measured one. The growth column is the card whose memory actually moved when a probe model loaded through the instance, against a threshold the file carries.

## The comparisons

#### `gemma4:26b` — this board against the 24 GB desktop rows

| configuration | cap, printed beside every figure | num_ctx | decode tok/s | median W (the boards) | J / 1,000 tokens | tok/s per 100 W | core temp max °C | source file |
|---|---|---|---|---|---|---|---|---|
| **this board — an NVIDIA GeForce RTX 5090 Laptop GPU 24 GB** · fixed · 95 W | **95 W (`enforced.power.limit`, read back at every stage boundary)** | 131,072 | **118.424** | 79.64 | 672.50 | 148.7 | 42.0 | `results/gpu-5090-laptop-24g-m5-one-95w-auto-ladder.json` |
| **this board — an NVIDIA GeForce RTX 5090 Laptop GPU 24 GB** · dynamic-boost · 95–175 W | **140–150 W (median 150, n=374)** | 131,072 | **155.251** | 122.57 | 787.76 | 126.7 | 44.0 | `results/gpu-5090-laptop-24g-m5-one-dynboost-auto-ladder.json` |
| **this board — an NVIDIA GeForce RTX 5090 Laptop GPU 24 GB** · shipped · boost + clock lock | **150 W, flat (n=385) — ⚠ THE LIMIT DID NOT MOVE, which is a FINDING and not a range: under this posture `nvidia-powerd` is supposed to be floating it** | 131,072 | **155.288** | 116.12 | 747.97 | 133.7 | 43.0 | `results/gpu-5090-laptop-24g-m5-one-shipped-auto-ladder.json` |
| a GeForce RTX 3090 24 GB — another box, another build | **350 W** | 131,072 | **136.618** | 240.49 | 1,758.13 | 56.8 | N/A | `references/desktop-3090-m5-one-350w-auto-ladder.json` |
| a GeForce RTX 3090 Ti 24 GB — another box, another build | **350 W — the shared-cap rung** | 131,072 | **147.823** | 282.64 | 1,909.96 | 52.3 | 57.0 | `references/desktop-3090ti-m5-one-350w-auto-ladder.json` |
| a GeForce RTX 3090 Ti 24 GB — another box, another build | **450 W — this board's own limit** | 131,072 | **147.560** | 283.01 | 1,915.13 | 52.1 | 57.0 | `references/desktop-3090ti-m5-one-450w-auto-ladder.json` |
| two GeForce RTX 3080 10 GB — another box, another build | **300 W each — 600 W between them** | 131,072 | **115.366** | 151.92 (one board of two; the file records no sum) | 1,326.84 (one board of two; the file records no sum) | — WITHHELD (one board of two) | N/A | `references/desktop-2x3080-m5-split-forced-ladder.json` |

> **a GeForce RTX 3090 Ti 24 GB — another box, another build** — **two rows, one board.** The `450 W — this board's own limit` row is this board at the limit it ships with. The `350 W — the shared-cap rung` row above it is the rung every board of this ladder shares — which is 78 % of this board's own limit, while the 3090's 350 W is 100 % of its. Both rows are printed so neither reading can be mistaken for the other: they are ONE board at two caps, not two boards, and each names the file it was read from.
> **two GeForce RTX 3080 10 GB — another box, another build** — this rig needed `options.num_gpu 99` to hold the model at all; its `auto` ladder held no rung whole. The forced ladder is the row, and the forcing is the finding.

> **This is not a cap-keyed table and it cannot be made into one.** This board's whole envelope — 95 W default, 175 W maximum, both READ off the card — sits below the lowest cap column the comparison piece has, which is 250 W. So the columns here are READINGS and the cap is printed beside every figure. Every reference row was measured in ANOTHER box on ANOTHER software stack: a ratio between two rows of this table is a box-and-board ratio, never a card-to-card one.
> `tok/s per 100 W` is derived from the two columns to its left, and both are printed so the derivation can be checked.

> **⚠ two GeForce RTX 3080 10 GB — another box, another build — the watt and joule cells of that row are ONE BOARD of two, and they say so.** `references/desktop-2x3080-m5-split-forced-ladder.json` carries no summed `*_all_cards` reduction at the rung this row reads: its `derived.power_mean_w` is card 0's own mean and `derived.energy_j_per_1k_tokens` is built from it. **The rig's draw IS recoverable from that file** — every run's `power.per_gpu` block holds both boards — but recovering it here would mean this reducer re-reducing another bench's raw per-card samples, and the legs that own the file already publish the result: **345.7 W** median over the whole configuration in `../gpu-3090-desktop-2026-09-17/RESULTS.md` and in the 3090 Ti leg's copy of the same table, 340.8 W mean in the pair bench's own fold. So the one-board readings are printed here because they are real, labelled because they are not the rig, and `tok/s per 100 W` is WITHHELD rather than derived: dividing this row's tokens by one board's watts would make a two-board rig look about 2× as efficient as it is. The other model's row of this same pair DOES carry the summed reduction, which is why one row of this table can name both boards and the other cannot.

#### `mistral-small3.2:24b` — this board against the 24 GB desktop rows

| configuration | cap, printed beside every figure | num_ctx | decode tok/s | median W (the boards) | J / 1,000 tokens | tok/s per 100 W | core temp max °C | source file |
|---|---|---|---|---|---|---|---|---|
| **this board — an NVIDIA GeForce RTX 5090 Laptop GPU 24 GB** · fixed · 95 W | **95 W (`enforced.power.limit`, read back at every stage boundary)** | 65,536 | **28.839** | 90.31 | 3,131.56 | 31.9 | 45.0 | `results/gpu-5090-laptop-24g-m4-one-95w-auto-ladder.json` |
| **this board — an NVIDIA GeForce RTX 5090 Laptop GPU 24 GB** · dynamic-boost · 95–175 W | **140–150 W (median 150, n=374)** | 65,536 | **44.897** | 135.76 | 3,019.40 | 33.1 | 47.0 | `results/gpu-5090-laptop-24g-m4-one-dynboost-auto-ladder.json` |
| **this board — an NVIDIA GeForce RTX 5090 Laptop GPU 24 GB** · shipped · boost + clock lock | **150 W, flat (n=385) — ⚠ THE LIMIT DID NOT MOVE, which is a FINDING and not a range: under this posture `nvidia-powerd` is supposed to be floating it** | 65,536 | **44.702** | 129.90 | 2,904.35 | 34.4 | 46.0 | `results/gpu-5090-laptop-24g-m4-one-shipped-auto-ladder.json` |
| a GeForce RTX 3090 24 GB — another box, another build | **350 W** | 65,536 | **46.880** | 273.35 | 5,842.92 | 17.2 | N/A | `references/desktop-3090-m4-one-350w-auto-ladder.json` |
| a GeForce RTX 3090 Ti 24 GB — another box, another build | **350 W — the shared-cap rung** | 65,536 | **53.354** | 338.05 | 6,343.81 | 15.8 | 56.0 | `references/desktop-3090ti-m4-one-350w-auto-ladder.json` |
| a GeForce RTX 3090 Ti 24 GB — another box, another build | **450 W — this board's own limit** | 65,536 | **57.396** | 418.76 | 7,295.49 | 13.7 | 57.0 | `references/desktop-3090ti-m4-one-450w-auto-ladder.json` |
| two GeForce RTX 3080 10 GB — another box, another build | **300 W each — 600 W between them** | 49,152 | **44.702** | 509.01 (both boards) | 11,385.40 (both boards) | 8.8 | N/A | `references/desktop-2x3080-m4-split-forced-ladder.json` |

> **a GeForce RTX 3090 Ti 24 GB — another box, another build** — **two rows, one board.** The `450 W — this board's own limit` row is this board at the limit it ships with. The `350 W — the shared-cap rung` row above it is the rung every board of this ladder shares — which is 78 % of this board's own limit, while the 3090's 350 W is 100 % of its. Both rows are printed so neither reading can be mistaken for the other: they are ONE board at two caps, not two boards, and each names the file it was read from.
> **two GeForce RTX 3080 10 GB — another box, another build** — this rig needed `options.num_gpu 99`; the forcing is the finding.

> **This is not a cap-keyed table and it cannot be made into one.** This board's whole envelope — 95 W default, 175 W maximum, both READ off the card — sits below the lowest cap column the comparison piece has, which is 250 W. So the columns here are READINGS and the cap is printed beside every figure. Every reference row was measured in ANOTHER box on ANOTHER software stack: a ratio between two rows of this table is a box-and-board ratio, never a card-to-card one.
> `tok/s per 100 W` is derived from the two columns to its left, and both are printed so the derivation can be checked.

#### The context ceiling — the largest window each configuration held WHOLE

| model | configuration | largest window held whole | how the ladder ended | layer policy | source file |
|---|---|---|---|---|---|
| `gemma4:26b` | this board · fixed · 95 W | **131,072** | the ladder ran out of rungs before the card ran out of memory | layers left to ollama's own planner (`auto`) | `results/gpu-5090-laptop-24g-m5-one-95w-auto-ladder.json` |
| `gemma4:26b` | this board · dynamic-boost · 95–175 W | **131,072** | the ladder ran out of rungs before the card ran out of memory | layers left to ollama's own planner (`auto`) | `results/gpu-5090-laptop-24g-m5-one-dynboost-auto-ladder.json` |
| `gemma4:26b` | this board · shipped · boost + clock lock | **131,072** | the ladder ran out of rungs before the card ran out of memory | layers left to ollama's own planner (`auto`) | `results/gpu-5090-laptop-24g-m5-one-shipped-auto-ladder.json` |
| `gemma4:26b` | a GeForce RTX 3090 24 GB — another box, another build · 350 W | **131,072** | the ladder ran out of rungs before the card ran out of memory | layers left to ollama's own planner (`auto`) | `references/desktop-3090-m5-one-350w-auto-ladder.json` |
| `gemma4:26b` | a GeForce RTX 3090 Ti 24 GB — another box, another build · 350 W — the shared-cap rung | **131,072** | the ladder ran out of rungs before the card ran out of memory | layers left to ollama's own planner (`auto`) | `references/desktop-3090ti-m5-one-350w-auto-ladder.json` |
| `gemma4:26b` | a GeForce RTX 3090 Ti 24 GB — another box, another build · 450 W — this board's own limit | **131,072** | the ladder ran out of rungs before the card ran out of memory | layers left to ollama's own planner (`auto`) | `references/desktop-3090ti-m5-one-450w-auto-ladder.json` |
| `gemma4:26b` | two GeForce RTX 3080 10 GB — another box, another build · 300 W each — 600 W between them | **131,072** | the ladder ran out of rungs before the card ran out of memory | layer count FORCED to 99 via options.num_gpu | `references/desktop-2x3080-m5-split-forced-ladder.json` |
| `mistral-small3.2:24b` | this board · fixed · 95 W | **65,536** | spilled at 98,304 | layers left to ollama's own planner (`auto`) | `results/gpu-5090-laptop-24g-m4-one-95w-auto-ladder.json` |
| `mistral-small3.2:24b` | this board · dynamic-boost · 95–175 W | **65,536** | spilled at 98,304 | layers left to ollama's own planner (`auto`) | `results/gpu-5090-laptop-24g-m4-one-dynboost-auto-ladder.json` |
| `mistral-small3.2:24b` | this board · shipped · boost + clock lock | **65,536** | spilled at 98,304 | layers left to ollama's own planner (`auto`) | `results/gpu-5090-laptop-24g-m4-one-shipped-auto-ladder.json` |
| `mistral-small3.2:24b` | a GeForce RTX 3090 24 GB — another box, another build · 350 W | **65,536** | spilled at 98,304 | layers left to ollama's own planner (`auto`) | `references/desktop-3090-m4-one-350w-auto-ladder.json` |
| `mistral-small3.2:24b` | a GeForce RTX 3090 Ti 24 GB — another box, another build · 350 W — the shared-cap rung | **65,536** | spilled at 98,304 | layers left to ollama's own planner (`auto`) | `references/desktop-3090ti-m4-one-350w-auto-ladder.json` |
| `mistral-small3.2:24b` | a GeForce RTX 3090 Ti 24 GB — another box, another build · 450 W — this board's own limit | **65,536** | spilled at 98,304 | layers left to ollama's own planner (`auto`) | `references/desktop-3090ti-m4-one-450w-auto-ladder.json` |
| `mistral-small3.2:24b` | two GeForce RTX 3080 10 GB — another box, another build · 300 W each — 600 W between them | **49,152** | N/A | layer count FORCED to 99 via options.num_gpu | `references/desktop-2x3080-m4-split-forced-ladder.json` |

> **This table has no watt column on purpose.** A context ceiling is a memory fact, not a power fact: it is the one question a 24 GB mobile board and a 24 GB desktop board answer under identical conditions, so it is the one comparison here that a power posture cannot move.

## The absences

#### What this bench did NOT measure, named so nobody quotes a gap as a result

| the absence | the reason, stated rather than left to be discovered |
|---|---|
| **wall watts** | There is no `upsc` on this box: NUT is not installed, the battery exposes no `power_now`/`energy_now`, and `/sys/class/powercap/*/energy_uj` is root-only. **Board watts are the measured quantity** and every watt in this document is one. The same absence is the one the desktop bench box published its whole 24 GB ladder under, so the two are comparable in this respect. The receipt is read and quoted below, not merely produced. |
| **arm T, the trainer** | Its preprocessed tensor set lives on benchbox, which has been unreachable since 2026-09-21T15:03Z. Arm T is a SECOND TRAIN and legs 1 and 2 never depended on it. It is **absent, not failed**: no figure anywhere in this document is affected by it and none is estimated in its place. |
| **the fan column** | **a declared non-figure.** `fan.speed` reads `[N/A]` on this board — the driver reports no fan for it — so no result file can carry a fan column and none is filled from another board's row. The operator held the laptop's own maximum-fan key by hand for the whole of this pass (`OPERATOR-CONDITIONS.md`, 2026-09-21T17:30Z and 19:10Z); that is a condition of the measurement and not a reading of it. |
| **the memory-die temperature** | `temperature.memory` reads `[N/A]` on this board, as on every board of this ladder. Every thermal figure here is CORE temperature and the driver's own thermal-slowdown reasons are the rest of the thermal instrument. |
| **the peak-bandwidth ceiling** | **WITHHELD.** This is a GDDR7 part and the harness's two-bits-per-clock formula is a GDDR6/GDDR6X fact with no receipt for GDDR7 signalling (`INSTRUMENT-DELTA.md` §2). The memory CLOCK is a reading and prints in the ladder tables; `peak GB/s`, `decode GB/s` and `% of peak` are `null` in every result file rather than computed wrong. |
| **a control seat** | **This card is soldered.** Nothing comes out of this machine and nothing can go in, so there is no board that can sit in the same seat and be measured by the same instrument. The fourteen CPU-only rows this same laptop has published were measured by another bench on another day with the card parked, and folding them in as this leg's control would be the substitution this bench refuses everywhere else. |

| posture | the whole-box power receipt, READ |
|---|---|
| fixed · 95 W | WITHHELD — no `upsc` on this box: NUT is not installed and nut-server/nut-monitor are inactive (read 2026-09-21T14:52Z). Whole-box watts and joules-per-1,000-tokens are WITHHELD with this reason; board watts, which is what every table in this ladder prints, are read normally from the card at 2 Hz. (asked `pr1500@localhost` at 2026-09-21T18:43:15Z; `results/gpu-5090-laptop-24g-ups-gate-95w.txt`) |
| dynamic-boost · 95–175 W | WITHHELD — no `upsc` on this box: NUT is not installed and nut-server/nut-monitor are inactive (read 2026-09-21T14:52Z). Whole-box watts and joules-per-1,000-tokens are WITHHELD with this reason; board watts, which is what every table in this ladder prints, are read normally from the card at 2 Hz. (asked `pr1500@localhost` at 2026-09-21T21:01:52Z; `results/gpu-5090-laptop-24g-ups-gate-dynboost.txt`) |
| shipped · boost + clock lock | WITHHELD — no `upsc` on this box: NUT is not installed and nut-server/nut-monitor are inactive (read 2026-09-21T14:52Z). Whole-box watts and joules-per-1,000-tokens are WITHHELD with this reason; board watts, which is what every table in this ladder prints, are read normally from the card at 2 Hz. (asked `pr1500@localhost` at 2026-09-22T00:10:41Z; `results/gpu-5090-laptop-24g-ups-gate-shipped.txt`) |


---

*Every table above was generated by `summarize_5090laptop.py` from the JSON receipts in `results/` (and, for the comparison tables, from three sibling benches' receipts — one 3090, one 3090 Ti and the two-3080 rig — each row naming its own file).*

<!-- GENERATED:END -->

---

*Every table above the END marker was generated by `summarize_5090laptop.py` from the JSON
receipts in `results/` and, for the comparison tables, from three sibling benches' receipts —
one 3090, one 3090 Ti and the two-3080 rig — each row naming its own file. The prose around them
is written and is never touched by the generator. Regenerate with `./report.sh`; prove it current
with `./report.sh --check`.*
