PASS — the laptop's 5090, posture SHIPPED (Dynamic Boost running AND the clock lock ON) started 2026-09-22T00:10:35Z on this laptop harness /workshop/bench-laptop-5090-2026-09-21/harness-llm render /workshop/bench-laptop-5090-2026-09-21/harness-render log /workshop/bench-laptop-5090-2026-09-21/logs/rung-shipped.log ⚠ this pass has NO CAP. The board refuses -pl, so there is nothing to set and nothing to wait for: nvidia-powerd floats the enforced limit between this board's default (95 W) and its maximum (175 W) against the CPU's draw. Every figure's cap cell is a RANGE out of the 2 Hz trace, and a single wattage typed into one would be a fabrication. The per-stage ranges land in results/*-cap-posture-shipped-*.json. ⚠ AND THE CLOCKS RUN LOCKED HERE, which is the whole point and is the one thing that separates this pass from the dynamic-boost one. This box boots with ai-perf.service applying a 1200,2550 MHz lock, so a locked board is what its owner actually uses. The lock is READ FROM THE CARD and this pass REFUSES without it (driver_block.sh's GATE 2, just after --stop-all, where a card at rest can answer). If it refuses, the paste is: sudo nvidia-smi -lgc 1200,2550 and nothing needs undoing afterwards — that is the state this box boots in. ⚠ THESE ROWS SIT BESIDE THIS MACHINE'S OTHER TWO POSTURES AND NOTHING ELSE. benchbox, which produced the desktop rows this bench compares against, runs no clock lock at all; PREREG §11.2 stands, so no cell from this pass is offered for those tables. PREREG.md Amendment 4 registers all of that, and the ruling that asked for it. receipt: nvidia-powerd is active — Dynamic Boost IS floating this board's limit, which is the whole point of this posture. Its cap cell is a RANGE, never a number. envelope at this pass's open: 150 95 175 W (enforced, default, max) memory guard: memshed.timer was active at the window's open (remembered in /workshop/bench-laptop-5090-2026-09-21/.memshed-was) memory guard: memshed.timer is now inactive; MEMSHED_DRY_RUN=1 (STOPPED, never masked: the unit file is a symlink into ~/estate/systemd and 'mask --force' would destroy it. glimmer-check.timer is left alone — it next fires Tue 14:43Z, outside any daytime window, and masking a timer this bench does not need to touch is one more thing to forget to restore.) ================================================================================ PASS shipped — posture shipped — LEG 1 (the eleven-arm language bank), then LEG 2 card GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff (NVIDIA GeForce RTX 5090 Laptop GPU) class gpu-5090-laptop-24g cap arg shipped start 2026-09-22T00:10:35Z ================================================================================ receipt: nvidia-powerd is active — Dynamic Boost IS floating this board's limit, which is the whole point of this posture. Its cap cell is a RANGE, never a number. envelope at the pass's open: enforced/default/max = 150 95 175 W the enforced limit FLOATS in this posture. Its cell is a range from the 2 Hz trace. clock lock [by the CARD]: ON-BY-CARD evidence: the card was AT REST for the whole sample (utilization.gpu read 0 % on all 12 reads) and clocks.sm never left the lock's floor: min 1192 MHz, max 1192 MHz, against a declared floor of 1200 MHz and the driver's 25 MHz quantization tolerance (the driver snaps 1200 to 1192 on this board, measured 2026-09-21T23:13Z). Unlocked and at rest this board reads 232 MHz or less (232 MHz at 17:30Z, 180 MHz at 18:11Z), so a resting card pinned at the floor is the lock holding it there. clocks.max.sm reads 3090 MHz, which is this board's own unlocked maximum with OR without the lock and is therefore not part of this verdict. the unit ai-perf.service reads active — NOT EVIDENCE: it is Type=oneshot, so it reads 'active (exited)' for the whole uptime whether or not 'nvidia-smi -rgc' has released the lock since. On 2026-09-21 a receipt that trusted it published 'ran with the lock ACTIVE' beside clocks that proved otherwise. clock-lock receipt -> /workshop/bench-laptop-5090-2026-09-21/logs/clock-lock-shipped.txt ---- LEG 1: /workshop/bench-laptop-5090-2026-09-21/harness-llm/driver_block.sh shipped ---- card name read back: 'NVIDIA GeForce RTX 5090 Laptop GPU' (arm one; exact equality, not a substring) === the board, before anything is started (this is the shipped block) === 00:10:41Z index, uuid, name, power.limit [W], enforced.power.limit [W], power.default_limit [W], power.max_limit [W], persistence_mode, pcie.link.width.current, clocks.current.sm [MHz], clocks.current.memory [MHz], clocks.max.sm [MHz], temperature.gpu, fan.speed [%] 0, GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff, NVIDIA GeForce RTX 5090 Laptop GPU, [N/A], 150.00 W, 95.00 W, 175.00 W, Enabled, 8, 1192 MHz, 810 MHz, 3090 MHz, 31, [N/A] receipt: nvidia-powerd is active — Dynamic Boost IS floating this board's limit, which is the whole point of this posture. Its cap cell is a RANGE, never a number. posture [block shipped open]: shipped — enforced/default/max = 150 95 175 W the enforced limit FLOATS here. Its cell is a RANGE from the 2 Hz trace, never a number. memory guard: memshed.timer is inactive; MEMSHED_DRY_RUN=1 === the UPS at the door (expected: absent on this box) === 00:10:41Z UPS UNREADABLE at block start -- whole-box watts and joules-per-1,000-tokens are WITHHELD with this reason, and board watts are read normally: no `upsc` on this box: NUT is not installed and nut-server/nut-monitor are inactive (read 2026-09-21T14:52Z). Whole-box watts and joules-per-1,000-tokens are WITHHELD with this reason; board watts, which is what every table in this ladder prints, are read normally from the card at 2 Hz. receipt -> results/gpu-5090-laptop-24g-ups-gate-shipped.txt clock lock [by the CARD]: ON-BY-CARD evidence: the card was AT REST for the whole sample (utilization.gpu read 0 % on all 12 reads) and clocks.sm never left the lock's floor: min 1192 MHz, max 1192 MHz, against a declared floor of 1200 MHz and the driver's 25 MHz quantization tolerance (the driver snaps 1200 to 1192 on this board, measured 2026-09-21T23:13Z). Unlocked and at rest this board reads 232 MHz or less (232 MHz at 17:30Z, 180 MHz at 18:11Z), so a resting card pinned at the floor is the lock holding it there. clocks.max.sm reads 3090 MHz, which is this board's own unlocked maximum with OR without the lock and is therefore not part of this verdict. the unit ai-perf.service reads active — NOT EVIDENCE: it is Type=oneshot, so it reads 'active (exited)' for the whole uptime whether or not 'nvidia-smi -rgc' has released the lock since. On 2026-09-21 a receipt that trusted it published 'ran with the lock ACTIVE' beside clocks that proved otherwise. clock-lock receipt -> results/gpu-5090-laptop-24g-clock-lock-ON-BY-CARD-shipped.txt receipt: the clock lock is IN FORCE (ON-BY-CARD), read from the card — which is this posture's declared condition, not a confound it tolerates. the card was AT REST for the whole sample (utilization.gpu read 0 % on all 12 reads) and clocks.sm never left the lock's floor: min 1192 MHz, max 1192 MHz, against a declared floor of 1200 MHz and the driver's 25 MHz quantization tolerance (the driver snaps 1200 to 1192 on this board, measured 2026-09-21T23:13Z). Unlocked and at rest this board reads 232 MHz or less (232 MHz at 17:30Z, 180 MHz at 18:11Z), so a resting card pinned at the floor is the lock holding it there. clocks.max.sm reads 3090 MHz, which is this board's own unlocked maximum with OR without the lock and is therefore not part of this verdict. === instance: one card, kv q8_0, parallel 1 === 00:10:49Z "persistence_mode_at_start": "0, Enabled;", "pcie_link_width_at_start": "0, 8;" } next: /workshop/bench-laptop-5090-2026-09-21/harness-llm/setup_laptop_instance.sh --shape one --status pass --instance-env-file /workshop/bench-laptop-5090-2026-09-21/harness-llm/results/instance-env-one-kvq8_0-p1.json to every bencher so the result files carry it posture [card-mapping open]: shipped — enforced/default/max = 150 95 175 W the limit is FLOATING by design; this gate records it and never stops on a change. clocks at [card-mapping] open: 1590,14001,3090 (sm,mem,max.sm MHz) cap witness [card-mapping]: every 5s -> results/cap-witness-shipped-card-mapping.csv (pid 1782890) === card mapping: watch which physical card's memory rises === 00:10:53Z one: MATCHES: declared soldered (mobile; no socket — link width is traced under load, never assumed) at 00000000:01:00.0 - vendor unread - GPU-edff232c, and that is the card whose memory grew every shape checked runs on the card it declares -> /workshop/bench-laptop-5090-2026-09-21/harness-llm/results/gpu-5090-laptop-24g-card-mapping-shipped.json clocks at [card-mapping] close: 1597,14001,3090 (sm,mem,max.sm MHz) posture [card-mapping close]: shipped — enforced/default/max at open 150 95 175 W, at close 150 95 175 W enforced.power.limit over this stage: min 150 W, median 150 W, max 150 W, mean 150.0 W, n=2 the board envelope it floated inside: default 95.0 W, max 175.0 W cap cell for this stage: 150 W, flat ⚠ THE LIMIT DID NOT MOVE over this stage. That is a finding, not a range: check that nvidia-powerd is actually floating the TGP. stage envelope -> results/gpu-5090-laptop-24g-cap-posture-shipped-card-mapping.json posture [armA-ladder-gemma4 open]: shipped — enforced/default/max = 150 95 175 W the limit is FLOATING by design; this gate records it and never stops on a change. clocks at [armA-ladder-gemma4] open: 1590,14001,3090 (sm,mem,max.sm MHz) cap witness [armA-ladder-gemma4]: every 5s -> results/cap-witness-shipped-armA-ladder-gemma4.csv (pid 1783283) === armA-ladder-gemma4 === 00:11:01Z run 2: decode=155.269 tok/s ttft=358.89 ms W=104.46 mem/card={'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 18845.0} run 3: decode=155.493 tok/s ttft=374.09 ms W=97.64 mem/card={'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 18845.0} --- num_ctx 98,304 --- fits: fully resident on the card; per-card delta {'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 500.0} per-card memory.used (MiB): {'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 19345.0}; runner RSS 1.48 GiB run 1: decode=155.529 tok/s ttft=401.06 ms W=102.31 mem/card={'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 19345.0} contention gate: busy: busiest card mean 17.8% > 5.0% -- waiting (attempt 1/6) run 2: decode=155.368 tok/s ttft=367.29 ms W=100.81 mem/card={'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 19345.0} run 3: decode=155.232 tok/s ttft=380.4 ms W=103.31 mem/card={'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 19345.0} --- num_ctx 131,072 --- fits: fully resident on the card; per-card delta {'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 500.0} per-card memory.used (MiB): {'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 19845.0}; runner RSS 1.51 GiB run 1: decode=155.246 tok/s ttft=369.68 ms W=116.12 mem/card={'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 19845.0} contention gate: busy: busiest card mean 18.4% > 5.0% -- waiting (attempt 1/6) run 2: decode=155.288 tok/s ttft=374.81 ms W=115.76 mem/card={'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 19845.0} contention gate: busy: busiest card mean 13.8% > 5.0% -- waiting (attempt 1/6) run 3: decode=155.43 tok/s ttft=371.01 ms W=146.42 mem/card={'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 19845.0} the biggest context this arm holds for gemma4:26b is 131,072 tokens -> /workshop/bench-laptop-5090-2026-09-21/harness-llm/results/gpu-5090-laptop-24g-m5-one-shipped-auto-ladder.json clocks at [armA-ladder-gemma4] close: 1597,14001,3090 (sm,mem,max.sm MHz) posture [armA-ladder-gemma4 close]: shipped — enforced/default/max at open 150 95 175 W, at close 150 95 175 W enforced.power.limit over this stage: min 150 W, median 150 W, max 150 W, mean 150.0 W, n=99 the board envelope it floated inside: default 95.0 W, max 175.0 W cap cell for this stage: 150 W, flat ⚠ THE LIMIT DID NOT MOVE over this stage. That is a finding, not a range: check that nvidia-powerd is actually floating the TGP. stage envelope -> results/gpu-5090-laptop-24g-cap-posture-shipped-armA-ladder-gemma4.json arm A's headline rung at shipped, read from its own result file: num_ctx 131072 posture [armA2-kv-axis open]: shipped — enforced/default/max = 150 95 175 W the limit is FLOATING by design; this gate records it and never stops on a change. clocks at [armA2-kv-axis] open: 1590,14001,3090 (sm,mem,max.sm MHz) cap witness [armA2-kv-axis]: every 5s -> results/cap-witness-shipped-armA2-kv-axis.csv (pid 1802401) === A2 instance kv=f16 === 00:19:21Z } next: /workshop/bench-laptop-5090-2026-09-21/harness-llm/setup_laptop_instance.sh --shape one --status pass --instance-env-file /workshop/bench-laptop-5090-2026-09-21/harness-llm/results/instance-env-one-kvf16-p1.json to every bencher so the result files carry it === armA2-kvf16-at131072 === 00:19:43Z prompt sha256=90eedd0c53f9554ae3837674504fcb7090013d9a432a653352183c0f25a7ce5c bytes=2101 base=http://127.0.0.1:11470 arm=one live=False runs=3 ### gpu-5090-laptop-24g m5 gemma4:26b -- arm one (the one NVIDIA GeForce RTX 5090 Laptop GPU 24 GB soldered to this machine's board, everything ollama's planner will put on it) ### memory-temperature sensor: the memory die's temperature is NOT readable on these cards through any of the three paths asked, so the memory-temperature stop condition could not arm and NO memory temperature is reported anywhere in this bench. The core temperature stop and the driver's own thermal-slowdown reasons are the thermal instrument instead. UPS before arm: load=None% status=None keep-alive: bench instance: held for the arm's duration /api/ps: size=17477150965 size_vram=17477150965 -> fits: fully resident on the card per-card VRAM delta (MiB): {'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 20696.0} run 1: decode=165.991 tok/s (wall 165.044) ttft=424.8 ms W=120.75 J/1k=727.45 temp=44.0C fan=None% clk=2070.0-2257.0MHz mem=14001.0MHz NEITHER cap: the SM clock varied with draw at 80.5% of the cap and the card at 44 C. On a memory-bound decode the clock follows the work, and nothing here was limiting it contention gate: busy: busiest card mean 23.5% > 5.0% -- waiting (attempt 1/6) run 2: decode=166.369 tok/s (wall 165.424) ttft=452.09 ms W=103.8 J/1k=623.92 temp=43.0C fan=None% clk=2092.0-2250.0MHz mem=14001.0MHz NEITHER cap: the SM clock varied with draw at 69.2% of the cap and the card at 43 C. On a memory-bound decode the clock follows the work, and nothing here was limiting it contention gate: busy: busiest card mean 23.5% > 5.0% -- waiting (attempt 1/6) run 3: decode=165.967 tok/s (wall 164.972) ttft=452.9 ms W=104.05 J/1k=626.93 temp=43.0C fan=None% clk=2092.0-2250.0MHz mem=14001.0MHz NEITHER cap: the SM clock varied with draw at 69.4% of the cap and the card at 43 C. On a memory-bound decode the clock follows the work, and nothing here was limiting it spread gate: spread 0.2% of the median within the 15% gate -> /workshop/bench-laptop-5090-2026-09-21/harness-llm/results/gpu-5090-laptop-24g-m5-one-shipped-kvf16-at131072.json wrote 1 file(s) kv=f16 holds num_ctx 131072 at shipped -- that is this KV type's largest window === A2 instance kv=q4_0 === 00:21:00Z } next: /workshop/bench-laptop-5090-2026-09-21/harness-llm/setup_laptop_instance.sh --shape one --status pass --instance-env-file /workshop/bench-laptop-5090-2026-09-21/harness-llm/results/instance-env-one-kvq4_0-p1.json to every bencher so the result files carry it === armA2-kvq4_0-at131072 === 00:21:36Z prompt sha256=90eedd0c53f9554ae3837674504fcb7090013d9a432a653352183c0f25a7ce5c bytes=2101 base=http://127.0.0.1:11470 arm=one live=False runs=3 ### gpu-5090-laptop-24g m5 gemma4:26b -- arm one (the one NVIDIA GeForce RTX 5090 Laptop GPU 24 GB soldered to this machine's board, everything ollama's planner will put on it) ### memory-temperature sensor: the memory die's temperature is NOT readable on these cards through any of the three paths asked, so the memory-temperature stop condition could not arm and NO memory temperature is reported anywhere in this bench. The core temperature stop and the driver's own thermal-slowdown reasons are the thermal instrument instead. UPS before arm: load=None% status=None keep-alive: bench instance: held for the arm's duration /api/ps: size=17747851344 size_vram=17747851344 -> fits: fully resident on the card per-card VRAM delta (MiB): {'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 19116.0} run 1: decode=155.889 tok/s (wall 154.852) ttft=352.58 ms W=135.35 J/1k=868.25 temp=44.0C fan=None% clk=2265.0-2332.0MHz mem=14001.0MHz no clock drop during decode (draw held at 90.2% of the 150 W cap) run 2: decode=155.534 tok/s (wall 154.675) ttft=367.21 ms W=98.31 J/1k=632.08 temp=44.0C fan=None% clk=1747.0-2332.0MHz mem=14001.0MHz NEITHER cap: the SM clock varied with draw at 65.5% of the cap and the card at 44 C. On a memory-bound decode the clock follows the work, and nothing here was limiting it run 3: decode=156.131 tok/s (wall 155.22) ttft=373.35 ms W=124.9 J/1k=799.97 temp=44.0C fan=None% clk=2272.0-2332.0MHz mem=14001.0MHz no clock drop during decode spread gate: spread 0.4% of the median within the 15% gate -> /workshop/bench-laptop-5090-2026-09-21/harness-llm/results/gpu-5090-laptop-24g-m5-one-shipped-kvq4_0-at131072.json wrote 1 file(s) kv=q4_0 holds num_ctx 131072 at shipped -- that is this KV type's largest window clocks at [armA2-kv-axis] close: 1987,14001,3090 (sm,mem,max.sm MHz) posture [armA2-kv-axis close]: shipped — enforced/default/max at open 150 95 175 W, at close 150 95 175 W enforced.power.limit over this stage: min 150 W, median 150 W, max 150 W, mean 150.0 W, n=39 the board envelope it floated inside: default 95.0 W, max 175.0 W cap cell for this stage: 150 W, flat ⚠ THE LIMIT DID NOT MOVE over this stage. That is a finding, not a range: check that nvidia-powerd is actually floating the TGP. stage envelope -> results/gpu-5090-laptop-24g-cap-posture-shipped-armA2-kv-axis.json === restore the q8_0 instance (the seat posture) === 00:22:33Z next: /workshop/bench-laptop-5090-2026-09-21/harness-llm/setup_laptop_instance.sh --shape one --status pass --instance-env-file /workshop/bench-laptop-5090-2026-09-21/harness-llm/results/instance-env-one-kvq8_0-p1.json to every bencher so the result files carry it posture [armA3-filled open]: shipped — enforced/default/max = 150 95 175 W the limit is FLOATING by design; this gate records it and never stops on a change. clocks at [armA3-filled] open: 1192,810,3090 (sm,mem,max.sm MHz) cap witness [armA3-filled]: every 5s -> results/cap-witness-shipped-armA3-filled.csv (pid 1810779) === armA3-filled === 00:23:09Z prompt sha256=90eedd0c53f9554ae3837674504fcb7090013d9a432a653352183c0f25a7ce5c bytes=2101 base=http://127.0.0.1:11470 arm=one live=False runs=3 memory-temperature sensor: the memory die's temperature is NOT readable on these cards through any of the three paths asked, so the memory-temperature stop condition could not arm and NO memory temperature is reported anywhere in this bench. The core temperature stop and the driver's own thermal-slowdown reasons are the thermal instrument instead. ### gpu-5090-laptop-24g m5 gemma4:26b -- FILLED WINDOW at num_ctx 131,072 (75% target = 98,304 tokens) ### fit at the window: fits: fully resident on the card per-card memory.used (MiB): {'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 19845.0} filler round 1: 183 copies -> 97992 tokens (target 98304) window filled: 97,992 tokens of 131,072 (74.8% occupied) COLD prefill: 2610.692 tok/s over 98025 tokens, ttft 38696.9 ms contention gate: busy: busiest card mean 38.0% > 5.0% -- waiting (attempt 1/6) run 1: prefill=1300335.726 tok/s (97992 tokens) ttft=745.29 ms decode=72.221 tok/s W=100.55 contention gate: busy: busiest card mean 59.1% > 5.0% -- waiting (attempt 1/6) run 2: prefill=1352546.584 tok/s (97992 tokens) ttft=672.85 ms decode=72.364 tok/s W=108.96 contention gate: busy: busiest card mean 19.0% > 5.0% -- waiting (attempt 1/6) run 3: prefill=1323518.686 tok/s (97992 tokens) ttft=678.53 ms decode=72.39 tok/s W=112.83 prompt tokens: 97992 of 131072 (0.7476) prefill: 1323518.686 tok/s ttft: 678.53 ms decode at depth: 72.364 tok/s -> /workshop/bench-laptop-5090-2026-09-21/harness-llm/results/gpu-5090-laptop-24g-m5-one-shipped-filled.json clocks at [armA3-filled] close: 1620,14001,3090 (sm,mem,max.sm MHz) posture [armA3-filled close]: shipped — enforced/default/max at open 150 95 175 W, at close 150 95 175 W enforced.power.limit over this stage: min 150 W, median 150 W, max 150 W, mean 150.0 W, n=34 the board envelope it floated inside: default 95.0 W, max 175.0 W cap cell for this stage: 150 W, flat ⚠ THE LIMIT DID NOT MOVE over this stage. That is a finding, not a range: check that nvidia-powerd is actually floating the TGP. stage envelope -> results/gpu-5090-laptop-24g-cap-posture-shipped-armA3-filled.json posture [armD-ladder-mistral open]: shipped — enforced/default/max = 150 95 175 W the limit is FLOATING by design; this gate records it and never stops on a change. clocks at [armD-ladder-mistral] open: 1590,14001,3090 (sm,mem,max.sm MHz) cap witness [armD-ladder-mistral]: every 5s -> results/cap-witness-shipped-armD-ladder-mistral.csv (pid 1818321) === armD-ladder-mistral === 00:25:59Z contention gate: busy: busiest card mean 14.8% > 5.0% -- waiting (attempt 1/6) run 2: decode=44.765 tok/s ttft=202.19 ms W=129.37 mem/card={'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 18865.0} contention gate: busy: busiest card mean 10.2% > 5.0% -- waiting (attempt 1/6) run 3: decode=44.751 tok/s ttft=191.25 ms W=134.81 mem/card={'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 18865.0} --- num_ctx 65,536 --- fits: fully resident on the card; per-card delta {'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 1440.0} per-card memory.used (MiB): {'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 20305.0}; runner RSS 0.88 GiB contention gate: busy: busiest card mean 55.0% > 5.0% -- waiting (attempt 1/6) run 1: decode=44.702 tok/s ttft=206.28 ms W=133.71 mem/card={'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 20305.0} contention gate: busy: busiest card mean 14.8% > 5.0% -- waiting (attempt 1/6) run 2: decode=44.68 tok/s ttft=191.65 ms W=129.41 mem/card={'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 20305.0} contention gate: busy: busiest card mean 9.9% > 5.0% -- waiting (attempt 1/6) run 3: decode=44.726 tok/s ttft=184.56 ms W=129.9 mem/card={'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 20305.0} --- num_ctx 98,304 --- does not fit: only part of the model is on the card (92.2% VRAM / 7.8% RAM); per-card delta {'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 2550.0} per-card memory.used (MiB): {'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 22855.0}; runner RSS 14.17 GiB ladder stops at num_ctx 98,304: does not fit: only part of the model is on the card (92.2% VRAM / 7.8% RAM) the biggest context this arm holds for mistral-small3.2:24b is 65,536 tokens -> /workshop/bench-laptop-5090-2026-09-21/harness-llm/results/gpu-5090-laptop-24g-m4-one-shipped-auto-ladder.json clocks at [armD-ladder-mistral] close: 1590,14001,3090 (sm,mem,max.sm MHz) posture [armD-ladder-mistral close]: shipped — enforced/default/max at open 150 95 175 W, at close 150 95 175 W enforced.power.limit over this stage: min 150 W, median 150 W, max 150 W, mean 150.0 W, n=104 the board envelope it floated inside: default 95.0 W, max 175.0 W cap cell for this stage: 150 W, flat ⚠ THE LIMIT DID NOT MOVE over this stage. That is a finding, not a range: check that nvidia-powerd is actually floating the TGP. stage envelope -> results/gpu-5090-laptop-24g-cap-posture-shipped-armD-ladder-mistral.json posture [armD-forced-probe open]: shipped — enforced/default/max = 150 95 175 W the limit is FLOATING by design; this gate records it and never stops on a change. clocks at [armD-forced-probe] open: 1590,14001,3090 (sm,mem,max.sm MHz) cap witness [armD-forced-probe]: every 5s -> results/cap-witness-shipped-armD-forced-probe.csv (pid 1856992) === D-probe: the planner SPILLED at num_ctx 98304 -- re-running it forced === 00:34:39Z === armD-forced-at98304 === 00:34:39Z prompt sha256=90eedd0c53f9554ae3837674504fcb7090013d9a432a653352183c0f25a7ce5c bytes=2101 base=http://127.0.0.1:11470 arm=one live=False runs=3 ### gpu-5090-laptop-24g m4 mistral-small3.2:24b -- arm one (the one NVIDIA GeForce RTX 5090 Laptop GPU 24 GB soldered to this machine's board, everything ollama's planner will put on it) ### memory-temperature sensor: the memory die's temperature is NOT readable on these cards through any of the three paths asked, so the memory-temperature stop condition could not arm and NO memory temperature is reported anywhere in this bench. The core temperature stop and the driver's own thermal-slowdown reasons are the thermal instrument instead. UPS before arm: load=None% status=None keep-alive: bench instance: held for the arm's duration /api/ps: size=23109042175 size_vram=23109042175 -> fits: fully resident on the card per-card VRAM delta (MiB): {'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 22908.0} run 1: decode=44.801 tok/s (wall 44.615) ttft=203.02 ms W=135.48 J/1k=3024.03 temp=48.0C fan=None% clk=1642.0-1657.0MHz mem=14001.0MHz no clock drop during decode (draw held at 90.3% of the 150 W cap) contention gate: busy: busiest card mean 14.8% > 5.0% -- waiting (attempt 1/6) run 2: decode=44.813 tok/s (wall 44.62) ttft=208.19 ms W=130.17 J/1k=2904.72 temp=47.0C fan=None% clk=1642.0-1657.0MHz mem=14001.0MHz no clock drop during decode contention gate: busy: busiest card mean 24.8% > 5.0% -- waiting (attempt 1/6) run 3: decode=44.767 tok/s (wall 44.581) ttft=217.01 ms W=129.82 J/1k=2899.91 temp=46.0C fan=None% clk=1635.0-1657.0MHz mem=14001.0MHz no clock drop during decode spread gate: spread 0.1% of the median within the 15% gate -> /workshop/bench-laptop-5090-2026-09-21/harness-llm/results/gpu-5090-laptop-24g-m4-one-shipped-forced-at98304.json wrote 1 file(s) === D-probe: the forced rung FIT -- climbing one more to find the first that does not === 00:36:11Z === armD-forced-at131072 === 00:36:11Z prompt sha256=90eedd0c53f9554ae3837674504fcb7090013d9a432a653352183c0f25a7ce5c bytes=2101 base=http://127.0.0.1:11470 arm=one live=False runs=3 ### gpu-5090-laptop-24g m4 mistral-small3.2:24b -- arm one (the one NVIDIA GeForce RTX 5090 Laptop GPU 24 GB soldered to this machine's board, everything ollama's planner will put on it) ### memory-temperature sensor: the memory die's temperature is NOT readable on these cards through any of the three paths asked, so the memory-temperature stop condition could not arm and NO memory temperature is reported anywhere in this bench. The core temperature stop and the driver's own thermal-slowdown reasons are the thermal instrument instead. UPS before arm: load=None% status=None keep-alive: bench instance: held for the arm's duration /api/ps: size=None size_vram=None -> refused: /api/ps reported no size for this model per-card VRAM delta (MiB): {'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 0.0} -> /workshop/bench-laptop-5090-2026-09-21/harness-llm/results/gpu-5090-laptop-24g-m4-one-shipped-forced-at131072.json wrote 1 file(s) clocks at [armD-forced-probe] close: 1590,14001,3090 (sm,mem,max.sm MHz) posture [armD-forced-probe close]: shipped — enforced/default/max at open 150 95 175 W, at close 150 95 175 W enforced.power.limit over this stage: min 150 W, median 150 W, max 150 W, mean 150.0 W, n=21 the board envelope it floated inside: default 95.0 W, max 175.0 W cap cell for this stage: 150 W, flat ⚠ THE LIMIT DID NOT MOVE over this stage. That is a finding, not a range: check that nvidia-powerd is actually floating the TGP. stage envelope -> results/gpu-5090-laptop-24g-cap-posture-shipped-armD-forced-probe.json posture [armE-concurrency open]: shipped — enforced/default/max = 150 95 175 W the limit is FLOATING by design; this gate records it and never stops on a change. clocks at [armE-concurrency] open: 1590,14001,3090 (sm,mem,max.sm MHz) cap witness [armE-concurrency]: every 5s -> results/cap-witness-shipped-armE-concurrency.csv (pid 1860759) === E instance: parallel 4 === 00:36:23Z next: /workshop/bench-laptop-5090-2026-09-21/harness-llm/setup_laptop_instance.sh --shape one --status pass --instance-env-file /workshop/bench-laptop-5090-2026-09-21/harness-llm/results/instance-env-one-kvq8_0-p4.json to every bencher so the result files carry it === E: concurrency 1 / 2 / 4 at shipped === 00:36:58Z ### concurrency level 2 ### contention gate: busy: busiest card mean 9.1% > 5.0% -- waiting (attempt 1/6) level 2: sum 255.69 tok/s, wall 200.455 tok/s over 2.5542s, W=104.13 temp=45.0C fan=None% per stream: [123.823, 131.867] level 2: sum 258.544 tok/s, wall 202.328 tok/s over 2.5305s, W=101.94 temp=45.0C fan=None% per stream: [129.297, 129.247] level 2: sum 259.142 tok/s, wall 204.545 tok/s over 2.5031s, W=128.14 temp=46.0C fan=None% per stream: [129.543, 129.599] ### concurrency level 4 ### level 4: sum 351.036 tok/s, wall 288.784 tok/s over 3.5459s, W=125.25 temp=46.0C fan=None% per stream: [85.866, 88.425, 88.359, 88.386] level 4: sum 354.932 tok/s, wall 277.294 tok/s over 3.6928s, W=121.92 temp=46.0C fan=None% per stream: [88.776, 88.748, 88.718, 88.69] level 4: sum 354.05 tok/s, wall 284.359 tok/s over 3.6011s, W=115.56 temp=47.0C fan=None% per stream: [88.506, 88.505, 88.506, 88.533] one-shipped takes 4 parallel streams at 354.1 tok/s aggregate (sum of streams) / 284.4 tok/s over wall -> /workshop/bench-laptop-5090-2026-09-21/harness-llm/results/gpu-5090-laptop-24g-concurrency-one-shipped.json clocks at [armE-concurrency] close: 1890,14001,3090 (sm,mem,max.sm MHz) posture [armE-concurrency close]: shipped — enforced/default/max at open 150 95 175 W, at close 150 95 175 W enforced.power.limit over this stage: min 150 W, median 150 W, max 150 W, mean 150.0 W, n=34 the board envelope it floated inside: default 95.0 W, max 175.0 W cap cell for this stage: 150 W, flat ⚠ THE LIMIT DID NOT MOVE over this stage. That is a finding, not a range: check that nvidia-powerd is actually floating the TGP. stage envelope -> results/gpu-5090-laptop-24g-cap-posture-shipped-armE-concurrency.json posture [g2-doorman open]: shipped — enforced/default/max = 150 95 175 W the limit is FLOATING by design; this gate records it and never stops on a change. clocks at [g2-doorman] open: 1590,14001,3090 (sm,mem,max.sm MHz) cap witness [g2-doorman]: every 5s -> results/cap-witness-shipped-g2-doorman.csv (pid 1866237) === G2: the doorman on mistral at num_ctx 65536 (1 caller, then 4) === 00:39:10Z memory-temperature sensor: the memory die's temperature is NOT readable on these cards through any of the three paths asked, so the memory-temperature stop condition could not arm and NO memory temperature is reported anywhere in this bench. The core temperature stop and the driver's own thermal-slowdown reasons are the thermal instrument instead. fit: SPLIT and scored as one: 56.2% VRAM / 43.8% RAM contention gate: busy: busiest card mean 9.6% > 5.0% -- waiting (attempt 1/6) concurrency 1: 10 calls, median 430.85 ms, p95 448.2 ms, max 448.2 ms, W/card {'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 32.37} concurrency 4: 8 calls, median 885.73 ms, p95 4452.8 ms, max 4452.8 ms, W/card {'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 70.93} a doorman call on mistral-small3.2:24b costs 430.85 ms at the median and 448.2 ms at p95, one caller at a time -> /workshop/bench-laptop-5090-2026-09-21/harness-llm/results/gpu-5090-laptop-24g-g2-doorman-shipped.json clocks at [g2-doorman] close: 1590,14001,3090 (sm,mem,max.sm MHz) posture [g2-doorman close]: shipped — enforced/default/max at open 150 95 175 W, at close 150 95 175 W enforced.power.limit over this stage: min 150 W, median 150 W, max 150 W, mean 150.0 W, n=13 the board envelope it floated inside: default 95.0 W, max 175.0 W cap cell for this stage: 150 W, flat ⚠ THE LIMIT DID NOT MOVE over this stage. That is a finding, not a range: check that nvidia-powerd is actually floating the TGP. stage envelope -> results/gpu-5090-laptop-24g-cap-posture-shipped-g2-doorman.json === G3/G4 instance: back to parallel 1 === 00:40:14Z next: /workshop/bench-laptop-5090-2026-09-21/harness-llm/setup_laptop_instance.sh --shape one --status pass --instance-env-file /workshop/bench-laptop-5090-2026-09-21/harness-llm/results/instance-env-one-kvq8_0-p1.json to every bencher so the result files carry it posture [g3-vision open]: shipped — enforced/default/max = 150 95 175 W the limit is FLOATING by design; this gate records it and never stops on a change. clocks at [g3-vision] open: 1590,14001,3090 (sm,mem,max.sm MHz) cap witness [g3-vision]: every 5s -> results/cap-witness-shipped-g3-vision.csv (pid 1868603) === G3: a vision seat (TEXT path only -- no image is read or sent by this bench) === 00:40:20Z === g3-vision === 00:40:20Z prompt sha256=90eedd0c53f9554ae3837674504fcb7090013d9a432a653352183c0f25a7ce5c bytes=2101 base=http://127.0.0.1:11470 arm=one live=False runs=3 ### gpu-5090-laptop-24g g3 minicpm-v4.5:latest -- arm one (the one NVIDIA GeForce RTX 5090 Laptop GPU 24 GB soldered to this machine's board, everything ollama's planner will put on it) ### memory-temperature sensor: the memory die's temperature is NOT readable on these cards through any of the three paths asked, so the memory-temperature stop condition could not arm and NO memory temperature is reported anywhere in this bench. The core temperature stop and the driver's own thermal-slowdown reasons are the thermal instrument instead. UPS before arm: load=None% status=None keep-alive: bench instance: held for the arm's duration /api/ps: size=5596931685 size_vram=5596931685 -> fits: fully resident on the card per-card VRAM delta (MiB): {'GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff': 6696.0} run 1: decode=121.65 tok/s (wall 120.421) ttft=127.47 ms W=99.24 J/1k=815.78 temp=44.0C fan=None% clk=1792.0-1830.0MHz mem=14001.0MHz no clock drop during decode run 2: decode=121.581 tok/s (wall 120.436) ttft=140.8 ms W=98.15 J/1k=807.28 temp=43.0C fan=None% clk=1800.0-1800.0MHz mem=14001.0MHz no clock drop during decode contention gate: busy: busiest card mean 14.6% > 5.0% -- waiting (attempt 1/6) run 3: decode=121.145 tok/s (wall 119.955) ttft=153.19 ms W=73.29 J/1k=604.98 temp=43.0C fan=None% clk=1800.0-1800.0MHz mem=14001.0MHz no clock drop during decode spread gate: spread 0.4% of the median within the 15% gate -> /workshop/bench-laptop-5090-2026-09-21/harness-llm/results/gpu-5090-laptop-24g-g3-one-shipped-g3-vision.json wrote 1 file(s) clocks at [g3-vision] close: 1762,14001,3090 (sm,mem,max.sm MHz) posture [g3-vision close]: shipped — enforced/default/max at open 150 95 175 W, at close 150 95 175 W enforced.power.limit over this stage: min 150 W, median 150 W, max 150 W, mean 150.0 W, n=14 the board envelope it floated inside: default 95.0 W, max 175.0 W cap cell for this stage: 150 W, flat ⚠ THE LIMIT DID NOT MOVE over this stage. That is a finding, not a range: check that nvidia-powerd is actually floating the TGP. stage envelope -> results/gpu-5090-laptop-24g-cap-posture-shipped-g3-vision.json posture [g4-embed open]: shipped — enforced/default/max = 150 95 175 W the limit is FLOATING by design; this gate records it and never stops on a change. clocks at [g4-embed] open: 1597,14001,3090 (sm,mem,max.sm MHz) cap witness [g4-embed]: every 5s -> results/cap-witness-shipped-g4-embed.csv (pid 1870916) === G4: embeddings === 00:41:28Z memory-temperature sensor: the memory die's temperature is NOT readable on these cards through any of the three paths asked, so the memory-temperature stop condition could not arm and NO memory temperature is reported anywhere in this bench. The core temperature stop and the driver's own thermal-slowdown reasons are the thermal instrument instead. keep-alive: bench instance: held for the arm's duration batch 1: 247.88 ms, 258.185 texts/s, W=22.47 batch 2: 249.2 ms, 256.826 texts/s, W=22.54 batch 3: 256.16 ms, 249.84 texts/s, W=22.53 batch 4: 255.33 ms, 250.655 texts/s, W=22.65 batch 5: 257.69 ms, 248.358 texts/s, W=22.57 one-shipped: 250.7 texts/s on a batch of 64 (255.33 ms per batch) -> /workshop/bench-laptop-5090-2026-09-21/harness-llm/results/gpu-5090-laptop-24g-embed-one-shipped.json clocks at [g4-embed] close: 1702,14001,3090 (sm,mem,max.sm MHz) posture [g4-embed close]: shipped — enforced/default/max at open 150 95 175 W, at close 150 95 175 W enforced.power.limit over this stage: min 150 W, median 150 W, max 150 W, mean 150.0 W, n=5 the board envelope it floated inside: default 95.0 W, max 175.0 W cap cell for this stage: 150 W, flat ⚠ THE LIMIT DID NOT MOVE over this stage. That is a finding, not a range: check that nvidia-powerd is actually floating the TGP. stage envelope -> results/gpu-5090-laptop-24g-cap-posture-shipped-g4-embed.json posture [idle open]: shipped — enforced/default/max = 150 95 175 W the limit is FLOATING by design; this gate records it and never stops on a change. clocks at [idle] open: 1590,14001,3090 (sm,mem,max.sm MHz) cap witness [idle]: every 5s -> results/cap-witness-shipped-idle.csv (pid 1871708) === idle receipt at shipped === 00:41:53Z memory-temperature sensor: the memory die's temperature is NOT readable on these cards through any of the three paths asked, so the memory-temperature stop condition could not arm and NO memory temperature is reported anywhere in this bench. The core temperature stop and the driver's own thermal-slowdown reasons are the thermal instrument instead. empty: no model on the cards cards empty, nothing loaded: per card soldered (mobile; no socket — link width is traced under load, never assumed) at 00000000:01:00.0 - vendor unread - GPU-edff232c 11.57 W | pair 11.57 W | UPS whole box None W (None %) loading gemma4:26b (num_ctx=131072, num_gpu=None) fit: fits: fully resident on the card model resident, not generating: per card soldered (mobile; no socket — link width is traced under load, never assumed) at 00000000:01:00.0 - vendor unread - GPU-edff232c 14.42 W | pair 14.42 W | UPS whole box None W (None %) with gemma4:26b resident and not generating, the card draws 14.42 W (board power), against 11.57 W with the cards empty -- a resident cost of 2.85 W -> /workshop/bench-laptop-5090-2026-09-21/harness-llm/results/gpu-5090-laptop-24g-idle-gemma4-shipped.json clocks at [idle] close: 1192,14001,3090 (sm,mem,max.sm MHz) posture [idle close]: shipped — enforced/default/max at open 150 95 175 W, at close 150 95 175 W enforced.power.limit over this stage: min 150 W, median 150 W, max 150 W, mean 150.0 W, n=20 the board envelope it floated inside: default 95.0 W, max 175.0 W cap cell for this stage: 150 W, flat ⚠ THE LIMIT DID NOT MOVE over this stage. That is a finding, not a range: check that nvidia-powerd is actually floating the TGP. stage envelope -> results/gpu-5090-laptop-24g-cap-posture-shipped-idle.json === shipped block finished === 00:43:29Z stopping bench-5090laptop-cpu-box --- receipt: units --- --- receipt: ports --- no bench port listening (1147[01]) -- clean index, uuid, power.limit [W], enforced.power.limit [W], persistence_mode, temperature.gpu, fan.speed [%], memory.used [MiB], clocks.current.sm [MHz], clocks.current.memory [MHz], clocks.max.sm [MHz] 0, GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff, [N/A], 150.00 W, Enabled, 33, [N/A], 15 MiB, 1192 MHz, 14001 MHz, 3090 MHz posture [block shipped close]: shipped — enforced/default/max = 150 95 175 W (the envelope at the block's open was 150 95 175 W. The enforced limit moving between the two IS the measurement; the DEFAULT or the MAX moving would not be, and cap_stage_close_posture refuses on that at every stage boundary.) the UPS is absent on this box -- see results/gpu-5090-laptop-24g-ups-gate-shipped.txt for the withheld-figure receipt written at the block's open. Board watts are the measured quantity in every table this bench fills. LEG 1 at shipped exited 0 envelope between the legs: 150 95 175 W ---- LEG 2: /workshop/bench-laptop-5090-2026-09-21/harness-render/run_rung.sh shipped ---- == render rung shipped · started 2026-09-22T00:43:29Z == card expected NVIDIA GeForce RTX 5090 Laptop GPU comfy root /workshop/bench-laptop-5090-2026-09-21/ComfyUI-0.21.1 comfy python /workshop/ComfyUI/.venv/bin/python comfy version 0.21.1 (26515acd) receipt: nvidia-powerd is active — Dynamic Boost IS floating this board's limit, which is the whole point of this posture. Its cap cell is a RANGE, never a number. card reads 0, 00000000:01:00.0, NVIDIA GeForce RTX 5090 Laptop GPU, [N/A], 150.00 W, 24463 MiB, 8, 16, 1522 MHz, 14001 MHz posture shipped — enforced/default/max = 150 95 175 W. The enforced limit FLOATS while these arms render; every result file's trace reduces it to min/median/max and THAT is the cap cell. clock lock [by the CARD]: ON-BY-CARD evidence: the card was AT REST for the whole sample (utilization.gpu read 0 % on all 12 reads) and clocks.sm never left the lock's floor: min 1192 MHz, max 1522 MHz, against a declared floor of 1200 MHz and the driver's 25 MHz quantization tolerance (the driver snaps 1200 to 1192 on this board, measured 2026-09-21T23:13Z). Unlocked and at rest this board reads 232 MHz or less (232 MHz at 17:30Z, 180 MHz at 18:11Z), so a resting card pinned at the floor is the lock holding it there. clocks.max.sm reads 3090 MHz, which is this board's own unlocked maximum with OR without the lock and is therefore not part of this verdict. the unit ai-perf.service reads active — NOT EVIDENCE: it is Type=oneshot, so it reads 'active (exited)' for the whole uptime whether or not 'nvidia-smi -rgc' has released the lock since. On 2026-09-21 a receipt that trusted it published 'ran with the lock ACTIVE' beside clocks that proved otherwise. clock lock ON-BY-CARD by the CARD (when the lock is in force this board cannot clock below 1200 MHz and a power posture means something different from what it means on the unlocked desktop cards this leg compares to). The unit state is not evidence; see above. receipt: the clock lock is IN FORCE (ON-BY-CARD), read from the card — which is this posture's declared condition, not a confound it tolerates. the card was AT REST for the whole sample (utilization.gpu read 0 % on all 12 reads) and clocks.sm never left the lock's floor: min 1192 MHz, max 1522 MHz, against a declared floor of 1200 MHz and the driver's 25 MHz quantization tolerance (the driver snaps 1200 to 1192 on this board, measured 2026-09-21T23:13Z). Unlocked and at rest this board reads 232 MHz or less (232 MHz at 17:30Z, 180 MHz at 18:11Z), so a resting card pinned at the floor is the lock holding it there. clocks.max.sm reads 3090 MHz, which is this board's own unlocked maximum with OR without the lock and is therefore not part of this verdict. port band 18190-18199 free; this box's own comfyui.service is inactive -- arm inventory · suffix inventory-shipped · 2026-09-22T00:43:35Z -> results/inventory-inventory-shipped.json -- arm inventory exited 0 · 2026-09-22T00:43:36Z posture [after inventory at shipped]: enforced/default/max = 150 95 175 W -- arm r1 · suffix r1-shipped · 2026-09-22T00:43:36Z card available at 2026-09-22T00:43:36Z · the image service holds 0 MiB · quiet: mean 0.0% <= 5.0% cap 150.0 W per card · UPS load unreadable R1 klein-4b-8st on index 0 (an NVIDIA GeForce RTX 5090 Laptop GPU 24 GB, SOLDERED to this machine — no slot, no partner, and no card that can replace it for a control, soldered (mobile; no socket — link width is traced under load)) r1-r1-shipped-klein-4b-8st-c0-b3-warm[0] 9.618s exec / 3 img = 3.206s per image r1-r1-shipped-klein-4b-8st-c0-b3-timed[0] 3.613s exec / 3 img = 1.204s per image r1-r1-shipped-klein-4b-8st-c0-b3-timed[1] 3.870s exec / 3 img = 1.290s per image R1 klein-4b-32st on index 0 (an NVIDIA GeForce RTX 5090 Laptop GPU 24 GB, SOLDERED to this machine — no slot, no partner, and no card that can replace it for a control, soldered (mobile; no socket — link width is traced under load)) r1-r1-shipped-klein-4b-32st-c0-b3-warm[0] 16.098s exec / 3 img = 5.366s per image r1-r1-shipped-klein-4b-32st-c0-b3-timed[0] 13.849s exec / 3 img = 4.616s per image r1-r1-shipped-klein-4b-32st-c0-b3-timed[1] 14.203s exec / 3 img = 4.734s per image R1 klein-4b-8st-pixel4 on index 0 (an NVIDIA GeForce RTX 5090 Laptop GPU 24 GB, SOLDERED to this machine — no slot, no partner, and no card that can replace it for a control, soldered (mobile; no socket — link width is traced under load)) r1-r1-shipped-klein-4b-8st-pixel4-c0-b3-warm[0] 5.591s exec / 3 img = 1.864s per image r1-r1-shipped-klein-4b-8st-pixel4-c0-b3-timed[0] 3.584s exec / 3 img = 1.195s per image r1-r1-shipped-klein-4b-8st-pixel4-c0-b3-timed[1] 3.858s exec / 3 img = 1.286s per image -> results/r1-r1-shipped.json -- arm r1 exited 0 · 2026-09-22T00:46:08Z posture [between R1 and R4 at shipped]: enforced/default/max = 150 95 175 W -- arm r4 · suffix r4-shipped · 2026-09-22T00:46:08Z card available at 2026-09-22T00:46:08Z · the image service holds 0 MiB · quiet: mean 0.0% <= 5.0% cap 150.0 W per card · UPS load unreadable r4-c0-warm[0] 3.579s exec / 1 img = 3.579s per image r4-c0-warm[1] 1.493s exec / 1 img = 1.493s per image r4-c0-warm[2] 1.474s exec / 1 img = 1.474s per image -> results/r4-r4-shipped.json -- arm r4 exited 0 · 2026-09-22T00:56:32Z posture [render rung shipped close]: enforced/default/max = 150 95 175 W == render rung shipped · done 2026-09-22T00:56:33Z (R1 exit 0, R4 exit 0) == results/r1-shipped.json and results/r4-shipped.json the UPS is absent on this box: no `upsc` on this box: NUT is not installed and nut-server/nut-monitor are inactive (2026-09-21T14:52Z). Whole-box watts are WITHHELD with this reason; board watts are read normally. the fan is unreadable on this board: fan.speed reads [N/A] on this board: the driver reports no fan for it. The fan-drawing cell is a declared non-figure with this reason — arm R4 measures ten minutes of sustained drawing and cannot report a fan curve for it. LEG 2 at shipped exited 0 envelope at the pass's close: 150 95 175 W artifact receipt -> /workshop/bench-laptop-5090-2026-09-21/RUNG-SHIPPED-RECEIPT.txt PASS shipped — COMPLETE posture shipped cap argument shipped written 2026-09-22T00:56:33Z box_class gpu-5090-laptop-24g card GPU-edff232c-7dbf-2bac-07fb-921a7c9eecff (NVIDIA GeForce RTX 5090 Laptop GPU) cap cell A RANGE, never a number — this board refuses -pl, so a floating posture has no cap to name. Per-stage ranges are in /workshop/bench-laptop-5090-2026-09-21/harness-llm/results/*-cap-posture-shipped-*.json and in the 2 Hz trace of every scored result file. Do not type a wattage here. envelope now 150 95 175 (enforced,default,max W) clocks 1590,14001,3090 (sm,mem,max.sm MHz) clock lock ON-BY-CARD, read FROM THE CARD the card was AT REST for the whole sample (utilization.gpu read 0 % on all 12 reads) and clocks.sm never left the lock's floor: min 1192 MHz, max 1192 MHz, against a declared floor of 1200 MHz and the driver's 25 MHz quantization tolerance (the driver snaps 1200 to 1192 on this board, measured 2026-09-21T23:13Z). Unlocked and at rest this board reads 232 MHz or less (232 MHz at 17:30Z, 180 MHz at 18:11Z), so a resting card pinned at the floor is the lock holding it there. clocks.max.sm reads 3090 MHz, which is this board's own unlocked maximum with OR without the lock and is therefore not part of this verdict. (the unit ai-perf.service reads active, which is NOT evidence: it is a oneshot and stays 'active' for the uptime after -rgc.) nvidia-powerd active LEG 1 exit 0 LEG 2 exit 0 PASS exit 0 LEG 1 result files carrying shipped in their name: bc7ab4e5c1bef28e cap-witness-shipped-armA2-kv-axis.csv 843ae2492926e7af cap-witness-shipped-armA3-filled.csv 21c3d51d9295d826 cap-witness-shipped-armA-ladder-gemma4.csv 09329e08d2c38e2b cap-witness-shipped-armD-forced-probe.csv bcc1107cd12317b0 cap-witness-shipped-armD-ladder-mistral.csv 9ae9a6c75a3c1380 cap-witness-shipped-armE-concurrency.csv c44c20648cc5ba4f cap-witness-shipped-card-mapping.csv 63acbb07db72e43a cap-witness-shipped-g2-doorman.csv 7f96a18002868fc8 cap-witness-shipped-g3-vision.csv 8a9c4e8dc018490b cap-witness-shipped-g4-embed.csv 0b6a7b81f4f3272c cap-witness-shipped-idle.csv f1509b7c18560448 gpu-5090-laptop-24g-cap-posture-shipped-armA2-kv-axis.json 962ca219445627a2 gpu-5090-laptop-24g-cap-posture-shipped-armA3-filled.json 2205a4c78ceaeeaa gpu-5090-laptop-24g-cap-posture-shipped-armA-ladder-gemma4.json 3f04599b56d07cdc gpu-5090-laptop-24g-cap-posture-shipped-armD-forced-probe.json 370150a115d718af gpu-5090-laptop-24g-cap-posture-shipped-armD-ladder-mistral.json 1d5983d31f6095a5 gpu-5090-laptop-24g-cap-posture-shipped-armE-concurrency.json 133c986b65a69f7d gpu-5090-laptop-24g-cap-posture-shipped-card-mapping.json aa9ef47cf444fb7f gpu-5090-laptop-24g-cap-posture-shipped-g2-doorman.json 43794f45a0a4b01a gpu-5090-laptop-24g-cap-posture-shipped-g3-vision.json 0d592599177c53c4 gpu-5090-laptop-24g-cap-posture-shipped-g4-embed.json bbe7f041de0eb4a9 gpu-5090-laptop-24g-cap-posture-shipped-idle.json f3619e238c7bd969 gpu-5090-laptop-24g-card-mapping-shipped.json ad015c13397b018b gpu-5090-laptop-24g-clock-lock-ON-BY-CARD-shipped.txt d62b7bd1790950bd gpu-5090-laptop-24g-concurrency-one-shipped.json cfd986ad3f93dbfb gpu-5090-laptop-24g-embed-one-shipped.json f7bf223307ea5515 gpu-5090-laptop-24g-g2-doorman-shipped.json fe5da48be411ceb3 gpu-5090-laptop-24g-g3-one-shipped-g3-vision.json 7029810e98d112ed gpu-5090-laptop-24g-idle-gemma4-shipped.json 49696337859e0a8c gpu-5090-laptop-24g-m4-one-shipped-auto-ladder.json 015dfc6590feefeb gpu-5090-laptop-24g-m4-one-shipped-forced-at131072.json 2166f2aacfcf6845 gpu-5090-laptop-24g-m4-one-shipped-forced-at98304.json 93f20d168508793e gpu-5090-laptop-24g-m5-one-shipped-auto-ladder.json 9906e60bab7fcb30 gpu-5090-laptop-24g-m5-one-shipped-filled.json bd7778a29624e39a gpu-5090-laptop-24g-m5-one-shipped-kvf16-at131072.json ab002d244b37a33e gpu-5090-laptop-24g-m5-one-shipped-kvq4_0-at131072.json 26562c3cb5908120 gpu-5090-laptop-24g-ups-gate-shipped.txt LEG 2 result files carrying shipped in their name: 7188d785ccc4face inventory-inventory-shipped.json 5df858e09f9549dd r1-r1-shipped.json 68d85142f20b7fb3 r4-r4-shipped.json stated absences at this pass: UPS no `upsc` on this box: NUT is not installed, nut-server and nut-monitor are inactive, the battery exposes no power_now/energy_now, and /sys/class/powercap/*/energy_uj is root-only (all read 2026-09-21T14:52Z). Whole-box watts are WITHHELD with this reason. Board watts — which is what every table in this ladder prints — are read normally at 2 Hz. fan fan.speed reads [N/A] on this board (2026-09-21T14:52Z): the driver reports no fan for it. The fan column is a declared non-figure with this reason, not a gap to be filled later. memory temp temperature.memory reads [N/A], as on every board of this ladder. The core-temperature stop and the driver's own thermal-slowdown reasons are the thermal instrument. arm T the training arm's preprocessed tensor set lives on benchbox, which was unreachable from 2026-09-21T15:03Z. Arm T is a SECOND TRAIN and legs 1 and 2 never depended on it. pass shipped finished rc=0 at 2026-09-22T00:56:33Z memory guard restored: memshed.timer is active (it was active); MEMSHED_DRY_RUN unset free receipt: journalctl --user -u memshed --since '2026-09-22T00:10:35Z' | grep SHEDDING — every line there is a shed this bench WOULD have taken, with its signature.