{
 "schema_version": "1.0",
 "page_kind": "article",
 "pack_kind": "generate",
 "slug": "a-laptop-asked-the-desktops-questions",
 "title": "A laptop, asked the desktop's questions",
 "dek": "One RTX 5090 Laptop part, 24 GB, soldered into a gaming laptop, put through the same language and render bench this shelf ran on a 3090, a 3090 Ti and a pair of 3080s. It was measured three times: held to its 95-watt default; with the maker's Dynamic Boost left to float the budget, which it held between 140 and 150 W; and the way the machine boots, boost and a clock lock together, which changed nothing under load and four watts at idle. The finding is a ratio, not a deficit: a part allowed 95 watts, where the desktop cards beside it are allowed 350 to 450, does most of the work on the model this workshop runs, loses a third of it on a dense one, and gets the third back the moment it is allowed 150. Every watt on the page is the card's own draw and not the wall's, the fan was held at maximum by hand, and nine things the page cannot say are written down.",
 "published": "2026-09-22",
 "series": [
  "bench"
 ],
 "licence": "CC BY 4.0",
 "status": "pending-judge",
 "pack_note": "no judge is seated, so every chip is held",
 "notice": "items with published:false are held and are not this page's published words",
 "source": {
  "url": "https://research.strata2signal.com/a-laptop-asked-the-desktops-questions/",
  "md_url": "https://research.strata2signal.com/a-laptop-asked-the-desktops-questions/index.md",
  "md_sha": "bc68457271dcd6641214cd1e6e1894b4794b62d714415f7ae2630016402a1c9c",
  "html_sha": "35b229e5b7e311ce28ff339648c8bc61a79be9995a21a169eba3011fcd05a84a",
  "anchors_sha": "6a34bbcb536d172f07550eeaa8f2fc18439bb35b0d49c2bcf6ba60308534eea3"
 },
 "short": {
  "paragraph": "yes, a laptop GPU is good for this. Held to 95 W, this part wrote at 118.4 tokens a second on the mixture-of-experts model this workshop actually runs, where a 3090 at 350 W wrote at 136.6 - 87 per cent of the speed on a third of the power the 3090 drew - and held the same 131,072-token window. On the dense 24-billion-parameter model it lost outright: 28.8 tokens a second against 46.9. In ten minutes of continuous drawing it made 294 images to the 3090's 311, peaking at 58 °C. With Dynamic Boost running, which is how a laptop like this arrives, the board's own limit read 140 to 150 W across the pass (n=374 at five-second intervals; 150 W in 335 of them) and never approached its 175 W maximum; on that budget the part wrote at 155.3 tokens a second on the mixture model - past the 3090 and past a 3090 Ti at its own 450 W limit (147.6) - brought the dense model to 44.9 against the 3090's 46.9, and drew 378 images in ten minutes to the 3090 Ti's 348 at 450 W, at between a third and two-fifths of the energy an image. A third pass put the boot-time clock lock this laptop normally runs with back on, and nothing moved under load; the lock costs about four watts at idle. The desktop comparator in the headlines is a 3090 at its own limit; a 3090 Ti is in every table, and it is ahead at both of its rungs on the dense model.",
  "from_dek": false,
  "counts": {
   "words": 7606,
   "minutes": 35,
   "tables": 11,
   "kit": false
  },
  "bullets": [
   {
    "text": "an operator found that the board's maximum draw was 175 W",
    "figure": "175 W",
    "cite": "the-board-and-the-two-things-it-will-not-do",
    "quote": "The bench was planned as a two-rung ladder, 175 W then 95 W, the way every desktop board on this shelf was measured.",
    "span": [
     7634,
     7640
    ]
   },
   {
    "text": "the core temperature reached a peak of 71",
    "figure": "71",
    "cite": "what-the-extra-watts-bought",
    "quote": "The board drew 145.9 W median through the ten minutes and its core plateaued at 68 °C, peaking at 71 - thirteen degrees hotter at the peak than at 95 W, and three under a 3090's 74 °C at 350 W, which drew 311 images in its ten minutes against this board's 378.",
    "span": [
     17416,
     17419
    ]
   },
   {
    "text": "the mixture model at a 131,072-token window wrote at 118.424",
    "figure": "118.424",
    "cite": "the-tables",
    "quote": "This board beside the desktop rows - gemma4:26b at a 131,072-token window Configuration Power setting Tokens/s First token, ms Board W (the three runs' means) J per 1,000 tokens Tokens/s per 100 W This board, fixed 95 W 118.424 372.49 79.64 (76.95 · 79.64 · 94.22) 672.50 148.7 This board, dynamic boost 140-150 W (n=374; 150 W in 335) 155.251 372.92 122.57 (98.58 · 146.10 · 122.57) 787.76 126.7 This board, its boot state 150 W, flat (n=385) 155.288 371.01 116.12 (116.12 · 115.76 · 146.42) 747.97 133.7 A GeForce RTX 3090 24 GB 350 W, its own limit 136.618 552.75 240.49 (275.82 · 240.49 · 222.39) 1,758.13 56.8 A GeForce RTX 3090 Ti 24 GB 350 W, the shared-cap rung 147.823 541.49 282.64 (284.40 · 254.06 · 282.64) 1,909.96 52.3 A GeForce RTX 3090 Ti 24 GB 450 W, its own limit 147.560 539.43 283.01 (229.38 · 302.00 · 283.01) 1,915.13 52.1 Two GeForce RTX 3080 10 GB, the model split across them, every layer forced onto the cards 300 W each, as the file records it 115.366 571.40 346, both boards, as the pair's own page prints it 2,994, as that page prints it 33.3 The three runs' means are the spread the median hides, and on this model it is wide on every board: a two-second run gives the 2 Hz trace three to seven busy samples.",
    "span": [
     22937,
     22947
    ]
   }
  ]
 },
 "sections": [
  {
   "id": "six-words-this-page-leans-on",
   "heading": "Six words this page leans on",
   "level": 2,
   "span": [
    2902,
    5264
   ],
   "chunks": [
    [
     2902,
     5264
    ]
   ],
   "chars": 2362,
   "digest": "this section defines key terms used throughout the page, including mixture of experts (MoE), dense models, the power budget, and Dynamic Boost. it also explains the context window, KV cache, the planner, first token, board power, and an arm.",
   "digest_skipped": null
  },
  {
   "id": "if-you-own-one-of-these-in-plain-words",
   "heading": "If you own one of these, in plain words",
   "level": 2,
   "span": [
    5264,
    6757
   ],
   "chunks": [
    [
     5264,
     6757
    ]
   ],
   "chars": 1493,
   "digest": "for the mixture model, a laptop 5090 on a 95-watt budget writes at 118.4 tokens a second. letting Dynamic Boost run raised the board's draw by about 43 W and put the mixture model ahead of both desktop cards. the dense model writes at 28.8 tokens a second against 46.9.",
   "digest_skipped": null
  },
  {
   "id": "the-board-and-the-two-things-it-will-not-do",
   "heading": "The board, and the two things it will not do",
   "level": 2,
   "span": [
    6757,
    10498
   ],
   "chunks": [
    [
     6757,
     10498
    ]
   ],
   "chars": 3741,
   "digest": "the NVIDIA GeForce RTX 5090 Laptop GPU has 24 GB of memory. it will not take a power cap, as the first command an operator ran showed it is not supported. it also will not report its fan, as the driver answers [N/A].",
   "digest_skipped": null
  },
  {
   "id": "five-things-the-95-w-pass-found",
   "heading": "Five things the 95 W pass found",
   "level": 2,
   "span": [
    10498,
    13922
   ],
   "chunks": [
    [
     10498,
     13922
    ]
   ],
   "chars": 3424,
   "digest": "on the mixture model, a quarter of the power allowed buys most of the speed. the dense model is where the mobile part actually loses. the context ceiling is identical to the 3090's. the planner left the card on the table, and a 16-bit KV cache is faster than the 8-bit default.",
   "digest_skipped": null
  },
  {
   "id": "what-the-extra-watts-bought",
   "heading": "What the extra watts bought",
   "level": 2,
   "span": [
    13922,
    18582
   ],
   "chunks": [
    [
     13922,
     18582
    ]
   ],
   "chars": 4660,
   "digest": "the mixture model's speed stopped following the power. the dense model is still power-bound, and every extra watt the second posture allowed went into speed. the renders went from slightly slower than the desktop cards to faster than both. the first token did not pay for the extra watts.",
   "digest_skipped": null
  },
  {
   "id": "the-way-it-boots-measured",
   "heading": "The way it boots, measured",
   "level": 2,
   "span": [
    18582,
    20900
   ],
   "chunks": [
    [
     18582,
     20900
    ]
   ],
   "chars": 2318,
   "digest": "the third pass ran the machine with the boot-time clock lock in force. under load the lock changed nothing, but at idle the card cannot clock down below 1,200 MHz, so it drew 11.57 W empty where the unlocked board drew 7.14 W. its idle link reads PCIe x8 at generation 2.",
   "digest_skipped": null
  },
  {
   "id": "the-tables",
   "heading": "The tables",
   "level": 2,
   "span": [
    20900,
    35885
   ],
   "chunks": [
    [
     20900,
     35885
    ]
   ],
   "chars": 14985,
   "digest": "the tables provide median values of three scored runs. they compare this board against desktop rows for mixture and dense models, show the largest window each configuration held whole, detail ten minutes of continuous drawing, and list various scored arms and recipes.",
   "digest_skipped": null
  },
  {
   "id": "what-this-page-cannot-say",
   "heading": "What this page cannot say",
   "level": 2,
   "span": [
    35885,
    39482
   ],
   "chunks": [
    [
     35885,
     39482
    ]
   ],
   "chars": 3597,
   "digest": "this page does not cover wall watts, noise, battery, or heat where hands are. it does not report an ordinary fan curve, a control seat, a memory-bandwidth ceiling, or how much the extra watts would stop paying. it also does not include a desktop 5090 or a price.",
   "digest_skipped": null
  },
  {
   "id": "how-it-was-measured",
   "heading": "How it was measured",
   "level": 2,
   "span": [
    39482,
    43048
   ],
   "chunks": [
    [
     39482,
     43048
    ]
   ],
   "chars": 3566,
   "digest": "the language bank ran against ollama 0.32.13 and the render arms ran against ComfyUI 0.21.1. every rung is three scored runs of one frozen prompt. board power, clocks, temperature and the enforced power limit were sampled at 2 Hz through every scored run.",
   "digest_skipped": null
  },
  {
   "id": "how-to-check-our-work",
   "heading": "How to check our work",
   "level": 2,
   "span": [
    43048,
    43963
   ],
   "chunks": [
    [
     43048,
     43963
    ]
   ],
   "chars": 915,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "the-rest-of-the-seminar",
   "heading": "The rest of the seminar",
   "level": 2,
   "span": [
    43963,
    45031
   ],
   "chunks": [
    [
     43963,
     45031
    ]
   ],
   "chars": 1068,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "who-ran-this-and-thanks",
   "heading": "Who ran this, and thanks",
   "level": 2,
   "span": [
    45031,
    46446
   ],
   "chunks": [
    [
     45031,
     46446
    ]
   ],
   "chars": 1415,
   "digest": null,
   "digest_skipped": "credits"
  }
 ],
 "chips": [
  {
   "id": "c-2bd7c6b3",
   "q": "how do mixture and dense models differ in weight usage?",
   "a": "A mixture of experts model consults only a slice of its weights for each token, allowing it to stay fast on modest power. In contrast, a dense model reads every one of its weights for every token, which is where a power budget becomes more apparent.",
   "cites": [
    "six-words-this-page-leans-on"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-6e250df2",
   "q": "what happens to performance when dynamic boost is active?",
   "a": "Enabling dynamic boost on the mixture model raised the board's draw by about 43 W at the headline window, closing the performance gap to 4 per cent and putting the mixture model ahead of both desktop cards.",
   "cites": [
    "if-you-own-one-of-these-in-plain-words"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-38c131c4",
   "q": "why can't the power limit be manually adjusted on this board?",
   "a": "The board will not take a power cap. An attempt to change the power management limit resulted in a message stating that changing the limit is not supported for the GPU.",
   "cites": [
    "the-board-and-the-two-things-it-will-not-do"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-808bd4ab",
   "q": "how much speed is gained by using more power on the mixture model?",
   "a": "A quarter of the power allowed buys most of the speed on the mixture model. At a 131,072-token window, the model achieved 118.424 tokens a second at 95 W, compared to 136.618 at 240.49 W on a 3090.",
   "cites": [
    "five-things-the-95-w-pass-found"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-11bf956a",
   "q": "how does extra wattage affect the dense model's speed?",
   "a": "The dense model remains power-bound, and every extra watt allowed by the second posture went into speed. It decoded at 44.8 to 44.9 tokens a second, which was 56 to 68 per cent faster than at 95 W.",
   "cites": [
    "what-the-extra-watts-bought"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-7bd9f4c8",
   "q": "how does a clock lock affect idle power draw?",
   "a": "The clock lock increases idle power draw because the card cannot clock down below 1,200 MHz. At idle, the card drew 11.57 W with the lock, compared to 7.14 W when the board was unlocked.",
   "cites": [
    "the-way-it-boots-measured"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-8f0a587a",
   "q": "what information is missing regarding environmental factors?",
   "a": "The article does not provide data on wall watts, noise, battery performance, or how the machine behaves on an ordinary fan curve. It also does not report temperatures at the keyboard, palm rest, or underside.",
   "cites": [
    "what-this-page-cannot-say"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-69529d9e",
   "q": "how were the language and render results verified?",
   "a": "Language results were run against a separate copy of ollama with keep-alive at zero, while render arms used a separate ComfyUI checkout. Every rung consisted of three scored runs of one frozen prompt.",
   "cites": [
    "how-it-was-measured"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  }
 ],
 "related": [
  {
   "slug": "two-used-3080s-priced",
   "why": "shares ground with § Five things the 95 W pass found · § How fast does each one write?"
  },
  {
   "slug": "one-3090-against-one-3090-ti",
   "why": "shares ground with § What this page cannot say · § What this page does not measure"
  },
  {
   "slug": "a-rig-your-friend-already-owns",
   "why": "shares ground with § The tables · § The table that answers the question"
  }
 ],
 "thanks": "## Who ran this, and thanks {#who-ran-this-and-thanks}\n\nThe board is an **NVIDIA GeForce RTX 5090 Laptop GPU 24 GB** in an **MSI Raider 16 Max HX**, and **NVIDIA** gets the credit for the board, for the driver every watt, clock and temperature on this page was read from, and for Dynamic Boost, the mechanism the second pass measured; **MSI** for the chassis that carried it and the max-fan key an operator leaned on. The language work stands on **ollama** and, beneath it, **llama.cpp** and **NVIDIA's CUDA**; the pictures on **ComfyUI**. The models are Google's **Gemma 4**, **Mistral Small 3.2** from Mistral AI, **MiniCPM-V 4.5** from OpenBMB, **nomic-embed-text** from Nomic, and Black Forest Labs' **FLUX.2 Klein 4B** with its decoder, which drew every picture timed here, and **Qwen** from Alibaba Cloud as the text encoder in those recipes. None of them owed us anything, and every one of them ran on a laptop.\n\nA small human team owns the laptop, set the rules and the refusals before the runs, held the fan and signed the numbers; a fleet of AI agents ran the harness and did the arithmetic under that team's rulings. Thanks to the people asking, on every forum this workshop reads, whether a laptop GPU is any good for this — that question is the whole of this page.\n\n*Measured on 2026-09-21 and 2026-09-22 by one instrument on one laptop, with the fan held by hand and the wattage beside every figure.*",
 "kit": null,
 "seat_class": {
  "writer": "gemma-class",
  "judge": null,
  "writer_runtime": "vllm"
 },
 "bench": {
  "bullets_written": 3,
  "bullets_kept": 3,
  "digests_written": 9,
  "digests_kept": 9,
  "chips_written": 8,
  "chips_kept": 8,
  "chips_grounded": 0,
  "chips_published": 0,
  "dropped_by": {
   "new_noun": 0,
   "figure": 0,
   "length": 0,
   "cite": 0,
   "judge": 0,
   "quote": 0,
   "directive": 0,
   "redaction": 0,
   "profanity": 0,
   "identifier": 0,
   "bare_figure": 0,
   "house_voice": 0
  },
  "new_noun_tokens_checked": [
   "0.21.1",
   "0.32.13",
   "072-token",
   "1",
   "11.57",
   "118.4",
   "118.424",
   "131",
   "136.618",
   "16-bit",
   "175",
   "2",
   "200",
   "24",
   "240.49",
   "28.8",
   "3090",
   "3090's",
   "4",
   "43",
   "44.8",
   "44.9",
   "46.9",
   "5090",
   "56",
   "68",
   "7.14",
   "71",
   "8-bit",
   "95",
   "95-watt",
   "Boost",
   "ComfyUI",
   "Dynamic",
   "GB",
   "GPU",
   "GeForce",
   "Hz",
   "KV",
   "Laptop",
   "MHz",
   "MoE",
   "N/A",
   "NVIDIA",
   "PCIe",
   "RTX",
   "W",
   "x8"
  ],
  "new_noun_tokens_withheld": 0,
  "figure_definition": "v3",
  "figure_boundary": "guarded",
  "drops": []
 },
 "generated_utc": "2026-09-22T02:11:48Z",
 "meta_refreshed_utc": "2026-09-22T03:12:07Z"
}
