# PREREG — the power ladder on the 96 GB card: what extra watts buy the image lane

**Registered 2026-08-27, before the first scored render of any arm.** Every timestamp in
this document and in every artefact it governs is UTC. Clock receipt: the harness machine and the
card's host both answered `date -u` = `2026-08-27T16:04:25Z` at the same instant, so the
harness clock and the card clock are the same clock to the second.

*Note on the dating of the governing amendments: Amendments 1–3 arrived carrying
approximate stamps (`~17:5xZ`, `16:4xZ`) that sit ahead of the measured UTC on both boxes.
This document uses only stamps it read from `date -u`. The amendments' CONTENT is binding
as written; their stamps are recorded as approximate.*

**The question (the operator):** how much performance do we get from extra power for images —
"I think it might be more noticeable [than serving]". This registers the design; it
answers nothing.

---

## 0. What is registered here, and what is not

Registered before any scored call: the arms and their order, the exact cell set (a file,
`sample.json`, whose sha256 is below), the measure, the gates, the drift/REDO rule, the
UNMEASURABLE rule, the VOID rule, and the exclusions. Not registered, because it is not
yet known: any number.

| artefact | sha256 |
|---|---|
| `sample.json` | `cc6a571cb422f06e0085de479cdff82d420adf26e030e4695b575001790e6c60` |
| `graphs/*.json` | *`graphs/SHA256SUMS`, sealed before the first render* |
| `the power bench's sealed prompt file (p512, sha-pinned above)` | `90eedd0c53f9554ae3837674504fcb7090013d9a432a653352183c0f25a7ce5c` (verified against its own `SHA256SUMS` at load; the probe refuses to run on a mismatch) |

---

## 1. Arms and their order

**500 W → 450 W → 600 W (last).** Walked in that order per Amendment 1: the operator had
already set 500 W before rung 0 finished building, and re-walking up to 600 first would
have spent an extra cap change for nothing.

The 600 W arm runs LAST and is the **control**: every published ratio in RESULTS.md has
the fresh, same-day 600 W arm as its denominator (Amendment 3), never a sealed number.

The cap is the **operator's hand**. This lane never sets it. Every cap reading in this
bench comes from a read-only `nvidia-smi --query-gpu=power.limit` over ssh. The runner
refuses to start an arm whose measured cap is not the arm's nominal cap
(`--expect-limit-w`), and reads the cap again at the arm's close.

**The VOID rule (Amendment 1).** Persistence mode is **OFF** on the card (measured:
`persistence_mode = Disabled`, `power.default_limit = 600.00 W`, `power.min_limit =
150.00 W`), so the limit can revert on a driver idle-unload. `power.limit` is therefore a
**sampled column on every power sample**, not a one-off check. Any cell whose render
window is covered by samples showing a limit other than its arm's nominal cap is **VOID
for that rung**, named in RESULTS.md. A rung is not silently rescued by a re-run: the void
cells are stated, then re-run, and both facts are printed.

---

## 2. The render set — the full sealed intersection, no sampling

Per Amendment 2 (the operator: *"run this full set on 450, 500, and 600w — to get a clean full
set"*), there is **no sampling**. The set is every sealed cell whose model the card's
ComfyUI holds, at its **sealed seeds** and **verbatim prompts**, **identical at every
rung**. The subjects are real cove content — the locked bake-off law; nothing is invented.

The card's ComfyUI offers exactly three UNETs (`flux-2-klein-4b`, `flux1-dev`,
`flux1-schnell`) and exactly two checkpoints (`albedobaseXL_v21`, `sd_xl_base_1.0`),
verified read-only against `/object_info` on 2026-08-27. All five run. Every node class
and every weight file the five graphs name is present at the endpoint (verified,
`validate.py --online`).

### 2.1 How the cells were chosen — a law, not a judgement

**Tier A** — the three lanes with a sealed 600 W history. The selection law is the
**rematch bench's own `pick_cells()`**, imported and re-run rather than re-implemented:
res-ladder (2 seeded scenes × 3 widths × 3 seeds), steps-ladder (4 spread points),
subjects-768 (3 seeded cast subjects × 3 seeds), with the standard kits at all three sizes
for the albedobase arm — a branch that driver already carries. **Receipt:** re-running it
reproduces the rematch manifest's 31 scored klein rows exactly. That is what licenses
extending the same law to the two control arms the rematch dropped for **licence** reasons
(`ctrl-flux1-dev`, `ctrl-albedobase`) — both fully sealed in the 08-23 bench.

**Tier B** — the two shipped lanes the card holds that have **no sealed history anywhere**
(`flux1-schnell`, `sd_xl_base_1.0`). They are not substituted for anything and they do not
borrow anyone's numbers: they inherit their sealed sibling's **cell definitions** (subject,
verbatim prompt, width, sealed seed) and run at their **own shipped locked** step/cfg
values. The steps-ladder axis is dropped for them because those step counts encode the
sibling's posture, not theirs. Every Tier B row carries `baseline_seconds: null` and is
published in its own block, **never merged into a Tier A table**.

Any lane whose intersection were empty would be **absent and stated, never substituted**.
None is empty; both Tier B lanes are stated as baseline-less rather than dropped, and this
paragraph is the ruling that put them in. *If the orchestrator rules otherwise, Tier B is
removed by one flag (`--lanes`) and the arm re-runs; the ruling belongs before the first
scored render, not after.*

### 2.2 The set as registered

**139 scored renders per rung. 175 submissions per rung including the unscored ones.**
Identical at 500 W, 450 W and 600 W.

| lane | weights | tier | cells | widths (n) | steps | sealed-baseline rows | shipped posture |
|---|---|---|---|---|---|---|---|
| klein-4b | `flux-2-klein-4b` | A | 31 | 512(6) 768(19) 1024(6) | 2, 4, 6 | 31 | the limner; vendor-default 4 steps, cfg 1.0 (distilled) |
| flux1-dev | `flux1-dev` | A | 30 | 512(6) 768(18) 1024(6) | 16, 24, 32, 40 | 30 | the HERO incumbent; 32 steps, guidance 3.5 |
| flux1-schnell | `flux1-schnell` | B | 24 | 512(6) 768(12) 1024(6) | 4 | 0 | the shipped fast lane; 4 steps, no guidance node |
| sdxl-albedobase | `albedobaseXL_v21` | A | 27 | 512(9) 768(9) 1024(9) | 32 | 27 | the painted-real incumbent; dpmpp_2m/karras, cfg 6.0 |
| sdxl-base | `sd_xl_base_1.0` | B | 27 | 512(9) 768(9) 1024(9) | 12 | 0 | the shipped LOCKED sketch tier; dpmpp_2m/karras, cfg 6.5 |

Six real subjects carry the set: two coastal scenes (Sorrowmoor Cove — The First Run;
Cleave Headland — The Point), two NPCs (Halar Kestrel, Salla Merrow), one avatar (the
newcomer), one foe (the Drowned-Hand). Seeds are the sealed three-per-timing-group
families (`Nxxxx`, `7Nxxxx`, `9Nxxxx`) exactly as sealed.

---

## 3. The measure

**Render seconds are ComfyUI's OWN `execution_start` → `execution_success` timestamps,
read from `/history`.** Never a wall clock on the harness side. The harness clock is
recorded per row as context and is never the measure. This is the harness law and the
method the 08-23 bench and the contention bench used, so the numbers are of a kind.

**One unscored warm render per lane per rung**, at the sealed warm seed `777333`, before
that lane's first scored cell. Loads are not the subject: the card cannot hold every
lane's weights at once beside the pinned serving seats, so ComfyUI will swap, and
ComfyUI's execution window **includes** a weight load. The arm therefore runs
**model-major** — each lane loads once, during its warm render, and stays resident for its
own block. Warm rows carry `warm: true` and are excluded from every table.

**The cache trap is guarded three ways.** Re-submitting a byte-identical graph is a full
ComfyUI cache hit that returns in milliseconds (the 08-23 driver's own note; render.py
calls it the orphan-resume trap). (1) The client **refuses** to submit a graph whose
fingerprint equals the previous submission's. (2) The submission order is proved
collision-free before the run: 175 submissions, 144 distinct graphs, **no two adjacent
submissions share a fingerprint** (`validate.py`). (3) Every row records the saved
filename and ComfyUI's own `execution_cached` node count, so a cache hit is visible in the
data and not only in the guard.

**Failure handling.** Two attempts per render, six seconds apart (the rematch's rule). A
render that fails twice is written as a row with `comfy_exec_seconds: null` and its
error — an empty cell recorded honestly, never omitted.

---

## 4. Power, clocks and the throttle flag

A read-only `nvidia-smi` stream from the harness machine for the whole of each arm:

```
power.draw, power.limit, clocks.sm, clocks.mem, utilization.gpu, temperature.gpu, memory.used
--format=csv,noheader,nounits -lms 500
```

one file per arm (`results/power-<arm>.csv`), **UTC-stamped by the harness clock line by
line**. The remote command carries its own `timeout` ceiling so a dropped local process can
never leave a polling loop alive on a live box. `nvidia-smi -q -d PERFORMANCE` snapshots at
each arm's open and close (`results/perf-<arm>.txt`).

**The armP lesson stands: the FLAG is evidence, the counter is not.** The `SW Power Cap`
clocks-event-reason flag is the throttle evidence. `Clocks Event Reasons Counters` are
recorded verbatim and are **never** used as a measure.

**Sampling-resolution caveat, registered now rather than discovered later.** The interval
is 500 ms per the brief. The sealed 600 W medians put the klein lane at 0.26–0.96 s a
render, so a 512 px klein render will be covered by **one sample or none**. Per-render
power attribution for the klein lane is therefore **not claimed**; the klein lane's power
column is reported at lane granularity (the samples spanning that lane's block), and the
per-render power column is claimed only for lanes whose sealed seconds exceed 2 s
(flux1-dev, sdxl-albedobase). Render SECONDS are unaffected — they come from ComfyUI, not
from the sampler.

---

## 5. The LLM probe — the serving lane's reading of the same cap

Per rung, per seat: **one unscored warm call, then 6 timed calls**.

- Seats: `gemma4:26b` on port the-answerer-seat, `mistral-small3.2:24b` on port the-classifier-seat, on the card's
  host, addressed **directly, never through the proxy**.
- **`num_ctx` = 32768 on every call** — the size both seats are forever-pinned at
  (verified read-only via `/api/ps`: both resident, `context_length` 32768). A mismatch
  forces a ~5.3 s reload; `load_duration` is recorded per call as the receipt.
- **`keep_alive: -1` on every call.** Omitting it would demote a production pin to the
  5-minute default and the seat would idle-unload later. The probe leaves the seats exactly
  as it found them.
- **Only those two model names are ever sent to those ports.** A third name at either port
  can evict a forever-pinned production seat. Residency is verified read-only before each
  rung's scored calls and the reading is written into the rows.
- `think:false`; `num_predict` 128; `temperature` 0; `seed` 20260827; `stream:false`.
- Prompt: the **frozen p512** from the FP8 bench — a real rulesage ruling body, house
  content, sha256 above, verified at load.
- **Measure: decode tok/s = `eval_count` / `eval_duration`**, from the response's own
  counters. Never a wall clock. Prefill tok/s recorded the same way, alongside.

---

## 6. Gates, pre-registered

1. **The headline** is the per-model render-seconds ratio **500/600** and **450/600**,
   with the drift column beside it. Denominator: the **fresh same-day 600 W arm**
   (Amendment 3). **Published per model, never averaged across models.** Tier B publishes
   in its own block.
2. **UNMEASURABLE.** More than **two** failed renders on a lane at a rung ⇒ that lane's row
   is **UNMEASURABLE at that rung**, printed as UNMEASURABLE with the failure count, never
   backfilled and never quietly dropped.
3. **VOID.** §1's rule. Cells covered by samples at the wrong cap are void for that rung
   and named.
4. **The drift / REDO gate is INTRA-ARM** (Amendment 3, superseding Amendment 1's
   sealed-baseline clause). Two instruments, both inside each arm:
   - **Sentinels** — two cells per lane (deterministically: the lane's widest width, one
     per subject, the clean high seed of each timing group), run at that lane's **open**
     and again at its **close**. Gives a per-model shift over the lane's own block.
   - **The arm-close re-check** — the arm's **first 10 scored cells**, same seeds, re-run at
     the very end of the arm behind an unscored warm render.

   **A per-model median shift greater than 5 % on either instrument means the ladder is a
   REDO, not a result.** Stated as a redo; no partial rescue.
5. **Self-refutation is honoured.** A rung that fails its own gate gets no threshold table
   and no invented numbers.

## 6.1 What is explicitly NOT a gate

**Sealed-versus-fresh 600 W is a cross-environment OBSERVATION, not a redo gate**
(Amendment 3). Published as its own per-model delta, with its confounds named, and **a
delta of ~0 is as much a finding as a large one.** The confounds, all measured or read
this morning and all stated up front:

- the card sat in a **different host** for the sealed sitting (the 08-23 bench ran against
  a different ComfyUI instance on a different machine);
- the card has since moved onto a **sine-wave UPS** — a different power environment;
- the **ComfyUI version differs** between the sittings (sealed: 0.33.3; today: 0.21.1);
- the sealed sitting did not share the card with two forever-pinned serving seats, which
  today hold ≈43 GB of the card's 96 GB resident throughout.

Because of these, the sealed numbers are never a denominator and never a redo trigger.

---

## 7. Exclusions and hygiene

- **Foreign work is excluded by allowlist, not by filtering.** Only prompt ids this
  harness submitted are ever read from `/history`. The live cove shares this card; the
  queue is inspected before every submission and any foreign prompt id seen is recorded
  with the time it was first seen (`results/arm-<arm>.json`), so contention is visible in
  the data rather than assumed absent.
- **Read-only everywhere else.** ssh to the card's host is nvidia-smi query forms only —
  never `-pl`, never sudo. No writes to any live service. The LLM probe's only side effect
  is the one it must have: preserving the seats' existing pin.
- **`results/` is written publish-safe by construction** — no host names, no addresses.
  The harness SOURCE carries the private-network address and the ssh alias; **anything kit-bound
  must be sanitized at the pen** before it leaves.
- Long runs go in their own systemd-run unit; logs under
  `~/projects/sb-lane-logs/bench-cap-ladder/`.

---

## 8. Outputs

`~/rs-capladder-0827/` — this file; `sample.json` (the registered cell set);
`graphs/` + `graphs/SHA256SUMS`; `results/rows.jsonl` (one row per render: arm, role,
lane, tier, cell id, weights, seed, resolution, steps, ComfyUI execution seconds and its
own start/end ms, cached-node count, output filename, prompt id, harness UTC start/end,
queue state at submit, sealed baseline where one exists); `results/power-<arm>.csv`;
`results/perf-<arm>.txt`; `results/llm-rows.jsonl`; `results/arm-<arm>.json` (cap at open
and close, ComfyUI version, VRAM, failures by lane, foreign ids seen); and `RESULTS.md`
with the per-model tables.

---

## 9. Rung-0 receipts (read-only, before any scored render)

| receipt | value | when (UTC) |
|---|---|---|
| cap | 600.00 W → **500.00 W** (the operator's paste landed between two reads) | 16:04:11Z → 16:04:25Z |
| persistence mode | **Disabled** (the VOID rule's reason) | 16:12:54Z |
| default / min / max limit | 600.00 / 150.00 / 600.00 W | 16:12:54Z |
| card idle | 28.5 W, 180 MHz SM, 29 °C | 16:12:5xZ |
| power sampler | 10 samples in 5 s at 500 ms, every line carrying `power_limit = 500.00` | 16:12:54–59Z |
| throttle flag | `SW Power Cap: Not Active` (card idle) | 16:13:06Z |
| endpoint | ComfyUI 0.21.1, Python 3.14.4, 101,971,001,344 B VRAM total, 34,738,955,700 B free | 16:13:xxZ |
| queue | 0 running, 0 pending, 0 foreign | 16:13:xxZ |
| seats | both resident, `context_length` 32768, both pinned (`expires_at` 2318) | 16:13:xxZ |
| graph check | 5 graphs, all node classes and all 5 weight files present at the endpoint | 16:1xZ |
| order check | 175 submissions, 144 distinct graphs, **0 adjacent fingerprint collisions** | 16:1xZ |

---

*Registered by the cap-ladder bench lane, 2026-08-27, before the first scored render.
Amendments 1, 2 and 3 are folded into the body above; where Amendment 3 conflicts with
Amendment 1 (the redo gate), Amendment 3 governs, as it says.*

---

## AMENDMENT 4 — the drift gate resolved to ONE instrument
**Stamped 2026-08-27T16:34:45Z, after the 500 W arm completed and BEFORE the 450 arm's first scored call.**
Ruled by the orchestrator on the 500 W arm's own measured trace. This amendment changes
NO scored cell: the 139 scored cells and the 5 warm cells are byte-identical to the set
registered above (proved by hash — see §A4.4).

### A4.1 The gate is the n=10 paired arm-close re-check. Only that.
Amendment 3's registered REDO instrument was always *"at the END of the arm, re-run its
FIRST 10 CELLS (same seeds); >5 % per-model median shift = REDO."* That instrument stands
unchanged as **THE gate at every arm**. On the 500 W arm it reads **+1.4 % and PASSES**.

### A4.2 The per-lane sentinel bracket is DEMOTED to a reported diagnostic
The lane-open → lane-close sentinel bracket was an *additional* instrument this lane added
at rung 0. The 500 W arm proved it **confounded by construction**: a lane's sentinel-OPEN
is structurally the coldest and least power-limited render of that lane's block. Measured
cause, from the arm's own power CSV (20 s buckets, busy samples only):

| minute | lane | temp | draw | clocks.sm |
|---|---|---|---|---|
| 16:19 | klein | 47 °C | 358.9 W | 2625 MHz |
| 16:20 | flux1-dev opens | 60 °C | **499.4 W — at the cap** | 2392 MHz |
| 16:22 | flux1-dev | 76 °C | 499.9 W | **1927 MHz** |
| 16:23 | flux1-dev closes | 77 °C | 499.7 W | 2010 MHz |

The card settles from 47 °C to 77 °C across the arm; under a binding cap its SM clock is
pulled from 2392 MHz down to 1927 MHz as it heats. The open→close bracket therefore spans
the settling transient **by construction**, and does so identically at every rung. The
**scored** window is stable through it: first half vs second half of every
(lane, width, steps) group with n ≥ 4 shifts by a median of +1.5 % / +0.7 % / -0.2 % /
-0.2 % / -3.0 % for the five lanes, flux1-dev's own max being +1.8 %.

**RULED:** sentinels are reported at every rung as a DIAGNOSTIC and **never trigger a
REDO**. Both the gate number and the diagnostic number are always printed.

### A4.3 Sentinels widen to n = 6 per lane (diagnostic only)
n = 2 was measured too thin to read a 5 % shift: klein's two cells split **+13.4 % and
-8.7 %** — quantisation noise at sub-second renders. Widened to **6 per lane**, drawn
deterministically from the lane's longest-running width, round-robined across subjects for
cast coverage, preferring the high seed of each timing group. Cost: 60 bracket renders per
arm instead of 20.

### A4.4 The statistic: PAIRED ratios are the verdict; ratio-of-medians rides beside
A paired instrument is read with the **median of PAIRED ratios**. Ratio-of-medians is
unstable on a bimodal cell set — the 500 W arm's klein re-check spans 512 px (~0.29 s) and
768 px (~0.53–0.68 s), and one changed value flipped which mode the median landed in,
producing a phantom **-13.8 %** where nine of ten pairs had moved +0.0 % to +2.9 %. (The
tenth was the arm's first scored render carrying weight-load overhead, 0.42 s → 0.291 s on
re-run — the first-of-group effect the sealed bench documents.) Median of paired ratios:
**+1.4 %**. Both statistics are printed at every rung; neither is hidden.

### A4.5 Cell-set continuity, proved not asserted
| artefact | sha256 |
|---|---|
| the 139 scored cells (`cells`), as registered above | `615e84aa16e4ee3ead360704d1333e3ceb4d6aaa455001d5023f613e0888b812` |
| the 139 scored cells, after this amendment | `615e84aa16e4ee3ead360704d1333e3ceb4d6aaa455001d5023f613e0888b812` — **identical** |
| `sample.json` before (kept as `sample-v1-500W.json`) | `cc6a571cb422f06e0085de479cdff82d420adf26e030e4695b575001790e6c60` |
| `sample.json` after | `1f3e700d1721c888f388ec6e0b1ae9df8971c794eb2fea76e22ec371be1789bf` |

The warm set is likewise identical, and every sentinel is a member of the scored set. The
500 W arm therefore remains directly comparable to the 450 W and 600 W arms.

### A4.6 One more instrument change, same rung-0 spirit
The throttle-reason FLAG is only evidence while the card is BUSY: an open/close snapshot
catches it idle and always reads `Not Active`. `run_arm.py` now takes one
`nvidia-smi -q -d PERFORMANCE` snapshot **inside each lane's own block** at a fixed
position, so arms stay comparable. On the 500 W arm the equivalent probes were taken live
per lane and read **`SW Power Cap: Active`, P1** during flux1-dev.
