# The noise floor these verdicts are read against

**Read this before you read any cross-cap difference in this kit.** The bench
pre-registered the rule before it ran: *a cross-cap difference smaller than the
same-cap repeat spread is not a result.* This table is that spread. The sparse
(MoE) sustained probe was interleaved — 600 W, 450 W, 600 W, 450 W — so that the
**same cell at the same limit** was measured twice, and the two readings could be
differenced against each other. Everything below is R1 minus R2 **at one limit**:
no cap changed between the two columns. Whatever moves here is the instrument,
not the cap.

Rungs: `rungs/vllm-fp8-gemma-sustR1-600W-p512-c16` vs `…sustR2-600W…`, and
`…sustR1-450W…` vs `…sustR2-450W…`. Every figure recomputes from those four
directories' `summary.json` and `power.csv`.

| cap   | c  | metric        | R1      | R2      | delta  | %      |
|-------|----|---------------|---------|---------|--------|--------|
| 600 W | 16 | TTFT p50 ms   | 129.8   | 130.3   | +0.5   | +0.4%  |
| 600 W | 16 | TTFT p95 ms   | 156.6   | 135.7   | -20.9  | -13.3% |
| 600 W | 16 | tok/s /stream | 170.33  | 170.15  | -0.18  | -0.1%  |
| 600 W | 16 | tok/s agg     | 2429.21 | 2454.37 | +25.16 | +1.0%  |
| 600 W | 16 | W mean        | 402.9   | 415.4   | +12.5  | +3.1%  |
| 600 W | 16 | W busy        | 407.8   | 417.8   | +10    | +2.5%  |
| 600 W | 16 | W peak        | 452.7   | 461.1   | +8.4   | +1.9%  |
| 600 W | 16 | T peak C      | 66      | 70      | +4     | +6.1%  |
| 600 W | 16 | SM MHz mean   | 2806    | 2788    | -18    | -0.6%  |
| 600 W | 16 | tok/s/W       | 6.029   | 5.908   | -0.121 | -2.0%  |
| 600 W | 16 | tok/s/W busy  | 5.957   | 5.875   | -0.082 | -1.4%  |
| 450 W | 16 | TTFT p50 ms   | 130.8   | 130.7   | -0.1   | -0.1%  |
| 450 W | 16 | TTFT p95 ms   | 140     | 138.5   | -1.5   | -1.1%  |
| 450 W | 16 | tok/s /stream | 170.06  | 170.11  | +0.05  | +0.0%  |
| 450 W | 16 | tok/s agg     | 2382.72 | 2426.38 | +43.66 | +1.8%  |
| 450 W | 16 | W mean        | 414.8   | 414.7   | -0.1   | -0.0%  |
| 450 W | 16 | W busy        | 418.8   | 419.2   | +0.4   | +0.1%  |
| 450 W | 16 | W peak        | 456.4   | 455.7   | -0.7   | -0.2%  |
| 450 W | 16 | T peak C      | 70      | 69      | -1     | -1.4%  |
| 450 W | 16 | SM MHz mean   | 2796    | 2798    | +2     | +0.1%  |
| 450 W | 16 | tok/s/W       | 5.744   | 5.851   | +0.107 | +1.9%  |
| 450 W | 16 | tok/s/W busy  | 5.689   | 5.788   | +0.099 | +1.7%  |

## The three numbers this table licenses

- **Aggregate throughput repeats to 1.0 % at 600 W and 1.8 % at 450 W.** The
  cross-cap aggregate difference on this arm is 1.5 % — *inside* that band, and
  therefore **not a result**. It is reported, and it is reported as noise.
- **Per-stream decode repeats to 0.1 %.** That is the only throughput column tight
  enough to carry a verdict on this arm, which is why the sparse verdict
  (170.24 → 170.09 tokens a second per stream, −0.09 %) is stated as *no
  measurable change* rather than as a loss.
- **Busy-mean power repeats to 10.0 W at 600 W** (407.8 vs 417.8). Any watt
  difference smaller than that, in either direction, is this instrument's own
  spread — including the sparse arm reading ~6 W *higher* under the *lower* cap.

## What the table does not cover

The dense (Q4) arm has **no same-cap repeat** — it was measured once at each
limit. Its 2.0 % per-stream difference is therefore read against the *sparse*
arm's floor, which is a borrowed denominator. That is a weaker claim than the
sparse arm's, and the exhibit says so in its own words rather than burying it.

The 3-wave burst rungs (`…-p512-c1`, `…-p512-c16` without `sust` in the name) are
7–9 seconds of card time each. They are kept for comparability with the morning
serving bench and **must not be read as power results**: their own c=16 pair
differs by 9.1 % with identical SM clocks at both limits, and neither leg reached
450 W. That difference is scheduling noise.
