Exhibit twenty-four · the close-up on the painter's dial
Fifty-Seven Milliseconds
exhibit twenty-four The bench
Published 2026-08-25 (UTC)
a small (human) team and a fleet of AI agents
One picture, drawn rung by rung up the step dial, on two painters that could not be more different. Timed to the millisecond and measured to the pixel, twice, a month apart — and the two runs say two different things, which turns out to be the finding.
Every diffusion pipeline has a dial labelled steps, and we’ve always turned it the way everyone seems to: up when the picture disappoints, down when the queue is slow. The folklore says more steps means more quality, the way more oven-time means more roast. We had priced the dial in seconds before; we had never checked the folklore against the same image drawn again and again.
So we did the small, boring version of the experiment — twice, a month apart, on two painters that could not be more different. The first run: six subjects from Sorrowmoor Cove, our game’s fishing village — two foes, two portraits, two scenes, real content the game renders — drawn by the painter the game’s in-play sketcher ships by default, at seven step counts from 20 to 40, everything else frozen: same sampler, same 512×512 canvas, same seed. Forty-two images.
The second run, on the season’s new painters: six of the game’s real sketch prompts — the exact assembled text the in-play sketcher sends — at 768×768, at 4, 8, 12, 16, 20 and 32 steps, on four painters, one of them a model distilled to finish in four. Every render carries its seconds, measured from the render graph’s own execution timestamps, and every cell on the first ladder and the new one opens a specimen sheet with its full provenance. This piece is the close-up on what those receipts say — and they say two different things, which turns out to be the finding.
What a step is, in one paragraph
A diffusion model doesn’t paint a picture; it un-ruins one. It starts from pure noise and, in a loop, asks a neural network: “what noise would I remove to make this look more like the prompt?” — then removes a little of it. Each pass of that loop is a step, and each step does the same amount of work as the last. That predicts something checkable — for a fixed sampler (the algorithm that walks the loop) and configuration, if steps are the only thing you change, time should be a straight line with a slope. It is rare, in benchmarks, for anything to be as clean as its theory. This one is — on both painters.
Finding one: a step is a fixed-price good
Fit the shipped painter’s per-rung medians and the line is seconds = 0.238 + 0.0569 × steps. Drop the one misbehaving render session — the wobble we get to below — and it is seconds = 0.216 + 0.0569 × steps. Both fits ship, because the headline survives either: a step costs 56.9 milliseconds at 512² on the 24G VRAM rig, and the slope holds to three decimal places whether we fit the medians, all forty-two renders individually, or the post-exclusion set (the intercept — the fixed overhead of prompt encoding, decoding, plumbing — moves with the slice: 0.216 s on the post-exclusion set, 0.238 s on the per-rung medians, 0.257 s fitting all forty-two renders individually; the slope does not budge).
| steps | 20 | 24 | 28 | 30 | 32 | 36 | 40 |
|---|---|---|---|---|---|---|---|
| median seconds | 1.36 | 1.56 | 1.81 | 2.08 | 2.06 | 2.23 | 2.50 |
Measured 2026-07-10 by the archive’s own clock, on the 24G VRAM rig · AlbedoBase XL v2.1 · dpmpp_2m/karras · 512×512 · seed 424242 · each cell the median of six subjects at that step count · full receipts: receipts.json.
The new ladder, on a different card and a bigger canvas, draws the same straight line with different numbers: the distilled painter — FLUX.2 [klein] 4B, a model distilled to finish in four steps — costs 73 milliseconds a step over a quarter-second floor (0.25 + 0.073 × steps), and its rungs read 0.62, 0.79, 1.10, 1.40, 1.71 and 2.62 seconds at 4 through 32. The shipped painter on the same rig and canvas costs 99 ms a step over a 0.32 s floor. Different slopes, same shape: budget steps like seconds, because they are seconds.
Read that first table again and you’ll spot the untidy part: the 30-step rung is slower than the 32-step rung — by 0.010 s in the raw medians, which the rounded row inflates to twice that — and we’re showing it to you anyway, because the house law here is that you show the wobble rather than smooth it. The explanation is in the timestamps: that kit ran in four separate render sessions across a few hours, and the 30-step rung ran alone in its own short session — it also has the widest spread in the kit (1.89 to 2.48 s). Drop that one session — not the rung we dislike, the session that ran differently — and the fit’s R² (the share of the variance the line explains) goes from 0.976 to 0.997 while the slope doesn’t move at all. A slope that survives every way of slicing the data is the trustworthy kind. An R² that needs a footnote is the honest kind.
Finding two: the first turn of the dial is the big one — and how big depends on the painter
Seconds are easy to measure. “Better” is not — better is a visual judgement, and we won’t pretend a number can make it for you; the galleries exist so you can make it yourself. What a number can say is how much the picture changes, so we measured that: for each pair of images, the mean absolute per-pixel difference, averaged over the three colour channels (0–255 scale). It is the bluntest instrument on the shelf, on purpose — perceptual metrics like SSIM or LPIPS answer “how similar does it look”, which smuggles a judging model into the measurement; we wanted a number anyone can recompute in three lines. The rule is printed in full below and reproduces on both archives’ lossless masters.
On the shipped painter (the 512² run), the 20→24 rung moves the picture about 1.6× as much as the typical later rung (a median change of 14.7 on the 0–255 scale, against 8.6 to 9.9 for the rungs after). On the distilled painter, starting from the bottom of the dial, the effect is far steeper: the 4→8 rung moves the picture 2.6× a later rung (a median of 19.9 against 7.7), and the biggest single move in either experiment — 32.7 — is a foe going from 4 to 8 steps. Look at the fisherman on the cove floor in the new gallery: at 4 steps there is no bird and not much of a floor; 8 and 16 draw the birds in and the detail with them, the figure squatting on the floor the whole way — fine, but squatting; and 32 stands him up on his legs. Each rung is a visibly different picture, and to our eyes each is better than the last — the gallery is there so you can disagree. The first turn of the dial is not just the biggest; on a painter built for few steps, it is most of the picture.
Finding three: the ladder settles on one painter and never on the other
Here is where the two runs — the 512² ladder from July and the 768² one from August — stop agreeing, and it is the finding we did not expect.
Measure each rung against the 32-step image with the same per-pixel rule as above — mean absolute difference across the three colour channels, 0–255 — and ask whether the picture is approaching it. On the distilled painter, every one of the six subjects walks monotonically toward its final image — the fisherman reads 25.6, 17.4, 13.4, 11.0, 9.7 at 4, 8, 12, 16 and 20 steps, each rung strictly closer than the last, and the other five subjects do the same without exception. Six of six. The ladder converges: the image at 32 really is what the image at 4 was becoming.
On the shipped painter, measured at 512² from 20 to 40 against its 40-step image, the trend points the same way in aggregate (13.98 at 20 steps, falling to about 9 on the same scale) — but not one of the six subjects approaches it monotonically. Every single one steps away from its own 40-step image at least once on the way up. Zero of six. And remember the sampler is deterministic and the seed is fixed: nothing here is random. Above that painter’s finishing point, the right picture of the ladder is not a photograph developing toward its final form — it’s a shelf of siblings, each rung a slightly different answer to the same prompt, drifting closer in aggregate without ever walking a straight line there.
Put the two ladders side by side and the dial has two regimes. Below a painter’s finishing point, each rung buys a visibly different and better picture, and the picture is converging. Above it, each rung buys a sibling. The shipped painter’s ladder started at 20 — already past its crossing, which the new run puts near 12 to 16 (at 768², the same six prompts are mush at 4, composed at 8, and finished for every subject by 16). The distilled painter crosses near its native 4 to 8 and keeps improving in smaller steps to 32. The honest way to set a step count is to find where your painter crosses from the first regime into the second, and stop there. In neither case is 40 buying what the folklore promised.
Finding four: the argument with our own settings — and what it changed
Our game rendered its in-play sketches at 512² and 32 steps — a tier chosen so a sketch lands in about two seconds and keeps pace with the conversation it illustrates, priced, at the time, in seconds alone. The first ladder raised an eyebrow at that: the largest rung-to-rung change was already behind the dial by 24, so a 24-step sketch gave up one typical rung of drift for half a second of pace (0.50 s by the measured medians: 2.06 s at 32 against 1.56 s at 24). We were careful then not to move a production setting off six subjects and one seed; that is how superstitions start.
The second run is what moved it. On the 96G VRAM workstation the sketch tier will render on, at 768² the shipped painter’s rungs read 0.77, 1.08, 1.51, 1.89 and 3.52 seconds at 4, 8, 12, 16 and 32 — and 768² costs only 0.04–0.08 s more than 512² at every step count, so the canvas was nearly free and the whole budget was the dial. The operators read the ladders and ruled: the jump from 8 to 12 is where the picture arrives; 12 to 16 is polish nobody sees at sketch scale, not worth the extra half second. The engine’s shipped sketch tier moved to 768² at 12 steps — 1.51 s against 3.48 s at the old setting on that card (512² at 32 steps, measured there; the 768² column reads 3.52 s at the same step count): more picture and less time, on the same painter. To be precise about what happened: the old tier was never wrong — it met the brief it was set. The receipts found a cheaper way to meet it, and the setting followed the receipts. That’s the research loop working as intended.
One more receipt, and it cuts against us: the distilled painter is the sharpest and fastest cell in the whole second run (0.62 s, finished, at 4 steps) and it did not take the tier — it paints a legible cursive signature on two of the six subjects — an inland figure and a foe — in every step column, because its template zeroes the negative conditioning that the no-signature guard rides on. A defect, not a taste call, and a research rung for another week. The fisherman above is one of the four it leaves unsigned.
What to take with you
- A denoising step is a fixed-price good: 56.9 ms at 512² on the 24G VRAM rig, 73 ms for a distilled painter at 768² on the 96G VRAM workstation and 99 ms for the shipped one on the same canvas — invariant to how you slice the data. Budget steps like seconds, because they are seconds.
- The biggest change your dial will ever buy is its first turn up from the floor — 1.6× a later rung on a conventional painter, 2.6× on a distilled one.
- The dial has two regimes. Below a painter’s finishing point the picture converges, rung by rung, toward its final form — six of six subjects, monotonically. Above it, rungs are siblings: zero of six approached the top monotonically. Find your painter’s crossing (near 12–16 for ours at 768², 4–8 for a distilled one) and stop there.
- When a rung misbehaves (our 30), show it, explain it, and publish the fit both ways. The slope that survives is the finding.
How to check our work — and see it live
Both grids are published with every cell captioned and every specimen sheet carrying the verbatim prompt: the first ladder — 42 images, albedobaseXL_v21 · dpmpp_2m/karras · <n> steps · 512×512 · seed 424242 · <seconds>s — and the new arrivals’ ladders — six real sketch prompts × six step counts on three painters whose licences allow the pictures, with the shipped painter’s two ladders counted, not shown, because its licence text has never been found. That the step count is the only thing moving is a verified property of each manifest — exactly one distinct value for every other field — not a promise. The first run’s renders date to 2026-07-10 by the archive’s own clock, on the 24G VRAM rig; the second to 2026-08-24 on the 96G VRAM workstation (machines here are named by VRAM class, not model); within-rig comparisons are exact, extrapolating absolute seconds to your card is approximate. All three fits are straight lines through the stated points; a rung’s median is the median of its six subjects. Both receipt sets ship as machine-readable JSON beside their galleries (the first, the second).
The first painter is AlbedoBase XL v2.1 — the engine’s shipped default for in-play sketches; the live world we run pins its own painters for bulk pre-generation, which is why the census in the walk next door gives this one a small share of all renders — whose creator’s permission flags — read first-hand from the creator’s listing on 2026-08-11 and recorded on our licence ledger — permit publishing outputs with credit: painted by albedobase xl 2.1, by albedobond. (Our newer bench harness holds itself to a stricter bar — the licence text itself, on disk — and until that text is found it withholds this painter’s images from new pages. Two evidence bars, both ours; the first gallery ships under the first, the new one under the second, and the ledger names both.) The pixel-difference rule, in full:
from PIL import Image, ImageChops, ImageStat
d = ImageChops.difference(Image.open(a).convert('RGB'), Image.open(b).convert('RGB'))
mad = sum(ImageStat.Stat(d).mean) / 3.0
Run it on either archive’s lossless masters (the published pages serve compressed derivatives, which add a small floor to every number) and every figure above reproduces.
See it live. Play RealKeep — free, no account — and the limner sketches at the tier this piece moved: 768² at 12 steps, about a second and a half a picture, while you read the room it drew.
The page carries no images; the galleries carry each painter’s credit, and every cell opens its full provenance.
One more disclosure: the prompts behind the first grid share a style clause our early art direction wrote, which names a commercial game as a style reference. The gallery publishes those prompts verbatim with a notice — they’re records of what we typed, not claims of association — and de-naming that clause product-side was already in progress before this piece was written.
The rest of the seminar
The week’s opener walked the whole pipeline — the file, the dice, the painter, and a dog. This was the close-up on the painter’s dial, and it ends with the dial moved. The bench that supplied the second ladder — the season’s open-weight painters sitting for the same small village — released the day before this piece. Two neighbours on the shelf are this piece’s other halves: Reading is fast, writing is slow is the other place where time turned out to be physics rather than theatre, and The free speed wasn’t free is the other measured slope. Still to come this week: how a vision model sees.
The whole shelf holds the benches behind the claims we publish, failures included. If there is a piece of the machinery you want opened next, say so — the suggestion box is read.
Licence: CC BY 4.0 for the text and tables — name the source and link to it; images are licensed individually and the galleries carry each painter’s credit.