The bench — what this page times, on which machines and which days, and what it never measured
Four homes for one reranker
exhibit fifty-six The bench
Published 2026-09-21 (UTC)
A small (human) team and a fleet of AI agents.
An operator at a game table waited nearly twenty seconds for the first word of a rules answer, and thought the app was broken. It was not. On the morning of 2026-09-17, one small stage of every answer from RuleSage, our board-game rules helper — the careful reader that re-reads the shortlist before the answer is written — was running on a rented server whose processors could not do that stage's arithmetic the fast way. That stage's own time to re-read a shortlist of 32: 1,792 ms there at one thread, against 30.1 ms on the card it runs on now. This page follows it across four machines in six weeks, three of them in five days, with the gate before the last move, the traps that bit on the way and what is still unmeasured.
ask about this page → assistant.strata2signal.com · in beta, still being tested
hardware on this page: RTX PRO 6000 Blackwell (96 GB) · RTX 3090 (24 GB) · Intel Core i5-12600H · the mini PC → research.strata2signal.com/hardware/ · the roster is in beta, still being tested
the short version
On the morning of 2026-09-17 a rules answer took nearly twenty seconds to start, and an operator at a game table thought RuleSage, our board-game rules helper, was broken. It was not. One small stage of the answer — the reranker, a cross-encoder that re-reads the shortlist before the answer is written — had moved to a rented server whose processors could not do that stage's arithmetic the fast way: 1,792 ms to re-read a shortlist of 32 at one thread, against 329 ms on the machine it came from. It has since moved to an RTX PRO 6000 Blackwell (96 GB) already in the house, where the same stage takes 30.1 ms. Measured on eight asks of our own, sent down the public path half an hour after the move, the first word now arrives in about a second and a half (1.56 s median, on the five that answered). Four homes in six weeks; every millisecond, with the bench file it came from, is on the page.
- The card's power under a 60 s saturated burst reached a median of 168.61 W (see What the seat costs, where we can say)
11,479 words, about 52 minutes to read.
The summary is this page’s own; the receipt lines were drafted by a model on this workshop’s network and every figure in them is in the article, checked before this page went out — what was dropped, and why, is in this page’s receipt file.
If you're new here: strata→signal is a small workshop (plus a friendly dog with a white patch) that builds on its own machines and writes up what it measures. Every figure below was read on 2026-09-20 (UTC) from the bench file, ledger row or live record named beside it. Machines are named by their role, never their hostname; graphics cards are named in full; a number we could not source is a dash with its reason.
Six words this page leans on. A passage is a bite-sized piece of a rulebook, usually a rule or two long. A reranker is a model that re-reads a shortlist of passages and puts it in a better order. A cross-encoder reads your question and one passage together, as a pair, instead of comparing them as two points on a map: slower per passage and much harder to fool, and ours is one. An instruction set is the vocabulary of operations a processor understands. VNNI is one of those operations, a shortcut for multiplying small whole numbers in batches: the one this model's arithmetic was built around, and the one the new server lacked. A thread is one worker inside a program; how many run at once is a setting, and on the new server it was still set for the old one's size.
A question on game night
Say it is a Friday night, six people around a board, and someone asks whether the longest road breaks when another player builds a settlement in the middle of it. Somebody types it into RuleSage, this workshop's board-game rules helper, which answers out of the rulebook's own words with the page cited. On the morning of 2026-09-17 the answer did not come. An operator at a table waited, and stopped waiting. The row they wrote in this workshop's own decision ledger that morning is the plainer version: "it takes FOREVER to answer." The app was two days into a new rented server.
It was not broken, and the app's own debug panel said what had not changed. The model that writes the answer never moved, and it was writing at the same speed for asks from either server. The search that finds candidate passages took about the same time on both. The files were identical on the two machines; we checked.
What had changed was the home of one stage in between: the reranker, the careful reader at the end of the funnel, which reads the question beside each shortlisted passage and decides which eight reach the answering model. On the old server that stage took 329 ms for a batch of 32 short candidates. On the new one it took 1,792 ms: 5.4 times longer at one thread, the posture both boxes were in that morning, on a server with twice the processors.
The first-token figures on this page — the bands above and in the table below — were read once off the app's own debug panel that morning and the rows were not kept, so every "before" here is a band and never a middle reading.
Four homes in six weeks, three of them in five days
The reranker was adopted on the first VPS on 2026-08-10 and lived there thirty-six days. Between the move on 2026-09-15 and the evening of 2026-09-20 it had three more addresses: the public host, a mini PC it was only measured on and the card it runs on now.
| Where it lived | Its own time to re-read a shortlist of 32 | The table's wait for the first word |
|---|---|---|
| The first VPS, a rented server with VNNI, 2026-08-10 → 2026-09-15 | 329 ms (2026-09-17, 1 thread) | 0.75–2.1 s (stated; rows not preserved) |
| The public host, a rented server without VNNI, on arrival 2026-09-15 | 1,792 ms (2026-09-17, 1 thread) | 2.5–19 s (stated; rows not preserved) |
| The public host after one setting was corrected, 2026-09-17 17:25Z | 522.3 ms (2026-09-18, 4 threads) | 3.8–5.1 s (three asks) |
| The mini PC on the workshop's bench, with VNNI, measured 2026-09-19 and never moved in | 118.1 ms (4 threads) | — (never served a visitor) |
| An RTX PRO 6000 Blackwell (96 GB) already in the house, since 2026-09-20 21:02Z | 30.1 ms (n = 5) | 1.56 s median, five answered of eight, on a clock that starts when the question leaves the asking machine |
Every cell is the reading as that stop happened, at the short shape, the only shape measured at the time; the long shape is in the builder's tables below. The card is read twice on this page: 30.1 ms in the afternoon's run of five and 28.9 ms in the evening's run of twenty, each printed with its own run. Sources: the workshop's decision ledger for the incident morning and its benches of 2026-09-18, -19 and -20, named in the builder's half.
The first VPS was never the slow one. The problem was not the app and not the model: it was one machine that could not do one kind of arithmetic the fast way.
We will not write "twenty seconds to thirty milliseconds", because those are two different clocks: one was the whole answer's wait at a table, the other is one stage's busy time on a card a network hop away. The honest pair is this. The reranker's share of the first token fell from seconds to tens of milliseconds. And the first word itself, on the five asks that answered, arrived in about a second and a half, on a clock that starts earlier than the one the twenty seconds were read on. A second and a half is about as long as it takes to pick the dice back up.
What a reranker does at a rules desk
RuleSage finds candidate passages with three searches: one for exact words, one for meaning and one that never leaves the game's own books. It fuses their ranked lists by position and cuts the fused list to a shortlist. A cross-encoder then reads the question and each shortlisted passage together and scores how well that passage answers that question. The eight best passages go to the answering model in a deliberate order, and the answer quotes them by page; the longer version is Three librarians and a careful reader.
The cross-encoder is the expensive stage by design, because it is the only one that reads pairs; on the first VPS it was very nearly the whole of an ask's time. That is why moving it moves everything, and why a server that runs it five times slower shows up as a table waiting.
That is the whole story, and this is a fine place to stop; thanks for reading it. If you came here from a game table, the page to read next is this one's prequel, Three librarians and a careful reader: how RuleSage finds the right page in the first place.
Everything below is the same story with its receipts open, for anyone who runs a search-and-answer stack of their own: where to put a reranker, what each class of machine measured, how it got here and what to check before you move one.
Welcome to the receipts, and first the numbers a practitioner reads for. The same int8 graph, checksum-matched across boxes, 32 candidates at about 133 tokens, four threads: 138.5 ms on a rented server with AVX-512 VNNI, 386.4 ms on one without it and 123.0 ms on a mini PC with AVX-VNNI; at 512 tokens, 734.5, 2,690.6 and 962.0 ms. In fp32 on an RTX PRO 6000 Blackwell (96 GB) already in the house: 28.9 and 67.2 ms.
Re-exporting the graph for AVX2 was slower in five cells of six. Before the flip, overlap@8 against the old graph was 0.9625: past our 0.90 bar, under our own 0.97 band, on a substitute fixture. The honest gap: the app pays about 100 ms, unpooled, to reach the door, about 195 ms an ask for the whole leg to the seat, and nobody has measured where production's pairs fall between those two shapes.
Every table below names its bench, its box, its date and its n; every ratio divides two cells of the same table; and where a number is missing there is a dash and a reason. Five more words first.
Five more words, for the tables below. A seat is one model, on one machine, reachable over the network, doing one job. The door is the workshop's own small gateway that every product here calls instead of calling a machine: you ask it for role:reranker and it finds whichever seat holds that role today. An int8 model has weights squeezed down to 8-bit whole numbers: small and fast where VNNI exists, much slower where it does not. fp32, or full precision, is the 32-bit form they were squeezed from, and the form the card runs. The two shapes are the two batches every table measures: 32 candidates of about 133 tokens each, the incident morning's shape, and 32 of 512 tokens each, the most the model reads of any pair.
Where should your reranker live?
The table a builder wants, by class of machine. Every rerank cell is the same model, the same two shapes, the same passages and the same evening (2026-09-20, n = 20, 4 threads on the CPUs), so any row divides by any other. How every cell was taken: one code path copied to the three CPU boxes, and the same synthetic passages generated on each box and digested there, so all four scored identical bytes. Twenty timed repetitions after a warm-up, one fresh process per CPU cell and the card timed over the wire through the product's own client.
| Class of machine | Our instance | Rerank, 32 × ~133 tok | Rerank, 32 × 512 tok | The lever, or the catch |
|---|---|---|---|---|
| A rented vCPU without VNNI | The public host, 8 vCPUs | 386.4 ms | 2,690.6 ms | Threads, and they run out |
| A rented vCPU with AVX-512 VNNI | The first VPS, 4 vCPUs | 138.5 ms | 734.5 ms | int8's best case on a CPU |
| A mini PC with AVX-VNNI | The mini PC, a mobile 12th Gen Intel Core i5-12600H | 123.0 ms | 962.0 ms | Fine at the short shape, slow at the long |
| A GPU already in the house | RTX PRO 6000 Blackwell (96 GB), 910 MiB used | 28.9 ms (2.8 ms of it network) | 67.2 ms (2.8 ms of it network) | Plus the hop from your app to it: about 195 ms an ask here, door included |
| A rented GPU | — | — | — | Not measured here |
| A hosted reranking API | — | — | — | Not measured here; your candidates leave the building |
The same-night pass, its own section below; the wait at each stop, on the clock each was read with, is in the wait table at the end of that section. The last two rows exist so that their blanks are visible.
Two projections exist and neither is in the table: the mini-PC bench's own arithmetic put a CPU seat there at 1.75–3.05 s and the plan's promise for the card at 0.81–2.14 s; the measured 1.32–1.57 s is on a different clock. Filling either blank row needs a price, not a millisecond: an hourly rate or a per-thousand fee, set against the fifteen calls RuleSage sent the new seat in the whole of its first day.
Every rerank cell compared above is 32 candidates, and a live ask's fused shortlist is wider: the adoption bench's mean was 52.5 and the fusion cuts at eighty, so the synthetic cells are a floor of unknown height. The eight live asks at the end of this page are the only readings taken at production's own width: 29–133 ms for the whole door-and-seat leg, median 94 ms. The card's own k-sweep is the only thing here that says how the cost grows with width: 31.0 ms at twenty candidates, 51.2 ms at forty, from the G5 seat bench's own sweep.
The one-line answer: a reranker belongs on the cheapest machine whose processor carries VNNI, and if there is already a GPU in the house, on the GPU. On our cells that is 138.5 ms on a VNNI vCPU, 123.0 ms on a mini PC and 28.9 ms on the card, against 386.4 ms on a rented vCPU without it; at the long shape the same four read 734.5, 962.0, 67.2 and 2,690.6 ms. If your processor has VNNI, the thread count is the lever you have left; if it does not, no export setting recovers it, and the lever is a different machine.
If a rented VPS is all you have. Check for VNNI in the processor flags before anything else: the same int8 graph read 329 ms on a VPS that had it and 1,792 ms on one that did not. If it is missing, do not re-export for AVX2: on the box with the problem, that graph was slower in five cells of six. And do not expect fp32 to save you; on that box int8 still beats it by 1.5× at the short shape, and at the long shape the two collapse onto each other.
The one free lever is the thread count, worth 2.7–3.1× if it is still sized for an older box's cores, and it runs out there. After that the options are a VNNI-capable instance, a mini PC or a card. Which one depends on what shape your candidates actually are: at 133 tokens a tuned VPS without VNNI is four tenths of a second; at 512 it is nearly three.
Four seats, one night, the same passages
On the evening of 2026-09-20, twenty minutes after the flip, the same cross-encoder was timed on all four seats it has ever had, with one passage generator and one summarizer shared by its two harnesses, so that the four compare like for like.
The deviations, stated. The bench ran master's loader on each CPU box rather than the installed one, since the first VPS's copy predates the thread and target knobs. The public host's cells ran niced and the other boxes' did not. The fp32 arm's torch build differs on the mini PC. And the public host had no fp32 weights installed, so they were staged in a home directory for the run. No unit was touched and nothing was deployed.
| Seat | Precision | 32 × ~133 tok, 1 thread | 32 × ~133 tok, 4 threads | 32 × 512 tok, 4 threads |
|---|---|---|---|---|
| The first VPS (AVX-512 VNNI) | int8 | 253.6 ms (p95 258.5) | 138.5 ms (p95 192.0) | 734.5 ms (p95 806.3) |
| The first VPS | fp32 | — (not run) | 419.4 ms (p95 474.0) | 1,825.6 ms (p95 2,439.6) |
| The public host (no VNNI) | int8 | 1,325.8 ms (p95 1,359.1) | 386.4 ms (p95 419.1) | 2,690.6 ms (p95 2,799.5) |
| The public host | fp32 | — (not run) | 587.2 ms (p95 658.3) | 2,591.1 ms (p95 2,798.7) |
| The mini PC (AVX-VNNI) | int8 | 379.2 ms (p95 391.5) | 123.0 ms (p95 129.6) | 962.0 ms (p95 971.8) |
| The mini PC | fp32 | — (not run) | 255.5 ms (p95 290.4) | 1,191.1 ms (p95 1,223.7) |
| The card, dialled direct, over the wire | fp32 (CUDA) | — (a seat has no thread count) | 28.9 ms (p95 29.9) | 67.2 ms (p95 81.7) |
The rerank-journey bench of 2026-09-20, Table 1: 21:22–21:32Z for the CPUs, 21:43–21:44Z for the card; n = 20 in every cell, p95 nearest-rank. The one-thread anchor was run for int8 at the short shape only, and no thread count but 1 and 4 was run. The card's cells sit under the four-thread columns, the posture they are compared against. Peak RSS: 840–882 MiB at the short shape and 988–1,075 MiB at the long one on the two VPSes, 522–747 MiB on the mini PC.
Measured again on one night with the same passages, the same graph cost 2.79× more at the short shape and 3.66× more at the long one on the public host than on the first VPS. Both are at four threads, against the 5.4× the incident read at one. The card now beats even the first VPS by 4.78× and 10.93×.
| Contrast, at four threads | 32 × ~133 tok | 32 × 512 tok |
|---|---|---|
| The migration's cost: the first VPS → the public host, int8 | 2.79× slower | 3.66× slower |
| The flip's gain: the public host → the card | 13.35× faster | 40.04× faster |
| The flip's gain on compute alone (the card's wall minus its 2.8 ms floor) | 14.8× faster | 41.8× faster |
| Against the original home: the first VPS → the card | 4.78× faster | 10.93× faster |
| The mini PC → the card | 4.25× faster | 14.32× faster |
Every ratio is arithmetic on two medians from the table above. The card's short-shape median is 28.95 ms before rounding (its two middle laps are 28.9 and 29.0), so the short-shape ratios are 386.4 / 28.95 = 13.35, 138.5 / 28.95 = 4.78 and 123.0 / 28.95 = 4.25; dividing the displayed 28.9 reads 13.37, 4.79 and 4.26. Every ratio understates the card, whose median carries a network floor the CPU cells do not.
Two of the same-night cells replicate the earlier readings closely: the mini PC's (123.0 against 118.1 ms; 962.0 against 963.6 ms). The public host's do not. Its CPU read 1.06–1.35× faster on 2026-09-20 than on 2026-09-18 on the same graph bytes and thread counts; the one-thread cell barely moved, the multi-threaded cells moved most, and a confirmation read at 22:04Z came in 6.8% faster again. Two things differed between the two runs: the later one was niced, and it ran the library's master tree rather than the installed one. So this page quotes 522.3 ms when it retells the 2026-09-18 diagnosis and 386.4 ms when it compares seats measured the same night, and never averages the two.
And the wait itself, by instrument:
| Stage | First token | How it was measured | n |
|---|---|---|---|
| Before the move, the first VPS | 0.75–2.1 s | The app's own debug rows, which start the clock after the search; before 2026-09-15; rows and window not preserved | — (stated) |
| After the move, before the thread fix | 2.5–19 s | The same rows, on the public host, 2026-09-17 04:17–14:10Z, not preserved; the asks whose passages filled the model's 512-token window sat in the upper part, 7–19 s | — (stated) |
| After the thread fix | 3.8–5.1 s (three asks) | The same rows, three real asks after the 17:25Z restart; the ledger's hotfix row, 2026-09-17 | 3 |
| After the flip to the card | 1.56 s median (1.32–1.57 s) | Client-side, from the moment the question leaves the asking machine to the first word of the answer; eight asks through the public reader path, 2026-09-20 21:35Z | 5 |
Different instruments: the first three rows are the server's own clock, started after the search; the last row's starts when the question leaves the asking machine, so it contains the search, one internet hop and any queueing the server-side figure never sees. Because RuleSage says the book is silent rather than guessing, three of the eight had no first word to time, so its n is five. The before was not re-measured; the product was not rolled back.
The model and the two shapes behind every latency table
| Thing | Value | Source |
|---|---|---|
| Model | cross-encoder/ms-marco-MiniLM-L6-v2, revision c5ee24cb16019beea0893ab7796b1df96625c6b8 | G1 bench, 2026-08-10 |
| Input | Query × candidate pairs, one score per pair; the pipeline keeps the top 8 | G1 bench; the prequel |
| Quantized form, the CPU seats | int8, dynamic, weights-only, per-channel, no calibration set, exported with AutoQuantizationConfig.avx512_vnni | G1 bench |
| Size on disk, int8 | 22.06 MiB | G1 bench |
| Size on disk, fp32 | 86.75 MiB | G1 bench |
| Full-precision form, the card | fp32 through the same loader, device=cuda, max_length=512 | The retrieval library's loader change; the ledger row of 2026-09-20 |
| Runtime, the CPU seats | ONNX Runtime 1.28.0 on all three boxes | The same-night pass, its instrument note |
| Runtime, the fp32 CPU arms | PyTorch 2.13.0+cu130 with sentence-transformers 5.7.0 on both VPSes; 2.14.0+cpu with 6.1.0 on the mini PC | Same |
| Batch size inside the seat | 8 | The shipped loader |
| The short shape | 32 candidates × about 133 tokens, 65 words each (the shape of the 2026-09-17 incident readings; the ledger rounded it to about 130) | G2, G3 and G5 benches; the same-night digests |
| The long shape | 32 candidates × 512 tokens, 600 words each (the model's max_length, the most it reads of any pair) | Same |
| Per-candidate text budget | The deployed engine's own configuration, read on 2026-09-20: a 1,200-character budget per passage, head and tail, with a 900-character floor; whether the live unit overrides it was not read | An operator's read of the deployed engine's own config.py on the public host, 2026-09-20 22:00Z; the unit's env override is root-owned and was not read |
| Live corpus passages, median length | 1,497 characters (mean 1,446; p90 2,300; max 30,351) | G1 bench |
| Live corpus passages over 1,200 characters | 61.3% of 29,505 | G1 bench |
The 1,200-character default was ruled in on 2026-08-12, on a Monopoly turn-question specimen. At 1,200 the int8 graph scored an overlap@8 of 0.9458 against 900's 0.9250 (the share of the top eight that two rankings agree on, the measure the gate section below uses). It cost a search p50 of 1.236 s against 0.834 s on the first VPS, and was accepted for reader quality; the 600-character budget had been rejected the day before, 2026-08-11, at 0.8542.
Two cuts happen, in order: the engine trims each passage to its 1,200-character budget, about 595 characters of the opening, the final 600 and a five-character joiner, as the prequel describes; the model then truncates each question-and-passage pair at 512 tokens. A median live passage, at 1,497 characters, is trimmed by the first cut, and so are 61.3% of all of them.
Which of the two shapes production's pairs resemble decides every claim in this half of the page. The short shape is what an operator had to hand on the morning of the incident. The long shape is the model's ceiling, not the typical pair. A 1,200-character passage with a question in front of it runs to roughly 300 tokens at about four characters a token — between the two shapes, and nearer the long one in ratio, at 2.3× the short shape's tokens. Nobody has read where the live pairs fall; the open-items table carries the unrun query. So what is left is arithmetic.
Correcting one thread setting saved 0.9 s at the short shape and 6.5 s at the long one on the public host, and production's wait fell by several seconds, a fall only the long cell is the right order of magnitude for. The bench states the caveat on its own argument: the before and after were different asks, so this compares magnitudes and is not a paired measurement. Every ratio below is labelled with the shape it was measured at.
The move to the public host and the slowness
On the first VPS, in the adoption bench of 2026-08-10, the rerank was 93.6% of a search's time and 95.6% of an ask's. On 2026-09-15 an operator moved RuleSage from that box to a new rented VPS that now serves most of what this workshop runs in public. Two days later the morning's asks looked like this.
| What was measured | The first VPS | The public host | Source |
|---|---|---|---|
| Rerank, 32 × ~133 tok, 1 thread | 329 ms | 1,792 ms | The ledger's incident row, 2026-09-17; tabled in the G2 bench |
| Rerank, 8 × ~120 tok, 1 thread | 96 ms | 434 ms | Same row |
| Virtual CPUs | 4 | 8 | Same row; the public host's flags, read that morning |
The row also records what did not change — search about 200 ms and generation about 195 tokens a second on both — and what did: the first token after search at 0.75–2.1 s against 2.5–19 s. The two rerank readings came from one script run on both boxes that morning, with no rep count recorded, so neither carries an n; until 2026-09-20 they were the only readings of the reranker's own milliseconds at this page's two shapes taken on the first VPS.
Everything that was the same ruled out everything that was not the problem. The diagnosis at 14:56Z was the processor: the new host's vCPUs are Haswell-class, with AVX2 and no VNNI in any width, and the shipped model graph had been quantized for avx512_vnni. The first VPS's vCPUs carried that instruction; the new host's did not. The deploy had verified the bytes, and the bytes were not the thing that mattered. Inside three hours the retrieval library learned which box it was on, with three environment knobs: the reranker's thread count (default 1), its batch size (default 8) and its graph target (default avx512_vnni). At 17:25Z an operator turned the first of them.
One thread was the whole free lever
The reranker had been running with intra_op_num_threads = 1, and it had been correct to. On the first VPS four concurrent callers, each with one ONNX Runtime thread, exactly filled four vCPUs. That arithmetic did not travel to a host with eight. Setting the thread count to 4 was one line in the service's environment file and a restart.
| Shape | 1 thread | 4 threads | Speed-up | Source |
|---|---|---|---|---|
| 32 × ~133 tok | 1,410.5 ms | 522.3 ms | 2.7× | G3 bench, 2026-09-18 |
| 8 × ~120 tok | 340.0 ms | 130.7 ms | 2.6× | Same |
| 32 × 512 tok | 9,589.7 ms | 3,063.3 ms | 3.1× | Same |
All six cells are the shipped graph on the public host, from the G3 run of 2026-09-18 21:25:35–21:32:36Z, pre-registered before it ran, five laps a cell; each speed-up divides two cells in its own row.
G3's own 1-thread reading at the short shape, 1,410.5 ms, is about 21% under the 1,792 ms of the incident the day before, on the same box at the same nominal shape. That gap is the ordinary spread of a shared vCPU, and the reason every cell here carries its median, its box and its date.
The thread knob helped more at the long shape than the short one, 3.1× against 2.7×. That is backwards from what thread scaling usually does, and it is the good direction here, since the long shape is the one that hurts. Memory did not move between one thread and four (845–853 MiB at the short shape, 1,069–1,075 MiB at the long one), so the lever cost nothing on a host with 24 GB, a figure from the workshop's inventory rather than from any bench here.
Production was measured the same evening, on three real asks after the restart at 17:25Z: first token after search 5.1 s, 4.3 s and 3.8 s (n = 3), from the 2.5–19 s window before. Usable, and nowhere near good. Three minutes after the receipt an operator ruled the next step: the rerank would leave the CPU for a card, through the remote-rerank seam the library already had.
A quantized graph targets an instruction set
The obvious fix was the other lever the release had added: export the graph for AVX2 instead of AVX-512 VNNI, and run the graph the box can actually execute. We benched it on three boxes. The dev laptop turned out to have avx_vnni and so disqualified itself as a control; the public host ran it pre-registered, six cells, every lap recorded; and the mini PC ran it later, where the avx2 export lost too.
| Shape | Threads | The shipped graph | The avx2 export | Speed-up |
|---|---|---|---|---|
| 32 × ~133 tok | 1 | 1,410.5 ms | 1,764.0 ms | 0.80× |
| 32 × ~133 tok | 4 | 522.3 ms | 619.8 ms | 0.84× |
| 8 × ~120 tok | 1 | 340.0 ms | 382.6 ms | 0.89× |
| 8 × ~120 tok | 4 | 130.7 ms | 120.9 ms | 1.08× |
| 32 × 512 tok | 1 | 9,589.7 ms | 10,021.7 ms | 0.96× |
| 32 × 512 tok | 4 | 3,063.3 ms | 3,538.7 ms | 0.87× |
The G3 bench, on the public host, 2026-09-18 21:25:35–21:32:36Z, five laps a cell; the shipped graph is the avx512_vnni export. Slower in five cells of six; the one cell above 1.0× is inside the noise, its laps overlapping the shipped graph's. The pre-registration named its outcomes before the run; the one that fired reads: the GPU seat is the only lever.
The two exports are structurally identical: 157 Constant nodes, 125 Mul, 50 DynamicQuantizeLinear, 50 MatMulInteger in each. They differ in one thing. The avx512_vnni export stores 76 weight tensors as signed 8-bit integers and 6 as unsigned, so its matrix multiplies are U8S8; the avx2 export stores all 82 unsigned, U8U8. That is a choice of which kernel family inside ONNX Runtime's MLAS the CPU will reach, and what each family costs is not read off the flag.
The two exports tied to within 2% on one laptop that has VNNI, ran 1.23–1.80× apart on the mini PC that also has it, and on the public host, which has none, the avx2 export was slower in five cells of six.
The bench that found the split named its candidates and refused to choose between them: the microarchitecture (both hybrid Intel parts from different generations, the mini PC's Alder Lake against the laptop's Arrow Lake-class 24-core), the runtime build each reading used and the thread counts each box was swept at. So the graph was not slow "because of AVX2", and the avx2 export was not the cure: every box that has run it has measured a loss or a wash.
Two int8 exports of the same model do not even agree with each other on scores. On twelve synthetic queries of forty passages on the public host their top eights agreed at 0.875, not one full order of forty matched, and the largest single score gap was 0.668684 against a mean of 0.172166. Any surface that displays a rerank score would show a different number under a different export.
What the missing instruction costs shows in the int8 graph's margin over fp32 on the same box, and the same-night pass measured that margin on all three CPUs, four threads, n = 20.
| Box, at the short shape (32 × ~133 tok) | int8 | fp32 | int8's margin |
|---|---|---|---|
| The first VPS (AVX-512 VNNI) | 138.5 ms | 419.4 ms | 3.0× |
| The mini PC (AVX-VNNI) | 123.0 ms | 255.5 ms | 2.1× |
| The public host (no VNNI) | 386.4 ms | 587.2 ms | 1.5× |
| Box, at the long shape (32 × 512 tok) | int8 | fp32 | int8's margin |
|---|---|---|---|
| The first VPS (AVX-512 VNNI) | 734.5 ms | 1,825.6 ms | 2.5× |
| The mini PC (AVX-VNNI) | 962.0 ms | 1,191.1 ms | 1.2× |
| The public host (no VNNI) | 2,690.6 ms | 2,591.1 ms | 0.96× (fp32 ahead by 1.04×) |
Both tables above are the same-night pass of 2026-09-20, four threads, n = 20 a cell; all four of each box's cells are from the one night, so the drift described there cannot move a margin. The fp32 arm's torch build differs on the mini PC from the two VPSes; the int8 arm is the same ONNX Runtime on all three.
So int8 is not slower than fp32 on the AVX2 box: at the incident's own shape it still wins by 1.5×, and at the full window the two collapse onto each other with fp32 marginally ahead. On the box the graph was built for, the same graph is 3.0× and 2.5× ahead; the difference between a 3× margin and none is the whole bill for a missing instruction set, and no setting on the host pays it.
A reader can ask whether the second box was simply slower at everything, and this page's own fp32 arm answers it. The fp32 arm needs no VNNI, and on the same torch build the two VPSes are 1.40× apart at the short shape (419.4 against 587.2 ms) and 1.42× at the long one (1,825.6 against 2,591.1 ms). The int8 graph is 2.79× and 3.66× apart.
Divide one by the other and the missing instruction is worth about 2.0× and 2.6×, with a generally slower box accounting for the remaining 1.4×. The control runs on a different runtime from the int8 arm, so it bounds the box's general speed rather than measuring ONNX Runtime's; the public host's cells were niced and the first VPS's were not, a per-box factor that cancels in a ratio of ratios.
The mini PC, measured and not chosen
Before the card there was a cheaper candidate: a mini PC already on the workshop's bench, whose mobile 12th-generation Intel CPU carries AVX-VNNI. The same graph, the same loader, the same two shapes, measured on 2026-09-19 between 13:27:54Z and 13:34:11Z.
| Shape | Threads | The mini PC | The public host | The mini PC's lead | Source |
|---|---|---|---|---|---|
| 32 × ~133 tok | 4 | 118.1 ms [116.8–131.0] | 522.3 ms | 4.42× | G5 CPU bench; G3 bench |
| 32 × 512 tok | 4 | 963.6 ms [952.6–971.7] | 3,063.3 ms | 3.18× | Same |
| 8 × ~120 tok | 4 | 27.9 ms | 130.7 ms | 4.68× | Same |
Brackets are the min–max of five laps. The public host's column is its four-thread cell of the day before, 2026-09-18; the incident's 1,792 ms, a one-thread reading, is not a fair divisor, and the third row is not the cell the decision turned on. Measured like for like on 2026-09-20, the first two rows read 3.14× and 2.80× (386.4 / 123.0 and 2,690.6 / 962.0).
A rerank of 118 ms is a fine number for a rules desk (eight threads over four bought 2–4%, 113.5 and 943.5 ms); a rerank of 964 ms is almost a second on every question. The bench's pre-registered rule fired for the mini PC, under 150 ms at the shape the rule named. The decision went the other way on a caveat in the same report: at the 512-token window the same box reads 964 ms, which the rule would have failed. The number that decided it was the one the rule had not named. An operator ruled for the card that afternoon, and the mini PC is documented as the CPU fallback.
The same model on a card already in the house
The inference box already carries the model that writes RuleSage's answers and a much larger generator engine for another product, both on the same 96 GB card the cross-encoder seat joined. A second, smaller card, a 3090 (24 GB), holds an embedder and a 17 GB language model that serves RuleSage's query classifier among other things. A seat there adds no machine and no new failure domain: if that box is down, RuleSage cannot answer anyway.
The seat is the same model at the same pinned revision, loaded in full precision on CUDA with a 512-token window, through the same loader with two arguments moved. The plan's own parity requirement, same model, same revision, same tokenization, is therefore true by construction.
The seat itself is this workshop's own small HTTP service around sentence-transformers: one process, a health route, one port, started by a user-level unit, not a dedicated inference server. The plan weighed a dedicated reranking server against it on paper and chose the service because the parity above holds by construction. No serving layer was benched against it, which leaves the serving layer the one part of this stack the page cannot put a number on.
| Thing | Reading | Source |
|---|---|---|
| The card | RTX PRO 6000 Blackwell (96 GB), 32,405 MiB free when the seat was placed | The ledger's seat-placement row, 2026-09-20 |
| Why not a 3090 (24 GB) | It already held 18,087 of 24,576 MiB, a 17 GB language model and an embedder, with 6,489 MiB free; a model reload there would have pushed the seat out of memory under Restart=always | The ledger's tenancy rows, 2026-09-19 and 2026-09-20 |
| Memory the seat holds | 910 MiB after 25 requests including 32 × 512-token batches (the plan predicted 0.4–0.8 GB; ruled under 2 GB) | The G5 seat bench's install record |
| Time to warm | 5.025 s | The seat's own health route, recorded in gate.json |
| Rerank, 32 × ~133 tok, direct (2026-09-20 17:10Z, n = 5) | 30.1 ms [28.2–31.0] | seat-direct.json |
| Rerank, 32 × 512 tok, direct (same run) | 72.3 ms [65.0–82.7] | Same |
| Rerank, 8 × ~120 tok, direct | 10.4 ms | Same |
| Rerank, 8 × 512 tok, direct | 20.2 ms | Same |
| The card's power, settled idle, every tenant resident | 17.78 W | The same-night pass, Table 5; the cost section below carries the samples and the phases |
| The card's power under a 60 s saturated burst | 168.61 W median | Same, Tables 5 and 6 |
| The seat's own first reading of the card's power, 2026-09-20, on its install run | 85.56 W of a 420 W cap at 35 °C, one sample moments after the gate's 25 requests: a true reading of the post-work state, not an idle level | The G5 seat bench's install record |
The G5 seat bench of 2026-09-20; the 30.1 and 72.3 ms cells were five calls each through the product's own HTTP client, not from the public host.
What a visitor's rerank costs is the seat's cell plus the hop. The public host to the door was measured that night at a median of 100.6 ms per call (p95 101.6, range 99.3–101.8, n = 20); then the door to the seat, which no instrument here timed on its own; the seat bench's own network floors are in the table below. The client is the standard library's urlopen, one call an ask, no session and no connection reuse. The link itself answers a ping in about 48 ms (read on 2026-09-18 and again on 2026-09-19), so 100.6 ms is a fresh connection's cost and not a slow wire.
The plan had predicted about 48 ms for a pooled contract; the client that shipped pools nothing. Whether anything pools between RuleSage and the door is the cheapest open item on this page.
Added up, that is a visitor's whole rerank leg from the app's side: 100.6 ms to reach the door, plus the door's own median of 94 ms on the eight live asks. That is about 195 ms an ask, of which the card's own work is 29 to 67 ms.
Against this box doing it on its own CPU, 386.4 ms at the short shape and 2,690.6 ms at the long one, the leg wins, and wins by 13.8× where it matters. Against a box that did carry VNNI it would not: 138.5 ms locally beats 195 ms at the short shape, and loses to it by 3.8× at the long one. Most of the 195 ms is a hop nothing pools.
An operator installed the seat on 2026-09-20, and found four things by starting and probing it rather than reading it. A TasksMax that starved one library's thread pool. A warm probe that outlived systemd's default start timeout. A refusal path that raised inside itself. And a HEAD response carrying a JSON body. All four are in the checklist below; a refusal path that has only been read is a draft.
Why a door stands in front of the seat
The seat could have been dialled directly from the public host. It is not. Every rerank RuleSage makes is one HTTP call to a role name, role:reranker, on the door, the workshop's own gateway, which its builders call quartermaster. The door runs on the mini PC, a third machine, so a rerank crosses two links: the public host to the door, and the door to the card. It relays the call to whichever member holds that role and writes one record per call: the app, the kind of call, the outcome, the wall time and which seat answered.
A caller that knows a port knows a machine; a caller that knows a role knows nothing that changes when the machine does. The model behind a role swaps with one line in the door's configuration and no change to the app, and the record is the reason this page can print live numbers at all.
Inside the client the degradation is designed and named, and it fails open: when the seat is unreachable the pipeline still answers, with the shortlist it already had. The DegradingReranker catches exactly one exception, RemoteRerankUnavailable, which is transport failure, and serves list(candidates[:k]), the fused top eight in fusion order. That is byte for byte the same dense-plus-fusion top-k that a deployment with reranking switched off would serve, with one line on stderr. A malformed 200 is a contract breach and propagates loudly; a missing seat URL refuses to build the pipeline at all.
The door's budget on every relayed call is 9,000 ms with an attempt count of 1, no retry, so a seat that is down costs an ask up to nine seconds through the door before the shortlist is served unranked. The client's own timeout, which is what governed the eleven inert minutes below (those calls never reached the door), was not read.
Nothing counts a degraded serve today: the only signals are one line on stderr and a silence in the door's rows, with no counter and no alarm. A counter on that one exception path is the cheapest alarm this page can recommend, and it is not built. The rollback is one environment file, kept beside the live one, and the band it returns to was measured before the flip rather than guessed: 3.8–5.1 s on the thread-fixed CPU path, on three real asks.
The door was built for language-model runtimes, and three of its habits were fatal for this seat. It rewrote the body's model field to the role's model name, and this seat refuses any top-level key it does not know, so every relayed call would have returned 400. Its liveness probe hit a path the seat does not have, so the seat would have read as dead. And it assumed the runtime's own URL paths. The fix made the dialect a declared object rather than a single string: each member of the door's role table declares its paths, its probe path, whether the door writes the model field and any legacy keys it accepts.
| Shape | Direct to the seat (17:10:26Z) | Through the door (17:42:11Z) | The difference | Source |
|---|---|---|---|---|
| 32 × ~133 tok | 30.1 ms [28.2–31.0] | 53.9 ms [46.5–65.3] | +23.8 ms | seat-direct.json, seat-through-door.json |
| 32 × 512 tok | 72.3 ms [65.0–82.7] | 62.8 ms [57.9–72.5] | −9.5 ms | Same |
| 8 × ~120 tok | 10.4 ms | 16.4 ms | +6.0 ms | Same |
| 8 × 512 tok | 20.2 ms | 38.1 ms | +17.9 ms | Same |
| Network floor of the bench | 3.1 ms | 5.0 ms | +1.9 ms | G5 seat bench README |
Five reps a cell in both columns. The difference column subtracts two cells taken 32 minutes apart, on a card shared with a large generator engine and two other seats, so it compares two runs and is not a paired measurement. That is why the 512-token row reads backwards, and it is not a pooling win. The door's own wall on the eight live asks later that evening ran 29–133 ms (the table at the end of this page): real questions with real candidate sets, where the synthetic cells above are one fixed shape.
The gate before the flip
No seat goes live here on speed alone. The G5 gate ran on 2026-09-20 before the flip, with three arms so the device and the quantization could be told apart. A is the int8 graph on a CPU: the same graph production runs, executed on the workshop's dev laptop, whose CPU has the VNNI instruction the public host lacks. A′ is the fp32 model on that same CPU. B is the fp32 model on the card.
Thirty queries, forty candidates each (the fixture's width, not a live ask's): the twenty-four best by a word-match ranking plus sixteen at random. The score is overlap@8, the fraction of each arm's top eight the other arm also puts in its top eight, with Spearman rank correlation beside it.
| Pair | What it isolates | overlap@8 mean | Min | Spearman | Bar | Result |
|---|---|---|---|---|---|---|
| Arm A (int8, CPU) vs arm A′ (fp32, CPU) | The quantization | 0.9625 | 0.875 | 0.9899 | — | — |
| Arm A (int8, CPU) vs arm B (fp32, card) | The gate: production's graph against the new seat | 0.9625 | 0.875 | 0.9900 | ≥ 0.90 | Passed, under its own expected band |
| Arm A′ (fp32, CPU) vs arm B (fp32, card) | The device control | 0.9917 | 0.875 | 0.9999 | ≥ 0.99 | Passed |
gate.json, 2026-09-20: verdict: PASS, failures: [], on a substitute fixture (prose this workshop could read without privileges, standing in for the live corpus it could not dump), which licenses a flip and never a graduation; queries_with_identical_scores = 0 on all three pairs, and notes[0] reading "EXPLANATION REQUIRED" for the missed 0.97 band.
The three Min cells read the same because 0.875 is seven of eight, the smallest disagreement overlap@8 can show, so the worst query on every pair differed by exactly one passage. The arm wall medians at forty candidates were 1,888.6 ms for A and 1,705.9 ms for A′ on the laptop's CPU, and 50.2 ms for B on the card: the gate's own timings on its own fixture, not comparable to the serving cells above.
Read the miss, not only the pass. The threshold was 0.90 and the reading 0.9625, so it passed; the gate's own expected band was 0.97, and 0.9625 is under it, so the file demands an explanation.
The three arms supply it. The quantization pair reads exactly the same 0.9625 as the gate pair, and the device control reads 0.9917 with a Spearman of 0.9999. The whole of the gap between the int8 graph and the new seat is the quantization; putting the same fp32 model on a card instead of a CPU moves almost nothing. The old seat was the one further from full precision. That is more useful than a clean 0.97 would have been, because it says which of the two things we changed carries the difference.
For int8 against fp32 we measured order, not values, and no score-delta figure exists for that pair. Order against order is all this gate grades: no arm here carries relevance labels, so nothing on this page says the new seat ranks better than the old one, or than the fusion it re-reads, only that moving it changed almost nothing. Both seats return raw logits, and the batch size alone moves a score by about 0.0000017, which is why the gate grades order.
Four limits. The gate leans on scores being a property of the graph, not of the CPU running it. That is very likely true and it is not evidenced here: arm A ran on a CPU with VNNI, and the fixture carries no score recorded on the public host's own seat, so gate.json marks its cross-box parity check unavailable at queries_checked = 0. The fixture is a substitute: real prose from this workshop's own documentation, 1,200 candidate slots (1,077 distinct passages) drawn from the 7,651 passages that run to 1,800 characters or more. It is not the live rules corpus, because dumping that needs privileges the bench did not take and a disclosure ruling nobody had made that evening.
And the fixture was cut at the 900-character floor rather than the shipped 1,200, so the gate scored slightly shorter passages than a live ask sends. The two budgets are worth 0.9458 against 0.9250 on a different contrast, which is a caution and not a correction factor. Its queries are section headings drawn from the same documents, a seeded sample of thirty, not the question-shaped asks the model was trained on and the app receives. So the gate compares two graphs and two devices on one input distribution, and rehearses neither the corpus nor the query shape a player sends.
The flip
The flip itself was a handful of pastes by an operator on the evening of 2026-09-20: an access key on the public host, two environment lines, the app's own row in the door's table on the mini PC, a door restart and one announced restart of RuleSage at 20:51:24Z. Every receipt was green. Two real asks then took a long first token, and the door had recorded zero rerank rows from the app: no calls, no refusals, one health probe.
RuleSage's service unit carries IPAddressDeny=0.0.0.0/0 ::/0, with an allow list covering loopback, Docker's own network and the box itself. Every connection the app made toward the door was dropped at the unit's cgroup filter, and the client failed open as designed: it waited out its timeout, served the fused shortlist unreranked, and wrote one line to stderr. The only symptom was a slow answer, the symptom the whole project existed to remove.
An operator diagnosed it from outside the box, without privileges, at 20:57Z. The first fix was a drop-in adding the door's address to the allow list, and it was a no-op: the unit's own later drop-in sorts after it alphabetically and opens with an empty IPAddressAllow=, which resets the list.
The second drop-in was given a zz- name so it sorts after all the others, and the unit was restarted at 21:02:27Z, eleven minutes after the flip. The first live row landed 24 seconds later, at 21:02:51.055Z by the door's own day file: kind rerank, outcome ok, wall 173 ms (the seat's warm-up on the first call after a restart), dialled from the door straight to the seat. Four more followed between 21:05:43Z and 21:06:01Z at 38, 51, 52 and 51 ms. At 21:04Z the CPU reranker left the live path, and its environment file was kept for the rollback.
The lesson a builder should carry: a fail-open path turns an outage into latency, and latency is the one symptom nobody escalates; what caught it was a second ledger the app does not own. Every move of a seat here now ends by reading the door's day file, not the app's own green.
Before you move a seat
| Before you move a seat | What bit us | How to check it |
|---|---|---|
| Read the new box's CPU flags, and check what your graph was compiled for | An avx512_vnni graph on a box with AVX2 and no VNNI: 5.4× slower at one thread, 2.79–3.66× at four, same bytes; the avx2 export did not recover it | lscpu, then the Flags line for avx512_vnni or avx_vnni |
| Re-do the thread arithmetic for the new core count | intra_op = 1 sized for 4 vCPUs and 4 callers filled four of eight vCPUs and left the other four idle; 2.7–3.1× from one line | nproc, against your concurrency |
| Check the card's free memory against a reload, not just a load | A 24 GB card already held 18,087 of 24,576 MiB; under Restart=always a reload would have pushed the seat out of memory | nvidia-smi, free memory against the model's footprint twice over |
| Check the seat unit's TasksMax against every library's thread pool | 64 on a 64-core box killed the warm-up inside the tokenizer's rayon pool, and only warm_error on the health route showed it; pinning OMP_, OPENBLAS_, MKL_ and RAYON_NUM_THREADS to 4 fixed it | systemctl show <unit> -p TasksMax |
| Check the start timeout against the warm probe | A 120 s ExecStartPost under a 90 s default reads as start-post operation timed out; TimeoutStartSec=300 | systemctl show <unit> -p TimeoutStartSec |
| Read the client's network wall, not just the seat's | IPAddressDeny on the app dropped every call to the door; fail-open made it a slow answer | systemctl show <unit> -p IPAddressAllow -p IPAddressDeny |
| Read the unit's drop-ins in load order | The first fix was a no-op: a later drop-in opens with an empty IPAddressAllow=, which resets the list. A drop-in that must win is named zz- | systemctl cat <unit> prints them in the order they apply |
| Make the door and the seat agree on a dialect | An unconditional model rewrite would have 400'd every call; a runtime-shaped probe would have marked the seat dead | The member's declared paths, probe path and key handling |
| Probe the refusal path with a raw socket | AttributeError inside the refusal; a HEAD with a JSON body | One line of garbage, and a bare HEAD |
| Measure the hop your app will actually make, on the client it actually builds | Ours reuses no connection: 100.6 ms a call to the door, about twice the pooled figure the plan assumed | One GET /health from the app's box, twenty times, on a fresh connection |
| Read the budget and the attempt count on the degrading path | The door's budget is 9,000 ms with one attempt, so a dead seat costs an ask up to nine seconds; the client's own timeout was never read | The door's own rows carry budget_ms and attempt; the client's timeout is in its source |
| Score the old seat against the new one before you flip, on order and not on values | The graph and the device changed in the same move; three arms told them apart, and the pass missed its own expected band | overlap@8 on your own shortlists, plus a device-control arm |
| Count a degraded serve | Nothing here counts one: the only signals are a stderr line and a silence in the door's rows | A counter on the fail-open exception, and an alarm on rerank rows falling to zero while asks continue |
| Read the door's rows after the flip, not the app's green | Fail-open hid a dropped-packet wall as latency for eleven minutes; nothing errored | The day file, for the rows that should exist |
The incidents are dated in the sections above: the instruction set and the threads on 2026-09-17, the card's tenancy on 2026-09-19, the rest on 2026-09-20. The IPAddressDeny row is now a law in the workshop's ledger, which is the most expensive way to learn it; the budget-and-attempt row and the count-a-degraded-serve row are recommendations this page makes and RuleSage has not yet adopted.
What the seat costs, where we can say
Nothing in this section is a per-ask cost, and that is the receipts' own caveat, not a hedge: the card draws what it draws for every tenant it holds, and the reranker's own share of that is not separable. A card's cost is watts; a CPU seat's cost is the box itself, for the duration of the call. The card's board power was sampled at 1 Hz on the evening of 2026-09-20 with every tenant resident: idle for 30 s, then through a 60 s burst of back-to-back reranks at the long shape. Neither VPS exposes a watt, and the mini PC's draw was not read.
| Seat | What a rerank at the long shape holds | Settled idle | Under a 60 s burst |
|---|---|---|---|
| The card, the seat | 67.2 ms of a card shared with other seats; 79.5 ms a call when saturated, because the seat queues behind itself | 17.78 W (n = 19, range 17.35–22.92) | 168.61 W (p95 177.06) |
| The first VPS's CPU | 4 vCPUs of 4 for 734.5 ms, all of the box | — (no watt reading on a rented vCPU) | — |
| The public host's CPU | 4 vCPUs of 8 for 2,690.6 ms, half the box | — (no watt reading on a rented vCPU) | — |
| The mini PC's CPU | 4 vCPUs of 16 for 962.0 ms, a quarter of the box | — (not read during its bench) | — |
The same-night pass, Tables 5 and 6: whole-card board power, every tenant resident; the burst is n = 70 samples over 775 calls at 12.92 calls a second, VRAM flat. One derived figure: 168.61 W × 60 s ÷ 775 calls = 13.1 J a rerank, 3.63 Wh per thousand, or 11.7 J and 3.24 Wh net of the settled idle; whole-card, at saturation, not a per-ask cost. No burst ran at the short shape.
In the unit a practitioner sizes with, that burst is 775 × 32 = 24,800 pairs in 60 seconds, about 413 pairs a second at the full 512-token window, on a card already holding a generator engine and two other seats. The idle has three phases: about 85 W for at least ten seconds after work (right-censored: nothing here measures where that state ends), a fall over about seven more, and the settled 17.78 W in the table. That tail, not the saturated joule figure, is what a fifteen-calls-a-day seat actually pays, and it is the one number this pass cannot size: the sampler stopped while the card was still high.
Two facts make the table honest. The card was shared with other tenants throughout, so the watts are the card's, not the seat's. And 12.92 calls a second is a saturation rate, not a duty cycle: the plan guessed a sub-100-ms burst a few times a minute, and the door's day file, as read at 22:05Z, holds fifteen rerank calls from RuleSage in the whole of 2026-09-20.
The 3.24 Wh figure is the closest thing here to a marginal cost, and it is still whole-card; the burst's individual laps were not saved, so that row has no distribution, and no dollar figure follows from any of this. The seat is idle almost all the time, which is why the card's settled draw of about 18 W with every tenant resident matters more than its busy draw. That figure is not the seat's alone; no measurement here divides it.
What is still unmeasured
| Open item | State at 2026-09-20 22:05Z | What closes it | Where it is written |
|---|---|---|---|
| The server-side first token after the flip | The client-side reading exists (median 1,559.6 ms, n = 5); the server's own first_delta_ms - search_ms lives behind a login this measurement did not have | The pre-registered query over the app's debug rows, p50/min/max, n ≥ 5 | The ledger's owed-measurement row, 2026-09-20; the same-night pass, §6 |
| Box parity for the gate's arm A | Arm A ran on a CPU with VNNI; the fixture carries no recorded score from the public host's own seat, so queries_checked = 0 | Re-dump the fixture with scores recorded on the public host | gate.json; the ledger's gate row, 2026-09-20 |
| The live-corpus fixture for the gate | The gate ran on a substitute fixture from repository prose, cut at the 900-character floor, with section headings for queries | A dump of the live rules corpus at the shipped budget, with question-shaped queries, which needs privileges and a disclosure ruling | gate.json; the ledger's gate row, 2026-09-20 |
| A relevance measure for the reranker | The gate grades agreement between arms; no arm carries labels, so nothing here says the reranker ranks better than the fusion it re-reads | A labelled set of real asks, and nDCG or MRR over it against the fused order | This page |
| Where production's pairs fall between the two shapes | Settled by arithmetic only; the one SELECT over the live fused set has no recorded result | Run it, and read what the fused set actually contains | The G5 CPU bench, its U-8 note |
| The live unit's per-candidate budget | The deployed engine's default is 1,200 characters with a 900 floor (an operator's read, 22:00Z); whether the live unit overrides it was not read | One read of the unit's environment | An operator's read of the deployed engine's configuration on the public host, 2026-09-20 22:00Z; the unit's environment is root-owned |
| Whether anything pools connections between RuleSage and the door | Unmeasured; 100.6 ms a call is the unpooled floor from the public host, on the standard library's urlopen | One read of the client's connection handling in production, or a session-reusing client and a re-measure | The same-night pass, §4 |
| A degraded serve is counted nowhere | No counter, no alarm; stderr and the door's silence are the only signals | A counter on the RemoteRerankUnavailable path, and an alarm on rerank rows falling to zero while asks continue | This page |
| The pre-flip first-token bands | Stated with their windows; no rows preserved, so no median or percentile can be given retroactively | Cannot be closed; the post-flip band has its median and range, and a week of asks would give it a p95 | The ledger's incident row, 2026-09-17 |
| The concurrent-caller shape and inter_op_num_threads | Open since the adoption bench; every seat figure on this page is one caller at a time | A bench with four callers at once against the seat | The G1 bench |
| A batch-size and candidate-count sweep on the CPU seats | The seat's batch size is 8 and its k-sweep has two points on the card; no CPU sweep of either knob was run | A sweep of the batch-size knob and of k on one CPU box, same passages | This page |
| Other reranker models | The model predates this page; no alternative cross-encoder was benched against it | A bench of two or three current rerankers on the same fixture and the same seats | This page |
| Heat as the card's tenants change | The burst reached 45 °C with the tenants the card had on 2026-09-20 | A thermal reading during a generation with the seat warm | The plan, §1.6 |
The state column is as of 2026-09-20 22:05Z, the end of the same-night measurement window, except the rows this page itself opened, which carry the page's own date. A row closed by a later measurement is marked here with its date, and the reading is added at the foot of the page.
Friday night, after the move
Ask it again now, on the same app. The question leaves RuleSage, crosses to the door and on to the inference box, is read there against the whole fused shortlist by a cross-encoder on a card, and comes back. The adoption bench counted a mean of 52.5 candidates on an ask, at most eighty, and nobody has counted since. On the eight questions asked through the public path on the evening of 2026-09-20, the door clocked that leg at 29–133 ms, median 94 ms. That is for a stage that cost between half a second and three seconds on a rented CPU.
What the player at the table actually waits for is the first word. On the five of those eight that answered it arrived at a median of 1.56 s, on a clock that starts when the question leaves the machine and so includes the search and the trip. Three of the eight abstained, which streams no word; the book was silent, or the shortlist missed, and this page cannot tell which.
The eight were invented questions of this page's shape, not the longest-road question itself, which is one paste away on the live app. The stage that made the table wait is no longer the stage worth measuring, and a second and a half is about as long as it takes to pick the dice back up.
What to take with you
- A quantized graph targets an instruction set, and a box move can lose it. The same bytes ran 5.4× slower at one thread and 2.79–3.66× at four on a host without VNNI. The avx2 export was slower still. Check the flags before the bytes.
- The thread pool is free and it runs out. One line took the public host's long shape from 9,589.7 to 3,063.3 ms and production from 2.5–19 s to 3.8–5.1 s: usable, not good.
- A card already in the house is the cheap seat. The same model in full precision reads 28.9 ms and 67.2 ms at the two shapes and holds 910 MiB. It draws 17.78 W settled, and it adds no failure domain the product did not already have. Every seat figure here is one caller at a time; the seat queues behind itself, and the concurrent shape is unmeasured.
- Gate on order with three arms, and print the miss. overlap@8 read 0.9625: it passed the bar and missed the band, and the arms show the whole gap is the int8 graph, not the device. The fixture was a substitute and arm A ran on the wrong CPU, so this was a flip and not a graduation.
- Fail-open hides a wall as latency. The flip was inert for eleven minutes behind a unit's IPAddressDeny, 20:51:24Z to 21:02:27Z, and nothing errored. Read the second ledger.
How to check our work
The machine this page describes is live. Ask RuleSage a rules question, free and without an account, and open the wrench on the ruling: it shows the search's own timing and the cited pages, and the rerank behind it is the seat this page just moved. The eight asks measured here went down the same public path a reader's does, because the measurement route needs a login this measurement did not have. So they sit on the desk's public record as ordinary rulings: invented for the measurement, unattributed, five answered with citations and three abstained. No visitor's words were sent or read.
The eight asks, paired one to one by timestamp with the door's own rows for them, which is the best evidence on this page that the door writes one record per call:
| Ask | Outcome | First streamed word | Total, question to done | The door's wall |
|---|---|---|---|---|
| 1 | Answered | 1,328.5 ms | 1,831.9 ms | 85 ms |
| 2 | Abstained | — | 1,576.6 ms | 133 ms |
| 3 | Answered | 1,317.9 ms | 1,570.1 ms | 29 ms |
| 4 | Abstained | — | 1,319.2 ms | 36 ms |
| 5 | Answered | 1,569.8 ms | 1,821.9 ms | 103 ms |
| 6 | Answered | 1,569.1 ms | 2,322.3 ms | 34 ms |
| 7 | Answered | 1,559.6 ms | 2,062.3 ms | 131 ms |
| 8 | Abstained | — | 1,563.7 ms | 113 ms |
The same-night pass, Table 3, 21:35:32–21:36:43Z; the first ask's stamp was taken on its way out, so the window opens at its derived start. Timings start when the question leaves the asking machine; the door's wall is its own clock, without the hop. An abstention streams no first word. The eight are contiguous in the day's rows (counted in the cost section); the nearest others are 9 m 43 s before and 13 m 28 s after.
Across the eight, the first word came at a median of 1,559.6 ms (n = 5, 1,317.9–1,569.8 ms), the whole answer at 1,699.2 ms (n = 8, interpolated, 1,319.2–2,322.3 ms) and the door's own leg at 94 ms (n = 8, 29–133 ms).
The benches, the ledger rows and the plan sections named beside the figures above are this workshop's own records and are not published; they are named so a later correction can point at the same reading, not as something a reader can open. What a reader can open is the live app, and the whole of this page's arithmetic is in its own cells. The benches, in the order they ran; the numbering is the plan's, and G4 was opened as a pre-registration for the card and holds no measurement, so it is not cited:
- G1, g1-rerank-int8-2026-08-11: the adoption bench on the first VPS, 2026-08-10; the tag carries the report's date, the day after the run.
- G2, g2-rerank-avx2-2026-09-17: the dev laptop, which disqualified itself as a control.
- G3, 2026-09-18: the AVX2 A/B on the public host, in the repository of the workshop's own status page; that directory is named for the box it ran on, and this shelf does not publish box names.
- G5 CPU, 2026-09-19: the mini PC, in the same measurements, named the same way.
- G5 seat, g5-rerank-seat-2026-09-20: gate.json, seat-direct.json, seat-through-door.json, README.md.
- The same-night pass, rerank-journey-2026-09-20: all four seats at n = 20, the eight asks, the hop, the power samples and a summarize.py that recomputes every median and p95 from the raw laps by one rule.
The code lives in repositories this workshop has not published. The commits a later correction may need are 10a65c8, 776c3ee, 3b71598, e6a7a18 and 2b27f39 in the retrieval library and f53bae6 in the door.
The rest of the seminar
- Three librarians and a careful reader — the prequel: how RuleSage finds the right page, and the reranker's place at the end of the funnel, with the shipped numbers.
- Where your question goes — the same journey for the person asking: which machine sees what, door to door.
- Two new frontier models at the rules desk — the answering model's own exam, and the two frontier models that did not clear it.
- A short history of Mistral — the other model at this workshop's doors, and its history.
The whole shelf holds the rest: the benches behind the claims we publish, failures included. If there is a measurement you want next, say so; the suggestion box is read.
Who ran this, and thanks
The model is cross-encoder/ms-marco-MiniLM-L6-v2, published on Hugging Face by the Sentence-Transformers project, at a pinned revision. On the CPU seats it ran as an int8 graph exported through Hugging Face Optimum and executed by ONNX Runtime. On the card it runs in full precision through sentence-transformers over PyTorch and CUDA (NVIDIA, proprietary, under the CUDA Toolkit EULA), the one closed piece in the stack. The open pieces' licences were not re-read for this page; the ones this shelf has read are on its licences page.
A small human team asked for this page, chose what it would and would not claim, and signed the numbers. A fleet of AI agents read every bench file, ledger row and live record listed above, re-derived every ratio from two cells of the same table and drafted it under that team's rulings.
Corrections and later measurements will be added below, each dated (UTC) with a window at both ends where one applies, each saying in plain words what it counts, and each comparing itself in one sentence to the reading it replaces.