# Arm C, the crossover: results (Ollama vs vLLM, release 1)

*Cold scroll: what this file is.* The readings of arm C of the Ollama vs vLLM bench, as registered in `../PREREG.md`
Part 1 (PLAN-v2 §2) and Amendment 5. Arm C is the crossover arm: vLLM 0.30.0 against Ollama 0.34.4 (with 16 slots,
and at its one-slot default), on two prompts at 1, 2, 4, 8 and 16 requests at once, on one EVGA GeForce RTX 3090 24 GB
at 300 W with persistence on. **Its reading of record is run `ovv-20261001T133502Z-C`**, window 2026-10-01T13:35:02Z
→ 14:29:21Z (UTC; `rows/C.json` `utc_start`, `utc_end`), on kit `1d1794cfa34f` (kit tree `cd0faba414db`) under
Amendment 5's PREREG (sha256 `withheld-2026-10-07`, `rows/C.json` `prereg.prereg_sha256`). Amendment 5 was pushed at
13:30:15Z, before the run's first model call (the day report, DAY B.3). Five boots, 30 cells and 181 scored runs. Arm
C's first run, refused at its second boot on 2026-09-30 (`ovv-20260930T201255Z-C`), is kept in
`refused-2026-09-30/rows/` as the finding it is, never as a result (the last section). Arms D, P, Q, A and G ran on
2026-09-30, a day earlier, on another kit commit ("Two dates, two kits", below).

*Where the figures come from.* Every figure here is read from `rows/C.json` by the field named beside it, or from
`rows/RECEIPT`. The full tables are `../TABLES.md`, printed from the same file by `../tools/r1_tables.py` (no figure
there is typed by hand). Two readings come from other row files: 2026-09-28's cold-compile P-BOOT and C-2
(`../dtt/rows/D.json`, landed in #191), and P-6's verdict, which sets the label every speed row carries
(`../qag/rows/Q.json`, in this PR). A condition that only the private half shows (the serve logs) is named with its
source, the hand-read prep `D-C-FLAGS.md` in the bench archive, and never supplies a figure. Every file was read
with its gates first: all 10 gates `ok`, `halted`, `refused` and `held` null, nothing `withheld`, not a rehearsal,
and no `voided_by_gate` field anywhere in `C.json`, `parity.json` or `servelines.json`.

## The line

**P-3 reads UNDECIDED: the crossover read 2, where the bet was 4 ("undecided if 2 or 8").** At sixteen requests at
once, vLLM 0.30.0 totals 506.829 tok/s and Ollama 0.34.4 with 16 slots 142.327 (`cells[4]`, `cells[16]`
`medians.aggregate_wall_tok_s.median`). **P-1** reads CONFIRMED (0.9728), **P-2** REFUTED (2.0745), **P-5** REFUTED on
both halves and **C-1** CONFIRMED at every level. **P-BOOT and C-2** read CONFIRMED on this run's warm compile cache,
and REFUTED on 2026-09-28's first boot, which compiled cold; both conditions print on every P-BOOT and C-2 line this
file and `../TABLES.md` write (OF-RECORD, D-20261001-067). The RECEIPT lines below are the kit's own, quoted
verbatim, so the first line's `pboot` and `c2` carry the warm reading alone. No level is VOID, and no cell is voided,
flagged or thin.

```
RECEIPT ovv C ovv-20261001T133502Z-C boots=5/5 cells=30 scored=181 void_levels=0 refused_prompts=none p1=CONFIRMED p2=REFUTED p3=UNDECIDED p5=REFUTED c1=CONFIRMED pboot=CONFIRMED c2=CONFIRMED aba_drift_L1=-0.0067 aba_drift_L16=0.0025 window=2026-10-01T13:35:02Z/2026-10-01T14:29:21Z halted=no
RECEIPT ovv chain ovv-20261001T133502Z-C arms=1/1 void_levels=0 halted=no
```

P-6 read REFUTED in arm Q (`../qag/RESULTS.md`), so every speed row here carries the kit's words **answers differ at
this level** (`../qag/rows/Q.json` `sixway.below_words`; PREREG Part 1 §2.8). The two servers do not give the same
answers on this model, so a speed row compares servers that answer differently.

## The verdicts

| Id | Verdict | The reading | Fields (`rows/C.json`) |
|---|---|---|---|
| P-1 | **CONFIRMED** | One request, decode: vLLM 82.693 tok/s, Ollama with 16 slots 85.007; ratio 0.9728, within ±0.1 | `verdicts[0]` `a`, `b`, `why` |
| P-2 | **REFUTED** | One request, first token: Ollama with 16 slots 441.866 ms over vLLM 212.999 ms is 2.0745, over 2 | `verdicts[1]` `num`, `den`, `why` |
| P-3 | **UNDECIDED** | The crossover read 2 (the bet: 4; undecided if 2 or 8). Level 2's ratio is 1.2077, with a spread of 0.0544 | `verdicts[2]` `crossover`, `why`, `per_level[1]` |
| P-3, long-write (printed beside) | **UNDECIDED** | Level 1's spread, 0.2499, is wider than the band, 0.2, before any decided crossover | `verdicts[3]` `why` |
| P-5 | **REFUTED** | Long-write at 16, each run's p95 first token: vLLM 17,325.908 to 17,333.515 ms (bar 1,500); Ollama with 16 slots 107,307.61 to 107,685.667 ms (bar 4,000) | `verdicts[4].parts` |
| C-1 | **CONFIRMED** | Ollama's default is a queue: its total at 2, 4, 8 and 16 over its total at 1 reads 0.9880, 0.9876, 0.9873 and 0.9869 | `verdicts[5].parts.<level>.why` |
| P-BOOT | **CONFIRMED** warm; **REFUTED** cold | 17.07× on this run's headline boot (warm compile cache); 14.78× on 2026-09-28's first boot (cold) | `verdicts[6]`; `../dtt/rows/D.json` `verdicts[1]` |
| C-2 | **CONFIRMED** warm; **REFUTED** cold | 0.9157 of the card at rest (22,505 MiB) warm; 0.8483 (20,847 MiB) cold | `verdicts[7]`, `engines[0].memory.at_rest_mib`; `../dtt/rows/D.json` `verdicts[2]`, `engines[0].memory.at_rest_mib` |

- **P-1 is read on the versions benched, beside arm S's reading of record**, never in its place. Arm S read P-1
  CONFIRMED on the proven pair, vLLM 0.27.1 and Ollama 0.32.13 (`../s/S-RESULTS.md`). P-1 also prints the kit's
  `beside` (`verdicts[0].beside`): "≈7.6 GiB (vLLM) vs ≈6.5 GiB (Ollama) per single-stream step, +17 %
  (CRITIQUE-method M-6)".
- **P-2 reads the grid, never the chat row.** The 16-slot chat row (`cells[22]`, variant `chat`) is a different cell
  at the same posture, prompt and level, and it is not P-2's input (B.6 item 2, below).
- **P-3 decides at level 2 first.** The kit walks the levels in order, and level 2 is the first whose ratio is at or
  over 1.2 with its spread inside the band. Level 2's ratio sits close to 1.2; read either side of 1.2, P-3 stays
  UNDECIDED, because the walk would then reach level 4, which is undecided (D-C-FLAGS §4). So the clause is "the
  crossover read 2", with 1.2077 beside it.
- **TT is not held.** C reads TT's verdict through arm D (`d_reading.tt`): "of record: max 3.111 <= 50
  (ovv-20260928T172031Z-D); replicated in D: max 3.119 <= 50". So C's first tokens are scored
  (`cells[i].ttft_scored` true on all 30 cells), and `calibration_blocks_C` is empty.

## Every gate

| Gate | Result | Reading (`rows/C.json`) |
|---|---|---|
| tree | pass | estate `1d1794cfa34f`, kit tree `cd0faba414db`, 114 files, clean, with this kit's own passing stage receipt (`gates[0].reading`) |
| units | pass | no other bench unit active |
| disk | pass | 91.47 GiB free, against a 15 GiB floor |
| ups | pass | load 6 % at the start |
| R-1 | pass | the one card, `EVGA GeForce RTX 3090 24 GB`, read by its label |
| caps | pass | 300 W, persistence Enabled (its default limit 350 W) |
| R-2, vLLM checkpoint | pass | every checkpoint file equal to its pin |
| R-2, Ollama tag | pass | manifest `38044be4f923…`, 4 layers |
| R-4 | pass | the kit's three ports free |
| R-5 | pass | no tenant on the card |
| boots | 5 of 5 | each resident by growth, each stopped clean back to its 1 MiB baseline (`engines[i].gone.ok`); both Ollama boots kept all layers on the card (`engines[i].readback.all_layers`) |
| S-1, the serve lines read back | pass on every boot | the default-batch boot read 2,048 batched tokens from the engine line's compile ranges (`engines[1].readback.max_num_batched_tokens_read.from`), the read #204 added after S-1 refused this posture on 2026-09-30 |
| voids and flags | none | 30 of 30 cells `voided: false`, no `voids`, no `flags`, none `thin`; 0 of 211 runs `trial_void`; all 886 requests, warm-ups included, `ok` on `length` at 256 tokens with 0 hidden tokens |
| stop guard | no stop | no run stopped (`runs[r].stop` null on all 211); the UPS stop never tripped; the highest UPS load in a run was 40 % |

## The readings

The tables are in `../TABLES.md`: C1 to C7 for series-p512, C8 to C12 for the long prompt, C13 and C14 for the
rows outside the grid, C15 and C16 for memory and the two compile conditions, C17 and C18 for the A-B-A check and the
verdict rows. What they say, with each figure's field:

- **Total speed, series-p512** (C1; `cells[i].medians.aggregate_wall_tok_s.median`). vLLM 0.30.0 grows with every
  level, from 77.653 tok/s at one request to 506.829 at sixteen. Ollama with 16 slots reads 74.358 at one request and
  117.621 at two; from four on its medians read 147.955, 136.181 and 142.327, but its cells at 4 and 8 are wider than
  the 0.2 band (spreads 0.2973 and 0.2352), so those two levels are undecided (C2, C3, C4). Ollama at its one-slot
  default stays at its level-1 total: 74.545 at one request and 73.565 at sixteen (C-1).
- **Per request** (C5 to C7). At one request the decode rates are 82.693 (vLLM) and 85.007 tok/s (16 slots), and
  the first tokens 212.999 and 441.866 ms. At sixteen, vLLM decodes each request at 41.428 tok/s and Ollama with 16
  slots at 24.449; the median first token is 1,809.039 ms on vLLM and 18,179.306 ms on 16 slots
  (`medians.ttft_ms_p50.median`).
- **The long prompt** (C8 to C12). vLLM's total grows from 59.884 tok/s at one request to 181.513 at sixteen; Ollama
  with 16 slots reads 24.282 at one and 35.373 at sixteen. P-5's bars are far from both: the p95 first token at
  sixteen is 17,327.116 ms on vLLM and 107,377.256 ms on 16 slots (`medians.ttft_ms_p95.median`).
- **The default-batch row** (C13, C14; `cells[11]`, variant `default-batch`). vLLM at its own default batch of 2,048
  tokens, long-write at sixteen: 186.538 tok/s total and a p95 first token of 16,395.26 ms, with the longest queue
  wait of a scored request 13,593.772847001674 ms. Its KV cache holds 37,541 tokens (7.33 requests of 5,120 at once).
  It is one labelled row, never a headline field, and it prints with its condition (B.6 item 1, below).
- **Memory** (C15). At rest, the headline boots grew 22,505 MiB, the default-batch boot 22,475, Ollama with 16 slots
  16,389 and Ollama's default 8,761 (`engines[i].memory.at_rest_mib`). Under sixteen requests, from the 2 Hz board
  samples less each boot's 1 MiB baseline (arithmetic), the headline read 22,507, the 16-slot boot 16,393, the
  default 8,769 and the default-batch boot 22,769 MiB.
- **A-B-A** (C17; `aba.by_level`). The anchor boot repeats the first block's nonces at one and sixteen requests:
  drift -0.0067 at one and 0.0025 at sixteen. It is printed beside the readings and corrects nothing.
- **Joules per 1,000 tokens are not printed here.** They are arithmetic over `runs[r].power["EVGA GeForce RTX 3090 24
  GB"].busy.power_w.mean`, `runs[r].wall_s` and `runs[r].completion_tokens_total` (the kit's rule
  `rules.j_per_1k_busy_w_x_wall`), with `rows/C.power.csv` as the 2 Hz source. A page that prints them computes them
  with their provenance and labels them arithmetic, not registered, as arm S's row is.

## P-BOOT and C-2: two compile conditions, both printed (OF-RECORD)

| Condition | File | Start to ready | KV blocks | KV tokens | P-BOOT | At rest | C-2 |
|---|---|---:|---:|---:|---|---:|---|
| Cold: 2026-09-28, this posture's first boot, which compiled its graph | `../dtt/rows/D.json` | 144.35 s | 10,718 | 75,691 | 14.78×, **REFUTED** | 20,847 MiB, 0.8483 | **REFUTED** |
| Warm: 2026-10-01, this run's headline boot, on the compile that first boot saved | `rows/C.json` | 47.58 s | 12,377 | 87,407 | 17.07×, **CONFIRMED** | 22,505 MiB, 0.9157 | **CONFIRMED** |

Fields: `engines[0].boot_s`, `.readback.cache_config.num_gpu_blocks`, `.readback.boot.kv_cache_tokens` and
`.memory.at_rest_mib`, and the P-BOOT and C-2 rows of `verdicts`, in each file (`../TABLES.md` C16). The bars are 16×
for P-BOOT and 0.9 of the card for C-2.

- **The compile state decides both verdicts, not noise.** A first boot compiles the model's graph and saves it; a
  later boot loads the saved compile (the serve logs, read in D-C-FLAGS §1). The loading boot's memory profile peaks
  lower (0.85 GiB here, `engines[0].readback.boot.peak_activation_gib`, against 2.47 GiB on the cold boot), so vLLM
  gives the KV cache more blocks. Neither reading sits near its bar.
- **No reading of record is registered for either.** The lead's ruling OF-RECORD (D-20261001-067) prints both
  conditions on every P-BOOT and C-2 line, with the condition named. The cold reading is the kit's on 2026-09-28's
  arm D, which refused at R-F1 for a cause that does not touch vLLM's boot; it stands as the kit's reading
  (D-20260929-002). D's re-run on 2026-09-30 reads both again on a warm cache; it lands in PR-A, not here.
- **The kit's `cold_compile` flag is never the condition.** It reads false on both boots, because the shared cache
  already held files from earlier boots (`engines[0].internals.compile_cache_files_before`).
- **An at-rest figure from one condition never sits beside an under-load figure from the other.** 09-28's cold
  20,847 MiB beside this run's warm 22,507 under sixteen requests would read as memory grown under load; the gap is the
  compile condition (D-C-FLAGS §5).

## The hand-read's six flags (DAY B.6), as conditions

None of the six is a void, a refusal or a flag in the rows, and none moves a registered verdict (D-C-FLAGS, each
section's "Does it change a verdict?"). Each prints only with its condition.

1. **The default-batch boot loaded the compile that the refused 2026-09-30 boot of the same posture saved.** It is a
   warm load, so its KV pool is larger than that cold boot's. Its figures (C14) print with the label "a warm load of
   the compile the refused 2026-09-30 boot saved". On 12,355 blocks here and 12,377 on the headline boot
   (`readback.cache_config.num_gpu_blocks`), the gap between their KV tokens is the batch, not the pool.
2. **Ollama with 16 slots, the chat row's first token** (`cells[22]`): median 9,518.334 ms, falling with every scored
   run from 10,970.019 to 6,639.418 ms (C11), with a decode of 85.384 tok/s beside the grid's 85.007. It ran after the
   long-write cells on the same runner. It fills no marker, is never a headline field and is not P-2's input; this
   file does not attribute it.
3. **Ollama with 16 slots, long-write at one request** (`cells[17]`): the first token rose with every scored run, from
   6,460.522 to 9,262.838 ms (C11), and the cell's spread is 0.2499; at two requests (`cells[18]`) the spread is 0.2256.
   The median, 7,460.251 ms, never prints as a steady figure without its rise. It is the row P-3 prints beside.
4. **Ollama with 16 slots at 4 and 8 at once** (`cells[14]`, `cells[15]`): spreads 0.2973 and 0.2352, wider than the
   0.2 band, so P-3's per-level rows read `decided: false` there (C2). Their runs split into two decode speeds (C4).
   "Flat from four on" is about the medians, 147.955, 136.181 and 142.327, never a measured ceiling.
5. **Memory under sixteen requests is a different read from memory at rest** (C15): 2 Hz samples less the baseline,
   22,507 and 16,393 MiB for the headline and the 16-slot boots, against 22,505 and 16,389 at rest.
6. **Release 1 spans two UTC dates and two kits** (the next section).

## Two dates, two kits

| Rows | Ran | Kit commit | Kit tree | PREREG sha256 |
|---|---|---|---|---|
| Arms D, P, Q, A and G | 2026-09-30 | `ae95012965c2` | `1d89dacb9b00` | `withheld-2026-10-07` (Amendment 4) |
| Arm C, reading of record | 2026-10-01 | `1d1794cfa34f` | `cd0faba414db` | `withheld-2026-10-07` (Amendment 5) |

The fields are each row file's `kit.git_sha`, `kit.kit_tree` and `prereg.prereg_sha256`; the nonce seed reads 20260929
in all of them. P's, Q's and A's row files are in this PR. D's re-run lands in PR-A (#230,
`../dtt-2026-09-30/rows/D.json`) and G's in PR-B2 (`../qag/rows/G.json`); until they land, their fields here are read
from those files' archived copies. Between the two kit commits, the runtime changes are #204's (S-1 reads vLLM's own
batch from the lines that carry it) and #212's (a run refuses without a passing stage receipt for its own kit and
tree); the judges (`prereg.py`), the gates (`gates.py`) and the client (`client.py`) show no diff (`git diff ae950129
1d1794cf -- bench/kits/vllm-seats`; D-C-FLAGS §6). C's verdicts read only C's own cells, and TT through arm D. Any
table that mixes C's figures with D's, P's, Q's, A's or G's labels C's date.

## The refused run of 2026-09-30, as the finding it is

`refused-2026-09-30/rows/` holds the rows of run `ovv-20260930T201255Z-C`: window 2026-09-30T20:12:55Z → 20:24:33Z
(`C.json` `utc_start`, `utc_end`), kit `ae95012965c2`, under Amendment 4's PREREG (`withheld-2026-10-07`). Its receipt,
verbatim from `refused-2026-09-30/rows/RECEIPT`:

```
RECEIPT ovv C ovv-20260930T201255Z-C refused=S-1 halted=no
RECEIPT ovv chain ovv-20260930T201255Z-C arms=1/1 void_levels=0 halted=C exited rc 2 (its refusal or error is in rows/C.json)
```

- **The refusal, verbatim** (`refused`): "REFUSED: vllm-0.30.0-default-batch read back max_num_batched_tokens=None and
  the table registers 2048; vllm stopped, readings kept". Its mismatch is `["max_num_batched_tokens", null, 2048]`.
- **What ran.** Boot 1, the headline posture, ran its 11 cells; boot 2, the default-batch posture, read back as
  resident and was refused at S-1, the serve-line read-back, before any cell ran; boots 3 to 5 never started. Every
  gate passed, and no stream failed.
- **The finding.** The kit's S-1 read-back had no line to read vLLM 0.30.0's own default batch from, for a posture
  that leaves it unset. #204 made S-1 read it from the engine line's compile ranges, Amendment 5 registered the fix,
  and arm C ran again whole on the fixed kit: the reading of record above (Amendment 5 items 1 to 13).
- **What it is not.** No arm is re-run to a pass, and these rows are not C's reading. No figure from
  `refused-2026-09-30/rows/` is a result, and none fills a marker. Their 11 cells were cut from the run ledger before
  C ran again, with a copy kept (Amendment 5 item 12; DAY B.3).
- **Why it still matters to the reading of record.** Its boot 2 compiled cold and saved the compile that the
  reading of record's default-batch boot loaded (B.6 item 1).

## What this does not say

- One card at one cap, one model and the two versions benched. P-1 on the proven pair is arm S's reading, not this.
- P-3 is UNDECIDED. Nothing here says "vLLM pulls ahead at two"; the clause is "the crossover read 2".
- The chat rows and the default-batch row are labelled rows, never headline fields.
- Why the 16-slot runner's first tokens rose or fell (B.6 items 2 and 3) or its decode split (item 4): the rows count,
  and no line splits the time.
- P-BOOT and C-2 under one condition alone.
- Joules per 1,000 tokens (above).
- Arm A's private file was not opened for this file, and no arm A or arm G figure is read here.

## Files

- `rows/`: the kit's outward output for the reading of record, every byte through its pen (a card label and never a
  UUID or a bus id; box-free run ids; UTC). sha256:
  - `C.json` `withheld-2026-10-07` (7,418,878 B)
  - `C.power.csv` `daab6b0a88455d66c95ceabed609382bb438e309d441bf93743605de8276f279` (the 2 Hz board samples)
  - `parity.json` `7742ee04c1ce7745ee83319b41cdd90209fdae08c56d2c77fce3e3772588e8ac`
  - `servelines.json` `withheld-2026-10-07`
  - `RECEIPT` `fc88be6318fd20bdddff63a13e217227883ed8a3fb00cb069944f05a58c5bee3`
- `refused-2026-09-30/rows/`: the refused run's outward rows, kept as its finding. sha256:
  - `C.json` `withheld-2026-10-07`
  - `C.power.csv` `withheld-2026-10-07`
  - `RECEIPT` `withheld-2026-10-07`
- Each file is byte-equal to the copy archived off the repo after the run (the day report, DAY B.7 and §2.2). The
  private half (the card's UUID map, the raw serve logs, the run log, the ledger) stays off the repo.
- Sources named above, all in the bench archive's `plans-2026-09-28/ollama-vs-vllm/` (off the repo): the day report
  `DAY-2026-09-30.md` (its 2026-10-01 addendum, B.1 to B.11) and the hand-read prep
  `bench-recon-2026-10-01/D-C-FLAGS.md`. Rulings: D-20260929-002, D-20261001-041 and D-20261001-067 in `DECISIONS.md`.
