# Arm P, a shared prompt: results (Ollama vs vLLM, release 1)

*Cold scroll: what this file is.* The readings of arm P of the Ollama vs vLLM bench, as registered in `../PREREG.md`
Part 1 (PLAN-v2 §2) and Amendments 3 and 4. Arm P sends one shared prompt to three servers on one EVGA GeForce RTX
3090 24 GB at 300 W: vLLM 0.30.0 with prefix caching on, Ollama 0.34.4 with 16 slots, and Ollama 0.34.4 at its
one-slot default. It sends the prompt in two shapes, "the same question again" and "a new question behind the same
prompt", at 1, 4 and 16 requests at once, and reads how many tokens each server serves from its cache and how long the
first token takes. **Run `ovv-20260930T204352Z-P`**, window 2026-09-30T20:43:52Z → 21:19:31Z (UTC; `rows/P.json`
`utc_start`, `utc_end`), on kit `ae95012965c2` (kit tree `1d89dacb9b00`) under Amendment 4's PREREG (sha256
`withheld-2026-10-07`, `rows/P.json` `prereg.prereg_sha256`). Three boots, 18 cells, three scored runs a cell.

*Where the figures come from.* Every figure here is read from `rows/P.json` by the field named beside it, or from
`rows/RECEIPT`; the one exception, arm C's chat-row figure in P16's caveat, names `../c/rows/C.json`. The full tables
are `../TABLES.md` P1 to P7, printed from the same file by `../tools/r1_tables.py`. The file was read with its gates
first: all 10 gates `ok`, `halted`, `refused` and `held` null, nothing `withheld`,
not a rehearsal, and no `voided_by_gate` field in `P.json`, `parity.json` or `servelines.json`. A condition that only
the private half shows (the serve logs) is named with its source, the read-only check `P16-CHECK.md` in the bench
archive, and never supplies a figure.

## The line

**P-4a and P-4b read REFUTED. Both are first readings, as registered, and both are flagged as an instrument gap**
(IR-1 (a), which D-20261001-016 covers; D-20261001-067). Their second readings, judged at the cache step vLLM itself
reports, come with Amendment 6 and land in PR-B2, printed beside these and never in their place. Every vLLM request
was served 2,496 cached tokens of the 2,533-token shared prompt. Ollama, on both postures, was served 2,558 to 2,559
on the same question again (a 2,560-token prompt) and 2,533 to 2,538 on a new question (P2); on 16 slots its runner
wrote a checkpoint at the shared prompt's end (P6; `verdicts[1].checkpoint_at_prefix_end`). TT reads CONFIRMED of
record, so P's first tokens are scored. No level is VOID; the new question at 4 and 16 at once is FLAGGED on all three
servers, because some answers ended before 256 tokens (P7).

```
RECEIPT ovv P ovv-20260930T204352Z-P boots=3/3 cells=18 prefix_tokens=2533 p4a=REFUTED p4b=REFUTED tt=CONFIRMED void_levels=0 window=2026-09-30T20:43:52Z/2026-09-30T21:19:31Z halted=no
RECEIPT ovv chain ovv-20260930T204352Z-P arms=1/1 void_levels=0 halted=no
```

## The verdicts: first readings, flagged

| Id | Verdict | The reading, in the kit's words and fields | Fields (`rows/P.json`) |
|---|---|---|---|
| P-4a | **REFUTED**, first reading | The same question again, vLLM against Ollama with 16 slots: the kit's cached readings 2,496 and 2,559, against a prefix of 2,533 tokens; "short of the prompt: {'vllm': 2496}" | `verdicts[0]` `cached`, `prefix_tokens`, `why` |
| P-4b | **REFUTED**, first reading | A new question at 4 and 16 at once: the kit's cached readings, vLLM 2,496 and llama.cpp under Ollama with 16 slots 2,538, with a checkpoint at the prompt's end (`true`) | `verdicts[1]` `vllm_cached`, `llama_cached`, `checkpoint_at_prefix_end` |

- **The registered rule** (`verdicts[i].rule`): a cached count within one 16-token block of the shared prompt's
  length is a hit. P-4b reads only new-question requests never sent before on their boot.
- **Why the flag.** Both judges count within a fixed 16-token block, the block size vLLM's cache config reports
  (`engines[0].readback.cache_config.block_size`), and every one of vLLM's 126 scored requests read the same 2,496
  (P2). The instrument reads found that, for this model, vLLM 0.30.0 ends each cache hit on a coarser step that its
  own boot log prints, so a judge set to the 16-token block reads a hit as a miss (INSTRUMENT-READS §2.3; the boot
  line was cleared for use by the lead, D-20261001-067). Amendment 6 reads that step back from vLLM's own lines into
  the rows; nothing is re-run on a card.
- **The words a page may not use.** The kit's `why` for P-4b reads "vLLM missed; llama.cpp hit with a checkpoint at
  the prompt's end". A page does not say that vLLM "missed", or that vLLM does not reuse a shared prompt
  (INSTRUMENT-READS l.566-569). The first readings print as registered, with the flag beside them.

## Every gate

| Gate | Result | Reading (`rows/P.json`) |
|---|---|---|
| tree | pass | estate `ae95012965c2`, kit tree `1d89dacb9b00`, 112 files, clean |
| units, disk, ups | pass | no other bench unit; 98.21 GiB free against a 15 GiB floor; UPS load 7 % at the start |
| R-1, caps | pass | the one card, `EVGA GeForce RTX 3090 24 GB`, at 300 W with persistence Enabled |
| R-2, the weights | pass | the vLLM checkpoint files equal their pins; the Ollama tag's manifest `38044be4f923…`, 4 layers |
| R-4, R-5 | pass | the kit's three ports free; no tenant on the card |
| boots | 3 of 3 | each resident by growth and stopped clean (`engines[i].gone.ok`); both Ollama boots kept all layers on the card |
| raw parity and the serve lines | pass | the parity file and the serve-line read-back record no refused prompt and no mismatch (`parity.json`, `servelines.json`) |
| voids | none | 18 of 18 cells `voided: false`, no cell `flags`, none `thin`; 0 of 72 runs `trial_void`; all 504 requests `ok`, with 0 hidden tokens |
| finish flags | 6 cells | the new question at 4 and 16 on all three servers: some requests ended on `stop` (P7; PREREG Part 1 §2.8: printed, never dropped) |
| stop guard | no stop | no run stopped (`runs[r].stop` null on all 72); the UPS stop never tripped; the highest UPS load in a run was 41 % |

## The readings

The tables are in `../TABLES.md`: P1 (first token), P2 (tokens from cache), P3 (spreads), P4 (the 16-slot cells at
sixteen, run by run), P5 (the prompt-cache lines the rows count), P6 (the verdict rows) and P7 (the flagged cells).
No verdict reads P's first tokens: the shared prompt's rows feed the first-word section, never the crossover headline
(`cells[i].not_headline_why`).

- **vLLM, cache on.** The first token reads 48.102 ms on the same question again at one request and 57.887 ms on a
  new question; at sixteen at once, 313.503 and 344.658 ms (P1; `cells[i].medians.ttft_ms_p50.median`).
- **Ollama with 16 slots, the same question again.** 227.911 ms at one request and 838.946 ms at four. At sixteen at
  once it reads 78,270.429 ms, the median of three runs of 66,329.152, 78,270.429 and 78,606.231 ms (P4).
- **Ollama with 16 slots, a new question at sixteen at once** (`cells[11]`): its first token rose with every run,
  95,395.498, 120,241.433 and 211,619.782 ms (P4; `runs[r].ttft_ms_p50`), and the cell's total-speed spread reads
  0.5987 (`medians.spread`). It prints with its runs beside it, never as the steady figure 120,241.433. No verdict
  reads it.
- **Ollama at its default.** 224.416 ms on the same question again at one request and 432.838 ms on a new question;
  at sixteen at once, 25,584.2 and 20,697.68 ms (P1).

### P16: Ollama with 16 slots, a new question at one request, with its condition

This is the reading the lead named P16 (D-20261001-067). It prints only with its condition, only in the first-word
section and never in a headline (`cells[9].not_headline_why`), and the idle-slot saves print as a count, never as a
cause. Every figure is in `rows/P.json`:

On a new question behind the 2,533-token shared prompt, one request at a time, Ollama 0.34.4 with 16 slots took
10,970.201 ms to its first token (`cells[9].medians.ttft_ms_p50.median`), and its one-slot default took 432.838 ms
(`cells[15]`). Each is the median of three scored runs: 10,989.521, 10,970.201 and 10,955.703 ms, and 432.838, 430.199
and 435.44 ms (`runs[r].ttft_ms_p50`). Each run opened with a one-token warm request outside the timed window, and
each cell ran straight after the same server had answered the same question again sixteen at a time, on the same boot
with no restart between (the cells' run times, `runs[r].utc`, and P16-CHECK §1). Both servers served all 2,533 shared
tokens from cache on every scored request (P2; `runs[r].streams[s].cached_tokens`). Ollama's own clock spent 412.264
to 413.285 ms of the 16-slot wait on the prompt, and 419.212 to 424.476 ms on the default
(`runs[r].streams[0].server_clock.prompt_eval_duration_ms`); the log does not split the rest. The 16-slot server
logged 30 idle-slot saves to its prompt cache in each of those runs, the same count as in its same-question row at
one request earlier on that boot, which read 227.911 ms; the default logged none (P5;
`runs[r].slot_log.counts.prompt_cache_idle_saved`). This file counts the saves; it does not say why.

- **It prints only as 10,970.201 ms with the condition above, never as a lone figure:** arm C's 16-slot chat row read
  10,970.019 ms at its first run (`../c/rows/C.json`), a different reading. It may not be given a cause: not the
  saves, the prompt cache, the slot count or the memory. Every scored request was served 2,533 tokens from cache, and
  Ollama's prompt time sits within 12.212 ms of the default's (arithmetic on the clocks above), so it may not say the
  server missed its cache. The warm request's and the warm-up run's own times are never figures.
- **The same idle-slot count, 30 a run, stands beside 227.911 ms and beside 10,970.201 ms**, so the count does not
  separate the two 16-slot readings (P5). Printed alone beside 10,970.201 against 432.838, "30 against 0" would read
  as a cause, so it is never printed that way.

### Arm P's finish flags (P7)

On the new question at 4 and 16 at once, every server ended some requests on `stop` before 256 tokens: 6 of 12 at four
and 25 of 48 at sixteen, on vLLM, on Ollama with 16 slots and on Ollama's default alike
(`runs[r].finish_flags`). The kit flags each such request and never drops it. A first token comes before any stop, so
P1's first tokens for those cells stand, with the flag beside them; a run's total speed on those cells counts answers
that stopped early, so P3's spreads there are not speed spreads.

## What this does not say

- That vLLM missed its cache, or that it does not reuse a shared prompt (the first readings, flagged).
- Why Ollama with 16 slots takes longer on a new question than its default does; the rows count lines and carry no
  split of the time.
- The second readings of P-4a and P-4b: they land in PR-B2 with Amendment 6.
- One card, one model, one 2,533-token prompt, the versions benched.

## Files

- `rows/`: the kit's outward output, every byte through its pen (a card label and never a UUID or a bus id; box-free
  run ids; UTC). sha256:
  - `P.json` `withheld-2026-10-07` (3,329,198 B)
  - `P.power.csv` `b5d48abab439294cf2eeba5587ec592d2d623be0b8ac412389da6b4ca6edc0c8` (the 2 Hz board samples)
  - `parity.json` `1f43cf52e1b0ce65f13a2914834f2701d14c451f681cedf9ac2c3b8018e92345`
  - `servelines.json` `withheld-2026-10-07`
  - `RECEIPT` `98afb6dc7eeeb533217447a87454ada665878e931a4fa4aaaeb1b3f6476b5eb3`
- Each file is byte-equal to the copy archived off the repo after the day (the day report, DAY §2.3). The private
  half (the card's UUID map, the raw serve logs, the slot-log lines, the run log, the ledger) stays off the repo; the
  raw slot-log lines are never printed.
- Sources named above, all in the bench archive's `plans-2026-09-28/ollama-vs-vllm/` (off the repo): the day report
  `DAY-2026-09-30.md` (§2.3), the instrument reads `INSTRUMENT-READS-2026-09-30.md` and the P16 check
  `bench-recon-2026-10-01/P16-CHECK.md`. Rulings: D-20261001-016 and D-20261001-067 in `DECISIONS.md`.
