# Ollama and vLLM on one RTX 3090 — data kit

**What this directory holds.** The receipts behind the exhibit *Ollama and vLLM on one RTX 3090*
(https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/) that our rules could clear for
release, and the name of every file they could not. The exhibit is the bench that put Ollama 0.34.4 and
vLLM 0.30.0 on one EVGA GeForce RTX 3090 XC3 Ultra (24 GB) at 300 W with the same 4-bit Gemma 4 12B weights,
from one request at a time to sixteen released together, and the first reading on the older pair it was
proven on, vLLM 0.27.1 and Ollama 0.32.13.

The arms, with their windows at both ends (UTC):

| arm | what it read | opened | closed |
|---|---|---|---|
| S | one request at a time on the older pair, the first reading | 2026-09-28 11:29:59 UTC | 2026-09-28 11:33:35 UTC |
| D and TT | the first full pass on the versions benched: memory, start-up, five roads to the first token, and the load-step test | 2026-09-28 17:20:32 UTC | 2026-09-28 17:30:20 UTC |
| D's re-run and TT's replication | the same pass with the client fixed | 2026-09-30 20:00:02 UTC | 2026-09-30 20:07:18 UTC |
| P | a shared prompt, the same question again and a new one behind it | 2026-09-30 20:43:52 UTC | 2026-09-30 21:19:31 UTC |
| Q | the six-way: are the answers the same? | 2026-09-30 21:27:35 UTC | 2026-09-30 21:31:13 UTC |
| A | which documented routes answered without a key, cut to the twelve rows the page prints | 2026-09-30 21:31:19 UTC | 2026-09-30 21:35:27 UTC |
| C | one to sixteen requests at once, the reading of record | 2026-10-01 13:35:02 UTC | 2026-10-01 14:29:21 UTC |

Arm G (what left the machine) is not in this kit: its rows join the page in a dated update, and this README
will say then what joins the kit.

**Licence: CC BY 4.0.** Take these rows, re-plot them, check our arithmetic, publish what you find. Attribution:
strata→signal research, research.strata2signal.com. If you find an error in any of this, we want to hear about
it: hello@strata2signal.com.

Machine readers start at `index.json`, which lists every file in this kit but three, each with its byte
count and its sha256; the three it leaves out are itself, `provenance.json` and this directory's index page.
`provenance.json` is the other half: it names every rule applied to every file, with the count of times each
one fired, and says whether any rule touched each file. `KIT-NOTE.md` says what is not here and why.

## What is in here

| path | what it is |
|---|---|
| `kit/prompts/MANIFEST.json`, `kit/prompts/frame/gemma4-canonical.json`, `kit/prompts/long-prefill.txt`, `kit/prompts/long-write.txt`, `kit/prompts/prefix-questions.json`, `kit/prompts/series-p512.txt`, `kit/prompts/short-chat.txt` | The frozen prompts the bench sent, byte for byte, arm P's questions and the frame each prompt was rendered with, and the manifest that names each prompt and its sha256. |
| `kit/rows/S/ovv-s-20260928T1130Z-ollama-scored-board.csv`, `kit/rows/S/ovv-s-20260928T1130Z-ollama-scored.jsonl`, `kit/rows/S/ovv-s-20260928T1130Z-vllm-scored-board.csv`, `kit/rows/S/ovv-s-20260928T1130Z-vllm-scored.jsonl` | Arm S's rows for both servers, the harness's own output: every scored request's token counts, first-token wait, decode rate and the server's own timers where it reports them, with the board's power at 2 Hz beside them. |
| `kit/rows/C/C.power.csv`, `kit/rows/C/RECEIPT`, `kit/rows/C/parity.json`, `kit/rows/D-2026-09-28/D.power.csv`, `kit/rows/D-2026-09-28/RECEIPT`, `kit/rows/D-2026-09-28/parity.json`, `kit/rows/D-2026-09-30/D.power.csv`, `kit/rows/D-2026-09-30/RECEIPT`, `kit/rows/D-2026-09-30/parity.json`, `kit/rows/P/P.power.csv`, `kit/rows/P/RECEIPT`, `kit/rows/P/parity.json`, `kit/rows/Q/Q.power.csv`, `kit/rows/Q/parity.json` | For arms D (both runs), C, P and Q: each run's board readings at 2 Hz and its token-ID parity check, the check that stops a run whose prompts the two servers turned into different token IDs; for arms D, C and P, the `RECEIPT` line each run printed. |
| `kit/records/C-RESULTS.md`, `kit/records/D-2026-09-30-RESULTS.md`, `kit/records/P-RESULTS.md` | The results files of the arms they name, as the bench wrote them after each run, apart from the digests withheld in them. |
| `kit/rows/A/A.routes-cut.json` | Arm A's census, cut to the twelve route rows the page prints, in the page's order and labelled as a cut. The rest of arm A's census is withheld whole, by ruling. |
| `kit/client/client.py`, `kit/client/engine.py`, `kit/client/engine_llamaserver.py`, `kit/client/engine_ollama.py`, `kit/client/framing.py`, `kit/client/gates.py`, `kit/client/nonces.py`, `kit/client/parity.py` | Client modules from the harness that drove the servers: the timing sweep, the engine protocol and the drivers listed, the parity check, the framing that renders each prompt once, the gates that read every run after it ends and void a level that breaks one of our rules, and the seeded tag each request opens with, the same on every server, so that no prompt cache can answer one request from another's. They are a reading set, not a program you can run: they import modules this kit does not carry (`pen.py`, `prereg.py`, `seat.py`, `seatlib.py`, `servelines.py`). |
| `KIT-NOTE.md` | What is not here, and why. |

## Not in this kit

Nothing in this kit was edited by hand. A rule either cleared a file or it did not; a file no rule could
clear is absent whole, and this list names it. The serve lines are printed on the page as code.

`bench/kits/vllm-seats/engine_vllm.py` — absent whole: the rules could not clear it for release, and a hand-edited copy would not be a record.

`bench/ollama-vs-vllm-2026-09-28/dtt/rows/D.json` — absent whole: the rules could not clear it for release, and a hand-edited copy would not be a record.

`bench/ollama-vs-vllm-2026-09-28/dtt/rows/servelines.json` — absent whole: the rules could not clear it for release, and a hand-edited copy would not be a record.

`bench/ollama-vs-vllm-2026-09-28/dtt-2026-09-30/rows/D.json` — absent whole: the rules could not clear it for release, and a hand-edited copy would not be a record.

`bench/ollama-vs-vllm-2026-09-28/dtt-2026-09-30/rows/servelines.json` — absent whole: the rules could not clear it for release, and a hand-edited copy would not be a record.

`bench/ollama-vs-vllm-2026-09-28/c/rows/C.json` — absent whole: the rules could not clear it for release, and a hand-edited copy would not be a record.

`bench/ollama-vs-vllm-2026-09-28/c/rows/servelines.json` — absent whole: the rules could not clear it for release, and a hand-edited copy would not be a record.

`bench/ollama-vs-vllm-2026-09-28/p/rows/P.json` — absent whole: the rules could not clear it for release, and a hand-edited copy would not be a record.

`bench/ollama-vs-vllm-2026-09-28/p/rows/servelines.json` — absent whole: the rules could not clear it for release, and a hand-edited copy would not be a record.

`bench/ollama-vs-vllm-2026-09-28/qag/rows/Q.json` — absent whole: the rules could not clear it for release, and a hand-edited copy would not be a record.

`bench/ollama-vs-vllm-2026-09-28/qag/rows/RECEIPT` — absent whole: the rules could not clear it for release, and a hand-edited copy would not be a record.

`bench/ollama-vs-vllm-2026-09-28/qag/rows/servelines.json` — absent whole: the rules could not clear it for release, and a hand-edited copy would not be a record.

`bench/ollama-vs-vllm-2026-09-28/PREREG.md` — absent whole: the rules could not clear it for release, and a hand-edited copy would not be a record.

`bench/ollama-vs-vllm-2026-09-28/s/S-RESULTS.md` — absent whole: the rules could not clear it for release, and a hand-edited copy would not be a record.

`bench/ollama-vs-vllm-2026-09-28/dtt/RESULTS.md` — absent whole: the rules could not clear it for release, and a hand-edited copy would not be a record.

`bench/ollama-vs-vllm-2026-09-28/qag/RESULTS.md` — absent whole: the rules could not clear it for release, and a hand-edited copy would not be a record.

## Records the page names that sit outside this kit's record

Arm S's independent check (arm S: 2026-09-28, 11:29:59–11:33:35 UTC) — not in this kit: it sits outside the bench record this kit was taken from.

The bench records of 2026-08-25 (the serving bench on an RTX PRO 6000 Blackwell (96 GB)) — not in this kit: they sit outside the bench record this kit was taken from.

The bench records of 2026-09-25 (the 12 GB seating diagnosis on the two RTX 3080 Ti boards) — not in this kit: they sit outside the bench record this kit was taken from.

## Withheld digests

A digest a file in this kit prints of bytes the kit does not serve is withheld and marked `withheld-2026-10-07`
where it stood, and `provenance.json` counts them per file, with two exceptions. The model's revision is a
public identifier: it names the version of the model's public files the frame was read from, and it stays. A
digest you can reproduce from this kit's own files, or from a public file or with a public program named
beside it, stays too, and each one is listed below with what it digests and how to reproduce it. A digest of
a file as it stood before a rule rewrote it would let a reader check a guess at what was cut, so it never
ships. A rule withholds a sha256 it knows, whether a file prints it whole, shortened to its first eight or
more hex characters, or as a head and a tail with the middle left out: the digests of the files it rewrote,
and of the records a ruling named, the pre-registration among them. A short prefix of a digest it does not
know, the first eight to fifteen hex characters as a record printed them, still stands; among those that
stand are prefixes of commit and tree ids in the bench's own repository, of the manifest of the Ollama model
tag the bench pulled, and of the token-ID lists the parity check hashed. `index.json` carries the sha256 of
every file it lists, as served.

- `kit/prompts/frame/gemma4-canonical.json`, `template_sha256` (1 in this file): the sha256 of `chat_template.jinja` in the model repository this file names (`repo`), at the `revision` it names. Download that file at that revision and hash its bytes.
- `kit/prompts/frame/gemma4-canonical.json`, `tokenizer_config_sha256` (1 in this file): the sha256 of `tokenizer_config.json` in the same repository, at the same revision. Download that file at that revision and hash its bytes.
- `kit/rows/D-2026-09-28/parity.json`, the render digests of the routes `vllm /tokenize messages`, `ollama /api/chat think:false` and `ollama /v1/chat/completions reasoning_effort:none` (12 in this file, under `D.render.PROMPT.shas`, where PROMPT is the prompt file's name without `.txt`; the three routes give one value per prompt): each is the sha256 of the token IDs that server's chat template made of the prompt, the IDs written as a JSON list with no spaces and hashed as UTF-8 bytes (`ids_sha` in `client/parity.py`). To reproduce one, take the prompt's text from its file in `prompts/`. Make its ticket by the rule in `client/nonces.py` (the formula is in its docstring): `Ticket`, a space, eight digits, a full stop and a blank line, the digits being the first 8 bytes of the sha256 of `SEED|D|PROMPT|1|0|0`, read as one big-endian number, modulo 100,000,000, where SEED is the nonce seed `records/C-RESULTS.md` prints. Build the string `client/framing.py` builds: the `head` of `prompts/frame/gemma4-canonical.json`, then the ticket and the prompt's text joined and trimmed of surrounding whitespace, then its `tail`. Tokenize that string, adding no special tokens, with `tokenizer.json` from the model repository and revision the frame file names, and hash the list.
- `kit/rows/D-2026-09-28/parity.json`, the render digests of the route `llama-server /apply-template` (4 in this file, one per prompt): the same kind of digest, of the token IDs llama-server made of the prompt with its built-in ChatML template rather than the model's own, because Ollama 0.34.4 starts its bundled llama-server (build b11081) with `--no-jinja --chat-template chatml`; that is why this file's render check reads `differs` for this route. To reproduce one, build the same ticket and the user text the same way (the ticket and the prompt's text joined and trimmed). Then, with llama.cpp at commit `3cf03257`, write the string its `chatml` template makes of that text as one user message (the rule is in `src/llama-chat.cpp`): `<|im_start|>user`, a newline, the user text, `<|im_end|>`, a newline, then `<|im_start|>assistant` and a newline, with nothing after it. Tokenize that string with that commit's `llama-tokenize`: give it, with `-m`, the Gemma 4 vocabulary llama.cpp carries at that commit, `models/ggml-vocab-gemma-4.gguf` (git blob `03c49541`), and feed it the string on standard input with `--stdin --no-escape --ids`; the tool then adds the vocabulary's beginning-of-sequence token and parses special tokens, as it does by default. Hash the list as above.
- `kit/rows/D-2026-09-30/parity.json`, the render digests of the routes `vllm /tokenize messages`, `ollama /api/chat think:false` and `ollama /v1/chat/completions reasoning_effort:none` (12 in this file, under `D.render.PROMPT.shas`, where PROMPT is the prompt file's name without `.txt`; the three routes give one value per prompt): each is the sha256 of the token IDs that server's chat template made of the prompt, the IDs written as a JSON list with no spaces and hashed as UTF-8 bytes (`ids_sha` in `client/parity.py`). To reproduce one, take the prompt's text from its file in `prompts/`. Make its ticket by the rule in `client/nonces.py` (the formula is in its docstring): `Ticket`, a space, eight digits, a full stop and a blank line, the digits being the first 8 bytes of the sha256 of `SEED|D|PROMPT|1|0|0`, read as one big-endian number, modulo 100,000,000, where SEED is the nonce seed `records/C-RESULTS.md` prints. Build the string `client/framing.py` builds: the `head` of `prompts/frame/gemma4-canonical.json`, then the ticket and the prompt's text joined and trimmed of surrounding whitespace, then its `tail`. Tokenize that string, adding no special tokens, with `tokenizer.json` from the model repository and revision the frame file names, and hash the list.
- `kit/rows/D-2026-09-30/parity.json`, the render digests of the route `llama-server /apply-template` (4 in this file, one per prompt): the same kind of digest, of the token IDs llama-server made of the prompt with its built-in ChatML template rather than the model's own, because Ollama 0.34.4 starts its bundled llama-server (build b11081) with `--no-jinja --chat-template chatml`; that is why this file's render check reads `differs` for this route. To reproduce one, build the same ticket and the user text the same way (the ticket and the prompt's text joined and trimmed). Then, with llama.cpp at commit `3cf03257`, write the string its `chatml` template makes of that text as one user message (the rule is in `src/llama-chat.cpp`): `<|im_start|>user`, a newline, the user text, `<|im_end|>`, a newline, then `<|im_start|>assistant` and a newline, with nothing after it. Tokenize that string with that commit's `llama-tokenize`: give it, with `-m`, the Gemma 4 vocabulary llama.cpp carries at that commit, `models/ggml-vocab-gemma-4.gguf` (git blob `03c49541`), and feed it the string on standard input with `--stdin --no-escape --ids`; the tool then adds the vocabulary's beginning-of-sequence token and parses special tokens, as it does by default. Hash the list as above.
- `kit/rows/S/ovv-s-20260928T1130Z-ollama-scored.jsonl`, `text_sha256` (11 in this file): the sha256 of the same row's `text` field, as UTF-8 bytes. Hash that field and compare.
- `kit/rows/S/ovv-s-20260928T1130Z-vllm-scored.jsonl`, `text_sha256` (11 in this file): the sha256 of the same row's `text` field, as UTF-8 bytes. Hash that field and compare.
