# Ollama and vLLM on one RTX 3090

*Ollama and vLLM both run an open language model on your own computer. This page puts both on one EVGA GeForce RTX 3090 XC3 Ultra (24 GB) at 300 W, with the same 4-bit Gemma 4 12B weights, from one request at a time to sixteen, and prints every reading beside the bet we wrote down before it. At sixteen requests at once, vLLM 0.30.0 wrote 506.829 tokens a second across them and Ollama 0.34.4, set to serve sixteen at once, 142.327 tokens a second; one request at a time, once an answer had started, the two wrote within 10 per cent of each other's speed; and on the same test questions their answers do not always agree. Benched: Ollama 0.34.4 (released 2026-09-23) and vLLM 0.30.0 (released 2026-09-22).*

*Published 2026-10-07 (UTC) · A small (human) team and a fleet of AI agents.*

**the short version:** Ollama and vLLM are two ways to run an open model on your own computer; this page measures both on one RTX 3090 (24 GB) with the same 4-bit Gemma 4 12B weights. One request at a time, the two write at nearly the same speed once an answer has started: on the older pair we tested first, Ollama 0.32.13 wrote 85.2 tokens a second and vLLM 0.27.1 wrote 78.1 tokens a second, inside the 10 per cent band we had predicted, and the prediction held again on Ollama 0.34.4 and vLLM 0.30.0. Many at once is the job vLLM was built for, and it shows: at sixteen requests released together, vLLM 0.30.0 wrote 506.829 tokens a second across them and Ollama 0.34.4, set to serve sixteen at once, 142.327 tokens a second. vLLM's total first reached 1.2 times Ollama's at two requests at once, a ratio of 1.2077, by a margin smaller than the variation between repeated runs; we had predicted four at once, and had written down before any run that two or eight at once would settle nothing either way, so that prediction stands undecided. Two predictions about the wait before an answer starts were wrong: at one request Ollama's wait was more than twice vLLM's, and with sixteen long prompts at once neither server kept that wait inside the limit we had set for it, 1.5 s for vLLM or 4 s for Ollama set to serve sixteen at once. One prediction about the answers was wrong too: on 108 lines of dinner-party talk, each asking which of six guests said it, the two servers agreed on fewer of them than the 97 per cent we had predicted, so every speed row here carries the label "answers differ at this level". It is one card and one model, with requests released together, not people arriving over time.

11,733 words · about 53 minutes (at 220 words/min) · 20 tables · data kit: yes

https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/

---

![An EVGA GeForce RTX 3090 XC3 Ultra standing upright with its top edge to the camera, reading EVGA and GEFORCE RTX 3090, beside a white moth orchid in a lavender pot, its stem arching over with open white flowers and green buds, against a plain wall.](images/gpus-and-orchids-2026-10-03-1-1600.jpg)

**An EVGA GeForce RTX 3090 XC3 Ultra Gaming (24G-P5-3975-KR), bought in 2021, beside a white moth orchid. Taken 2026-10-03.**

*If you're new here: [strata→signal](https://strata2signal.com) is a small workshop (plus a friendly dog with a white patch) that builds on its own machines and writes up what it measures. Our products run on both servers: [the assistant on our pages](https://research.strata2signal.com/ask-about-this-page/) runs on vLLM, and the long table's chairs sit on both ([how the long table works](https://research.strata2signal.com/how-the-long-table-works/)). Every figure names its reading: the arm that took it and the day. This page is a working document, and a dated update will add [the readings listed near its end](#what-comes-next).*

## The words this page leans on {#the-words-this-page-leans-on}

- **Reading and writing.** A server reads the whole prompt, then writes the answer a token (a piece of a word) at a time. **First token**, which this page also calls the first word, is the wait before the answer starts; **decode speed**, how fast one answer is written after that; **total speed**, every token written across all requests over the time from sending to the last one finishing.
- **Window and slot.** For each request a server keeps the conversation so far on the card, beside the weights; the **window** is the most tokens one request may hold. Ollama serves as many requests at once as it has **slots**, one by default, and every slot reserves a whole window of memory.
- **Requests at once** is this page's unit: how many arrive together. Sixteen at once is not sixteen people. A family of four chatting with a home assistant almost never sends four in the same second; one person's coding agent can send several; a classroom of twenty pressing enter after the same instruction can send twenty.
- **Posture.** One server started with one set of settings. This page has four: vLLM 0.30.0 at the headline settings and at its default step, and Ollama 0.34.4 with sixteen slots and at its default; [what each server was allowed](#what-was-written-down-before-the-first-run) lists every flag.
- **Arm, prediction, verdict.** An arm is one timed job of the bench, named by a letter. Every prediction was written down and pushed before the runs it governs, and reads CONFIRMED, REFUTED or UNDECIDED; a reading that breaks one of our own rules is VOID, never a number. A level's **spread** is its highest run less its lowest, divided by the median, of the figure its prediction reads (decode speed for P-1, total speed for the rest), and a prediction whose band is narrower than a level's spread reads UNDECIDED at that level. Spreads print as fractions, so P-3's band of 0.2 is 20 per cent; the first table below prints its two spreads as per cents.

## Two names, one card {#two-names-one-card}

Say you have a graphics card with 24 GB of memory and an open model you want to run on it. You meet these two names early. Both load an open model's weights onto the card and answer requests over a local web address, so they look like two answers to one question. They are closer to answers to two different ones. Ollama asks how easy it can be for one person to run a model on the machine in front of them. vLLM asks how many requests one card can answer at once without wasting its memory. Each has a history of its own on this shelf, [A short history of Ollama](https://research.strata2signal.com/a-short-history-of-ollama/) and [A short history of vLLM](https://research.strata2signal.com/a-short-history-of-vllm/), and this page is the part both of them point at: the two on one card, the same model on both, measured.

The difference that matters here is how each one hands out the card's memory.

**Ollama gives each request a slot**, working memory reserved for a whole window. By default there is one, so a second request waits for the first. `OLLAMA_NUM_PARALLEL` adds slots, each reserving another whole window: with sixteen slots of 5,120 tokens, the runner Ollama started on our card was told to hold 81,920 tokens. Its own FAQ says the memory needed scales with `OLLAMA_NUM_PARALLEL` times `OLLAMA_CONTEXT_LENGTH`, so the two are set together. It loads a model on the first request for it and, by default, unloads it after five idle minutes. Since Ollama 0.30 the program that runs a GGUF model (llama.cpp's file format) under it is llama.cpp's own `llama-server`, started as a separate process; Ollama 0.34.4 ran build b11081 of it.

**vLLM batches continuously.** It hands out working memory in small pages as requests grow, moves every request in flight forward together on each step, and lets new ones join between steps. To do it, it reserves most of the card at start, 92 per cent by default, and keeps it.

**Both reuse a repeated prompt, differently.** vLLM keeps a shared cache of prompt pages that any later request with the same beginning can reuse. `llama-server`, under Ollama, keeps each slot's last prompt; at the build Ollama 0.34.4 ships (b11081) it also keeps up to 8,192 MiB of prompts in the computer's memory, saves idle slots into it, and sends a new request to the slot whose last prompt looks most like it. We read those defaults from the bundled program; Ollama sets none of them, and its one patch to llama.cpp does not touch them. Which helps, and when, we measured: [the first word](#the-first-word).

The longer account of each is in its history: [how Ollama serves a request](https://research.strata2signal.com/a-short-history-of-ollama/#how-it-serves-a-request) and [how vLLM does](https://research.strata2signal.com/a-short-history-of-vllm/#how-it-serves-a-request). What follows is the bench, results first. The card, the weights, the prompt, every posture and every prediction are written out further down, in [what was written down before the first run](#what-was-written-down-before-the-first-run), and every result links back to it.

## One request at a time {#one-request-at-a-time}

The first reading, arm S, ran on 2026-09-28 from 11:29:59 to 11:33:35 UTC, 17 minutes after the go-ahead, on the pair our bench kit had already proven on this computer: vLLM 0.27.1, and Ollama 0.32.13 with sixteen slots. One 536-token prompt (`series-p512`), 256 tokens out, one request at a time, a warm-up and ten scored runs per server, the same text to both, decode scored on our own clock for both. vLLM ran our kit's own serve line, not the headline posture [described below](#what-was-written-down-before-the-first-run): a 4,352-token window, its memory pool pinned at 0.8913 of the card, prefix caching off and 32 requests at most. The pre-registration names each departure and why it does not move one request's decode. Ollama ran with sixteen slots and `num_batch: 512`.

| Median of ten runs | vLLM 0.27.1 | Ollama 0.32.13, 16 slots |
|---|---:|---:|
| Decode, tokens a second (scored) | 78.126 | 85.216 |
| Spread of the ten runs | 0.50 % | 0.41 % |
| Decode, tokens a second (the server's own clock) | 78.118 | 85.553 |
| First token, ms (recorded, not scored) | 226.0 | 1,018.1 |
| Of which Ollama's load step, ms | — | 761.1 |
| Total speed of one answer, tokens a second (send to last token) | 73.364 | 63.865 |
| Board power while busy, W | 299.14 | 289.88 |
| Joules per 1,000 tokens, busy watts ÷ decode speed (arithmetic, not registered) | 3,829 | 3,402 |

*Answers differ at this level: each server returned the same text on all ten of its scored runs, and the two servers' texts differ from each other. On the versions benched, [are the answers the same?](#are-the-answers-the-same) says by how much.*

**P-1: CONFIRMED on the proven pair.** vLLM's decode ÷ Ollama's is 0.9168, 8.3 per cent below 1 (put the other way, Ollama 9.1 per cent faster): inside the ±10 per cent band by 1.7 points, with each server's ten runs spread 0.50 and 0.41 per cent. A bet that held, not a tie. The bytes point the same way: vLLM reads 17 per cent more per token, and the gap is about half that.

**P-1 on the versions benched: CONFIRMED again.** Arm C, the sweep from one to sixteen requests at once, ran on 2026-10-01 from 13:35:02 to 14:29:21 UTC and read one request first, on the headline postures: vLLM 0.30.0 decoded at 82.693 tokens a second and Ollama 0.34.4 with sixteen slots at 85.007, a ratio of 0.9728, inside the ±10 per cent band, where arm S's ratio on the proven pair read 0.9168. That reading sits beside arm S's, never in its place. S is the reading of record for one request on the proven pair; C's level 1 is the same question asked of the versions this page stamps. Answers differ at this level.

What this reading is not: a first-token comparison. Ollama's 1,018 ms is about 761 ms of its own load step, on a model already loaded (a step that reads about 2 ms on 0.34.4; [the first word](#the-first-word)), plus about 253 ms of reading the prompt. The total-speed row, where vLLM finished each answer sooner, carries the same wait. vLLM's board sat at its 300 W cap on every run and Ollama's just under it, so on that row's sum Ollama used about 11 per cent fewer joules per token.

## Requests at once {#requests-at-once}

This is the question vLLM was built for, and one our earlier two-server comparisons could not answer fairly: they ran Ollama at one slot, a queue by design. Arm C answers it, on 2026-10-01 from 13:35:02 to 14:29:21 UTC: five boots on the one card, thirty cells (one posture, one prompt and one level each), and every scored request wrote its 256 tokens. The prompt is `series-p512`, 536 tokens as arm S sent it and 548 on every server here, where each request opens on its own seeded tag so that no prompt cache flatters either server; the levels are one, two, four, eight and sixteen requests released together. Ollama ran twice over: with sixteen slots, and at its default of one.

**Total speed at one to sixteen requests at once**

*Tokens a second across all requests, the median of each level's scored runs: Gemma 4 12B, 4-bit, on one EVGA GeForce RTX 3090 XC3 Ultra (24 GB) at 300 W, read by arm C on 2026-10-01 (UTC); the 536-token prompt with each request's seeded tag, 256 tokens out per request. Ollama's default posture is not drawn: it stays flat, and its figures are in their own table. At four and eight requests the sixteen-slot runs did not agree with each other (the spreads are in the table). Answers differ at every level.*

Total speed at one to sixteen requests at once. 1 at once: vLLM 0.30.0 77.653 tok/s; Ollama 0.34.4, sixteen slots 74.358 tok/s. 2 at once: vLLM 0.30.0 142.05 tok/s; Ollama 0.34.4, sixteen slots 117.621 tok/s. 4 at once: vLLM 0.30.0 242.982 tok/s; Ollama 0.34.4, sixteen slots 147.955 tok/s. 8 at once: vLLM 0.30.0 377.228 tok/s; Ollama 0.34.4, sixteen slots 136.181 tok/s. 16 at once: vLLM 0.30.0 506.829 tok/s; Ollama 0.34.4, sixteen slots 142.327 tok/s.

Arm C, 2026-10-01, rows/C.json, cells 0 to 4 and 12 to 16, medians.aggregate\_wall\_tok\_s

**Each request's own writing speed**

*Decode, tokens a second per request, the median: Gemma 4 12B, 4-bit, on one EVGA GeForce RTX 3090 XC3 Ultra (24 GB) at 300 W, read by arm C on 2026-10-01 (UTC); the 536-token prompt with each request's seeded tag, 256 tokens out per request, one to sixteen requests at once. At four and eight requests the sixteen-slot runner's three runs decoded at two speeds, and each median drawn is one of them; both speeds are under the table. Answers differ at every level.*

Each request's own writing speed. 1 at once: vLLM 0.30.0 82.693 tok/s; Ollama 0.34.4, sixteen slots 85.007 tok/s. 2 at once: vLLM 0.30.0 79.677 tok/s; Ollama 0.34.4, sixteen slots 76.43 tok/s. 4 at once: vLLM 0.30.0 72.183 tok/s; Ollama 0.34.4, sixteen slots 51.713 tok/s. 8 at once: vLLM 0.30.0 58.778 tok/s; Ollama 0.34.4, sixteen slots 23.816 tok/s. 16 at once: vLLM 0.30.0 41.428 tok/s; Ollama 0.34.4, sixteen slots 24.449 tok/s.

Arm C, 2026-10-01, rows/C.json, cells 0 to 4 and 12 to 16, medians.decode\_tok\_s\_p50

*Two charts, two series each: vLLM 0.30.0 against Ollama 0.34.4 with sixteen slots, total speed by level and each request's own decode by level, tokens a second drawn from zero. Ollama's default posture is a flat line, and its own table below carries it; it is not drawn as a third series. Every value drawn is printed in the tables below. Answers differ at every level drawn.*

**What the tables say.** Every speed figure in them carries P-6's label, answers differ at this level ([why](#are-the-answers-the-same)).

**The crossover, and what sixteen slots did.** The crossover, the first level at which vLLM's total reached 1.2 times Ollama's with sixteen slots, read two requests at once, a ratio of 1.2077, and vLLM's total stood at 3.561 times Ollama's at sixteen; Ollama's sixteen-slot total read 147.955 tokens a second at four, 136.181 at eight and 142.327 at sixteen, flat from four on; answers differ at every level here. Read it with the labels the rows carry. At four and at eight requests, the sixteen-slot runner's three scored runs did not agree with each other: their spreads are wider than the 0.2 band P-3 was registered with, so P-3 is undecided at those two levels and the ratios there print as readings, not as verdicts. P-3 itself, the bet that vLLM's total would first reach 1.2 times Ollama's with sixteen slots at four requests, read UNDECIDED: the crossover came at a level the pre-registration had marked undecided before any run ("undecided at two or eight"), and its margin sits within run noise of the 1.2 bar. Had two fallen short, the next level, four, was too noisy to decide, so the word would have been the same. What one slot did is C-1's row: Ollama at its default is a queue, and its total stayed within 10 per cent of its one-request figure at every level, CONFIRMED.

**The long prompt, and P-5.** A second prompt of about 2,600 tokens, written to ask for a long answer and cut at 256 tokens out like every other, ran at the same five levels; its totals are in the table below. At sixteen the first token's 95th percentile, the median of three runs, read 17,327.116 ms on vLLM and 107,377.256 ms on Ollama with sixteen slots, against bars of 1,500 and 4,000 ms (P-5 reads every run, and even each server's best missed its bar); on the sixteen-slot runner at one request the first token rose from 6,460.522 to 9,262.838 ms across the ten scored runs, and that cell's total-speed spread read 0.2499; at two requests the spread read 0.2256. The sixteen-slot cells at one and two requests on that prompt ran straight after the `series-p512` cells on the same runner, and at one request the first token rose with every scored run; those two cells carry spreads wider than the band and print with that condition, never as steady figures. P-5, the bet that sixteen long prompts at once would see a first token within 1.5 s on vLLM and 4 s on Ollama with sixteen slots, read REFUTED on both halves. Answers differ at this level.

**Total speed, the long prompt of about 2,600 tokens, 256 tokens out per request, released together.** Arm C, 2026-10-01 (UTC). Tokens a second, the median of the level's scored runs: ten at one and two requests, three above.

| At once | vLLM 0.30.0 | Ollama 0.34.4, 16 slots | Answers |
|---:|---:|---:|---|
| 1 | 59.884 | 24.282 | Differ |
| 2 | 93.695 | 21.68 | Differ |
| 4 | 130.187 | 24.012 | Differ |
| 8 | 161.508 | 24.902 | Differ |
| 16 | 181.513 | 35.373 | Differ |

*At one and two requests the sixteen-slot spreads are wider than the band, as the paragraph above says, so those two rows print with that condition, never as steady figures.*

**vLLM at its default step.** vLLM at its default 2,048 tokens per step wrote 186.538 tokens a second across sixteen long requests (answers differ at this level), with a first-token 95th percentile, the median of three runs, of 16,395.26 ms and queue waits of up to 13,593.773 ms across its scored requests, on a working memory of 37,541 tokens (room for 7.33 requests at the full window). That boot loaded the compiled graph that [a refused boot](#what-was-written-down-before-the-first-run) of the same posture had saved on 2026-09-30, so its working memory came out larger than a cold boot's would have, and its figures print with that condition and no other.

**Drift, and energy.** Inside arm C's window (2026-10-01, 13:35:02 to 14:29:21 UTC), vLLM ran again at one and sixteen requests after the Ollama block, to check for drift: its total speed on that second pass read 0.67 per cent lower at one request and 0.25 per cent higher at sixteen than its first reading, which is printed beside the readings and corrects nothing. Board energy is arithmetic, not a registered reading, and the two tables below give it three ways (answers differ at every level here). The first row counts the card's busy watts over each run's whole wall time. The kit's own busy figure counts only the half-second samples in which the card was busy; where the card was busy throughout, those samples run past the end of the run, so it reads above the first row. At sixteen on Ollama with sixteen slots the card was busy in only 36 to 40 of each run's 55 samples, so the first row charges busy watts to idle seconds there, and the sixteen-request table's third row counts the whole run's mean watts instead; where the card was busy throughout, as on vLLM, that row reads the same as the first. Arm S's row above is a different sum, busy watts over decode speed, which does not compare with these, so the one-request table gives that sum its own row.

**Board energy at one request, joules per 1,000 tokens.** Arm C, 2026-10-01 (UTC); arithmetic, not registered; answers differ at this level.

| Sum | vLLM 0.30.0 | Ollama 0.34.4, 16 slots |
|---|---:|---:|
| Busy watts over the whole wall time | 3,849 | 3,905 |
| The kit's busy figure | 4,352 | 4,223 |
| Arm S's sum, busy watts over decode speed | 3,618 | 3,511 |

**Board energy at sixteen requests, joules per 1,000 tokens.** Arm C, 2026-10-01 (UTC); arithmetic, not registered; answers differ at this level.

| Sum | vLLM 0.30.0 | Ollama 0.34.4, 16 slots |
|---|---:|---:|
| Busy watts over the whole wall time | 590 | 1,638 |
| The kit's busy figure | 624 | 1,090 |
| Mean watts over the whole run | 590 | 1,384 |

**Total speed, the 536-token prompt `series-p512` with each request's seeded tag, 256 tokens out per request, released together.** Arm C, 2026-10-01 (UTC). Tokens a second, the median of the level's scored runs: ten at one and two requests, three above. Answers: P-6's label, [answers differ at this level](#are-the-answers-the-same), on every speed row.

| At once | vLLM 0.30.0 | Ollama 0.34.4, 16 slots | Answers |
|---:|---:|---:|---|
| 1 | 77.653 | 74.358 | Differ |
| 2 | 142.05 | 117.621 | Differ |
| 4 | 242.982 | 147.955 | Differ |
| 8 | 377.228 | 136.181 | Differ |
| 16 | 506.829 | 142.327 | Differ |

*A request that stopped short of 256 tokens would print flagged, never dropped: every request at every level wrote its 256 tokens, so each run wrote from 256 tokens at one request to 4,096 at sixteen, and no cell carries a flag.*

**Ollama 0.34.4 at its default of one slot, the same prompt and levels.** Arm C, 2026-10-01 (UTC). Tokens a second, the median of the level's scored runs. One slot is a queue by design, which is what C-1 bet on and read.

| At once | Ollama 0.34.4, its default | Answers |
|---:|---:|---|
| 1 | 74.545 | Differ |
| 2 | 73.649 | Differ |
| 4 | 73.621 | Differ |
| 8 | 73.597 | Differ |
| 16 | 73.565 | Differ |

**vLLM against Ollama with sixteen slots, level by level.** Arm C, 2026-10-01 (UTC). The ratio is vLLM's total over Ollama's. The spread is the wider of the two compared cells' at that level, (highest − lowest) ÷ median; a level whose sixteen-slot spread is wider than P-3's 0.2 band reads undecided there, whatever its ratio.

| At once | vLLM ÷ Ollama, 16 slots | Spread | P-3 at this level |
|---:|---:|---:|---|
| 1 | 1.0443 | 0.0099 | Decided |
| 2 | 1.2077 | 0.0544 | Decided; the crossover |
| 4 | 1.6423 | 0.2973 | Undecided: the 16-slot spread is wider than the band |
| 8 | 2.77 | 0.2352 | Undecided: the same |
| 16 | 3.561 | 0.001 | Decided |

*Every row: answers differ at this level, P-6's label.*

**What each request felt, the same prompt.** Arm C, 2026-10-01 (UTC). Decode per request, tokens a second: one answer's own writing speed once it has started; each run's median request, then the median of the level's scored runs.

| At once | vLLM 0.30.0 | Ollama 0.34.4, 16 slots | Answers |
|---:|---:|---:|---|
| 1 | 82.693 | 85.007 | Differ |
| 2 | 79.677 | 76.43 | Differ |
| 4 | 72.183 | 51.713 | Differ |
| 8 | 58.778 | 23.816 | Differ |
| 16 | 41.428 | 24.449 | Differ |

*At four and at eight the sixteen-slot runner's three scored runs decoded at two speeds (52.026, 51.713 and 32.831 tokens a second at four; 23.816, 32.485 and 23.814 at eight), and each median above is one of them.*

**The wait for the first token, 95th percentile, ms.** Arm C, 2026-10-01 (UTC). Each run's 95th-percentile request, the median of the level's scored runs (at one request a run holds one request, so its 95th percentile is that request's wait). These rows print because TT, our test of Ollama's load step, came out the same when it ran again ([the first word](#the-first-word)).

| At once | vLLM 0.30.0 | Ollama 0.34.4, 16 slots |
|---:|---:|---:|
| 1 | 212.999 | 441.866 |
| 2 | 417.849 | 1,068.577 |
| 4 | 822.939 | 2,244.823 |
| 8 | 1,641.068 | 5,365.722 |
| 16 | 3,333.177 | 20,935.24 |

*Every row: answers differ at this level, P-6's label.*

## The first word, and a shared prompt {#the-first-word}

Arm S, on 2026-09-28, recorded a wide first-token gap: 226 ms for vLLM against 1,018 ms for Ollama, 761 ms of it in a step Ollama's own timer calls `load_duration`, on a model already loaded. An earlier page of ours, [Reading the answer](https://research.strata2signal.com/reading-the-answer/), put about 870 ms of each short request down to "the server's per-request overhead", read on this same bench computer under the same version, 0.32.13. That step needed explaining before any first-token figure could mean anything, so it got a test of its own, TT: with the window set at the server and none in the request, Ollama 0.34.4's load step at one request is 50 ms or less.

**TT: CONFIRMED.** On 2026-09-28 between 17:20 and 17:30 UTC, all ten scored runs on Ollama 0.34.4 read a load step of 1.284 to 3.111 ms; a nearly identical request shape on 0.32.13 that morning read 722 to 833 ms. The shapes were close, not identical (0.32.13's request also asked to keep the model loaded and set `top_p` to 1), so the reading says the step is gone on 0.34.4, not which change removed it. When arm D ran again on 2026-09-30, from 20:00:02 to 20:07:18 UTC, with the client fixed, TT came out the same: CONFIRMED, 1.31 to 3.119 ms on all ten scored runs. That repeat is why this page prints its first-token rows. [Reading the answer](https://research.strata2signal.com/reading-the-answer/#the-870-ms-read-again) now carries a dated correction, added 2026-10-07 (UTC), with a pointer at each of the three places it attributes about 870 ms of each short request to the server's per-request overhead; its figures stand as read on 0.32.13.

On arm D's first pass, on 2026-09-28 from 17:20 to 17:30 UTC, we sent the one-request prompt five ways, ten scored runs each: (a) to Ollama's `/api/generate`, raw, with no window in the request; (b) the same, asking for the server's own window, 5,120; (c) the same, asking for a larger window, 8,192, which reloads the model, so its registered reading is that reload and its cells read VOID; (d) to Ollama's OpenAI-style `/v1/completions`, which templated the text again (563 prompt tokens against 548, every visible answer empty), VOID as hidden reasoning; (e) to the bundled `llama-server` alone, `/completion`.

| Road | First token, median, ms | Fastest to slowest, ms | Load step, ms |
|---|---:|---:|---:|
| (a) | 463.2 | 441.6–668.8 | 1.284–3.111 |
| (b) | 487.7 | 445.1–580.7 | 1.262–2.960 |
| (c) | VOID: reloaded | VOID | VOID |
| (d) | VOID: hidden reasoning | VOID | — |
| (e) | 440.3 | 438.1–575.9 | — |

Road (a) is TT's own request, so its load step is TT's first reading above. Two lines are worth keeping. Ollama's median sat 23 ms above `llama-server` alone (463.2 against 440.3), inside the spread of either set, so at one request Ollama's own layer adds little to the first word on this version. A request asking for a larger window than the loaded model's reloads it, because the window is set when the model loads: road (c)'s registered reading is that reload, whose warm-up spent 6,333.1 ms loading. If your client sends a window size, send the server's own.

vLLM's first-token rows from that pass are not printed, because of [the start-token mistake](#what-was-written-down-before-the-first-run). On the versions benched, arm C, on 2026-10-01, read the first word at one request on the headline postures: 212.999 ms for vLLM 0.30.0 and 441.866 ms for Ollama 0.34.4 with sixteen slots, a ratio of 2.0745. P-2, our bet that Ollama's first token would come within twice vLLM's, read REFUTED, and the margin holds at every pairing of the two servers' scored runs. The reading is arm C's one-request cell, the same request shape as every other level on this page; a chat-API row the same runner served later in its boot is not P-2's input and is not printed as a figure.

**A shared prompt.** A long system prompt reused on every request is an assistant's common case, and where the two caches part. In Gemma 4 12B, 40 of 48 attention layers look back only 1,024 tokens and 8 see the whole window, which changes what "the prompt is cached" means: a new question behind a long shared prompt can need work that a match on the prompt's beginning does not supply. `llama-server` at b11081 keeps up to 32 checkpoints of a slot's state, at least 8,192 tokens apart. Arm P, on 2026-09-30 from 20:43:52 to 21:19:31 UTC, sent a shared prompt of about 2,500 tokens in two shapes, at one, four and sixteen requests interleaved, and recorded what each server said it served from cache. The same question again is a 2,560-token request; a new question behind the same prompt shares its first 2,533 tokens.

**The same question again, first token, median, ms.** Arm P, 2026-09-30 (UTC).

| At once | vLLM 0.30.0 | Ollama 0.34.4, 16 slots |
|---:|---:|---:|
| 1 | 48.102 | 227.911 |
| 4 | 149.682 | 838.946 |
| 16 | 313.503 | 78,270.429 (runs below) |

**A new question behind the same prompt, first token, median, ms.** Arm P, 2026-09-30 (UTC).

| At once | vLLM 0.30.0 | Ollama 0.34.4, 16 slots |
|---:|---:|---:|
| 1 | 57.887 | 10,970.201 (condition below) |
| 4 | 166.486, flagged | 39,991.196, flagged |
| 16 | 344.658, flagged | 120,241.433, rising; flagged |

*Flagged, on a new question at four and at sixteen at once, on each server: on vLLM 0.30.0, 6 of 12 requests at four and 25 of 48 requests at sixteen, and on Ollama 0.34.4 with sixteen slots, 6 of 12 requests at four and 25 of 48 requests at sixteen, ended their answer on a stop before 256 tokens; the counts are the same on both, each read from its own server's cells. A first token comes before any stop, so these first tokens stand, flagged, never dropped. At sixteen at once on Ollama, run by run: the same question again read 66,329.152, 78,270.429 and 78,606.231 ms; a new question rose with every scored run, 95,395.498, 120,241.433 and 211,619.782 ms, so its median is never a steady figure (the cell's total-speed spread, 0.5987, is wider than P-3's 0.2 band, though on a flagged cell, where answers stopped early, a total-speed spread is not a speed spread).*

**Tokens served from cache, per request.** Arm P, 2026-09-30 (UTC).

| Shape | At once | vLLM 0.30.0 | Ollama 0.34.4, 16 slots |
|---|---:|---:|---:|
| Same question | 1 | 2,496 | 2,559 |
| Same question | 4 | 2,496 | 2,559 |
| Same question | 16 | 2,496 | 2,558–2,559 |
| New question | 1 | 2,496 | 2,533 |
| New question | 4 | 2,496 | 2,533–2,536 |
| New question | 16 | 2,496 | 2,533–2,538 |

*vLLM's count is the same on every request in both shapes, and it is the whole of this prompt a cache hit can reach: vLLM 0.30.0 keeps this model's cache in six groups, so a hit ends on a 64-token step, and the tail of the prompt past the last whole step is read again, with the question. Ollama's counts show `llama-server` serving the same question again from a slot's own last prompt, and a new question from a checkpoint it had saved at the prompt's end.*

**What each cache did, and where the requests landed.** P-4a and P-4b both read REFUTED on a first reading judged within one 16-token step, a step smaller than the 64-token step vLLM uses for this model; a second reading at vLLM's own step, from a corrected instrument, follows in a dated update and prints beside the first. One reading in these tables is a condition, not a cause, and it prints as one. On a new question behind the 2,533-token shared prompt, one request at a time, Ollama 0.34.4 with sixteen slots took 10,970.201 ms to its first token, and its one-slot default took 432.838 ms. Each is the median of three scored runs (10,955.703 to 10,989.521 ms, and 430.199 to 435.44 ms), each run opened with a one-token warm request outside the timed window, and both were read straight after the same server had answered "the same question again" sixteen at a time, with no restart between. Both served all 2,533 shared tokens from cache, and Ollama's own clock spent 412.264 to 413.285 ms of the sixteen-slot wait on the prompt (419.212 to 424.476 ms on the default); the log does not split the rest. So on the sixteen-slot runner a new question behind the shared prompt waited far longer for its first token than the same request did on Ollama's default posture of one slot, at one request, and the same question again waited far longer at sixteen at once than at one or four. The sixteen-slot runner's log for those runs counts idle slots being saved to its host prompt cache between a request's launch and its prompt, 30 a run, the same count as in its "same question again" row at one request earlier in that boot, which read 227.911 ms; the default posture's log counts none. The count is printed; the time it took is not in the log, so no cause is read here. The raw lines of that log stay in our records.

## What each holds, and how fast it starts {#what-each-holds}

Memory was read one way for every server: how much the card's used memory grew as it started and loaded the model, the median of three reads. For vLLM that figure depends on whether the boot compiled the model's graph or loaded a compile saved by an earlier boot, so each vLLM row names which.

| Server, as set up | Grown at rest, MiB | Share of the card, % | Reading |
|---|---:|---:|---|
| vLLM 0.30.0, its default memory share; its first boot on this computer, which compiled cold | 20,847 | 84.8 | D, 2026-09-28 |
| vLLM 0.30.0, its default memory share; a later boot that loaded the saved compile | 22,505 | 91.6 | D, 2026-09-30 |
| Ollama 0.34.4, 16 slots of 5,120 tokens | 16,387 | 66.7 | D, 2026-09-28 |
| `llama-server` alone, Ollama's settings | 16,377 | 66.6 | D, 2026-09-28 |
| Ollama 0.34.4, its default | 8,761 | 35.6 | C, 2026-10-01 |
| vLLM 0.27.1, the kit's pinned pool | 22,071 | 89.8 | Arm S, 2026-09-28 |
| Ollama 0.32.13, 16 slots of 5,120 tokens | 16,389 | 66.7 | Arm S, 2026-09-28 |

*Reading names the arm each row comes from: S, the first reading, on the proven pair; D, the first full pass on the new versions and its re-run; C, the sweep from one to sixteen requests at once. The share is arithmetic on the grown figure, against the card's 24,576 MiB. Ollama's rows include Gemma 4's image projector, which its runner also loads (a 167 MiB file); vLLM ran text only and built no image, audio or video parts. We have not measured vLLM with those parts built.*

Under load the read is a different one: the card's used memory in the 2 Hz board samples during the runs at sixteen requests, less the boot's own baseline. Arm C took it on 2026-10-01, beside what the same two boots had grown at rest; its vLLM boot loaded a saved compile.

| Server, as set up | Grown at rest, MiB | Used under sixteen requests, MiB |
|---|---:|---:|
| vLLM 0.30.0, the headline; compile cache warm | 22,505 | 22,507 |
| Ollama 0.34.4, 16 slots of 5,120 tokens | 16,389 | 16,393 |

Each reads a few MiB above what that same boot had grown at rest. Do not read either against the first row of the table above: that row is the cold boot, and the gap between the two vLLM conditions is the compile cache's, not the load's.

vLLM's reservation is taken at start; Ollama's follows its slots and window, and returns when the model unloads. Inside its reservation on arm D's first pass, vLLM 0.30.0 put 8.28 GiB of weights (loaded in 6.7 s), 0.29 GiB of captured graphs and 10.47 GiB of working memory: 75,691 tokens, room for 14.78 requests at the full window. Two registered expectations came in under their bars on that pass, which compiled cold: room for sixteen (P-BOOT, at 14.78 requests) and 90 per cent of the card (C-2, at 84.83 per cent); vLLM's log prints beside them `--gpu-memory-utilization=0.9200 is equivalent to --gpu-memory-utilization=0.8947 without CUDA graph memory profiling`. Our bench's verdict rows marked both REFUTED, but on a pass that stopped at another gate, so both were read again when arm D repeated on 2026-09-30, from 20:00:02 to 20:07:18 UTC, on a boot that loaded the compile the first pass had saved: room for 17.07 requests at the full window (87,407 tokens, 12.09 GiB of working memory), P-BOOT CONFIRMED, and 91.57 per cent of the card at rest, C-2 CONFIRMED. The two readings differ by condition, not by noise. On its first boot vLLM compiles the model's graph and saves the result; its memory profile during that boot peaks higher, so it gives the working memory less. Every later boot with the same serve line loads the saved compile, profiles a lower peak, and gives the working memory more. Neither reading is noise at its bar: every boot of the same condition in release 1's rows read the same figures exactly. Both are printed wherever either prediction is, cold REFUTED and warm CONFIRMED, and which one you meet depends on whether vLLM has already compiled that model, with that same serve line, on that computer.

| Server | What was timed | Seconds | Arm |
|---|---|---:|---|
| vLLM 0.30.0 | Launch to ready, its first start on this computer | 144.35 | D |
| vLLM 0.30.0 | Of that, compiling, by its own log | 56.95 | D |
| vLLM 0.30.0 | Launch to ready, a later start that loaded the saved compile | 53.79 | D |
| vLLM 0.27.1 | Launch to ready, compile cache warm from a 104.65 s start minutes earlier | 31.18 | Arm S |
| Ollama 0.34.4 | Launch to ready; the model loads on the first request | 4.06 | D |
| Ollama 0.34.4 | First request: loads the model, asks for 8 tokens; one start, not yet explained | 68.44 | D |
| Ollama 0.34.4 | Of that, the load, by Ollama's timer | 38.90 | D |
| Ollama 0.32.13 | The model's first load, by Ollama's timer | 4.94 | Arm S |
| `llama-server` b11081 | Launch to ready, model loaded, with Ollama's settings from 0.34.4's bundle | 2.92 | D |

*One start each, the model's files already in the computer's memory: arm S's on the morning of 2026-09-28 (UTC), arm D's on 2026-09-28 from 17:20 to 17:30 UTC, and the one later start of vLLM 0.30.0 on arm D's re-run of 2026-09-30; three cold starts each come in [the follow-up](#what-comes-next).*

A new vLLM version, or a serve line it has not compiled before, compiles on first start (0.30.0 logged 56.95 s of `torch.compile` inside 97.84 s of engine setup) and reuses the result after, which is why 0.27.1 started in 31 s; the reuse also moves the working memory it reserves, as the two readings above show. One oddity is Ollama's: 0.34.4's first load read 38.9 s by its own timer (68.4 s to the answer), where 0.32.13's first load that morning read 4.9 s by the same timer, and the `llama-server` from 0.34.4's own bundle started with the model loaded in 2.9 s. We have not traced why. Its second start, read by arm D on 2026-09-30, between 20:00:02 and 20:07:18 UTC, with the page cache warm and one start only, came in far shorter: 8.24 s by Ollama's own timer, 8.62 s to the first answer, after a 2.86 s launch. Both readings stand, the first start beside the second, until a later day reads three cold starts of each.

## Are the answers the same? {#are-the-answers-the-same}

Speed means little if the answers change. In arm S each server returned the same text on all ten scored runs, so temperature 0 repeated on both; the two servers' texts differ from each other, as the arithmetic differences in the weights make likely.

The registered test is arm Q, the six-way, on 2026-09-30 from 21:27:35 to 21:31:13 UTC: 108 lines from [our long table](https://research.strata2signal.com/a-dinner-party-for-the-dead/), each asking which of six guests said it, public since 2026-09-23 in [the kit for Reading the answer](https://research.strata2signal.com/reading-the-answer/data/). Both servers answer it on the same raw path at one request and at sixteen, scored strictly: a reply counts only if it opens with the letter. The bar: each above the six-way's informed floor of 24.0 per cent (what a guesser scores by crossing off the guests a line names; pure chance is 16.7), and the two agreeing on at least 97 per cent of items (P-6); below that, every speed row prints "answers differ at this level".

| Server | Right of 108, one request | Right of 108, sixteen at once | Items the two agree on, one request |
|---|---:|---:|---:|
| vLLM 0.30.0 | 51 (47.22 %) | 52 | 63 (58.33 %) |
| Ollama 0.34.4, 16 slots | 9 (8.33 %) | 10 | — |

*The agreement is one figure for the pair, printed once in vLLM's row. The agreement column is the registered count: an item on which neither reply opens with a letter counts as agreeing, and 54 of the 63 are such items. The right-answer columns are strict; read leniently, the letter found wherever it appears in the reply, vLLM was right on 81 of 108 (75 %) at one request and Ollama on 86 (79.63 %), printed beside and never in a verdict.*

**P-6: REFUTED**, on the strict scorer, as registered: the two agreed on 63 of 108 items (58.33 per cent) against the 97 per cent bet, by the registered count, which takes an item on which neither reply opens with a letter as agreeing: 54 of the 63 are such items, and on the other 9 both opened with the same letter; read leniently, on 98 (90.74 per cent); Ollama's replies open with "The correct answer is" on 99 of 108 and vLLM's on 51; on all 45 strict disagreements both named the same guest; vLLM's strict score, 47.22 per cent, clears the floor, and Ollama's, 8.33 per cent, does not; read leniently, beside and never in the floor, Ollama's is 79.63 per cent (86 of 108) and vLLM's 75 per cent (81 of 108); sixteen at once changed vLLM's strict choice on 1 of 108 items and Ollama's on 5. The lenient reading, which finds the letter wherever it appears in the reply, prints beside the strict one and never enters the ratio. The two differ because the strict rule counts only a reply that opens with the letter. We re-scored nothing. The label stays on every speed row, because the registered rule reads the strict scorer, and under it the two servers' answers differ.

## What was written down before the first run {#what-was-written-down-before-the-first-run}

**The pre-registration, an arm, a verdict.** Predictions and rules were pushed before the first model call, and each later amendment before the scored runs it governs; an arm is one timed job of the bench. A prediction reads CONFIRMED, REFUTED or UNDECIDED; a reading that breaks one of our rules is VOID, never a number. A refusal by one of our own gates is a finding, and no arm was re-run to a pass: each refusal below got a fix of its own, and the next run read the fixed kit.

**The card and the computer.** One EVGA GeForce RTX 3090 XC3 Ultra Gaming (24 GB), capped at 300 W with persistence on, alone in the bench computer's x16 slot: an AMD Ryzen 7 3700X, 30,990 MiB of memory, Ubuntu 26.04 LTS, NVIDIA driver 595.84. One server on the card at a time, the card read back to empty between them.

**The same weights, said exactly.** Both ran the same Gemma 4 12B weights, which Google trained to hold up at 4 bits (quantization-aware training, QAT): Ollama's `gemma4:12b-it-qat`, a 4-bit GGUF, and Google's `google/gemma-4-12B-it-qat-w4a16-ct` at revision `1d2c2d7f` for vLLM. The downloads differ: about 7.2 GB for Ollama's tag, its image part included, and about 10.3 GB for vLLM's checkpoint. The linear layers are 4-bit in blocks of 32 on both. The output layer differs, 6-bit in the GGUF and 16-bit in vLLM's checkpoint, as do the scale formats and the arithmetic: llama.cpp multiplies the 4-bit weights by 8-bit activations, vLLM's Marlin kernel by 16-bit ones. So for every token it writes, vLLM reads about 17 per cent more bytes from the card's memory, about 7.6 GiB against 6.5, nearly all of the difference in the output layer (1.88 GiB at 16 bits against 0.77 at 6).

**The same prompt, token for token.** Chat templates are where two-server comparisons often go wrong, so we took them out of the timed path: each prompt is rendered once with Google's own Gemma 4 template, reasoning off, and the same text goes to each server's raw endpoint (vLLM's `/v1/completions`; Ollama's `/api/generate` with `raw: true`). Before any scored run the bench checks that both servers turn it into the same token IDs, and stops if not. Asked the same chat questions with reasoning off (Ollama's `think: false` on `/api/chat`, or `reasoning_effort: "none"` on its OpenAI-style chat), vLLM 0.30.0 and Ollama 0.34.4 turned all four of our prompts into identical token IDs. Ollama's OpenAI-style `/v1/completions` applies the model's template again, as its `/api/generate` does unless `raw` is set, which is why road (d) above is void and not a speed. After every stream a gate counts generated tokens that never became visible text; any surplus voids the level as hidden reasoning (one closing end token is allowed on an answer that stopped itself; the proven pair's smoke run allowed two for a re-tokenizing edge, and read zero on every stream). Temperature 0 goes on every request, and in the sweeps each request opens with a tag from one seeded list, the same on both, so no prompt cache flatters either.

**What each server was allowed**, every flag registered before the scored runs it governs:

| Posture | What was set | Why |
|---|---|---|
| vLLM 0.30.0, the headline | A 5,120-token window and 512 tokens per step (labelled tuned); text only, so its image, audio and video parts are never built; the checkpoint's own 16-bit type and seed 0 pinned; loopback only; its FlashInfer sampler off (the bench computer has no compiler for it); cached-token and per-request timing reports on. The rest at its defaults: 0.92 of the card, prefix caching on | 5,120 is one Ollama slot; 512 is the step Ollama's runner uses; every prompt here is text |
| vLLM 0.30.0, its default step | The same at its default 2,048 tokens per step; one row, long prompt, sixteen requests | Whether its default queues there is worth knowing |
| Ollama 0.34.4, sixteen slots | Sixteen slots of 5,120 tokens, default cache type and attention, kept loaded, cloud calls and its Vulkan backend off; each request carries `num_batch: 512`; the model's image part loaded, as Ollama does by default | Without `num_batch`, on 0.32.13, sixteen slots of 5,120 tokens pushed part of the model off this card (46 of the 49 layers llama.cpp counts stayed on it, the rest in the computer's memory); 0.34.4 was not run without it |
| Ollama 0.34.4, its default | One slot and the window its memory tier picks; set only to stay loaded, with its cloud calls and Vulkan backend off | What you get if you change nothing that affects speed |

**Levels and runs.** One, two, four, eight and sixteen requests, released together; ten scored runs at one and two, three above, each level after a warm-up. A level's spread is (highest − lowest) ÷ median, and a prediction whose band is narrower reads UNDECIDED. vLLM runs again at one and sixteen after the Ollama block, to catch drift.

**What we predicted.** A refuted prediction stays on this page.

| ID | The prediction | Reading |
|---|---|---|
| P-1 | One request: decode within ±10 per cent between vLLM and Ollama with sixteen slots | CONFIRMED on the proven pair, vLLM 0.27.1 and Ollama 0.32.13; on the versions benched, CONFIRMED: vLLM 0.30.0 82.693 tokens a second, Ollama 0.34.4 with sixteen slots 85.007, ratio 0.9728; answers differ at this level |
| P-2 | One request: Ollama's first token within twice vLLM's | REFUTED: Ollama 0.34.4 with sixteen slots 441.866 ms over vLLM 0.30.0's 212.999 ms, a ratio of 2.0745 |
| P-3 | vLLM's total first reaches 1.2 times Ollama's with sixteen slots at four (undecided at two or eight) | UNDECIDED: the crossover read 2, a level marked undecided before any run, at 1.2077 on a spread of 0.0544; at four and eight the sixteen-slot spreads are wider than the band; the long-prompt row also reads UNDECIDED: vLLM's total there was over 1.2 times Ollama's at every level, but its spread at one request is wider than the band, so no crossover is decided; answers differ at every level here |
| P-4a | The same question again behind a shared prompt: both serve the prompt from cache | REFUTED on a first reading judged within one 16-token step, smaller than the 64-token step vLLM uses for this model; a second reading at that step, from a corrected instrument, follows in a dated update beside this one |
| P-4b | A new question behind it, four or more interleaved: vLLM serves the prompt from cache; `llama-server` reads it again unless a checkpoint lands at its end | REFUTED on a first reading judged within one 16-token step; a second reading at vLLM's own 64-token step follows in a dated update beside this one |
| P-5 | Sixteen requests, long prompt, first token 95th percentile: vLLM 1.5 s or less, Ollama with sixteen slots 4 s or less | REFUTED on both halves: vLLM's first-token 95th percentile read 17,325.908 ms at its best run against a bar of 1,500; Ollama's with sixteen slots 107,307.61 ms at its best run against 4,000 |
| P-6 | The two agree on at least 97 per cent of the six-way's 108 items at one request | REFUTED: 63 of 108 (58.33 per cent) on the strict scorer, by the registered count, under 97; read leniently 98 of 108 (90.74 per cent), printed beside and never in the ratio |
| TT | Window set at the server, none in the request: Ollama's load step 50 ms or less | CONFIRMED, 1.284 to 3.111 ms on 2026-09-28; the replication of 2026-09-30 is in [the first word](#the-first-word) |
| P-BOOT | vLLM's start-up reports room for sixteen requests at the full window | Arm D's first pass, 14.78 requests, on a boot that compiled cold: REFUTED. Arm D's re-run, on a boot that loaded the saved compile: 17.07 requests, CONFIRMED. Both conditions stand; [what each holds](#what-each-holds) explains the gap |
| C-1 | Ollama at its default is a queue: its total within ±10 per cent of its one-request figure at every level | CONFIRMED at two, four, eight and sixteen: ratios 0.9880, 0.9876, 0.9873 and 0.9869; answers differ at every level here |
| C-2 | vLLM holds at least 90 per cent of the card at rest | Arm D's first pass, 84.83 per cent, on a boot that compiled cold: REFUTED. Arm D's re-run, on a boot that loaded the saved compile: 91.57 per cent, CONFIRMED. Both conditions stand |
| C-3 | No outbound connection from either server with the switches set | Its first reading joins this page with arm G's rows, in a dated update; a second reading from a corrected instrument follows beside it |

*Readings by arm: P-1 on the proven pair is arm S's (2026-09-28); P-1 on the versions benched, P-2, P-3, P-5 and C-1 are arm C's (2026-10-01, the later build of the kit); TT, P-BOOT and C-2 are arm D's (2026-09-28, and its re-run on 2026-09-30); P-4a and P-4b are arm P's and P-6 is arm Q's (both 2026-09-30); C-3's reading is arm G's, which joins this page in a dated update.*

**Our own mistake, kept.** The first full pass on the new versions, arm D on 2026-09-28 from 17:20 to 17:30 UTC, stopped at two of our own gates: every text our client sent vLLM 0.30.0 arrived without the model's start token, `<bos>` (547 token IDs where the other two servers read 548). Our client stripped it on the rule that each server adds one back; this checkpoint's tokenizer adds none, and the stand-in servers we rehearsed against did, so the rehearsal missed it. The fix keeps the token in the text for vLLM, and the stand-ins now add none. D ran again first, on 2026-09-30: from 20:00:02 to 20:07:18 UTC; its receipt reads `RECEIPT ovv D ovv-20260930T200002Z-D frame=ok render=differs:llama-server raw_parity=368/368 nonce_counts=ok eligible=series-p512,long-write,long-prefill tt=CONFIRMED replan_c=re-planned calibration=ok egress=ready void_levels=1 halted=no` (`render=differs:llama-server` is `llama-server`'s own template render, the generic one Ollama starts it with and never uses for this model; the raw path the bench timed read the same IDs on all 368). On the first pass, Ollama and `llama-server` agreed ID for ID on all 364 texts, and this page's D readings from that pass come from them, from vLLM's start-up, or from vLLM's chat render, whose template writes the `<bos>` itself (548 IDs, the same as Ollama's); the missing token touches none of them.

**Arm C's own refusals, kept.** Arm C was refused by its own gates three times before it ran whole, and each refusal is a finding, not a void. On 2026-09-30 the bench refused the second boot, vLLM at its default step, before any of its cells ran: the kit read that posture's step back from the server's log, found no line stating it, and stopped rather than guess. The fix taught the kit the line vLLM 0.30.0 prints for a default it never states. On 2026-10-01 the kit's own test suite refused that fix twice on the bench computer, first on one test the suite did not name, then on one it did, a timing bound on the stand-in servers; each got a fix of its own, the suite then passed on every leg, and only then did C run, under a fifth amendment to the pre-registration pushed before any C model call.

**Two days, two builds of the kit.** This page's first release of readings, release 1, comes from two UTC dates; arm G's rows land in a second. Arm D's re-run and arms P, Q, A and G ran on 2026-09-30, on one build of our bench kit and under the pre-registration's fourth amendment. Arm C, the reading of record for every level above one request, ran on 2026-10-01, on the later build that the refusals above produced and under the fifth amendment. The judges and the measured path are the same in both builds, line for line; the diff between them touches the read-back that refused, the test suite and the receipts. Every table that mixes a C figure with a D, P, Q, A or G figure says so in its Reading column or its caption.

## What else the bench read {#what-else-the-bench-read}

The findings page, when it comes, draws its recommendations from this bench and links each one back here. These are the readings it will lean on that the sections above do not carry.

**Structured answers, and the chat path's token IDs.** Asked for a structured answer, a JSON schema alone and with reasoning off, ten requests each on both servers: how many of each server's answers parsed and validated, in four variants of ten, joins this page with arm G's rows in a dated update. The token-ID check on the chat path is in [the method](#what-was-written-down-before-the-first-run).

**Which routes answered without a key.** Arm A, on 2026-09-30 from 21:31:19 to 21:35:27 UTC, called each server's routes on the loopback address only, on the versions benched, over the routes each project's own documentation names, and recorded which of those each server answered without a key: on vLLM 0.30.0, 7 routes that its own security guide names as served without the key, each of which answered without it; on Ollama 0.34.4, the 5 routes in the table, and its own documentation says it has no key system; the table below lists each. The table claims nothing about any route it does not list.

*The routes in this table: for vLLM 0.30.0, routes its security guide at the tag names as served without the API key, each of which answered without it; the guide names others, which this table does not list. For Ollama 0.34.4, five of its API routes, none of which adds, changes or removes a model. Each was called on the loopback address by arm A on 2026-09-30. Yes means the route answered without a key; a dash marks a route that is not on that server's list.*

| Route | Ollama 0.34.4 | vLLM 0.30.0 |
|---|---|---|
| `POST /api/generate` | Yes | — |
| `GET /api/ps` | Yes | — |
| `POST /api/show` | Yes | — |
| `GET /api/tags` | Yes | — |
| `GET /api/version` | Yes | — |
| `POST /detokenize` | — | Yes |
| `GET /health` | — | Yes |
| `POST /invocations` | — | Yes |
| `GET /load` | — | Yes |
| `GET /ping` | — | Yes |
| `POST /tokenize` | — | Yes |
| `GET /version` | — | Yes |

**What left the machine.** Arm G, on 2026-09-30, started each server in a sandbox whose only network was the loopback address, with every system call recorded and every name lookup sent to a resolver on loopback that refused it. No name server and no listener ran, so the count is of attempts, and the window ran from launch to 120 s after ready, which is inside vLLM's ten-minute report interval and Ollama's four-hour one. C-3's first reading joins this page with arm G's rows, in a dated update, and a second reading from a corrected instrument follows beside it. The method is the one we used on [ONNX Runtime's telemetry on Linux](https://research.strata2signal.com/onnx-runtime-telemetry-on-linux/).

**What Ollama's default runner was started with.** Arm G read the arguments of the runner Ollama 0.34.4 started for this model at its defaults on this card, straight from the process table, and it sent each default instance a prompt longer than its window; both readings join this page with arm G's rows, in a dated update.

**What it took to get here, with receipts.** Ollama was one download, then `ollama pull` and a model name; on this card its server was ready in about 4 s, and the model loaded in about 5 s on 0.32.13 (on 0.34.4, 38.9 s on its first start, not yet explained, and 8.24 s on its second). It asked one change of us at sixteen slots, `num_batch: 512` with each request, so the whole model stayed on the card. vLLM asked more, and our own records show it. Seating this same Gemma 4 12B under vLLM 0.27.1 on a 12 GB card, the house's two RTX 3080 Ti boards, an NVIDIA GeForce RTX 3080 Ti Founders Edition (12 GB) and an EVGA GeForce RTX 3080 Ti XC3 Ultra (12 GB), one at a time, took 27 recorded experiments on 2026-09-25 before it would start with room to answer. Our lock installs vLLM 0.30.0 beside `transformers` 5.14.1, the version this model needed on the vLLM version we ran before it. The bench computer lacks the compiler toolkit one of vLLM's optional speed-ups needs, its FlashInfer sampler, so we switched that off. A first start on a new vLLM version compiles before it serves: 144.35 s from launch to ready on this card, 56.95 s of it compiling. Cold starts of vLLM 0.27.1 on our RTX PRO 6000 Blackwell (96 GB), with larger models, took about four to eleven minutes, by our serving bench's records of 2026-08-25. None of this is a flaw; it is the price of a server built to be tuned, mostly paid once per setup.

**What has shipped since.** Since the versions benched, Ollama 0.34.4 (released 2026-09-23) and vLLM 0.30.0 (released 2026-09-22), each project's GitHub release record, read on 2026-10-07 (UTC), lists three full releases of Ollama, the newest 0.40.0 on 2026-10-06, and one of vLLM, the newest 0.31.0 on 2026-10-05; pre-releases are not counted. Whether any default this page states moved in them is not read here: the versions benched are the ones every figure names.

## What it does not say {#what-it-does-not-say}

- One card, an RTX 3090 (24 GB); one model family, Gemma 4; one operating system, Linux. Not two or more cards, and not a model larger than the card. A later piece may measure a Mac; no date.
- Not other servers: not SGLang, not LM Studio, not Hugging Face's text-generation-inference, now archived.
- Requests released together are a burst, not people arriving over time.
- Not other request shapes: every speed reading came from requests capped at 256 tokens out at temperature 0, and no timed prompt ran past about 2,600 tokens; long answers, sampled answers and long windows were not measured.
- Not bit-identical weights: both servers ran Google's 4-bit Gemma 4 12B, but the two downloads differ in the output layer, the scale formats and the arithmetic ([the same weights, said exactly](#what-was-written-down-before-the-first-run)).
- Not any production traffic, ours included, and not the versions after those benched.
- Not a comparison of the two servers' answers beyond the six-way: P-6 says they differ, and how often; it does not say which is right.
- What the bench's own gates voided in release 1: none in arms C, P and Q, whose receipts each read void_levels=0, with none of C's thirty cells voided; arm A, a census with no levels, read a reply from every route the routes table lists; arm G's reading waits for its rows. Arm D's one void level is road (d)'s, a templated-again request that spent its tokens reasoning; road (c)'s cells read VOID too, because a reload is that road's registered reading, which the kit counts as no void level. The three refusals are named above, as findings.
- Two verdicts on this page are first readings only, each labelled so: the two shared-prompt verdicts, P-4a and P-4b; a third, C-3, prints with arm G's rows. The second readings of all three, from a corrected instrument, will print beside the first, each naming what it counts, dated at both ends of its window (UTC) and compared with the first reading in one sentence, and will not replace it.

## What comes next on this page {#what-comes-next}

A later day adds a dated block here, on the same card, changing nothing above: `llama-server` as Ollama bundles it, as its own series, so Ollama's layer reads at every level; Gemma 4 E4B at 16 bits on both, a control with no 4-bit arithmetic; Ollama's default `gemma4:26b` tag against a popular community 4-bit build for vLLM, fitting whole or refused, which is a comparison of two builds and will never be headlined as the same model on both; whole-computer watts under load and at idle; three cold starts of each. The block will open with what it counts, its dates at both ends, and how it compares with this release.

## What to take with you {#what-to-take-with-you}

- **One request at a time, the two are close, and that is the registered result twice over.** Decode fell inside our ±10 per cent band on the proven pair, Ollama 0.32.13 against vLLM 0.27.1, and again on the versions benched, Ollama 0.34.4 against vLLM 0.30.0 (P-1, CONFIRMED on both). The first word is another matter: Ollama's wait at one request was more than twice vLLM's on the versions benched (P-2, REFUTED), and on 0.32.13 most of a second of it was a load step that 0.34.4 no longer has (TT, CONFIRMED).
- **Many at once is where they part, and Ollama with sixteen slots does not keep up.** vLLM's total reached 1.2 times Ollama's with sixteen slots at the first level above one request, by a margin within run noise, and kept climbing to sixteen; Ollama's sixteen-slot total went flat from four requests on, with runs that disagreed with each other at four and at eight. The bet on where the crossover would fall read UNDECIDED, because it fell at a level the pre-registration had marked undecided. Ollama at its default of one slot is a queue, and reads as one at every level (C-1, CONFIRMED).
- **Sixteen long prompts at once blew through both first-token bars** (P-5, REFUTED): neither the 1.5 s we bet for vLLM nor the 4 s for Ollama with sixteen slots held, and on Ollama with sixteen slots the longest waits ran well past a minute. A shared prompt reused on every request is served from cache by both, differently, by each server's own count; the two registered verdicts on it, P-4a and P-4b, read REFUTED on a first reading judged within one 16-token step, smaller than the 64-token step vLLM uses for this model, and their second readings, from a corrected instrument, will print beside the first.
- **What each holds depends on a condition the page names every time.** vLLM reserves most of the card at start; whether its first boot on your computer leaves room for sixteen requests at the full window turns on whether vLLM has already compiled the model there with the same serve line (P-BOOT and C-2: REFUTED on the cold boot, CONFIRMED on the warm one this page prints). Ollama's reservation follows its slots and window and comes back when the model unloads.
- **The answers differ, so every speed row says so.** Scored strictly, the two servers agreed on fewer of the six-way's 108 items than the 97 per cent we bet (P-6, REFUTED); read leniently, on most of them, and on every strict disagreement both named the same guest.

## How to check our work — and see it live {#how-to-check-our-work-and-see-it-live}

- **The kit.** [The data kit](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/data/) sits beside this page, taken from the bench's own record by a fixed list and by rule: the frozen prompts with their sha256; arm S's per-request rows and board readings for both servers; the board readings and token-ID parity checks of arms D (both runs), C, P and Q, with the run receipts of arms D, C and P; the results files [its README](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/data/README.md) lists, as the bench wrote them apart from the digests withheld in them; arm A's twelve route rows, labelled as a cut; and the client modules the README lists, a reading set rather than a program you can run. The README names, file by file, every file the rules could not clear for release, among them the per-request rows of every arm but S, the serve-line files and the pre-registration; the serve lines are printed below as code. Row files carry a card label, never a serial or bus address.
- **The prompt** this page leans on is `series-p512`, sha256 `90eedd0c53f9554ae3837674504fcb7090013d9a432a653352183c0f25a7ce5c`, and the long prompt is `long-write`, sha256 `e7fb0b295f6de90744d01abceec0e2273f80a4c4a1ddb32df493203952d9f3db`, each read off the kit's own copy; the kit's parity files give each prompt's token count and the verdict of the check, made before any scored request, that both servers turned it into the same token IDs.
- **The six-way** is already public: [the kit for Reading the answer](https://research.strata2signal.com/reading-the-answer/data/kit/task_c.json), sha256 `61a0b1d6e7cf66fd91992c0c3a4572f6fbae90e6471048fd7dc1eb41ff2d7f39`.
- **The pre-registration** and its amendments, each pushed before the scored runs it governs (the first, found by a dry run, before any scored call), are not in the kit: the rules could not clear the file for release, a hand-edited copy would not be a record, and [the kit's README](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/data/README.md) names it absent. Every prediction, its bar and its reading are printed on this page, in [what we predicted](#what-was-written-down-before-the-first-run).
- **Run one level yourself.** Pull `gemma4:12b-it-qat` into Ollama and `google/gemma-4-12B-it-qat-w4a16-ct` at revision `1d2c2d7f` into vLLM, start each as below, send the same rendered text to each raw endpoint with the `<bos>` the template writes kept at its start, and time the first and last tokens on your own clock. Each Ollama request goes to `/api/generate` with `"raw": true` and `"options": {"num_batch": 512}`.
- **See both live.** [Ask About This Page](https://research.strata2signal.com/ask-about-this-page/) answers on vLLM; [the long table](https://research.strata2signal.com/a-dinner-party-for-the-dead/) seats its guests on both. Both rooms are open: the assistant's door is the "ask about this page" line at the top of every article on this shelf, and the long table's is at [longtable.strata2signal.com](https://longtable.strata2signal.com/).

The two serve lines below are the bench's own, kept as receipts: they carry the switches the bench set so that its outbound count could be read, and they start each server as this page measured it; the findings page is where running either server quietly is set out. Ollama 0.34.4, sixteen slots, as the bench started it (its Vulkan backend off, so the card is used through CUDA alone):

```
OLLAMA_NUM_PARALLEL=16 \
OLLAMA_CONTEXT_LENGTH=5120 \
OLLAMA_KEEP_ALIVE=-1 \
OLLAMA_NO_CLOUD=1 \
OLLAMA_VULKAN=0 \
ollama serve
```

vLLM 0.30.0, the headline posture (our lock installs it beside `transformers` 5.14.1; on a computer without the CUDA compiler, add `VLLM_USE_FLASHINFER_SAMPLER=0`, as we did):

```
VLLM_NO_USAGE_STATS=1 vllm serve \
  google/gemma-4-12B-it-qat-w4a16-ct \
  --revision 1d2c2d7f2466070e69d6fb3fd5ce9a7d75f2f6ee \
  --max-model-len 5120 \
  --max-num-batched-tokens 512 \
  --dtype bfloat16 --seed 0 \
  --limit-mm-per-prompt '{"image":0,"audio":0,"video":0}' \
  --enable-prompt-tokens-details \
  --enable-per-request-metrics \
  --host 127.0.0.1
```

## The rest of the seminar {#the-rest-of-the-seminar}

- [A short history of Ollama](https://research.strata2signal.com/a-short-history-of-ollama/): the project and the company behind one of these two servers, and every default of its we found that moved, dated.
- [A short history of vLLM](https://research.strata2signal.com/a-short-history-of-vllm/): the lab, the paging idea and the foundation behind the other, and its defaults that moved.
- The distilled findings and recs (coming soon): which to pick when, the settings that matter, and how to run both quietly, each recommendation linking its evidence on this page.
- [What 150 Watts Buys](https://research.strata2signal.com/what-150-watts-buys/) and [Two Hours at 12 tok/s — On Battery](https://research.strata2signal.com/two-hours-on-battery/): pages that promised a write-up of our serving bench of 2026-08-25, run on an RTX PRO 6000 Blackwell (96 GB) with other models; this page takes its place.
- [Reading the answer instead of writing it](https://research.strata2signal.com/reading-the-answer/): the six-way, and the first-token figure revisited here.
- [The free speed wasn't free](https://research.strata2signal.com/the-free-speed-wasnt-free/): a sampler default that moved.
- [The compressed photograph](https://research.strata2signal.com/the-compressed-photograph/): what the 4-bit tags mean.
- [RTX 5080 vs RTX 3080 Ti](https://research.strata2signal.com/rtx-5080-vs-rtx-3080-ti/): Ollama's planner deciding what fits.
- [A Dinner Party for the Dead](https://research.strata2signal.com/a-dinner-party-for-the-dead/) and [How the long table works](https://research.strata2signal.com/how-the-long-table-works/): both servers at work.
- [Four homes for one reranker](https://research.strata2signal.com/four-homes-for-one-reranker/): the gateway in front of our seats.
- [ONNX Runtime's telemetry on Linux](https://research.strata2signal.com/onnx-runtime-telemetry-on-linux/): an outbound call, measured, by the method arm G used.

## Who ran this, and thanks {#who-ran-this-and-thanks}

The six-way's 108 lines are our own, from [the long table](https://research.strata2signal.com/a-dinner-party-for-the-dead/), published with their scorer in the kit for Reading the answer. **Google** made Gemma 4 and its QAT weights and released them under the Apache 2.0 licence; **Hugging Face** hosts them and makes **transformers**, a library vLLM depends on. **Ollama**'s team and contributors built one server (MIT), on **llama.cpp** (MIT), begun by **Georgi Gerganov** and kept by the ggml authors, whose `llama-server` README states the defaults this page reads; the **vLLM** project (Apache-2.0), begun in UC Berkeley's **Sky Computing Lab** and hosted by the **PyTorch Foundation**, built the other, on **PyTorch** (BSD 3-clause). On this card both run over **CUDA** (NVIDIA, proprietary, the one closed piece in either stack). The card is an **EVGA** board on an **NVIDIA** chip. None of them owed us anything.

A small human team asked for this page, set the rules and the predictions before the first run, chose what it would and would not claim, and signed off on every figure; a fleet of AI agents wrote the pre-registration, ran and checked each arm, and wrote this page under that team's rulings.

## Sources {#sources}

Read 2026-09-28 to 2026-10-01 (UTC); every version and default line was read again on 2026-10-07 (UTC), the day this page went up.

- **Ollama** source at tag v0.34.4: `llm/llama_server.go` (the runner's flags, `cache_prompt: true` on every generation request, and the two chat paths, one of which starts `llama-server` with a generic template it never uses for this tag), `envconfig/config.go` (the defaults the postures above set or left), `server/routes.go` (the window tiers by card memory) and `server/sched.go` (the reload on a larger window); the `gemma4:12b-it-qat` tag's manifest (the GGUF, its image projector and its default thinking); `docs/faq.mdx` and `docs/context-length.mdx` at the tag (the slots and the window, set together).
- **llama.cpp** at build b11081, the one Ollama 0.34.4 bundles: `tools/server/README.md` (the prompt cache's defaults, its checkpoints and how a request picks a slot), read on the bench beside the bundled program's own behaviour.
- **vLLM** source at tag v0.30.0: `vllm/config/cache.py` (the 0.92 share), `vllm/engine/arg_utils.py` (the batch tiers by card memory, and prefix caching on by default), and the serve log's own lines for the six cache groups and the 64-token step on this model.
- **Google's checkpoint** at revision `1d2c2d7f`: `chat_template.jinja`, `tokenizer.json` and `tokenizer_config.json` (the template, and no `<bos>` added to raw text).
- **Our records:** the pre-registration and its five amendments; arm S (2026-09-28, 11:29:59–11:33:35 UTC) and its independent check; arm D and TT (2026-09-28, 17:20:32–17:30:20 UTC), whose results quote vLLM 0.30.0's compile times (56.95 s of `torch.compile` inside 97.84 s of engine setup) from that start's own serve log; arm D's re-run and TT's replication (2026-09-30, 20:00:02–20:07:18 UTC); arm C (2026-10-01, 13:35:02–14:29:21 UTC); arm P (2026-09-30, 20:43:52–21:19:31 UTC); arm Q (2026-09-30, 21:27:35–21:31:13 UTC); arm A (2026-09-30, 21:31:19–21:35:27 UTC); arm G's window joins with its rows; the bench records of 2026-08-25 (the serving bench on an RTX PRO 6000 Blackwell (96 GB)) and 2026-09-25 (the 12 GB seating diagnosis on the two RTX 3080 Ti boards). In the kit, as the bench wrote them apart from the digests withheld in them: the results of [arm D's re-run](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/data/kit/records/D-2026-09-30-RESULTS.md), [arm C](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/data/kit/records/C-RESULTS.md) and [arm P](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/data/kit/records/P-RESULTS.md); [the kit's README](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/data/README.md) names each other record above as absent, or links it where it is already published.
- **This workshop's own released pages, linked above:** [Reading the answer instead of writing it](https://research.strata2signal.com/reading-the-answer/) (the six-way and its kit; the 870 ms this page revisits), [What 150 Watts Buys](https://research.strata2signal.com/what-150-watts-buys/), [Two Hours at 12 tok/s — On Battery](https://research.strata2signal.com/two-hours-on-battery/), [ONNX Runtime's telemetry on Linux, measured](https://research.strata2signal.com/onnx-runtime-telemetry-on-linux/) (the method arm G used), [A Dinner Party for the Dead](https://research.strata2signal.com/a-dinner-party-for-the-dead/), [How the long table works](https://research.strata2signal.com/how-the-long-table-works/), [Ask About This Page](https://research.strata2signal.com/ask-about-this-page/), [A short history of Ollama](https://research.strata2signal.com/a-short-history-of-ollama/) and [A short history of vLLM](https://research.strata2signal.com/a-short-history-of-vllm/).

*First readings 2026-09-28 (UTC) on one EVGA GeForce RTX 3090 XC3 Ultra (24 GB) at 300 W, one server at a time; release 1's readings from arm D's re-run and arms P, Q and A on 2026-09-30, 20:00:02 to 21:35:27 UTC, with arm G's to join them, and from arm C on 2026-10-01, 13:35:02 to 14:29:21 UTC, on a later build of the bench kit.*

Corrections and later measurements will be added below, each dated (UTC), with a window at both ends where one applies and saying in plain words what it counts.

<!-- derived 2026-10-07 (UTC) by tools/derive_md.py from the pour source.
     source html sha256: 3b6bde653ded637d6b9193396d3d092dc8a2cbaef1f43965f75e363c11397a62
     derivation sha256:  296ac48df65ab94ab9cd167d101728fdd4350e8ec93f72f7f454bf507dd7de38
     the {#id} on each heading is the anchor that heading carries on the page. -->
