# Ollama or vLLM

*Ollama and vLLM are two open-source programs that run an open language model on your own computer's graphics card and answer requests over a local web address. This page is the advice page of a four-page set on the two: for the case you bring, which one to run, the two or three settings that decide it, who can reach it once it is running, and what each sends out of your machine by default. Every line points at the measurement on our bench page, the dated source on one of the two histories, or another of this workshop's pages named beside it; the safety paragraph reads vLLM's own repository on GitHub, at the tags and on the pages the last bullet under Sources names, and nothing here was measured or read anywhere else. Benched: Ollama 0.34.4 (released 2026-09-23) and vLLM 0.30.0 (released 2026-09-22), on one EVGA GeForce RTX 3090 XC3 Ultra (24 GB), with the same 4-bit Gemma 4 12B weights on both.*

*Published 2026-10-08 (UTC) · A small (human) team and a fleet of AI agents.*

**the short version:** Ollama and vLLM are two programs that run an open language model on your own graphics card, and which one to run depends mostly on how many requests reach it at once. Every figure here is our bench page's, from one RTX 3090 (24 GB) with the same 4-bit Gemma 4 12B weights on both, where each reading sits beside the prediction we wrote down before measuring it; speeds are in tokens a second, a token being a word or a piece of one, and Ollama 0.34.4 is set to serve sixteen requests at once in every comparison here (at its default it serves one at a time). One request at a time, run either: once an answer had started, vLLM 0.30.0 wrote it at 82.693 tokens a second and Ollama at 85.007 tokens a second, inside the 10 per cent we had predicted, though Ollama's wait before the answer started was more than twice vLLM's, 441.866 ms against 212.999 ms. Two to four at once, run vLLM, or Ollama set to serve more at once if another case below keeps you on it: vLLM was ahead on every speed and every wait the bench page prints at two and at four, its total across both requests 1.2077 times Ollama's at two, though at four Ollama's repeated runs disagreed widely. Eight to sixteen at once, run vLLM: at sixteen requests released together it wrote 506.829 tokens a second across them and Ollama 142.327 tokens a second, 3.561 times as much, and Ollama's total stayed flat from four requests on. Whichever you run, the two did not always give the same answer to the same test question, so check the answers on your own prompts. On this page, each of these cases links the section of our bench page, Ollama and vLLM on one RTX 3090, that holds its figures.

8,568 words · about 39 minutes (at 220 words/min) · 1 table · data kit: no

https://research.strata2signal.com/ollama-or-vllm-which-to-run/

---

*If you're new here: [strata→signal](https://strata2signal.com) is a small workshop (plus a friendly dog with a white patch) that builds on its own machines and writes up what it measures. Our products run on both servers: [the assistant on our pages](https://research.strata2signal.com/ask-about-this-page/) runs on vLLM, and the long table, our dinner-party exhibit, seats its guests on both ([how the long table works](https://research.strata2signal.com/how-the-long-table-works/)). This page gives advice; the evidence and the method are on [the bench page](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/), and every recommendation here links the section of it that the recommendation rests on.*

## Who this page is for {#who-this-page-is-for}

You have one graphics card, or are about to buy one, and you have met both names. You might be running a home assistant for the family, one person's coding agent, a classroom or an office of about twenty, or a small shop that wants chat and embeddings on one card. You are not running a cluster, and you do not need to read the other three pages first: this one sends you to them where a claim needs its evidence.

Some words this page leans on, in the bench page's sense. A **token** is a word or a piece of one. A **slot** is room for one request; Ollama serves as many requests at once as it has slots, one by default, and reserves a whole **window** of memory for each, the most tokens one request may hold. Ollama holds the model in a separate program it starts for it, its **runner** (llama.cpp's `llama-server`). vLLM instead reserves most of the card at start and hands working memory out in small pages as requests grow. Three speeds: the **first word** (the bench page's first token) is the wait before an answer starts; **decode** is how fast one answer is written after that; **total speed** is every token written across all the requests at once, over the time from sending to the last one finishing. [Two names, one card](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#two-names-one-card), on the bench page, says the rest in a screen.

And these from the bench's method, which [the words the bench page leans on](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#the-words-this-page-leans-on) gives in full. Each **prediction**, which this page also calls a bet (P-1, P-2, C-1 and the rest), was written down with its rule and put on record before the runs it governs, and reads CONFIRMED, REFUTED or UNDECIDED by that rule; together the predictions and their rules are the bench's **pre-registration**, and a figure a prediction is judged against is its **bar**. An **arm** is one timed job of the bench, named by a letter. A **level** is a number of requests released together: one, two, four, eight or sixteen. A prediction's **band** is the margin written with it: ±10 per cent for P-1, and 0.2, that is 20 per cent, for the bet on where vLLM's total would first reach 1.2 times Ollama's (the predictions table in [What was written down before the first run](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#what-was-written-down-before-the-first-run) carries P-1's, and [Requests at once](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#requests-at-once) the 0.2). A level's **spread** is how far its runs disagreed: its highest run less its lowest, divided by the median, printed as a fraction (0.2 is 20 per cent). Where a level's spread is wider than a prediction's band, the runs disagreed by more than the margin being tested, so that prediction reads UNDECIDED there, and the level's figures print as a **reading**: a measurement, with no verdict on the bet. A **posture** is one server started with one set of settings; the bench ran four, vLLM 0.30.0 at its headline settings and at its default step, and Ollama 0.34.4 with sixteen slots and at its default of one slot. A **first reading** is a verdict whose second reading, from a corrected instrument, will print beside it, not in its place.

## Which to pick when {#which-to-pick-when}

Each row is a case readers bring, with our lean and the one reading that decides it; the list under the table carries each case's figures, conditions, costs and the rest of its links. *Measured* means the bench page prints the reading, and the link goes to the section that holds it; *sourced* means one of the histories dates the fact from the project's own documentation or code, and the link goes there. Most readers fall in more than one row. Where two rows lean apart, read the platform row first: vLLM is a Linux server (on Windows, officially, only through the Windows Subsystem for Linux) and on NVIDIA needs the RTX 20 series or newer, so where it cannot run the other rows do not arise; after that, the list under the table gives each side's cost. Every speed reading on the bench page carries the label "answers differ at this level", because on the answer test we registered the two servers did not always give the same answer, and speed means little if the answers change; the note after the list says how far apart.

| Your case | Our lean, and how we know |
|---|---|
| One request at a time, on one machine (a family's assistant, most of the time) | **Either.** Decode is inside our ±10 per cent band twice over; Ollama's first word takes more than twice as long. *Measured:* [One request at a time](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#one-request-at-a-time), [The first word, and a shared prompt](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#the-first-word). |
| A few at once (two to four) | **vLLM, or Ollama with its slots and window set together.** vLLM's total led at both levels; Ollama's half is for the reader another row keeps there. *Measured:* [Requests at once](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#requests-at-once). |
| One person's coding agent | **Not measured at its shape.** It can send several at once; where its prompts or windows run past what the bench sent, nothing here measured it. See the row above for speed, and the window settings before raising either. *Not measured:* [What it does not say](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#what-it-does-not-say). |
| Many at once (eight to sixteen): an app, a team | **vLLM.** Its total at sixteen was 3.561 times Ollama's with sixteen slots, and Ollama's was flat from four on. *Measured:* [Requests at once](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#requests-at-once). |
| One long prompt reused all day (an assistant with a long system prompt) | **vLLM for the first word; both reuse the prompt, by each server's own count.** On the same question again at one request, 48.102 ms on vLLM against 227.911 ms on Ollama with sixteen slots. The two verdicts on that reuse are first readings, and a new question behind the prompt carries a condition; the list states both. *Measured:* [The first word, and a shared prompt](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#the-first-word). |
| I switch models often | **Ollama.** It loads a model on request and lets it go after five idle minutes; vLLM serves one model per process. *Sourced:* [how Ollama serves a request](https://research.strata2signal.com/a-short-history-of-ollama/#how-it-serves-a-request), [how vLLM does](https://research.strata2signal.com/a-short-history-of-vllm/#how-it-serves-a-request). |
| Chat and embeddings on one card | **Ollama.** One server loads both on demand; vLLM runs one model per process, so two processes share the card. *Sourced:* [what Ollama speaks](https://research.strata2signal.com/a-short-history-of-ollama/#what-it-speaks-and-where-it-runs), [what vLLM speaks](https://research.strata2signal.com/a-short-history-of-vllm/#what-it-speaks-and-where-it-runs). |
| Windows, an older NVIDIA card, an AMD card, Apple silicon | **Ollama; on an AMD card under Linux, either.** vLLM is a Linux server, and on NVIDIA needs the RTX 20 series or newer; on an AMD card under Linux both run, vLLM through ROCm and Ollama through ROCm or Vulkan, so the other rows decide; no Mac row here is measured. *Sourced:* [where Ollama runs](https://research.strata2signal.com/a-short-history-of-ollama/#what-it-speaks-and-where-it-runs), [where vLLM runs](https://research.strata2signal.com/a-short-history-of-vllm/#what-it-speaks-and-where-it-runs). |
| Other people or apps will reach it | **Either, behind something that checks keys.** Ollama checks none; vLLM's key covers part of its surface. [Who can reach it](#who-can-reach-it), below. |
| A classroom or office of about twenty pressing enter together (measured at sixteen) | **vLLM.** At sixteen the wait for the first word was far shorter on vLLM, and neither server met its long-prompt bar. *Measured:* [Requests at once](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#requests-at-once). |
| A model larger than the card; two or more cards | **Not measured here.** The bench page's [What it does not say](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#what-it-does-not-say) names both; the histories say what each server does by design ([Ollama](https://research.strata2signal.com/a-short-history-of-ollama/#what-it-speaks-and-where-it-runs), [vLLM](https://research.strata2signal.com/a-short-history-of-vllm/#what-it-speaks-and-where-it-runs)). No recommendation. |

**Each case, in more detail.**

- **One request at a time.** On the versions benched, vLLM 0.30.0 decoded at 82.693 tokens a second and Ollama 0.34.4 with sixteen slots at 85.007 tokens a second, vLLM's speed 0.9728 of Ollama's, inside our ±10 per cent band, where the bet had already held on the older pair we tested first (P-1, CONFIRMED twice); the first word at one request took 441.866 ms on Ollama with sixteen slots against 212.999 ms on vLLM, 2.0745 times as long, where we had bet on twice at most (P-2, REFUTED). On Ollama, the posture for this case is its default of one slot, which reserves one window rather than sixteen: at rest it grew the card's used memory by 8,761 MiB, against 16,389 MiB with sixteen slots, both read by arm C on 2026-10-01. On a new question behind a long shared prompt it waited 432.838 ms for its first word, where Ollama with sixteen slots waited 10,970.201 ms; both were read straight after the same server had answered sixteen at a time, and the bench page prints that as a condition and reads no cause into it (the shared-prompt case below). What you give up with vLLM: most of the card from the moment it starts, and a compile on the first start of a new version or a new serve line. *Measured:* [One request at a time](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#one-request-at-a-time); [The first word, and a shared prompt](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#the-first-word); vLLM's reservation and compile in [What each holds, and how fast it starts](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#what-each-holds).
- **A few at once (two to four).** At two, vLLM's total was 1.2077 times Ollama's with sixteen slots; at four, 1.6423 times, but there the wider of the two servers' spreads read 0.2973, wider than the bet's band of 0.2, so the bench reads that level as undecided. Per request at four, 72.183 tokens a second on vLLM against 51.713 tokens a second on Ollama with sixteen slots, where Ollama's three scored runs wrote at two different speeds, and that median is one of them. vLLM was ahead on every speed reading the bench page prints at two and at four. What you give up with Ollama: a whole window of memory for every slot, and speed at these levels. Pick Ollama here when another row keeps you on it, and set its slots to the requests you expect, with the window beside them; the bench ran Ollama with sixteen slots of 5,120 tokens and `num_batch: 512`, and at its default of one slot, and no count between. *Measured:* [Requests at once](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#requests-at-once); the two Ollama postures, in the postures table in [What was written down before the first run](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#what-was-written-down-before-the-first-run).
- **One person's coding agent.** It can send several requests at once, like the case above, and where its prompts or windows run longer than the bench ran, nothing here measured them: no timed prompt passed about 2,600 tokens, and the headline postures' window was 5,120 tokens. Read the case above for the speeds, and the window lines under [the settings that matter](#the-settings-that-matter) before you raise either server's window. *Not measured:* [What it does not say](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#what-it-does-not-say); the window the bench ran, in the postures table in [What was written down before the first run](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#what-was-written-down-before-the-first-run).
- **Many at once (eight to sixteen).** 506.829 tokens a second across sixteen on vLLM against 142.327 tokens a second on Ollama with sixteen slots, 3.561 times; at eight, 2.77 times, on a spread of 0.2352, wider than the band, so that level reads undecided; and Ollama's sixteen-slot total was flat from four on. The bet on where the crossover would fall reads UNDECIDED: it came at two, a level the pre-registration had marked undecided before any run. With long prompts, read the classroom case below: neither server met its first-token bar at sixteen. What you give up: Linux, one model per process, and a decision about who may reach it, since vLLM listens on every IPv4 address until told otherwise. *Measured:* [Requests at once](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#requests-at-once); *sourced:* [where vLLM runs](https://research.strata2signal.com/a-short-history-of-vllm/#what-it-speaks-and-where-it-runs).
- **One long prompt reused all day.** First reading. By each server's own count, the same question again at one request was served nearly whole from cache on both: 2,496 tokens of its 2,560-token request on vLLM, whose cache hit ends on a 64-token step, and 2,559 tokens on Ollama with sixteen slots. Our two bets on that reuse, P-4a (the same question again: both serve the prompt from cache) and P-4b (a new question behind it, at four or more at once: vLLM serves the prompt from cache, and Ollama's `llama-server` reads it again unless it saved the slot's state at the prompt's end), read REFUTED on a first reading: the rule they were judged by let the count served from cache fall short of the prompt by one 16-token step at most, while vLLM's cache for this model matches in whole 64-token steps. Their second readings, judged at vLLM's own step, will print beside them in a dated update. The first word was the shorter on vLLM on every row the bench page prints: on that request, 48.102 ms against 227.911 ms on Ollama with sixteen slots. The everyday shape of this case is a new question behind the same prompt, and there, one at a time, Ollama with sixteen slots waited 10,970.201 ms and its default of one slot 432.838 ms; both were read straight after the same server had answered the same question sixteen at a time, with no restart between, and the bench page prints that as a condition and reads no cause into it. *Measured:* [The first word, and a shared prompt](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#the-first-word).
- **I switch models often.** Ollama loads a model on the first request for it and, by default, lets it go after five idle minutes; vLLM serves one model per process and, by default, reserves 92 per cent of the card the moment it starts. Each switch costs a model load, and the bench page prints this model's loads, one of them not yet explained ([What each holds, and how fast it starts](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#what-each-holds)). *Sourced:* [how Ollama serves a request](https://research.strata2signal.com/a-short-history-of-ollama/#how-it-serves-a-request), [how vLLM does](https://research.strata2signal.com/a-short-history-of-vllm/#how-it-serves-a-request), [vLLM's defaults that moved](https://research.strata2signal.com/a-short-history-of-vllm/#the-defaults-that-moved).
- **Chat and embeddings on one card.** Ollama loads both on demand, several models at once up to its own limit; vLLM runs one model per process, so chat and embeddings are two processes, each taking a smaller share of the card. What you give up with Ollama: every embedding model is held at one slot, so embedding requests are served one at a time. If many people will also press enter at once, the many-at-once case applies to the chat model, and vLLM can serve the embedding model from a second process with its own smaller share of the card; the bench did not measure that pairing. *Sourced:* [what Ollama speaks](https://research.strata2signal.com/a-short-history-of-ollama/#what-it-speaks-and-where-it-runs), [how Ollama serves a request](https://research.strata2signal.com/a-short-history-of-ollama/#how-it-serves-a-request), [what vLLM speaks](https://research.strata2signal.com/a-short-history-of-vllm/#what-it-speaks-and-where-it-runs), [how vLLM serves a request](https://research.strata2signal.com/a-short-history-of-vllm/#how-it-serves-a-request).
- **Windows, an older NVIDIA card, an AMD card, Apple silicon.** vLLM is a Linux server (Windows, officially, only through the Windows Subsystem for Linux), needs an NVIDIA card of the RTX 20 series or newer, runs on AMD cards through ROCm, and on Apple silicon either processor-only or through a separate plugin for its graphics; Ollama runs on macOS, Windows and Linux, on NVIDIA cards back to compute capability 5.0 (the GTX 750, 900 and 10 series among them), on AMD cards through ROCm or Vulkan, and on Apple silicon through Metal or MLX (Apple's array library). No Mac row on this page is measured. *Sourced:* [where Ollama runs](https://research.strata2signal.com/a-short-history-of-ollama/#what-it-speaks-and-where-it-runs), [where vLLM runs](https://research.strata2signal.com/a-short-history-of-vllm/#what-it-speaks-and-where-it-runs).
- **Other people or apps will reach it.** Ollama's local server checks no key; vLLM's key guards four path prefixes, and its own guide says not to rely on the key alone. [Who can reach it](#who-can-reach-it), below, has the recipe. *Sourced:* [who can talk to Ollama](https://research.strata2signal.com/a-short-history-of-ollama/#who-can-talk-to-it), [who can talk to vLLM](https://research.strata2signal.com/a-short-history-of-vllm/#who-can-talk-to-it).
- **A classroom or office of about twenty.** At sixteen requests at once the wait for the first token, at the 95th percentile and the median of three runs, read 3,333.177 ms on vLLM and 20,935.24 ms on Ollama with sixteen slots; with sixteen long prompts at once neither met the bar we set it, 17,327.116 ms on vLLM against our 1,500 ms and 107,377.256 ms on Ollama with sixteen slots against our 4,000 ms, the medians again, and even each server's best run missed its bar (P-5, REFUTED on both halves). Twenty was not measured; sixteen is the nearest level. *Measured:* [Requests at once](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#requests-at-once).

*On the answers.* On our six-way test at one request, 108 lines each asking which of six dinner guests said it, the two servers agreed on 63 of 108 items (58.33 per cent) by the strict rule we registered, against the 97 per cent we bet (P-6, REFUTED). That count takes an item on which neither reply opens with a letter as agreeing, and 54 of the 63 are such items; read leniently, the letter found wherever it appears in the reply, they agreed on 98 (90.74 per cent). Scored strictly, vLLM 0.30.0 was right on 47.22 per cent and Ollama 0.34.4 with sixteen slots on 8.33 per cent. That gap is in how a reply opens, not in which guest it names: Ollama's replies open with "The correct answer is" on 99 of 108 and vLLM's on 51, the strict rule counts only a reply that opens with the letter, and on every strict disagreement both named the same guest. Read leniently, beside and never in the verdict, Ollama's is 79.63 per cent (86 of 108) and vLLM's 75 per cent (81 of 108). The same model in two downloads that differ in their arithmetic: [What was written down before the first run](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#what-was-written-down-before-the-first-run) says exactly how the two downloads differ, and [Are the answers the same?](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#are-the-answers-the-same) has the six-way's rows. Pick for your case, then check the answers on your own prompts.

## The settings that matter {#the-settings-that-matter}

Each is a default you can change, with the reason it matters here. Ollama's settings below are environment variables its server reads when it starts, except the fields a bullet names as travelling with each request (`num_batch`, `truncate`, `raw`, `think` and `reasoning_effort`); vLLM's are flags on its `vllm serve` command. The bench page's postures table lists every flag the bench set, in [What was written down before the first run](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#what-was-written-down-before-the-first-run), and its two serve lines, the starts it measured, close [How to check our work — and see it live](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#how-to-check-our-work-and-see-it-live).

**Ollama**

- **Slots and window are one memory budget; set them together.** Each slot reserves a whole window, so the runner is started with window × slots: with sixteen slots of 5,120 tokens, the runner on our card was told to hold 81,920 tokens. Ollama's own FAQ says the memory needed scales with `OLLAMA_NUM_PARALLEL` times `OLLAMA_CONTEXT_LENGTH`. Raise one and set the other beside it. *Measured and sourced:* [Two names, one card](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#two-names-one-card); [how it serves a request](https://research.strata2signal.com/a-short-history-of-ollama/#how-it-serves-a-request).
- **One slot by default is a queue, by design.** Since v0.10.0 Ollama serves one request at a time unless you say otherwise, so a second request waits for the first. On the bench its default posture wrote 74.545 tokens a second across one request and 73.565 tokens a second across sixteen released together: the same total speed, queued (C-1, CONFIRMED). *Measured:* Ollama at its default in [Requests at once](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#requests-at-once); *sourced:* [the defaults that moved](https://research.strata2signal.com/a-short-history-of-ollama/#the-defaults-that-moved).
- **At sixteen slots, with `num_batch: 512`, the whole model stayed on this card.** It is a request option, so each request carries it, as `"options": {"num_batch": 512}`. Without it, on the older version we tested first, sixteen slots of 5,120 tokens pushed part of the model off the card, the rest in the computer's memory; with it, the whole model stayed on the card. Ollama 0.34.4 was not run without it, so whether the newer version still needs it is not measured. If you raise the slots, check that `ollama ps` shows the model wholly on the card under PROCESSOR, not split between CPU and GPU. *Measured:* the postures table in [What was written down before the first run](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#what-was-written-down-before-the-first-run), and the request's form in the bench page's [How to check our work — and see it live](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#how-to-check-our-work-and-see-it-live); *sourced:* [see it on your own install](https://research.strata2signal.com/a-short-history-of-ollama/#how-to-check-our-work-and-see-it-live).
- **The window follows your card's memory, and a larger one reloads the model.** With `OLLAMA_CONTEXT_LENGTH` unset, Ollama picks the window from the memory of every card it can see, added up: 4,096 tokens below 23 GiB, 32,768 tokens from 23 GiB, 262,144 tokens from 47 GiB, so one 24 GB card gets 32,768 tokens. The window is fixed when the runner starts, so a request that asks for a larger one means a new runner: on the bench, the request that asked for a larger window reloaded the model before it answered, so the bench prints no speed for it. If your client sends a window size, send the server's own. *Measured:* [The first word, and a shared prompt](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#the-first-word); *sourced:* [how it serves a request](https://research.strata2signal.com/a-short-history-of-ollama/#how-it-serves-a-request), [the defaults that moved](https://research.strata2signal.com/a-short-history-of-ollama/#the-defaults-that-moved).
- **A long chat is trimmed quietly unless you say `truncate: false`.** On the chat route `truncate` defaults to true for GGUF models (the single-file format most of Ollama's library ships in): messages that overflow the window are dropped from the front, keeping the system messages and the newest, and the log records it only at debug level. A chat request can carry `truncate: false` to turn that off. *Sourced:* [how it serves a request](https://research.strata2signal.com/a-short-history-of-ollama/#how-it-serves-a-request).
- **`OLLAMA_HOST` is where it listens; leave it at its loopback default unless something that checks keys stands in front.** By default only the computer itself can reach it; its FAQ's answer to exposing it on your network is to change `OLLAMA_HOST`, and its official Docker image sets it to every interface itself. [Who can reach it](#who-can-reach-it) has the rest. *Sourced:* [who can talk to it](https://research.strata2signal.com/a-short-history-of-ollama/#who-can-talk-to-it).
- **A loaded model stays five minutes, then goes.** `OLLAMA_KEEP_ALIVE` sets that; the bench kept its models loaded. The queue for a model still loading holds 512 requests, then answers HTTP 503. Some model families are held at one slot whatever you set, eleven architectures at v0.34.4 and every embedding model; Gemma 4 is not among them. *Sourced:* [how it serves a request](https://research.strata2signal.com/a-short-history-of-ollama/#how-it-serves-a-request); the bench's serve line is a receipt in [How to check our work — and see it live](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#how-to-check-our-work-and-see-it-live) on the bench page.
- **The `gemma4` tag thinks by default, and Ollama's OpenAI-style completions route applies the template again.** Sent our already-formatted prompt, that route wrapped it in the model's template a second time with reasoning on, and every visible answer came back empty: the bench prints no speed for that way of sending it, because its tokens went to reasoning that never became visible text. Send text you have formatted yourself to `/api/generate` with `raw: true`; for chat, turn reasoning off with `think: false` on `/api/chat` or `reasoning_effort: "none"` on the OpenAI-style chat route. *Measured:* the fourth way the bench sent it, road (d), in [The first word, and a shared prompt](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#the-first-word); the same prompt, token for token, in [What was written down before the first run](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#what-was-written-down-before-the-first-run); *sourced:* [what it speaks](https://research.strata2signal.com/a-short-history-of-ollama/#what-it-speaks-and-where-it-runs), the `think` field.

**vLLM**

- **It takes 92 per cent of the card at start and keeps it; `--gpu-memory-utilization` sets the share.** If less than its share is free when it starts, because a desktop, a loaded Ollama model or an image generator is on the same card, it refuses and says why. Two vLLM processes share one card by each taking a smaller share. What that reservation held on our card depends on a condition the bench page names every time: on a boot that loaded a saved compile, started text only, room for 17.07 requests at the full window and 91.57 per cent of the card at rest (our bets of room for sixteen requests, P-BOOT, and of at least 90 per cent of the card, C-2, both CONFIRMED there; both REFUTED on the first boot, which compiled cold). *Measured:* [What each holds, and how fast it starts](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#what-each-holds); *sourced:* [how it serves a request](https://research.strata2signal.com/a-short-history-of-vllm/#how-it-serves-a-request).
- **The window is the model's own unless `--max-model-len` sets less, and a longer prompt is refused, not trimmed.** If the card cannot hold one request at the window, vLLM refuses to start rather than shrink it. The bench set 5,120 tokens, the window of one Ollama slot. *Sourced:* [how it serves a request](https://research.strata2signal.com/a-short-history-of-vllm/#how-it-serves-a-request); *measured:* the postures table in [What was written down before the first run](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#what-was-written-down-before-the-first-run).
- **The step, `--max-num-batched-tokens`, the most tokens vLLM works through in one pass across every request in flight.** On a 24 GB card vLLM's default is 2,048 tokens per step and up to 256 requests in a batch; the bench's headline posture set 512 tokens, the step Ollama's runner uses, so the two were compared like for like, and ran one row at the default step too, sixteen long requests at once. On that row the default read no lower on total speed, and no longer on the first token's 95th percentile, than the headline step at the same level, on a boot the bench page names as loading a saved compile. The same row prints how much working memory the default step left for requests and how long its requests waited in its queue; read both before you keep or change the step. The row is printed in the same section as the long-prompt totals, with the condition the bench page names. *Measured:* vLLM at its default step in [Requests at once](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#requests-at-once); *sourced:* [how it serves a request](https://research.strata2signal.com/a-short-history-of-vllm/#how-it-serves-a-request).
- **The first start compiles; every later start with the same serve line reuses it.** The compile is paid once per model and serve line on a computer, and it also moves how much working memory vLLM gives itself, which is why the bench page prints two conditions for what it holds. The seconds are on the bench page, not here. `--enforce-eager` turns off the compile and the CUDA graphs it captures on every start (graphs that record the model's steps for replay); what that costs, we have not measured. *Measured:* [What each holds, and how fast it starts](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#what-each-holds); *sourced:* [how it serves a request](https://research.strata2signal.com/a-short-history-of-vllm/#how-it-serves-a-request).
- **`--host` and `--allowed-origins` are the first two flags to decide.** By default vLLM listens on every IPv4 address the computer has, and lets a web page from any site call it. `--host 127.0.0.1` keeps it on the computer, as the bench ran it; `--allowed-origins` narrows the pages. [Who can reach it](#who-can-reach-it) has the rest. *Sourced:* [who can talk to it](https://research.strata2signal.com/a-short-history-of-vllm/#who-can-talk-to-it).
- **Gemma 4 has image, audio and video parts; the bench started vLLM without them.** `--limit-mm-per-prompt`, set to no images, audio or video, kept those parts from being built, so every vLLM reading on this page is text only; what the reservation holds with them built, we have not measured. Ollama's rows include the model's image part, which its runner loads by default. *Measured:* the postures table in [What was written down before the first run](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#what-was-written-down-before-the-first-run), the flag in the vLLM serve line in [How to check our work — and see it live](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#how-to-check-our-work-and-see-it-live), and [What each holds, and how fast it starts](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#what-each-holds).
- **Raw text sent to vLLM needs its own start token.** A model's text opens with a start token, `<bos>`. Google's checkpoint (the weights as downloaded) for vLLM adds none to raw text, and vLLM follows it; the GGUF under Ollama adds one. Our client stripped it on the rule that each server adds one back, and the first full pass stopped at two of our own checks; the fix keeps the token in the text for vLLM. If you send raw text to vLLM with this checkpoint, keep the `<bos>` the template writes at its start. Chat requests are not affected, because the template writes it. *Measured:* our own mistake, kept, in [What was written down before the first run](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#what-was-written-down-before-the-first-run).

**Both**

- **What each holds, read two ways.** At rest, Ollama at its default of one slot had grown the card's used memory by 8,761 MiB, 35.6 per cent of the card, and gives it back when the model unloads. Under sixteen requests at once, the card's used memory less the boot's own baseline read 22,507 MiB on vLLM 0.30.0, on a boot that loaded a saved compile, and 16,393 MiB on Ollama 0.34.4 with sixteen slots, each a few MiB above what that same boot had grown at rest, 22,505 MiB and 16,389 MiB. vLLM's reservation is taken at start; Ollama's follows its slots and window. *Measured:* [What each holds, and how fast it starts](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#what-each-holds).
- **A structured answer.** Both take a JSON schema: Ollama in the `format` field of a request, vLLM through its structured-output backends. How many of ten answers parsed and validated on each, in four variants, lands on the bench page with arm G's results. *Read, not yet printed:* [What else the bench read](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#what-else-the-bench-read); *sourced:* [what Ollama speaks](https://research.strata2signal.com/a-short-history-of-ollama/#what-it-speaks-and-where-it-runs), [what vLLM speaks](https://research.strata2signal.com/a-short-history-of-vllm/#what-it-speaks-and-where-it-runs).

## What each calls home {#what-each-calls-home}

Both make calls out of your machine by default, and both document a switch that stops those default calls; fetching a model is a separate call. The bench page's two serve lines are the receipts; this is the recipe.

Ollama, with its cloud calls off:

```
OLLAMA_NO_CLOUD=1 ollama serve
```

vLLM, with its usage report off:

```
VLLM_NO_USAGE_STATS=1 \
vllm serve <the model>
```

The other two documented spellings of vLLM's switch do the same: `DO_NOT_TRACK=1`, or an empty file named `do_not_track` in the `.config/vllm` folder under your home folder. Started as written, vLLM takes the model's own window and refuses to start if the card cannot hold one request at it; add the `--max-model-len` you need, as the settings above say. Ollama's switch has a second spelling too, a file rather than a variable, so it does not depend on the shell the server was started from: `disable_ollama_cloud` in its `server.json`, in the `.ollama` folder under the home folder of the account the server runs as. Its FAQ calls that "local only mode", and the log then reads `Ollama cloud disabled: true`. The switch also turns off Ollama's cloud models and its web search.

**What goes out by default.** On Ollama, unless the cloud features are off, the server fetches two things from ollama.com: at start, the catalogue of its cloud models, and at start and about every four hours after, a list of recommended models. Each request is signed with the key the installation generated, so calls from one installation carry the same public key. Nothing on that list is your prompt; its FAQ says it does not see your prompts or data when you run locally, and the code read at v0.34.4 agrees ([what leaves the machine](https://research.strata2signal.com/a-short-history-of-ollama/#who-can-talk-to-it), on the Ollama history, which dates each fetch to the release that added it). On vLLM, the usage report goes out once at start, under a random ID new for each process: the machine's details (a cloud provider if detected, processor, memory, operating system and kernel, graphics cards, CUDA version), the vLLM version, the model's architecture as a class name rather than the model's name, twenty-six settings of the run and five vLLM environment variables; then, every ten minutes, the same ID and the time, and at v0.30.0 nothing more. A copy of the report sits on your computer, in `usage_stats.json` beside the `do_not_track` file ([what leaves the machine](https://research.strata2signal.com/a-short-history-of-vllm/#who-can-talk-to-it), on the vLLM history). Started from a Hugging Face model ID, vLLM also reaches Hugging Face for the weights.

**What the bench counted.** Arm G, on 2026-09-30, started each server in a sandbox whose only network was the loopback address, with every system call recorded and every name lookup sent to a resolver on loopback that refused it; no name server and no listener ran, so the count is of attempts, and the window ran from launch to 120 s after ready, inside vLLM's ten-minute report interval and Ollama's four-hour one. The method is the one we used on [ONNX Runtime's telemetry on Linux](https://research.strata2signal.com/onnx-runtime-telemetry-on-linux/), and the bench page states it in [What else the bench read](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#what-else-the-bench-read). The counts land on the bench page with arm G's results, and this page will carry them in a dated update that says what it counted, its window at both ends, and how it compares with the code-read lists on the two histories.

The recommendation stands without a number: set the switches where the server's own process reads them (a switch set in the shell you type in does not reach a server that was started as an app or a service), then check: for Ollama, the log line above says the switch took; the vLLM history names no log line for it, and `usage_stats.json` in the `.config/vllm` folder under your home folder holds the copy of what vLLM sent while the report was on, field for field, the list the vLLM history describes.

## Who can reach it, and what to put in front {#who-can-reach-it}

A server on a network address answers anyone who can reach it, unless it checks. Here the two differ in kind, and each history gives the project's own words.

**Ollama's local server checks no key.** Its documentation says the local API does not require authentication, and the code at v0.34.4 agrees: two checks stand in front of every route, one on which web pages may call it from a browser and, while it is bound to the loopback address, one on the request's `Host` header, and neither asks who you are. The routes that pull, push, copy and delete models are as open as the ones that answer, and the server speaks plain HTTP. What keeps it private is the default: `OLLAMA_HOST` binds the loopback address, so only the computer itself can reach it. Ollama's official Docker image changes that default itself, binding every interface, and its Docker guide publishes it on the host. The maintainers' position is on record from February 2025: no key system in the server for now; use a VPN or a proxy. The engine underneath, `llama-server`, takes a key and serves TLS (the protocol behind https); Ollama starts it on the loopback address and passes neither through. *Sourced:* [who can talk to it](https://research.strata2signal.com/a-short-history-of-ollama/#who-can-talk-to-it).

**vLLM checks, over part of its surface.** Its key, from `--api-key` or `VLLM_API_KEY`, guards requests whose path begins with one of four prefixes; the keys are interchangeable, with no name, limit or record of their own. Its own security guide, read at v0.30.0, says what the key does not cover and puts it plainly: do not rely exclusively on the key, and deploy vLLM behind a reverse proxy, a web server that takes each request first and passes on only what it allows (the guide names nginx, Envoy or a Kubernetes Gateway), allowing the routes you mean to expose and blocking the rest. Among the routes the guide names as open is `/invocations`, which, in the guide's words, "routes to the same inference functions as `/v1` endpoints". By default vLLM listens on every IPv4 address the computer has; `--host 127.0.0.1` keeps it on the computer. *Sourced:* [who can talk to it](https://research.strata2signal.com/a-short-history-of-vllm/#who-can-talk-to-it).

**As measured.** Arm A called each server's routes on the loopback address only, on the versions benched, over the routes each project's own documentation names, and recorded which answered without a key: on vLLM 0.30.0, 7 routes that its own security guide names as served without the key, each of which answered without it; on Ollama 0.34.4, the 5 routes in the bench page's table, and its own documentation says it has no key system. The bench page's table claims nothing about any route it does not list. *Measured:* which routes answered without a key, in [What else the bench read](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#what-else-the-bench-read).

**The recommendation.** Put something in front of either server before anyone else can reach it, and keep the server itself on the loopback address behind it: `OLLAMA_HOST` at its default, `--host 127.0.0.1` for vLLM. Ollama will not check for you, and its maintainers point to a proxy or a VPN; vLLM checks a part, and its own guide says that part is not enough and names the proxy's job: allow the routes you mean to expose, block the rest, and add authentication, rate limiting and logging there ([who can talk to it](https://research.strata2signal.com/a-short-history-of-vllm/#who-can-talk-to-it)). The workshop runs a small gateway of its own in front of our seats (the model servers behind it), which [Four homes for one reranker](https://research.strata2signal.com/four-homes-for-one-reranker/) follows one call through: every product calls a role name on it rather than a machine. That is one pattern, in front of our seats, not the only way in; any reverse proxy in front of either server can carry it.

**Two more things from vLLM's own record.** Each is a behaviour of the server rather than a speed, read on 2026-10-07 (UTC) from vLLM's security guide at tag v0.30.0 and its source at each tag named here, which dates each to the release that changed it. A key given through `VLLM_API_KEY` was written in plain text into the compile cache. In the source at v0.27.1 and v0.28.0, `cache_key_factors.json`, the small file that records the settings a compile was keyed on, takes the value of every vLLM environment variable not on its list of exclusions, and the key's variable was not on that list; from v0.29.0 it is, marked there as a credential that "must not be persisted" in that file (the only settings it records in the clear are environment variables, so a key given on the command line with `--api-key` never entered it). And the engine's start-up handshake, a small store that PyTorch, the machine-learning library vLLM is built on, opens so the engine's parts can find each other: the security guide at v0.30.0 says vLLM uses PyTorch's distributed library, `torch.distributed`, "including when using vLLM on a single host", warns that this store, when started over the network, "by default, listens on all network interfaces", and adds that "any use of `torch.distributed` should be considered insecure by default"; in the source at v0.27.1 a run on one machine opened it over the network, and from v0.28.0 (a change merged on 2026-08-10) a run on one card at the default settings hands the handshake to a file in the temporary folder instead, unless a ROCm setting keeps the network form; some set-ups, data parallelism (several copies of the model at once) among them, still use the network form, and an open issue of 2026-08-17 asks the project to say which. The guide's advice covers both: restrict the cache folder's permissions so that only the account the server runs as can read or write it, and never share it with users you do not trust or mount it from storage you do not trust; and put a firewall around the machine that exposes only the smallest network surface it needs. Read the release notes before you upgrade across these versions.

## Getting it running, and keeping it running {#getting-it-running}

**Ollama** is one download, then `ollama pull` and a model name; it asked one change of us at sixteen slots, `num_batch: 512` with each request, so the whole model stayed on the card. Its server was ready in seconds on our card, and the model loaded in seconds on the older version we tested first; the newer version's one first load took longer and is not yet explained, and its second start came in far shorter. The change is in what it took to get here, in [What else the bench read](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#what-else-the-bench-read), and the seconds are in [What each holds, and how fast it starts](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#what-each-holds).

**vLLM** is a Python package for Linux, installed with pip or run from its official Docker image, and one command serves one model: `vllm serve` and a Hugging Face model ID. The packages pip installs are built for CUDA 13.0 since v0.20.0, by its release notes, which ask for an NVIDIA driver of the R580 branch or newer, so check the driver first. Its installation page at v0.30.0 still says CUDA 12.9; the vLLM history notes the disagreement. On Windows it runs, officially, only through the Windows Subsystem for Linux; on macOS the project's own support is experimental and processor-only, with Apple silicon graphics through a separate plugin. *Sourced:* [where it runs](https://research.strata2signal.com/a-short-history-of-vllm/#what-it-speaks-and-where-it-runs), [the defaults that moved](https://research.strata2signal.com/a-short-history-of-vllm/#the-defaults-that-moved). What it took us is on the bench page, with receipts, and we will not re-tell it here: the experiments it took to fit this model under vLLM 0.27.1 on a 12 GB card, each of our two RTX 3080 Ti boards in turn, a library version pinned beside it, a compiler the bench computer lacked, and the first start that compiles before it serves (what it took to get here, in [What else the bench read](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#what-else-the-bench-read)). None of it is a flaw; it is the price of a server built to be tuned, mostly paid once per setup.

**Keeping it running.** Both projects ship often; each history counts the releases in the twelve months before its read ([Ollama](https://research.strata2signal.com/a-short-history-of-ollama/#what-have-they-shipped), [vLLM](https://research.strata2signal.com/a-short-history-of-vllm/#the-releases-where-it-bent)), and each dates the defaults that moved under their users: on Ollama, the slot count, the window, the engine and the sampler; on vLLM, the share of the card, the engine, the model runner and the CUDA build. A default that holds at the version you set up may not at the next. Write down the version you installed, and read the release notes before you upgrade. Since the versions benched, Ollama 0.34.4 (released 2026-09-23) and vLLM 0.30.0 (released 2026-09-22), each project's GitHub release record, read on 2026-10-07 (UTC) for the bench page's own line in [What else the bench read](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#what-else-the-bench-read), lists three full releases of Ollama, the newest 0.40.0 on 2026-10-06, and one of vLLM, the newest 0.31.0 on 2026-10-05; pre-releases are not counted. Whether any default this page states moved in them is not read here: the versions benched are the ones every figure names.

## What to take with you {#what-to-take-with-you}

- **One request at a time, pick on everything but writing speed.** Decode fell inside our ±10 per cent band twice over (P-1); the first word came sooner on vLLM, Ollama's wait at one request being more than twice as long (P-2). [One request at a time](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#one-request-at-a-time); [The first word, and a shared prompt](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#the-first-word).
- **Two or more at once, vLLM was ahead on every speed reading, and the gap grew with the load.** Its total reached 1.2 times Ollama's with sixteen slots at two requests at once, past that bar by less than the runs varied, and kept climbing to sixteen; Ollama's sixteen-slot total went flat from four requests on, with runs that disagreed with each other at four and at eight, and at its default of one slot Ollama is a queue. The bet on where the crossover would fall reads UNDECIDED, because it fell at a level marked undecided before any run. [Requests at once](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#requests-at-once).
- **Stay on Ollama where the table leans to it:** models you switch between, chat beside embeddings, Windows, a Mac or an older card. [Which to pick when](#which-to-pick-when).
- **Set the slots and the window together, and check the model stayed on the card.** Each slot reserves a whole window; with `num_batch: 512` the whole model stayed on this card at sixteen slots. [The settings that matter](#the-settings-that-matter).
- **Set the switches for the calls out, keep each server on the loopback address, and put something in front.** `OLLAMA_NO_CLOUD=1` on one side; `VLLM_NO_USAGE_STATS=1`, `DO_NOT_TRACK=1` or the `do_not_track` file on the other; the count with them set lands with arm G's results on the bench page. Ollama checks no key, and vLLM's own guide says its key is not enough on its own. [What each calls home](#what-each-calls-home); [who can reach it](#who-can-reach-it).

## What it does not say {#what-it-does-not-say}

- One card, an RTX 3090 (24 GB); one model family, Gemma 4; one operating system, Linux. Not two or more cards, not a model larger than the card, and no Mac row is measured.
- Not other servers: not SGLang, not LM Studio, not Hugging Face's text-generation-inference, now archived.
- Requests released together are a burst, not people arriving over time. A classroom pressing enter together is the nearest real case to the sixteen-request level; a family's assistant almost never sends four in the same second.
- Not other request shapes: every speed reading came from requests capped at 256 tokens out at temperature 0, no timed prompt ran past about 2,600 tokens, and the headline postures' window was 5,120 tokens; long answers, sampled answers and long windows were not measured, and no coding agent's own requests were sent.
- Not Ollama at a slot count between one and sixteen, and not vLLM with Gemma 4's image, audio and video parts built: the bench started vLLM text only, while Ollama's rows include the model's image part.
- Not bit-identical weights: both servers ran Google's 4-bit Gemma 4 12B, but the two downloads differ in the output layer, the scale formats and the arithmetic (the same weights, said exactly, in [What was written down before the first run](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#what-was-written-down-before-the-first-run)).
- Two verdicts behind this page are first readings, each labelled so on the bench page: the two shared-prompt verdicts, P-4a and P-4b; a third, C-3, our bet that neither server calls out with its switches set, prints with arm G's results as a first reading. Their second readings, from a corrected instrument, will print beside the first and will not replace it.
- Not which server gives the right answers: our six-way says the two differ, and how often; it does not say which is right.
- Not the versions after those benched. The dated updates below will say what moved.

## How to check our work — and see it live {#how-to-check-our-work-and-see-it-live}

- **The kit, the prompts and the pre-registration** are on the bench page: its [How to check our work — and see it live](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/#how-to-check-our-work-and-see-it-live) section links the kit, the frozen prompts, the six-way and the predictions put on record before each run.
- **See both live.** [Ask About This Page](https://research.strata2signal.com/ask-about-this-page/) answers on vLLM; [the long table](https://research.strata2signal.com/a-dinner-party-for-the-dead/), our dinner-party exhibit, seats its guests on both. The assistant opens from the "ask about this page" line at the top of every article on this site, and the long table is at [longtable.strata2signal.com](https://longtable.strata2signal.com/).
- **On your own Ollama.** `ollama ps` prints the split under PROCESSOR when a model did not fit the card, which is the check behind the `num_batch` line above; with `OLLAMA_DEBUG=1` the log says when it is truncating a chat that outgrew its window.
- **On your own vLLM.** After a start with the usage report on, `usage_stats.json` in the `.config/vllm` folder under your home folder is the copy of what was sent, field for field, which is the check behind the list above.

## The rest of the seminar {#the-rest-of-the-seminar}

- [A short history of Ollama](https://research.strata2signal.com/a-short-history-of-ollama/): the project and the company behind one of these two servers, and every default of its we found that moved, dated.
- [A short history of vLLM](https://research.strata2signal.com/a-short-history-of-vllm/): the lab, the paging idea and the foundation behind the other, and its defaults that moved.
- [Ollama and vLLM on one RTX 3090](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/): the bench this page leans on, both servers on one card with the same model, from one request at a time to sixteen at once, every reading beside the bet we wrote down before it, with the method, the kit and the receipts.
- [Four homes for one reranker](https://research.strata2signal.com/four-homes-for-one-reranker/): the gateway in front of our seats, one call followed through it.

## Who ran this, and thanks {#who-ran-this-and-thanks}

This page measures nothing of its own: every reading is the bench page's, taken on one EVGA GeForce RTX 3090 XC3 Ultra (24 GB) capped at 300 W, and every dated fact but the safety paragraph's is one of the two histories', read from the project's own release notes, documentation and source at the tags named there. **Google** made Gemma 4 and its QAT (quantization-aware training) weights and released them under the Apache 2.0 licence; **Ollama**'s team and contributors built one server (MIT) on **llama.cpp** (MIT), begun by **Georgi Gerganov** and kept by the ggml authors; the **vLLM** project (Apache-2.0), begun in UC Berkeley's **Sky Computing Lab** and hosted by the **PyTorch Foundation**, built the other, on **PyTorch** (BSD 3-clause). On this card both run over **CUDA** (NVIDIA, proprietary, the one closed piece in either stack). None of them owed us anything.

A small human team asked for this page and for the bench it draws on, chose what each would and would not claim, and signed off on every figure; a fleet of AI agents ran the bench and wrote this page under that team's rulings.

## Sources {#sources}

Read 2026-10-02 to 2026-10-07 (UTC), from this workshop's own three pages on these servers, the other pages of ours listed below, and, for the safety paragraph only, vLLM's own repository at the tags and on the pages the last bullet names, as they stood then; every version and default line was read again on 2026-10-08 (UTC), from the bench page as released and the two histories, the day this page went up.

- **The bench page**, [Ollama and vLLM on one RTX 3090](https://research.strata2signal.com/ollama-and-vllm-on-one-rtx-3090/): every reading on this page is printed there first, in the section each link names, with its arm and its day, and its pre-registration and kit are linked there.
- **The two histories**, [A short history of Ollama](https://research.strata2signal.com/a-short-history-of-ollama/) and [A short history of vLLM](https://research.strata2signal.com/a-short-history-of-vllm/): every dated fact on this page but the safety paragraph's is one of theirs, read from Ollama's source and documentation at tag v0.34.4, llama.cpp's server README at build b11081, and vLLM's source and documentation at tag v0.30.0, each file linked at its tag under their own Sources.
- **This workshop's other pages, linked above:** [Four homes for one reranker](https://research.strata2signal.com/four-homes-for-one-reranker/), [ONNX Runtime's telemetry on Linux, measured](https://research.strata2signal.com/onnx-runtime-telemetry-on-linux/), [Ask About This Page](https://research.strata2signal.com/ask-about-this-page/), [A Dinner Party for the Dead](https://research.strata2signal.com/a-dinner-party-for-the-dead/) and [How the long table works](https://research.strata2signal.com/how-the-long-table-works/).
- **vLLM's own source and documentation, read on 2026-10-07 (UTC) for the safety paragraph only:** the [security guide at v0.30.0](https://github.com/vllm-project/vllm/blob/v0.30.0/docs/usage/security.md) ("Cache Directory Security", with its recommendations; "Security and Firewalls: Protecting Exposed vLLM Systems", "including when using vLLM on a single host"); [`envs.py`](https://github.com/vllm-project/vllm/blob/v0.30.0/vllm/envs.py) (`compile_factors`, where `VLLM_API_KEY` is on the list of exclusions as a credential that "must not be persisted in cache_key_factors.json") and [`backends.py`](https://github.com/vllm-project/vllm/blob/v0.30.0/vllm/compilation/backends.py) (the file written into the compile cache), `envs.py` read also at [v0.27.1](https://github.com/vllm-project/vllm/blob/v0.27.1/vllm/envs.py), [v0.28.0](https://github.com/vllm-project/vllm/blob/v0.28.0/vllm/envs.py) and [v0.29.0](https://github.com/vllm-project/vllm/blob/v0.29.0/vllm/envs.py), and `backends.py` at [v0.27.1](https://github.com/vllm-project/vllm/blob/v0.27.1/vllm/compilation/backends.py) and [v0.28.0](https://github.com/vllm-project/vllm/blob/v0.28.0/vllm/compilation/backends.py), for the version the change first shipped in; [`uniproc_executor.py`](https://github.com/vllm-project/vllm/blob/v0.30.0/vllm/v1/executor/uniproc_executor.py), read also at [v0.27.1](https://github.com/vllm-project/vllm/blob/v0.27.1/vllm/v1/executor/uniproc_executor.py), [`network_utils.py`](https://github.com/vllm-project/vllm/blob/v0.30.0/vllm/utils/network_utils.py) (the file-based start-up rendezvous, and the ROCm exception) and [`parallel_state.py`](https://github.com/vllm-project/vllm/blob/v0.30.0/vllm/distributed/parallel_state.py) (the network form kept for data parallelism), with [the pull request that made the change](https://github.com/vllm-project/vllm/pull/50999), merged 2026-08-10 and first shipped in v0.28.0 (GitHub's comparison finds it in v0.28.0 and not in v0.27.1), and [the open issue](https://github.com/vllm-project/vllm/issues/52638), opened 2026-08-17 and open when read; the release records for [v0.27.1](https://github.com/vllm-project/vllm/releases/tag/v0.27.1) (2026-08-11), [v0.28.0](https://github.com/vllm-project/vllm/releases/tag/v0.28.0) (2026-08-26), [v0.29.0](https://github.com/vllm-project/vllm/releases/tag/v0.29.0) (2026-09-09) and [v0.30.0](https://github.com/vllm-project/vllm/releases/tag/v0.30.0) (2026-09-22).

Corrections and later measurements will be added below, each dated (UTC), with a window at both ends where one applies and saying in plain words what it counts.

<!-- derived 2026-10-08 (UTC) by tools/derive_md.py from the pour source.
     source html sha256: a81349fcc15afd7d8564401b2333b4ae1ea02b2a33c7d85c650000791197afb5
     derivation sha256:  fc05a7de9eb0f91052ef6a1c224e9a73201e9c89d5e8f39db9f42903f9e31967
     the {#id} on each heading is the anchor that heading carries on the page. -->
