# A short history of vLLM

*The open-source server that lets one graphics card run a language model for many people at once: a Berkeley lab's borrowing from operating systems, the foundation and the company around it, and the defaults that moved under its users, dated.*

*Published 2026-10-02 (UTC) · A small (human) team and a fleet of AI agents.*

**the short version:** vLLM is free, open-source software that serves a large language model, the kind of AI inside a chatbot; on this hub it runs the model behind every article's "ask about this page" link. It began in a UC Berkeley lab and was announced on 2023-06-20, built to make one graphics card hold many conversations at once: like an operating system, vLLM hands out working memory in pages as each answer grows. It has been a PyTorch Foundation project since 2025-05-07; some of its creators announced a $150M-funded company beside it on 2026-01-22. This page also dates the defaults that moved under vLLM's users. The only speed figures here are the project's own claims, each dated, none re-run by us.

6,633 words · about 30 minutes (at 220 words/min) · 5 tables · data kit: no

https://research.strata2signal.com/a-short-history-of-vllm/

---

*If you're new here: [strata→signal](https://strata2signal.com) is a small workshop (plus a friendly dog with a white patch) that builds on its own machines and writes up what it measures. Every date below was read on 2026-09-30 and 2026-10-01 (UTC) from the source credited beside it — among them vLLM's own posts and release notes, its documentation and source code as of version 0.30.0, GitHub's record of the repository, the paper, the foundations' and companies' own announcements and pages, and one press report, read beside the company post it reports on; for Ion Stoica's other roles it is our only source. Where we found no source, the claim is not here.*

## One card, many questions at once {#one-card-many-questions}

Every article on this hub carries a small link at the top: ask the page a question, and a model answers from the page's own text ([Ask About This Page](https://research.strata2signal.com/ask-about-this-page/)). The model is open-weight, meaning the numbers it learned in training are published for anyone to download. The program that runs it, on a card in this building, is vLLM. The same program serves some of the chairs at [our long table](https://research.strata2signal.com/how-the-long-table-works/), where language models are seated as famous people from history.

Now picture two readers asking at once. A language model writes one token at a time (a token is a word or a piece of one), and for every token it keeps the conversation so far on the card beside the weights, in a working memory called the KV cache. Give each request a fixed slab of that memory, sized for the longest answer you allow, and most of the slab sits empty most of the time: the card fills up with reservations before it fills up with work.

That was the state of things in the spring of 2023, when a group at UC Berkeley needed to serve a great many strangers on not many cards. The engine they wrote for it is vLLM.

## Who are these people, and where did they come from? {#who-are-these-people}

vLLM's README, read at the v0.30.0 tag, gives the origin in one line: "Originally developed in the Sky Computing Lab at UC Berkeley". The launch post of 2023-06-20 is bylined **Woosuk Kwon** and **Zhuohan Li**, marked as equal contributors, with Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Yu, Joey Gonzalez, Hao Zhang and **Ion Stoica**; the paper that followed lists the same nine people in the same order, Cody Yu and Joey Gonzalez as Cody Hao Yu and Joseph E. Gonzalez. Stoica, a computer science professor, directs Berkeley's Sky Computing Lab and co-founded Databricks, per SiliconANGLE's report of January 2026.

The post also says where the engine had already been: "vLLM has been developed at UC Berkeley and deployed at Chatbot Arena and Vicuna Demo for the past two months" and "handling an average of 30K requests daily and a peak of 60K". LMSYS was "a small research team" that "developed the popular Vicuna chatbot models" that April, and Vicuna had since "been served in Chatbot Arena for millions of users". The post adds: "With vLLM, LMSYS was able to cut the number of GPUs used for serving the above traffic by 50%". Those are the project's numbers, not ours. The point they make is about timing: the engine was serving the public for two months before it was announced.

The repository was created on 2023-02-09, and its first GitHub release, v0.1.0, is dated 2023-06-20, the same day as the post (GitHub, read 2026-09-30).

## What the v stands for {#what-the-v-stands-for}

Issue #835 on the project's tracker, titled "vLLM full name", was opened on 2023-08-23. Woosuk Kwon answered on 2023-08-25, and his answer, quoted here without its greeting, is the only account of the name we found: "originally we used the 'v' to represent 'virtual', just like vCPU stands for virtual CPU. This is because our core technique was inspired by virtual memory (and paging) in OS. But we also think the 'v' could represent other ideas, like velocity and victory."

You will see "virtual large language model" spelled out elsewhere, Red Hat's "What is vLLM?" page among them. Nothing in the project's repository at v0.30.0 spells it out; the README's line under the logo is "Easy, fast, and cheap LLM serving for everyone".

## The idea: paging for a language model {#paging-for-a-language-model}

The paper is "Efficient Memory Management for Large Language Model Serving with PagedAttention", submitted to arXiv on 2023-09-12 and presented at SOSP 2023, the 29th ACM Symposium on Operating Systems Principles, in Koblenz, 2023-10-23 to 2023-10-26. An operating-systems conference for a language-model engine is the tell: the abstract says the algorithm is "inspired by the classical virtual memory and paging techniques in operating systems".

In plain words: a program does not get one solid slab of memory; it gets pages, handed out as it grows, from wherever there is room. PagedAttention does the same for each request's working memory: fixed-size blocks, allocated as the answer lengthens, so the only waste is whatever is left in a request's last block. The launch post put the before and after in its own terms: "existing systems waste 60% – 80% of memory due to fragmentation and over-reservation", against "a mere waste of under 4%".

With the memory no longer spoken for, many requests can share one card, and the engine moves all of them forward together, letting new arrivals join the ones already running rather than wait for a batch to finish. The README's name for that at v0.30.0 is "Continuous batching of incoming requests".

vLLM's launch claims of 2023 count throughput, how much work a server gets through in a given time across all its requests: "up to 24x higher throughput than HuggingFace Transformers" and "up to 3.5x higher throughput than TGI" (Hugging Face's Text Generation Inference, "the previous state of the art" in the post's words). The paper's figure is "2-4× with the same level of latency compared to the state-of-the-art systems, such as FasterTransformer and Orca". We have not re-run any of them, and Transformers, TGI and vLLM have all changed since.

The graphics-card code written for the paper, its kernels, did not last: v0.25.0 (2026-07-11) deleted them, and its notes say "PagedAttention has been removed". The paging outlived them; the README at v0.30.0 still credits PagedAttention, and the old kernel's design page now calls itself "a historical document".

## Who owns it, and who pays for it {#who-owns-it}

The licence is the shortest story on this page. vLLM is Apache-2.0, a licence that lets anyone use, change and redistribute it, for money or not, as long as the licence and its notices travel with it. The LICENSE file has exactly one commit in its history, "Add Apache-2.0 license (#102)", dated 2023-05-15, five weeks before the first release (GitHub, commits on that path, read 2026-09-30).

| Date | What happened | Who said so |
|---|---|---|
| 2023-08-30 | The Open Source AI Grant program of Andreessen Horowitz (a16z) names "Woosuk Kwon and Zhuohan Li (vLLM): library for high-throughput LLM inference" among its recipients: "grant funding (**not** an investment or SAFE note)" | [a16z](https://a16z.com/supporting-the-open-source-ai-community/) |
| 2024-07-25 | The project's governance post: "vLLM has started the incubation process into LF AI & Data Foundation" and "no one party will have exclusive control over the future of vLLM". It says the project was then maintained by "a consortium of groups such as UC Berkeley, Anyscale, AWS, CentML, Databricks, IBM, Neural Magic, Roblox, Snowflake, and others" | [vLLM](https://vllm.ai/blog/2024-07-25-lfai-perf) |
| 2024-10-28 | LF AI & Data announces vLLM as its latest incubation project | [LF AI & Data](https://lfaidata.foundation/blog/2024/10/28/lf-ai-data-announces-vllm-as-its-latest-incubation-project/) |
| 2025-01-13 | Red Hat completes its acquisition of Neural Magic, which its release calls "a leading contributor to vLLM, an open source project developed by UC Berkeley for open model serving"; terms not stated | [Red Hat](https://www.redhat.com/en/about/press-releases/red-hat-completes-acquisition-neural-magic-fuel-optimized-generative-ai-innovation-across-hybrid-cloud) |
| 2025-05-07 | The PyTorch Foundation, "Hosted by the Linux Foundation", announces its expansion into an umbrella foundation and accepts vLLM and DeepSpeed as its first two hosted projects. Its welcome page says the project was "Contributed by the University of California – Berkeley" | [PyTorch Foundation](https://pytorch.org/blog/press-release-pytorch-foundation-expands-welcomes-projects-vllm-deepspeed/), [its welcome page](https://pytorch.org/blog/pytorch-foundation-welcomes-vllm/) |
| 2026-01-22 | Inferact announces itself as "a startup founded by creators and core maintainers of vLLM", in a launch post signed by Simon Mo, Woosuk Kwon, Kaichao You, Roger Wang, Joseph Gonzalez and Ion Stoica "and the rest of founding members": it "has raised $150M at $800M valuation, led by a16z and Lightspeed". SiliconANGLE's report calls the round seed funding | [Inferact](https://inferact.ai/), [SiliconANGLE](https://siliconangle.com/2026/01/22/inferact-launches-150m-funding-commercialize-vllm/) |

So the answer to "who owns it" is: no single party, on purpose. Since 2025-05-07 it has been a project of the PyTorch Foundation, the home of the machine-learning library vLLM runs on, and the foundation's description of a hosted project is that maintainers "transfer assets to the Linux Foundation for independent stewardship and adopt an open governance model". The code's copyright stays with its contributors ("Copyright contributors to the vLLM project" heads almost every Python and Rust file). The PyTorch Foundation's announcements do not say what became of the LF AI & Data incubation of 2024; both foundations sit under the Linux Foundation.

Inferact has its own "Inferact Platform", which "is built on top of OSS vLLM" (its news of 2026-08-12), and says "The optimizations we develop flow back to the community". SiliconANGLE read Inferact's launch post as a hint that the company "plans to launch a paid serverless version of vLLM". Inferact does not hold the project.

Two counts say how big the thing has become. The PyTorch Foundation's welcome page, dated 2025-05-07, said "over 46,500 GitHub stars and over 1000 contributors". On 2026-09-30 GitHub's API reported 92,997 stars for the repository, and the README at v0.30.0 says "over 2000 contributors".

## The releases where it bent {#the-releases-where-it-bent}

| Date (UTC) | Release | What changed |
|---|---|---|
| 2023-06-20 | v0.1.0 | The first GitHub release, the day of the launch post |
| 2025-01-27 | v0.7.0 | "Minimum requirements for SageMaker compatibility" (#11576, "Implements `/ping` and `/invocations`"), the origin of the route that matters [below](#who-can-talk-to-it) |
| 2025-01-27 | V1 alpha | The rewritten engine, "up to 1.7x higher throughput compared to V0 (without multi-step scheduling)" by its own post; its opt-in was `VLLM_USE_V1=1` |
| 2025-03-18 | v0.8.0 | "We have now enabled V1 engine by default (#13726) for supported use cases" |
| 2025-12-03 | v0.12.0 | Model Runner V2, "Experimental", a rewrite of the layer that runs the model |
| 2026-06-29 | v0.24.0 | GGUF support, in vLLM's own code since v0.5.5 (2024-08-23), moves out into a plugin installed separately (#39612) |
| 2026-07-11 | v0.25.0 | Model Runner V2 becomes "the default for all dense models" (models that use all their weights for every token); and "PagedAttention has been removed" (#47361): the paper's kernels go, [the paging stays](#paging-for-a-language-model) |
| 2026-09-09 | v0.29.0 | "Model Runner V2 is now the default for all models"; the old runner "deprecated", with removal "targeting v0.32"; two queue limits that turn overflow into HTTP 503 |
| 2026-09-22 | v0.30.0 | "762 commits from 315 contributors (104 new)"; a helper process that keeps the weights on the card for fast restarts; the newest release when we read the sources for this page, eight days later |

Read the middle of that table and a pattern shows: twice in two years the project has rewritten its core, and each time the replacement ran beside the old one, opt-in, and then became the default. One old core is gone: the V0 engine stopped being the default on 2025-03-18 and had left the codebase by 2025-10-02, six and a half months later. The other, Model Runner V1, is the part inside the V1 engine that runs the model, not an engine of its own, and its removal is targeted for v0.32.

GitHub lists 106 releases as of 2026-09-30, 97 of them not marked pre-release, and 32 in the twelve months to that date, none of them a pre-release and 20 of them new minor versions, v0.11.0 to v0.30.0.

## What it speaks, and where it runs {#what-it-speaks-and-where-it-runs}

Programs talk to a server through an API, an agreed shape for a request and its answer. vLLM answers in several, most of them shapes first set by AI companies' own services; its README lists an "OpenAI-compatible API server, plus Anthropic Messages API and gRPC support". Read at v0.30.0 on 2026-09-30 and 2026-10-01.

| Feature | vLLM at v0.30.0 |
|---|---|
| OpenAI-style API | An "OpenAI-compatible API server" from v0.1.0 (2023-06-20), in its README's words then; now chat completions, completions, embeddings, models and audio routes, with the Responses API since v0.10.0 (2025-07-24) |
| Anthropic-style API | `/v1/messages`, since v0.11.1 (2025-11-18) |
| Cohere-style API | Rerank since v0.7.0 (2025-01-27), Embed v2 since v0.18.0 (2026-03-20), Chat v2 since v0.27.0 (2026-08-10), the last off unless `VLLM_ENABLE_COHERE_API=1` is set and the `cohere` package installed |
| Its own routes | Tokenize and detokenize, score and rerank; gRPC only when asked for (`--grpc`), and then instead of HTTP |
| A library with no server | The Python `LLM` class runs a batch of prompts offline, in your own code |
| Structured output | A JSON schema, a regex or a grammar the answer must follow, through xgrammar or guidance (outlines and lm-format-enforcer are also named); the default backend is `auto` |
| Tool calling | Yes, with a parser chosen per model family; a named or required tool uses structured output, and the docs warn the first such call compiles a grammar, with "several seconds of latency (or more)" |
| Embeddings, scoring, classification | Yes, as pooling models |
| LoRA adapters | "Efficient multi-LoRA support"; loading adapters at run time is off by default, and the security guide calls it "not a secure operation" |
| Several models | One model per server process; the FAQ says to run one instance per model and route in front of them |
| Getting a model | A Hugging Face model ID, downloaded on first start ("By default, vLLM downloads models from Hugging Face"), or a local folder; weights in the safetensors format, with quantised builds in FP8, INT8, INT4, GPTQ, AWQ, compressed-tensors and other formats the README lists |
| GGUF, llama.cpp's single-file format, which most of Ollama's library ships in | Through plugins installed separately: `vllm-gguf-plugin` since v0.24.0, and vllm-metal on Apple silicon for some dense models. vLLM's docs say "GGUF support in vLLM is highly experimental and under-optimized at the moment" |

Read from the installation pages and the README at v0.30.0 on 2026-09-30.

| Platform | vLLM at v0.30.0 |
|---|---|
| Operating system | Linux ("OS: Linux"), with an official Docker image. Windows: "vLLM does not support Windows natively", and its docs point to the Windows Subsystem for Linux or community forks. macOS: "experimental support for macOS with Apple Silicon", processor only, built from source |
| Python | 3.10 to 3.13 by the installation pages; 3.10 to 3.14 by the package's own `requires-python` |
| NVIDIA cards | "compute capability 7.5 or higher (e.g., T4, RTX20xx, A100, L4, H100, B200, etc.)": the RTX 20 series onward, so not the GTX 10 series |
| AMD cards | Yes, through ROCm |
| Intel graphics | Yes, through its Intel XPU backend |
| Processor only | x86, ARM and PowerPC in the README; the CPU installation page adds IBM Z |
| Other accelerators | Plugins: "Google TPUs, Intel Gaudi, IBM Spyre, Huawei Ascend, Rebellions NPU, Apple Silicon, MetaX GPU, and more" |
| Apple silicon graphics | Through vllm-metal, a separate package under the vllm-project organisation, built on Apple's MLX. The v0.30.0 docs call it "a community-maintained hardware plugin"; the project's blog of 2026-09-22 calls vllm-metal's v0.28.0, published on GitHub on 2026-09-01, "Our first official release" |
| Two or more cards | "Tensor, pipeline, data, expert, and context parallelism": it spreads one model, or copies of it, across cards by design |

## How it serves a request {#how-it-serves-a-request}

Read from the v0.30.0 source and documentation on 2026-09-30 and 2026-10-01; each line is what a fresh install does for a text-generating model until you change it.

- **It reserves most of the card at start, and keeps it.** It takes 92% of the card's memory (`gpu_memory_utilization`, 0.92), "a per-instance limit": two vLLM processes fit on one card by each taking a smaller share. If less than its share is free when it starts, it refuses and says why: "Decrease GPU memory utilization or reduce GPU memory used by other processes."
- **The context window is the model's own.** How many tokens one request may hold, prompt and answer together, is the longest the model's configuration allows unless `--max-model-len` sets less. If the card cannot hold one request that long, vLLM refuses to start rather than shrink the window. A longer prompt is refused, not trimmed.
- **Working memory is handed out in pages, and a page can be shared.** The pages are [the idea above](#paging-for-a-language-model). Prefix caching, on by default, lets a later request reuse the pages of an earlier one that began the same way, as when the question box at the top is asked many questions about one long article. Which models get it is decided as the engine starts: most text-generating ones yes, a few other kinds no.
- **Requests move together.** Chunked prefill is on by default, decided the same way: a long prompt is read in pieces sized to what a step has left, rather than in whole steps while everyone else waits. Scheduling is first come, first served.
- **Its batch sizes depend on the card.** Under 70 GiB of card memory, or on any A100, the API server batches up to 256 requests and 2,048 tokens per step; from 70 GiB, 1,024 requests and 8,192 tokens; from 160 GiB, 1,024 requests and 16,384 tokens (`get_batch_defaults`). A 24 GB card is in the first tier.
- **Its queue has no limit unless you set one.** Two limits arrived in v0.29.0, on requests waiting or running (`max_num_queued_reqs`) and on prompt tokens still being read; both default to none, and once one is reached, new requests are "rejected with HTTP 503", the web's code for a server too busy to answer.
- **The first start compiles; every start captures.** vLLM runs a model through `torch.compile` on its first start and keeps the output in `~/.cache/vllm/torch_compile_cache` ("the cache is enabled by default"). CUDA graphs, which record the model's steps for replay, are captured on every start. One flag turns both off (`--enforce-eager`): "This disables both `torch.compile` and CUDA graphs".

## Who can talk to it {#who-can-talk-to-it}

A server on a network address answers anyone who can reach it, unless it checks. vLLM checks, over part of its surface, and its security guide is the plainest description of that part. Read at v0.30.0 on 2026-09-30.

**The key.** vLLM takes its keys from `--api-key` (one or more since v0.10.1) or one from `VLLM_API_KEY`, and a client sends `Authorization: Bearer <key>`. The middleware, 62 lines in `authenticate.py`, keeps a SHA-256 digest of each key and compares in constant time; a request without a match gets HTTP 401. The keys are interchangeable: no name, limit or record attaches to any of them.

**What it guards.** Only requests whose path begins with one of four prefixes, named in the code as `GUARDED_PREFIX = ("/v1", "/v2", "/inference", "/cohere")`. The guide says the same and goes on: "Many other sensitive endpoints are exposed on the same HTTP server without any authentication enforcement." Its list of what stays open includes `/invocations`, the SageMaker-style route, which "routes to the same inference functions as `/v1` endpoints" and is "particularly concerning"; `/score` and `/rerank` in their non-`/v1` forms; `/pause`, which "causes denial of service"; `/tokenize`, `/health` and `/version`; and, if a development mode is switched on, `/collective_rpc`, "extremely dangerous". Above that list, the guide's overview says: "**Important:** Do not rely exclusively on `--api-key` for securing access to vLLM."

**What the guide says to do.** "The most effective approach is to deploy vLLM behind a reverse proxy (such as nginx, Envoy, or a Kubernetes Gateway)" that allowlists the routes you mean to expose, blocks the rest, and "Implements additional authentication, rate limiting, and logging at the proxy layer".

**A bypass, since fixed.** CVE-2026-48746 is a flaw on the public security record: the key check could be bypassed by an attacker able to "include special URL characters such as `/` or `?` in the `Host:` header", on versions from 0.3.0 up to but not including 0.22.0, which fixed it on 2026-05-29. The project's own advisory, GHSA-94f4-hr76-p5j6, published 2026-06-02, rates it critical, 9.1 of 10 on the CVSS 3.1 scale.

**Where it listens.** By default the launcher listens on every IPv4 address the computer has (`host` defaults to `None`), so any machine that can reach the computer over the network can reach its API; `--host 127.0.0.1`, the loopback address, keeps its API on the computer. The allowed browser origins default to `["*"]`, so any web page may call it until `--allowed-origins` narrows that. It can encrypt its own connections with TLS, the protocol behind https, with client certificates if you ask.

**What leaves the machine.** vLLM has sent a usage report by default since v0.4.0 (2024-03-30, "Usage statistics collection (#2852)"), and its page opened then as it opens now: "vLLM collects anonymous usage data by default to help the engineering team better understand which hardware and model configurations are widely used."

Once at start, under a random ID new for each process, vLLM sends the machine's details (cloud provider if detected, processor, memory, operating system and kernel, graphics cards, CUDA version), the vLLM version, the model's architecture (a class name such as `LlamaForCausalLM`, not the model's name), twenty-six settings of the run and five vLLM environment variables. Then, every ten minutes, it sends the same ID and the time, and at v0.30.0 nothing more. The report goes to `stats.vllm.ai`; a copy sits in `~/.config/vllm/usage_stats.json`. `VLLM_NO_USAGE_STATS=1`, `DO_NOT_TRACK=1` or an empty file at `~/.config/vllm/do_not_track` stops it (the code also honours `VLLM_DO_NOT_TRACK=1`). Started from a Hugging Face model ID, vLLM also reaches Hugging Face for the weights.

## The defaults that moved {#the-defaults-that-moved}

Some of what older forum threads say about vLLM was true when written and is not now. The dates below are the releases and commits that moved each default, read on 2026-09-30 and 2026-10-01.

| Date (UTC) | What moved | Before → after | Where it is written |
|---|---|---|---|
| 2023-10-07 | Where its API listens | `localhost` → unset (`None`), which then listened on every address the computer had, in a change titled "API server support ipv4 / ipv6 dualstack" | Commit 09ff7f10 (#1288); first shipped in v0.2.1 |
| 2025-01-27 | Prefix caching | Opt-in → on by default in the V1 engine: "Thanks to the near-zero overhead, we now enable prefix caching by default in V1" | The V1 alpha post |
| 2025-03-18 | The engine | V0 → V1, "for supported use cases"; `VLLM_USE_V1` stops being necessary | v0.8.0 release notes |
| 2025-10-02 | The V0 engine | A fallback → gone: "V1 is the only engine in the codebase now" | v0.11.0 release notes |
| 2026-04-21 | The share of the card reserved at start | 0.9 → 0.92, in the change that turned CUDA-graph memory profiling on by default | Commit 96a85c57 (#38284); first shipped in v0.20.0 |
| 2026-04-27 | The CUDA build `pip install vllm` gives | CUDA 12.9 → 13.0, and CUDA 13 "requires an R580 or newer driver" | v0.20.0 release notes; the CUDA installation page at v0.19.1 and at v0.30.0 |
| 2026-07-11 | The model runner, for dense models | Model Runner V1 → V2, "the default for all dense models"; `VLLM_USE_V2_MODEL_RUNNER=1`, the opt-in its announcement of 2026-03-24 gave, no longer needed for them | v0.25.0 release notes |
| 2026-09-09 | The model runner, for the rest | V1 → V2 "for all models", though "MRV1 remains in use for a few ROCm models and features MRV2 does not yet support" | v0.29.0 release notes |

The CUDA installation page at v0.30.0 still reads "vLLM's binaries are compiled with CUDA 12.9 and public PyTorch release versions by default", while the v0.30.0 release lists "PyPI (CUDA 13.0)" for the same command; the disagreement is the project's own.

Now picture two readers asking this article's question box at once. vLLM, by default, hands each answer its working memory a page at a time as it grows, and reads their shared opening once.

*A footnote on numbers. This page carries no measurement of our own, on purpose: every figure on it comes from the source named beside it. The workshop's own measurements of vLLM beside Ollama, another open-source model server, on one EVGA RTX 3090 XC3 Ultra (24 GB) — one request at a time and several at once, the same model on both — are coming soon, on a separate page.*

## What to take with you {#what-to-take-with-you}

- **The speed came out of a memory idea.** PagedAttention (arXiv 2023-09-12, SOSP 2023) put a request's working memory in fixed-size pages, as an operating system keeps a program's, so one card can hold many requests; the paging outlived the paper's kernels, deleted in v0.25.0 (2026-07-11). The v was first for virtual (Woosuk Kwon, 2023-08-25).
- **It was serving strangers before it was announced.** Two months at Chatbot Arena and the Vicuna demo before the post of 2023-06-20, by its own account.
- **One licence, never moved; no single party owns it, on purpose.** Apache-2.0 has been its licence since 2023-05-15, and the PyTorch Foundation its host since 2025-05-07. The company some of its creators announced on 2026-01-22 stands beside it, not over it.
- **It is a Linux server that reserves most of the card at start, and its defaults have moved.** Since v0.20.0 (2026-04-27) that share has been 92%. It compiles on first start, serves one model per process, and has run Model Runner V2 by default since 2026-07-11 for dense models and 2026-09-09 for nearly all the rest.
- **At v0.30.0 its API key guards four path prefixes, and its own guide says not to rely on it alone.** Its API listens on every IPv4 address until told otherwise; the guide's answer is a reverse proxy in front. A usage report goes out by default; three documented switches stop it.

## How to check our work — and see it live {#how-to-check-our-work-and-see-it-live}

- **Ask this page.** The "ask about this page" link at the top of every article here opens a question box whose model runs on vLLM; ask it what the v stands for and see whether it cites "What the v stands for".
- **Read the source at the tag.** Every file named here is linked at v0.30.0 under Sources, so a phone will do; at a computer, `git clone https://github.com/vllm-project/vllm && git -C vllm checkout v0.30.0`, then read the middleware's `GUARDED_PREFIX`, the 0.92 in `vllm/config/cache.py`, the batch tiers in `vllm/engine/arg_utils.py`, the reporter in `vllm/usage/usage_lib.py`, and `report_usage_stats` in `vllm/v1/utils.py`, which returns before it gathers anything when usage stats are off. `docs/usage/usage_stats.md` prints one example report, which it calls "an example as of v0.4.0".
- **Read the guide the project wrote.** `docs/usage/security.md` at the tag is the list of what the key does and does not cover, in the project's own words.

## The rest of the seminar {#the-rest-of-the-seminar}

- [A short history of Ollama](https://research.strata2signal.com/a-short-history-of-ollama/): the other model server in this set.
- [Ask About This Page](https://research.strata2signal.com/ask-about-this-page/): the question box behind the link at the top of this page, and how it works.
- [A Dinner Party for the Dead](https://research.strata2signal.com/a-dinner-party-for-the-dead/) and [How the long table works](https://research.strata2signal.com/how-the-long-table-works/): an exhibit whose guests are played by models that vLLM and Ollama serve.
- [What 150 Watts Buys](https://research.strata2signal.com/what-150-watts-buys/): what a lower power cap cost the two models our products serve, on an RTX PRO 6000 Blackwell (96 GB), one measured through vLLM and the other through Ollama.
- [ONNX Runtime's telemetry on Linux, measured](https://research.strata2signal.com/onnx-runtime-telemetry-on-linux/): an outbound call, and the method we use for one.
- [A short history of Mistral](https://research.strata2signal.com/a-short-history-of-mistral/): the same shape, for a company and its models rather than a server.

## Who ran this, and thanks {#who-ran-this-and-thanks}

Thanks to the **vLLM** project (Apache-2.0), whose launch post, blog, documentation, release notes and source are dated, detailed and still up, and whose maintainer answered a question about the name in two days in 2023; to the **PyTorch Foundation** and **LF AI & Data**, both under the **Linux Foundation**, for dated announcements; to **arXiv** and the **ACM**'s SOSP for the paper and its venue; to **a16z**, **Red Hat**, **Inferact** and **SiliconANGLE** for the money side; and to **GitHub**, whose record of the repository dates most of the rows above. The model this page opens with is served by vLLM, which runs over **PyTorch** (BSD 3-clause; the foundation calls vLLM "deeply integrated into PyTorch") and **CUDA** (NVIDIA, proprietary, the one closed piece in that stack). None of them owed us anything. A small human team asked for this page, chose what it would and would not claim, and signed off on it; a fleet of AI agents fetched and read every source listed below and drafted it under that team's rulings.

## Sources {#sources}

Every external link below was fetched and read between 2026-09-30 18:43 and 2026-10-01 05:33 (UTC); the version and default lines, and the sibling page's link, were read again on 2026-10-02 (UTC). Source files were read at the tags named; the tag-pinned URLs reproduce each read. Every GitHub release named here is dated by its release record's publication time.

- vLLM — [README at v0.30.0](https://github.com/vllm-project/vllm/blob/v0.30.0/README.md) ("Easy, fast, and cheap LLM serving for everyone"; "Originally developed in the Sky Computing Lab at UC Berkeley"; "over 2000 contributors"; "Efficient management of attention key and value memory with PagedAttention"; the feature, API, quantisation and hardware lists quoted above), and [at v0.1.0](https://github.com/vllm-project/vllm/blob/v0.1.0/README.md) ("OpenAI-compatible API server" among the features it lists)
- vLLM — [vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention](https://vllm.ai/blog/2023-06-20-vllm), 2023-06-20 (the byline of nine; "developed at UC Berkeley and deployed at Chatbot Arena and Vicuna Demo for the past two months"; the 30K / 60K requests and the 50% GPU figures; the 24x, 3.5x, 60% – 80% and under 4% claims; LMSYS as "a small research team" that "developed the popular Vicuna chatbot models"; Vicuna as having "been served in Chatbot Arena for millions of users"; TGI as "the previous state of the art")
- Kwon et al. — [Efficient Memory Management for Large Language Model Serving with PagedAttention](https://arxiv.org/abs/2309.06180), submitted 2023-09-12; "SOSP 2023" in its comments; "inspired by the classical virtual memory and paging techniques in operating systems"; "2-4×" against FasterTransformer and Orca
- ACM — [SOSP 2023](https://sosp2023.mpi-sws.org/), the 29th ACM Symposium on Operating Systems Principles, Koblenz, Germany, 2023-10-23 to 2023-10-26, whose [program](https://sosp2023.mpi-sws.org/program.html) lists the paper in Session 10 on 2023-10-26
- GitHub — the repository API for [`vllm-project/vllm`](https://github.com/vllm-project/vllm) (created 2023-02-09T11:23Z; licence Apache-2.0; 92,997 stars at 18:43 UTC on 2026-09-30), its releases API ([106 entries](https://github.com/vllm-project/vllm/releases) from v0.1.0 to v0.30.0, 9 of them marked pre-release, 32 in the twelve months to 2026-09-30, 20 of those new minor versions, v0.11.0 to v0.30.0; the publication times quoted above: v0.1.0 2023-06-20T06:28Z, v0.2.1 2023-10-16, v0.4.0 2024-03-30, v0.5.5 2024-08-23, v0.7.0 2025-01-27, v0.8.0 2025-03-18, v0.10.0 2025-07-24, v0.10.1 2025-08-18, v0.11.0 2025-10-02, v0.11.1 2025-11-18, v0.12.0 2025-12-03, v0.18.0 2026-03-20, v0.20.0 2026-04-27, v0.22.0 2026-05-29, v0.24.0 2026-06-29, v0.25.0 2026-07-11, v0.27.0 2026-08-10, v0.29.0 2026-09-09, v0.30.0 2026-09-22T05:20Z), and the [commits on its `LICENSE` path](https://github.com/vllm-project/vllm/commits/v0.30.0/LICENSE) (one: 89988ec8, 2023-05-15, "Add Apache-2.0 license (#102)")
- vLLM — the release notes for [v0.2.1](https://github.com/vllm-project/vllm/releases/tag/v0.2.1) ("API server support ipv4 / ipv6 dualstack"), [v0.4.0](https://github.com/vllm-project/vllm/releases/tag/v0.4.0) ("Usage statistics collection (#2852)"), [v0.5.5](https://github.com/vllm-project/vllm/releases/tag/v0.5.5) ("Support loading GGUF model (#5191)"), [v0.7.0](https://github.com/vllm-project/vllm/releases/tag/v0.7.0), [v0.8.0](https://github.com/vllm-project/vllm/releases/tag/v0.8.0), [v0.10.0](https://github.com/vllm-project/vllm/releases/tag/v0.10.0), [v0.10.1](https://github.com/vllm-project/vllm/releases/tag/v0.10.1), [v0.11.0](https://github.com/vllm-project/vllm/releases/tag/v0.11.0), [v0.11.1](https://github.com/vllm-project/vllm/releases/tag/v0.11.1), [v0.12.0](https://github.com/vllm-project/vllm/releases/tag/v0.12.0), [v0.18.0](https://github.com/vllm-project/vllm/releases/tag/v0.18.0), [v0.20.0](https://github.com/vllm-project/vllm/releases/tag/v0.20.0), [v0.22.0](https://github.com/vllm-project/vllm/releases/tag/v0.22.0), [v0.24.0](https://github.com/vllm-project/vllm/releases/tag/v0.24.0), [v0.25.0](https://github.com/vllm-project/vllm/releases/tag/v0.25.0) ("Model Runner V2 is now the default for all dense models"; "PagedAttention has been removed"), [v0.27.0](https://github.com/vllm-project/vllm/releases/tag/v0.27.0), [v0.29.0](https://github.com/vllm-project/vllm/releases/tag/v0.29.0) and [v0.30.0](https://github.com/vllm-project/vllm/releases/tag/v0.30.0) (each line quoted in the releases table is from the release, post or pull request named beside it)
- vLLM — [issue #835, "vLLM full name"](https://github.com/vllm-project/vllm/issues/835), opened 2023-08-23; Woosuk Kwon's answer of 2023-08-25, quoted above without its greeting
- vLLM — [vLLM V1: A Major Upgrade to vLLM's Core Architecture](https://vllm.ai/blog/2025-01-27-v1-alpha-release), 2025-01-27 ("up to 1.7x higher throughput compared to V0 (without multi-step scheduling)"; "we now enable prefix caching by default in V1"; `VLLM_USE_V1=1`)
- vLLM — [Model Runner V2: A Modular and Faster Core for vLLM](https://vllm.ai/blog/2026-03-24-mrv2), 2026-03-24 ("a ground-up re-implementation of the vLLM model runner"; `VLLM_USE_V2_MODEL_RUNNER=1`)
- vLLM — [Announcing vllm-metal: Concurrent Serving on Apple Silicon](https://vllm.ai/blog/2026-09-22-vllm-metal-v0-28-0), 2026-09-22 ("Our first official release, v0.28.0"; "vllm-metal builds on MLX and mlx_lm from Apple's MLX team"), and that [v0.28.0 release](https://github.com/vllm-project/vllm-metal/releases/tag/v0.28.0) on GitHub, published 2026-09-01T05:34Z; and [`docs/gguf.md`](https://github.com/vllm-project/vllm-metal/blob/v0.30.0/docs/gguf.md) at vllm-metal's v0.30.0, published 2026-09-23 ("vllm-metal supports dense decoder GGUF checkpoints through the MLX runtime", with a "Current scope" of four model families and three quantisation types)
- vLLM — [vLLM's Open Governance and Performance Roadmap](https://vllm.ai/blog/2024-07-25-lfai-perf), 2024-07-25 (the incubation sentence, "no one party will have exclusive control", and the consortium "such as" nine named groups "and others")
- LF AI & Data — [LF AI & Data Announces vLLM as its Latest Incubation Project](https://lfaidata.foundation/blog/2024/10/28/lf-ai-data-announces-vllm-as-its-latest-incubation-project/), dated 2024-10-28 and updated 2026-08-26
- PyTorch Foundation — [PyTorch Foundation Expands to Umbrella Foundation and Welcomes vLLM and DeepSpeed Projects](https://pytorch.org/blog/press-release-pytorch-foundation-expands-welcomes-projects-vllm-deepspeed/), dateline Paris, 2025-05-07 (the press release; "Hosted by the Linux Foundation"); [PyTorch Foundation Welcomes vLLM as a Hosted Project](https://pytorch.org/blog/pytorch-foundation-welcomes-vllm/), dated 2025-05-07 ("Contributed by the University of California – Berkeley"; "deeply integrated into PyTorch"; "over 46,500 GitHub stars and over 1000 contributors"); [PyTorch Foundation Expands to an Umbrella Foundation to Accelerate AI Innovation](https://pytorch.org/blog/pt-foundation-expands/), dated 2025-04-29 on its page (2025-04-30T03:34Z by its metadata) and updated 2025-05-03 (hosted projects "transfer assets to the Linux Foundation for independent stewardship and adopt an open governance model")
- a16z — [Supporting the Open Source AI Community](https://a16z.com/supporting-the-open-source-ai-community/), 2023-08-30 (the grant line naming Kwon and Li; "grant funding (**not** an investment or SAFE note)")
- Red Hat — [Red Hat Completes Acquisition of Neural Magic to Fuel Optimized Generative AI Innovation Across the Hybrid Cloud](https://www.redhat.com/en/about/press-releases/red-hat-completes-acquisition-neural-magic-fuel-optimized-generative-ai-innovation-across-hybrid-cloud), 2025-01-13 ("a leading contributor to vLLM, an open source project developed by UC Berkeley for open model serving"), and [What is vLLM?](https://www.redhat.com/en/topics/ai/what-is-vllm), shown as published 2026-05-12 ("vLLM, which stands for virtual large language model")
- Inferact — [its launch post](https://inferact.ai/), dated 2026-01-22 and carried on its home page when read ("a startup founded by creators and core maintainers of vLLM"; the six named signatories; "has raised $150M at $800M valuation, led by a16z and Lightspeed"; "The optimizations we develop flow back to the community"), and [its news of 2026-08-12](https://inferact.ai/news/kimi-k3-kvv) ("is built on top of OSS vLLM")
- SiliconANGLE — [Inferact launches with $150M in funding to commercialize vLLM](https://siliconangle.com/2026/01/22/inferact-launches-150m-funding-commercialize-vllm/), dated 2026-01-22 on its page and published 2026-01-23T01:13Z by its metadata ("$150 million in seed funding"; Stoica as a "computer science professor", a Databricks co-founder and the director of the Sky Computing Lab; "The blog post hints that Inferact plans to launch a paid serverless version of vLLM"; the one press report on this page, read beside the company's own post)
- vLLM — the source at tag v0.30.0: [`vllm/entrypoints/serve/middleware/authenticate.py`](https://github.com/vllm-project/vllm/blob/v0.30.0/vllm/entrypoints/serve/middleware/authenticate.py) (`GUARDED_PREFIX`, the SHA-256 digests, `secrets.compare_digest`, the 401), [`vllm/entrypoints/launchers/cli_args.py`](https://github.com/vllm-project/vllm/blob/v0.30.0/vllm/entrypoints/launchers/cli_args.py) (`host: str | None = None`; `allowed_origins` `["*"]`; `api_key: list[str] | None`; the `ssl_*` flags; `--grpc`, "Launch a gRPC server instead of the HTTP OpenAI-compatible server"), [`vllm/entrypoints/launchers/launcher.py`](https://github.com/vllm-project/vllm/blob/v0.30.0/vllm/entrypoints/launchers/launcher.py) (`sock_addr = (args.host or "", args.port)`), [`vllm/config/cache.py`](https://github.com/vllm-project/vllm/blob/v0.30.0/vllm/config/cache.py) (`gpu_memory_utilization` default 0.92, "a per-instance limit"), [`vllm/config/load.py`](https://github.com/vllm-project/vllm/blob/v0.30.0/vllm/config/load.py) (`load_format` `auto`: it "will try to load the weights in the safetensors format and fall back to the pytorch bin format"), [`vllm/config/scheduler.py`](https://github.com/vllm-project/vllm/blob/v0.30.0/vllm/config/scheduler.py) (`policy` fcfs; the two queue limits, on requests "in-flight (waiting or running)" and on prompt tokens "in the prefill phase", defaulting to `None`, "rejected with HTTP 503"), [`vllm/engine/arg_utils.py`](https://github.com/vllm-project/vllm/blob/v0.30.0/vllm/engine/arg_utils.py) (`get_batch_defaults`: the 160 GiB, 70 GiB and A100 branches; `_set_default_chunked_prefill_and_prefix_caching_args`: prefix caching and chunked prefill on for each model that supports them), [`vllm/config/model.py`](https://github.com/vllm-project/vllm/blob/v0.30.0/vllm/config/model.py) (`max_model_len`: "If unspecified, will be automatically derived from the model config"; `enforce_eager`: "This disables both `torch.compile` and CUDA graphs"; `is_prefix_caching_supported` and `is_chunked_prefill_supported`, the two tests the defaults follow), [`vllm/usage/usage_lib.py`](https://github.com/vllm-project/vllm/blob/v0.30.0/vllm/usage/usage_lib.py) (`uuid4()` per process; the fields collected once, the five environment variables of `_USAGE_ENV_VARS_TO_COLLECT` among them; the ten-minute loop; the local copy), [`vllm/v1/utils.py`](https://github.com/vllm-project/vllm/blob/v0.30.0/vllm/v1/utils.py) (`report_usage_stats`: it returns before collecting anything when usage stats are off, and adds twenty-six settings of the run to the first report), [`vllm/v1/core/kv_cache_utils.py`](https://github.com/vllm-project/vllm/blob/v0.30.0/vllm/v1/core/kv_cache_utils.py) (`_check_enough_kv_cache_memory`: "To serve at least one request with the model's max seq len"), [`vllm/v1/engine/input_processor.py`](https://github.com/vllm-project/vllm/blob/v0.30.0/vllm/v1/engine/input_processor.py) (`_validate_prompt_len`: a prompt "longer than the maximum model length" is an error), [`vllm/v1/worker/utils.py`](https://github.com/vllm-project/vllm/blob/v0.30.0/vllm/v1/worker/utils.py) (`request_memory`: "Decrease GPU memory utilization or reduce GPU memory used by other processes."), [`vllm/v1/worker/gpu_worker.py`](https://github.com/vllm-project/vllm/blob/v0.30.0/vllm/v1/worker/gpu_worker.py) (`compile_or_warm_up_model`: `capture_model()` on every start unless `enforce_eager`) and [`vllm/envs.py`](https://github.com/vllm-project/vllm/blob/v0.30.0/vllm/envs.py) (`VLLM_USAGE_STATS_SERVER` `https://stats.vllm.ai`; `VLLM_NO_USAGE_STATS`; `DO_NOT_TRACK` and `VLLM_DO_NOT_TRACK`; `VLLM_ENABLE_COHERE_API` off by default), [`vllm/entrypoints/cohere/api_router.py`](https://github.com/vllm-project/vllm/blob/v0.30.0/vllm/entrypoints/cohere/api_router.py) (`attach_router`: no Cohere chat route without the flag and the `cohere` package), [`pyproject.toml`](https://github.com/vllm-project/vllm/blob/v0.30.0/pyproject.toml) (`requires-python = ">=3.10,<3.15"`)
- vLLM — the docs at tag v0.30.0: [`docs/usage/security.md`](https://github.com/vllm-project/vllm/blob/v0.30.0/docs/usage/security.md) (the four prefixes; the guarded and unprotected lists; "Do not rely exclusively on `--api-key`"; "Deploy Behind a Reverse Proxy"; run-time LoRA "not a secure operation"), [`docs/usage/usage_stats.md`](https://github.com/vllm-project/vllm/blob/v0.30.0/docs/usage/usage_stats.md) ("collects anonymous usage data by default"; the three switches; "an example as of v0.4.0"; `tail ~/.config/vllm/usage_stats.json`) and [at v0.4.0](https://github.com/vllm-project/vllm/blob/v0.4.0/docs/source/serving/usage_stats.md), then at `docs/source/serving/usage_stats.md` (the same opening sentence), [`docs/usage/faq.md`](https://github.com/vllm-project/vllm/blob/v0.30.0/docs/usage/faq.md) (several models "not currently supported"; one instance per model), [`docs/getting_started/installation/gpu.md`](https://github.com/vllm-project/vllm/blob/v0.30.0/docs/getting_started/installation/gpu.md) ("OS: Linux"; Python 3.10 to 3.13; "does not support Windows natively"; WSL), [`gpu.cuda.inc.md`](https://github.com/vllm-project/vllm/blob/v0.30.0/docs/getting_started/installation/gpu.cuda.inc.md) ("compute capability 7.5 or higher"; "vLLM offers an official Docker image for deployment"; "compiled with CUDA 12.9 … by default"; "CUDA 13 requires an R580 or newer driver") and [at v0.19.1](https://github.com/vllm-project/vllm/blob/v0.19.1/docs/getting_started/installation/gpu.cuda.inc.md) (the same "compiled with CUDA 12.9 … by default", before the switch), [`gpu.apple.inc.md`](https://github.com/vllm-project/vllm/blob/v0.30.0/docs/getting_started/installation/gpu.apple.inc.md) ("a community-maintained hardware plugin that uses MLX as the compute backend"), [`cpu.apple.inc.md`](https://github.com/vllm-project/vllm/blob/v0.30.0/docs/getting_started/installation/cpu.apple.inc.md) ("experimental support for macOS with Apple Silicon"; build from source; processor only), [`cpu.md`](https://github.com/vllm-project/vllm/blob/v0.30.0/docs/getting_started/installation/cpu.md) (x86, ARM, Apple silicon and IBM Z tabs), [`docs/getting_started/quickstart.md`](https://github.com/vllm-project/vllm/blob/v0.30.0/docs/getting_started/quickstart.md) ("By default, vLLM downloads models from Hugging Face"), [`docs/features/quantization/gguf.md`](https://github.com/vllm-project/vllm/blob/v0.30.0/docs/features/quantization/gguf.md) ("highly experimental and under-optimized"; the plugin), [`docs/features/tool_calling.md`](https://github.com/vllm-project/vllm/blob/v0.30.0/docs/features/tool_calling.md) (`--enable-auto-tool-choice` and `--tool-call-parser`; "several seconds of latency" on the first named call), [`docs/features/structured_outputs.md`](https://github.com/vllm-project/vllm/blob/v0.30.0/docs/features/structured_outputs.md) (the backends; default `auto`), [`docs/design/paged_attention.md`](https://github.com/vllm-project/vllm/blob/v0.30.0/docs/design/paged_attention.md) ("This is a historical document based on the original paper for vLLM. It no longer describes the code used in vLLM today.", a warning the page has opened with since v0.10.1), [`docs/design/torch_compile.md`](https://github.com/vllm-project/vllm/blob/v0.30.0/docs/design/torch_compile.md) (the cache directory; "the cache is enabled by default") and [`docs/serving/offline_inference.md`](https://github.com/vllm-project/vllm/blob/v0.30.0/docs/serving/offline_inference.md) ("Offline inference is possible in your own code"; the library with no server)
- GitHub — [commit 96a85c57](https://github.com/vllm-project/vllm/commit/96a85c57), 2026-04-21T22:16Z, "[Startup][UX] Enable CUDAGraph memory profiling by default (#38284)", whose diff to `vllm/config/cache.py` changes `default=0.9` to `default=0.92`; its [pull request #38284](https://github.com/vllm-project/vllm/pull/38284) "correspondingly increases the default `--gpu-memory-utilization` to 0.92" as it turns the profiling on; first shipped in v0.20.0 (GitHub's comparison finds the commit in v0.20.0 and not in v0.19.1)
- GitHub — [GHSA-94f4-hr76-p5j6](https://github.com/advisories/GHSA-94f4-hr76-p5j6), CVE-2026-48746, published by the project on 2026-06-02 and in GitHub's advisory database on 2026-06-16 (critical, CVSS 3.1 score 9.1; affected ≥ 0.3.0 and < 0.22.0; patched in 0.22.0; "include special URL characters such as `/` or `?` in the `Host:` header")
- GitHub — [pull request #2852](https://github.com/vllm-project/vllm/pull/2852), "Usage Stats Collection", merged 2024-03-29, first shipped in v0.4.0; [pull request #47361](https://github.com/vllm-project/vllm/pull/47361), "Delete PagedAttention", merged 2026-07-02, which deletes `paged_attention_v1.cu`, `paged_attention_v2.cu` and their shared kernels, first shipped in v0.25.0 (GitHub's comparison finds it in v0.25.0 and not in v0.24.0)
- GitHub — [pull request #1288](https://github.com/vllm-project/vllm/pull/1288), "API server support ipv4 / ipv6 dualstack", merged 2023-10-07 (commit 09ff7f10), whose diff to `vllm/entrypoints/openai/api_server.py` changes the `--host` default from `"localhost"` to `None` (its description reports `localhost` failing on the IPv6 loopback address); first shipped in v0.2.1 (GitHub's comparison finds it in v0.2.1 and not in v0.2.0)
- GitHub — [pull request #11576](https://github.com/vllm-project/vllm/pull/11576), "[Misc] Minimum requirements for SageMaker compatibility", merged 2025-01-02 ("Implements `/ping` and `/invocations`"), the origin of the `/invocations` route
- This workshop's own released pages, linked above: [Ask About This Page](https://research.strata2signal.com/ask-about-this-page/) ("One line at the top of each released article leads to that article's room"; the seat: "a gemma-class open-weight writer on a vLLM runtime with FP8 weights"), [How the long table works](https://research.strata2signal.com/how-the-long-table-works/) ("the chairs are served by vLLM and Ollama"), [What 150 Watts Buys](https://research.strata2signal.com/what-150-watts-buys/) ("the two models our products actually serve": one measured through vLLM as "the candidate fast lane", the other through Ollama; its hardware line names the RTX PRO 6000 Blackwell (96 GB)), [A Dinner Party for the Dead](https://research.strata2signal.com/a-dinner-party-for-the-dead/), [ONNX Runtime's telemetry on Linux, measured](https://research.strata2signal.com/onnx-runtime-telemetry-on-linux/), [A short history of Mistral](https://research.strata2signal.com/a-short-history-of-mistral/), [A short history of Ollama](https://research.strata2signal.com/a-short-history-of-ollama/)

Corrections and later measurements will be added below, each dated (UTC), with a window at both ends where one applies and saying in plain words what it counts.

<!-- derived 2026-10-02 (UTC) by tools/derive_md.py from the pour source.
     source html sha256: c4e47e6e7acabb9cea35c2e4f0f4e7ef0cb283dcf819d558452db5f2322f491e
     derivation sha256:  53a5a6f2c3e849b95e0dfedf12ee9717922d27686b94fd736aaac89f26a7f980
     the {#id} on each heading is the anchor that heading carries on the page. -->
