{
 "schema_version": "1.0",
 "page_kind": "article",
 "pack_kind": "generate",
 "slug": "a-short-history-of-vllm",
 "title": "A short history of vLLM",
 "dek": "The open-source server that lets one graphics card run a language model for many people at once: a Berkeley lab's borrowing from operating systems, the foundation and the company around it, and the defaults that moved under its users, dated.",
 "published": "2026-10-02",
 "series": [
  "notes"
 ],
 "licence": "CC BY 4.0",
 "status": "pending",
 "pack_note": "5 chip(s) survived the gates and 4 of them were removed by hand, leaving 1; the floor is 2; no judge is seated, so every chip is held",
 "notice": "items with published:false are held and are not this page's published words",
 "source": {
  "url": "https://research.strata2signal.com/a-short-history-of-vllm/",
  "md_url": "https://research.strata2signal.com/a-short-history-of-vllm/index.md",
  "md_sha": "7d7851012e1a07f283eb215ba26def41d77286d5181873e817093981c8cd3f21",
  "html_sha": "beaadb5236b990c86f95978ad63441417855520a0c07ffafb76215407c5b738f",
  "anchors_sha": "a6003ca9df5183a79ebfeccc5be2df2ac43d5d94a56d2f875debce86bd34022e"
 },
 "short": {
  "paragraph": "vLLM is free, open-source software that serves a large language model, the kind of AI inside a chatbot; on this hub it runs the model behind every article's \"ask about this page\" link. It began in a UC Berkeley lab and was announced on 2023-06-20, built to make one graphics card hold many conversations at once: like an operating system, vLLM hands out working memory in pages as each answer grows. It has been a PyTorch Foundation project since 2025-05-07; some of its creators announced a $150M-funded company beside it on 2026-01-22. This page also dates the defaults that moved under vLLM's users. The only speed figures here are the project's own claims, each dated, none re-run by us.",
  "from_dek": false,
  "counts": {
   "words": 6633,
   "minutes": 30,
   "tables": 5,
   "kit": false
  },
  "bullets": []
 },
 "sections": [
  {
   "id": "one-card-many-questions",
   "heading": "One card, many questions at once",
   "level": 2,
   "span": [
    1876,
    3142
   ],
   "chunks": [
    [
     1876,
     3142
    ]
   ],
   "chars": 1266,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "who-are-these-people",
   "heading": "Who are these people, and where did they come from?",
   "level": 2,
   "span": [
    3142,
    4670
   ],
   "chunks": [
    [
     3142,
     4670
    ]
   ],
   "chars": 1528,
   "digest": "Originally developed at the Sky Computing Lab at UC Berkeley, vLLM was deployed at Chatbot Arena and Vicuna Demo for two months before its announcement. The project's numbers state that LMSYS was able to cut the number of GPUs used for serving traffic by 50%.",
   "digest_skipped": null
  },
  {
   "id": "what-the-v-stands-for",
   "heading": "What the v stands for",
   "level": 2,
   "span": [
    4670,
    5457
   ],
   "chunks": [
    [
     4670,
     5457
    ]
   ],
   "chars": 787,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "paging-for-a-language-model",
   "heading": "The idea: paging for a language model",
   "level": 2,
   "span": [
    5457,
    7644
   ],
   "chunks": [
    [
     5457,
     7644
    ]
   ],
   "chars": 2187,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "who-owns-it",
   "heading": "Who owns it, and who pays for it",
   "level": 2,
   "span": [
    7644,
    11718
   ],
   "chunks": [
    [
     7644,
     11718
    ]
   ],
   "chars": 4074,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "the-releases-where-it-bent",
   "heading": "The releases where it bent",
   "level": 2,
   "span": [
    11718,
    13996
   ],
   "chunks": [
    [
     11718,
     13996
    ]
   ],
   "chars": 2278,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "what-it-speaks-and-where-it-runs",
   "heading": "What it speaks, and where it runs",
   "level": 2,
   "span": [
    13996,
    17986
   ],
   "chunks": [
    [
     13996,
     17986
    ]
   ],
   "chars": 3990,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "how-it-serves-a-request",
   "heading": "How it serves a request",
   "level": 2,
   "span": [
    17986,
    20607
   ],
   "chunks": [
    [
     17986,
     20607
    ]
   ],
   "chars": 2621,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "who-can-talk-to-it",
   "heading": "Who can talk to it",
   "level": 2,
   "span": [
    20607,
    24433
   ],
   "chunks": [
    [
     20607,
     24433
    ]
   ],
   "chars": 3826,
   "digest": "vLLM uses API keys to guard specific path prefixes. The security guide advises deploying vLLM behind a reverse proxy to implement additional authentication and rate limiting. A critical flaw, CVE-2026-48746, was fixed on 2026-05-29.",
   "digest_skipped": null
  },
  {
   "id": "the-defaults-that-moved",
   "heading": "The defaults that moved",
   "level": 2,
   "span": [
    24433,
    27075
   ],
   "chunks": [
    [
     24433,
     27075
    ]
   ],
   "chars": 2642,
   "digest": "Several defaults in vLLM have changed over time, including the API server listening address, prefix caching, and the default engine. The share of the card reserved at start moved from 0.9 to 0.92 in a change shipped in v0.20.0.",
   "digest_skipped": null
  },
  {
   "id": "what-to-take-with-you",
   "heading": "What to take with you",
   "level": 2,
   "span": [
    27075,
    28493
   ],
   "chunks": [
    [
     27075,
     28493
    ]
   ],
   "chars": 1418,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "how-to-check-our-work-and-see-it-live",
   "heading": "How to check our work — and see it live",
   "level": 2,
   "span": [
    28493,
    29539
   ],
   "chunks": [
    [
     28493,
     29539
    ]
   ],
   "chars": 1046,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "the-rest-of-the-seminar",
   "heading": "The rest of the seminar",
   "level": 2,
   "span": [
    29539,
    30716
   ],
   "chunks": [
    [
     29539,
     30716
    ]
   ],
   "chars": 1177,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "who-ran-this-and-thanks",
   "heading": "Who ran this, and thanks",
   "level": 2,
   "span": [
    30716,
    31798
   ],
   "chunks": [
    [
     30716,
     31798
    ]
   ],
   "chars": 1082,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "sources",
   "heading": "Sources",
   "level": 2,
   "span": [
    31798,
    51631
   ],
   "chunks": [
    [
     31798,
     51631
    ]
   ],
   "chars": 19833,
   "digest": null,
   "digest_skipped": null
  }
 ],
 "chips": [
  {
   "id": "c-7e85c52c",
   "q": "what api formats does vllm support?",
   "a": "vLLM provides an OpenAI-compatible API server, Anthropic Messages API support, and gRPC support. It also offers Cohere-style APIs for Rerank, Embed, and Chat, though some require specific settings or packages to be enabled.",
   "cites": [
    "what-it-speaks-and-where-it-runs"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  }
 ],
 "related": [
  {
   "slug": "a-short-history-of-ollama",
   "why": "shares ground with § Who are these people, and where did they come from? · § The engine underneath: llama.cpp"
  },
  {
   "slug": "llms-txt",
   "why": "shares ground with § Who can talk to it · § The confession."
  },
  {
   "slug": "the-free-speed-wasnt-free",
   "why": "shares ground with § The defaults that moved · § The part that isn't about speed"
  }
 ],
 "thanks": "## Who ran this, and thanks {#who-ran-this-and-thanks}\n\nThanks to the **vLLM** project (Apache-2.0), whose launch post, blog, documentation, release notes and source are dated, detailed and still up, and whose maintainer answered a question about the name in two days in 2023; to the **PyTorch Foundation** and **LF AI & Data**, both under the **Linux Foundation**, for dated announcements; to **arXiv** and the **ACM**'s SOSP for the paper and its venue; to **a16z**, **Red Hat**, **Inferact** and **SiliconANGLE** for the money side; and to **GitHub**, whose record of the repository dates most of the rows above. The model this page opens with is served by vLLM, which runs over **PyTorch** (BSD 3-clause; the foundation calls vLLM \"deeply integrated into PyTorch\") and **CUDA** (NVIDIA, proprietary, the one closed piece in that stack). None of them owed us anything. A small human team asked for this page, chose what it would and would not claim, and signed off on it; a fleet of AI agents fetched and read every source listed below and drafted it under that team's rulings.",
 "kit": null,
 "seat_class": {
  "writer": "gemma-class",
  "judge": null,
  "writer_runtime": "vllm"
 },
 "bench": {
  "bullets_written": 3,
  "bullets_kept": 0,
  "digests_written": 12,
  "digests_kept": 3,
  "chips_written": 8,
  "chips_kept": 1,
  "chips_grounded": 0,
  "chips_published": 0,
  "dropped_by": {
   "new_noun": 3,
   "figure": 0,
   "length": 0,
   "cite": 0,
   "judge": 0,
   "quote": 0,
   "directive": 0,
   "redaction": 0,
   "profanity": 0,
   "identifier": 0,
   "bare_figure": 0,
   "house_voice": 0,
   "hand-edit": 16
  },
  "new_noun_tokens_checked": [
   "0.9",
   "0.92",
   "024",
   "1",
   "2023",
   "2025",
   "2025-10-02",
   "2026-05-29",
   "2026-09-30",
   "2026-10-01",
   "24x",
   "4%",
   "50%",
   "60%",
   "70",
   "80%",
   "92%",
   "AMD",
   "API",
   "APIs",
   "Anthropic",
   "Apache-2.0",
   "Apple",
   "Arena",
   "Berkeley",
   "CPU",
   "CVE-2026-48746",
   "Chat",
   "Chatbot",
   "Cohere-style",
   "Computing",
   "Demo",
   "Embed",
   "Foundation",
   "GPUs",
   "GiB",
   "GitHub",
   "Inferact",
   "Intel",
   "KV",
   "LMSYS",
   "Lab",
   "Linux",
   "May",
   "Messages",
   "Model",
   "NVIDIA",
   "OS",
   "OpenAI-compatible",
   "PagedAttention",
   "PyTorch",
   "Rerank",
   "Runner",
   "Silicon",
   "Sky",
   "UC",
   "V0",
   "V1",
   "Vicuna",
   "v0.20.0",
   "v0.30.0",
   "v0.32"
  ],
  "new_noun_tokens_withheld": 3,
  "figure_definition": "v3",
  "figure_boundary": "guarded",
  "drops": [
   {
    "kind": "chip",
    "reason": "new_noun",
    "item": null,
    "item_chars": 7,
    "withheld": "a string naming something the article does not",
    "detail": ""
   },
   {
    "kind": "chip",
    "reason": "new_noun",
    "item": null,
    "item_chars": 4,
    "withheld": "a string naming something the article does not",
    "detail": ""
   },
   {
    "kind": "chip",
    "reason": "new_noun",
    "item": null,
    "item_chars": 5,
    "withheld": "a string naming something the article does not",
    "detail": ""
   },
   {
    "kind": "bullet",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 32,
    "text_sha": "5e17d19e47ce",
    "figure": "24",
    "withheld": "one bullet removed from this pack by hand",
    "removed_utc": "2026-10-02",
    "why": "the line drops what the 24x measures and what it is against: the section gives vLLM's launch claim of 2023, up to 24x higher throughput than HuggingFace Transformers, says the workshop has not re-run it, and says Transformers, TGI and vLLM have all changed since",
    "detail": "removed from this pack by hand on 2026-10-02 (UTC), after the writer ran and after the gates had counted: the line drops what the 24x measures and what it is against: the section gives vLLM's launch claim of 2023, up to 24x higher throughput than HuggingFace Transformers, says the workshop has not re-run it, and says Transformers, TGI and vLLM have all changed since. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "bullet",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 37,
    "text_sha": "5bb340816ec0",
    "figure": "80",
    "withheld": "one bullet removed from this pack by hand",
    "removed_utc": "2026-10-02",
    "why": "the line states the launch post's claim of 2023 as a present fact and drops what is wasted: the section gives it in the post's own terms, that existing systems waste 60% to 80% of memory due to fragmentation and over-reservation, and the page carries no measurement of its own",
    "detail": "removed from this pack by hand on 2026-10-02 (UTC), after the writer ran and after the gates had counted: the line states the launch post's claim of 2023 as a present fact and drops what is wasted: the section gives it in the post's own terms, that existing systems waste 60% to 80% of memory due to fragmentation and over-reservation, and the page carries no measurement of its own. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "digest",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 288,
    "text_sha": "5a916f56c580",
    "section": "one-card-many-questions",
    "withheld": "one section digest removed from this pack by hand",
    "removed_utc": "2026-10-02",
    "why": "the digest inverts the section: it says vLLM gives each request a fixed slab of the KV cache, where the section describes that fixed slab, sized for the longest answer allowed and mostly empty, as the state of things in the spring of 2023 that vLLM was written to end; its first sentence also runs in a circle, vLLM running on a card using vLLM",
    "detail": "removed from this pack by hand on 2026-10-02 (UTC), after the writer ran and after the gates had counted: the digest inverts the section: it says vLLM gives each request a fixed slab of the KV cache, where the section describes that fixed slab, sized for the longest answer allowed and mostly empty, as the state of things in the spring of 2023 that vLLM was written to end; its first sentence also runs in a circle, vLLM running on a card using vLLM. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "digest",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 276,
    "text_sha": "18346be5cd22",
    "section": "paging-for-a-language-model",
    "withheld": "one section digest removed from this pack by hand",
    "removed_utc": "2026-10-02",
    "why": "the digest states the launch post's claim, memory waste cut from 60% to 80% to under 4%, as a fact of the engine; the section gives both figures in the post's own terms, and the page carries no measurement of its own",
    "detail": "removed from this pack by hand on 2026-10-02 (UTC), after the writer ran and after the gates had counted: the digest states the launch post's claim, memory waste cut from 60% to 80% to under 4%, as a fact of the engine; the section gives both figures in the post's own terms, and the page carries no measurement of its own. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "digest",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 229,
    "text_sha": "a1c0e15385ba",
    "section": "the-releases-where-it-bent",
    "withheld": "one section digest removed from this pack by hand",
    "removed_utc": "2026-10-02",
    "why": "the digest calls both rewrites new engines; the section says the second, Model Runner V1 and the V2 that replaced it, is the part inside the V1 engine that runs the model, not an engine of its own",
    "detail": "removed from this pack by hand on 2026-10-02 (UTC), after the writer ran and after the gates had counted: the digest calls both rewrites new engines; the section says the second, Model Runner V1 and the V2 that replaced it, is the part inside the V1 engine that runs the model, not an engine of its own. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "digest",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 252,
    "text_sha": "c8b461df2772",
    "section": "what-it-speaks-and-where-it-runs",
    "withheld": "one section digest removed from this pack by hand",
    "removed_utc": "2026-10-02",
    "why": "the digest says NVIDIA, AMD and Intel graphics are supported via plugins; the section's table supports NVIDIA cards of compute capability 7.5 or higher, AMD cards through ROCm and Intel graphics through vLLM's own Intel XPU backend, and names plugins only for other accelerators, Google TPUs and Intel Gaudi among them",
    "detail": "removed from this pack by hand on 2026-10-02 (UTC), after the writer ran and after the gates had counted: the digest says NVIDIA, AMD and Intel graphics are supported via plugins; the section's table supports NVIDIA cards of compute capability 7.5 or higher, AMD cards through ROCm and Intel graphics through vLLM's own Intel XPU backend, and names plugins only for other accelerators, Google TPUs and Intel Gaudi among them. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "digest",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 243,
    "text_sha": "9bd6aabae4e8",
    "section": "how-it-serves-a-request",
    "withheld": "one section digest removed from this pack by hand",
    "removed_utc": "2026-10-02",
    "why": "the digest gives 1,024 requests as the batch for a card over 70 GiB; the section says the API server batches up to that many from 70 GiB, and keeps one card family it names in the first tier whatever its memory, up to 256 requests and 2,048 tokens per step, so the digest is wrong for that family's cards over 70 GiB",
    "detail": "removed from this pack by hand on 2026-10-02 (UTC), after the writer ran and after the gates had counted: the digest gives 1,024 requests as the batch for a card over 70 GiB; the section says the API server batches up to that many from 70 GiB, and keeps one card family it names in the first tier whatever its memory, up to 256 requests and 2,048 tokens per step, so the digest is wrong for that family's cards over 70 GiB. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "digest",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 221,
    "text_sha": "7b935f14d2bc",
    "section": "sources",
    "withheld": "one section digest removed from this pack by hand",
    "removed_utc": "2026-10-02",
    "why": "the Sources list names real people (the paper's authors, the launch post's signatories, maintainers); a model does not paraphrase a named third party, and pack_corpus.NEVER_GENERATED does not match the heading",
    "detail": "removed from this pack by hand on 2026-10-02 (UTC), after the writer ran and after the gates had counted: the Sources list names real people (the paper's authors, the launch post's signatories, maintainers); a model does not paraphrase a named third party, and pack_corpus.NEVER_GENERATED does not match the heading. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "chip",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 270,
    "text_sha": "16150120c927",
    "chip_id": "c-1ac295cc",
    "withheld": "one chip removed from this pack by hand",
    "removed_utc": "2026-10-02",
    "why": "the answer states the launch post's claim, memory waste cut to under 4%, as a fact of the method; the section gives it in the post's own terms, and the page carries no measurement of its own",
    "detail": "removed from this pack by hand on 2026-10-02 (UTC), after the writer ran and after the gates had counted: the answer states the launch post's claim, memory waste cut to under 4%, as a fact of the method; the section gives it in the post's own terms, and the page carries no measurement of its own. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "bullet",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 55,
    "text_sha": "c8f10b349127",
    "figure": "0.92",
    "withheld": "one bullet removed from this pack by hand",
    "removed_utc": "2026-10-02",
    "why": "the default share is dated on the page (read from the v0.30.0 source; 92% since v0.20.0); the line drops the version, its condition, on a page that exists to date defaults that moved",
    "detail": "removed from this pack by hand on 2026-10-02 (UTC), after the writer ran and after the gates had counted: the default share is dated on the page (read from the v0.30.0 source; 92% since v0.20.0); the line drops the version, its condition, on a page that exists to date defaults that moved. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "digest",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 234,
    "text_sha": "acdbc3b3c2bb",
    "section": "what-to-take-with-you",
    "withheld": "one section digest removed from this pack by hand",
    "removed_utc": "2026-10-02",
    "why": "states the 92% default with no version; the page scopes it to the v0.30.0 source and since v0.20.0, and this page dates defaults that moved",
    "detail": "removed from this pack by hand on 2026-10-02 (UTC), after the writer ran and after the gates had counted: states the 92% default with no version; the page scopes it to the v0.30.0 source and since v0.20.0, and this page dates defaults that moved. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "chip",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 218,
    "text_sha": "c4ccbd3830f6",
    "chip_id": "c-b0041e8e",
    "withheld": "one chip removed from this pack by hand",
    "removed_utc": "2026-10-02",
    "why": "the 92% default stated with no version; the page dates it (v0.30.0 source; since v0.20.0)",
    "detail": "removed from this pack by hand on 2026-10-02 (UTC), after the writer ran and after the gates had counted: the 92% default stated with no version; the page dates it (v0.30.0 source; since v0.20.0). Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "chip",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 267,
    "text_sha": "3aa215b2c023",
    "chip_id": "c-4ab4d233",
    "withheld": "one chip removed from this pack by hand",
    "removed_utc": "2026-10-02",
    "why": "states as a fact about vLLM what the page gives as the foundation's own description of any hosted project; the attribution is lost",
    "detail": "removed from this pack by hand on 2026-10-02 (UTC), after the writer ran and after the gates had counted: states as a fact about vLLM what the page gives as the foundation's own description of any hosted project; the attribution is lost. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "digest",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 245,
    "text_sha": "62597ad5d30f",
    "section": "what-the-v-stands-for",
    "withheld": "one section digest removed from this pack by hand",
    "removed_utc": "2026-10-02",
    "why": "states as fact what the page gives as Woosuk Kwon's answer of 2023-08-25, the only account of the name it found; the attribution is lost",
    "detail": "removed from this pack by hand on 2026-10-02 (UTC), after the writer ran and after the gates had counted: states as fact what the page gives as Woosuk Kwon's answer of 2023-08-25, the only account of the name it found; the attribution is lost. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "chip",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 245,
    "text_sha": "9b5fa58a8775",
    "chip_id": "c-6bbc0096",
    "withheld": "one chip removed from this pack by hand",
    "removed_utc": "2026-10-02",
    "why": "states as fact what the page gives as Woosuk Kwon's answer of 2023-08-25, the only account of the name it found; the attribution is lost",
    "detail": "removed from this pack by hand on 2026-10-02 (UTC), after the writer ran and after the gates had counted: states as fact what the page gives as Woosuk Kwon's answer of 2023-08-25, the only account of the name it found; the attribution is lost. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "digest",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 200,
    "text_sha": "65d6d717ff07",
    "section": "who-owns-it",
    "withheld": "one section digest removed from this pack by hand",
    "removed_utc": "2026-10-02",
    "why": "says as fact that creators and core maintainers of vLLM founded Inferact; the cited section gives those words only as Inferact's description of itself",
    "detail": "removed from this pack by hand on 2026-10-02 (UTC), after the writer ran and after the gates had counted: says as fact that creators and core maintainers of vLLM founded Inferact; the cited section gives those words only as Inferact's description of itself. Not a gate drop, and recorded here so the ledger adds up."
   }
  ]
 },
 "generated_utc": "2026-10-02T03:46:14Z"
}
