# A short history of Ollama

*Two founders who had sold their first company to Docker in 2015, a one-command model server that began with llama.cpp compiled in and came back to llama.cpp's own server in 2026, and every default we found that moved under its users, dated.*

*Published 2026-10-02 (UTC) · A small (human) team and a fleet of AI agents.*

**the short version:** Most of the language-model readings this workshop has published were taken through Ollama, the program that turns an open-weight model into one command. It pulls the weights, loads them when asked, and by default answers over a local address and lets them go five minutes after the last request. This is the history of the project and the company behind it: a first release on 2023-07-08, llama.cpp underneath, then engines of its own, then llama.cpp's own server again from 2026, and $88M raised by July 2026. It is also a dated record of the defaults that moved under its users, among them a penalty on repeated words, the context window, the number of requests it serves at once, and the engine itself.

6,710 words · about 30 minutes (at 220 words/min) · 3 tables · data kit: no

https://research.strata2signal.com/a-short-history-of-ollama/

---

*If you're new here: [strata→signal](https://strata2signal.com) is a small workshop (plus a friendly dog with a white patch) that builds on its own machines and writes up what it measures. Every date below was read on 2026-09-30 and 2026-10-01 (UTC) from the source credited beside it — among them Ollama's own blog, release notes, documentation and source code at its release tags, GitHub's record of its commits and releases, two press reports where no primary exists, the security researchers' reports, and its maintainers' words in public threads. Where we found no source, the claim is not here.*

*Six words this page leans on.* A **token** is the unit a language model reads and writes: a word or a piece of one. A model's **context window**, the **window** for short, is how many tokens it can hold at once: the conversation so far and the answer it is writing, together. A **slot** is room for one request; Ollama serves as many requests at once as it has slots, and reserves a whole window of memory for each. A **tag** is a name in Ollama's library, like `gemma3:4b`, that pulls one particular build of a model. **safetensors** is a common file format for a model's weights; most of Ollama's library ships in GGUF, llama.cpp's single-file format, instead. The **loopback address** is the one a computer uses to talk to itself: a server that listens only there cannot be reached from another machine.

## The server under most of our numbers {#the-server-under-most-of-our-numbers}

On 2026-08-16 the busiest model on one of this workshop's machines started writing faster than it ever had, and nobody had touched the model, its settings or the card. Ollama had just been upgraded, from 0.32.9 to 0.32.13, and a default had moved with it: the penalty on repeated words, from 1.1 to off for models that set none. [The free speed wasn't free](https://research.strata2signal.com/the-free-speed-wasnt-free/) is that story.

Ollama also runs the doorman at our doors, the Mistral Small 3.2 model (Apache-2.0) that reads what strangers type into [the print lab](https://research.strata2signal.com/how-the-print-lab-works/) and at the rules desk. [A short history of Mistral](https://research.strata2signal.com/a-short-history-of-mistral/) credited Ollama in one line: every reading there was "taken through Ollama (MIT) over llama.cpp (MIT)". Most of this shelf's language-model numbers could say the same. This page is the history behind them, down to [every default we could find that moved](#the-defaults-that-moved).

## The engine underneath: llama.cpp {#the-engine-underneath}

On 2023-03-10 **Georgi Gerganov** created llama.cpp, a C and C++ program for running these models on ordinary computers, processor or graphics card. It is MIT-licensed, and on 2026-09-30 GitHub counts 129,980 stars on it. On 2023-08-21 it merged GGUF, the single-file model format most of Ollama's library ships in ([PR #2398](https://github.com/ggml-org/llama.cpp/pull/2398), opened 2023-07-26).

Ollama's first README, at v0.0.1 on 2023-07-08, put its debt in the first line: "Run large language models with `llama.cpp`", and under features, "Fast inference server written in Go, powered by llama.cpp". Jeffrey Morgan, one of its two founders, said the same on the day he showed it to Hacker News, 2023-07-20: "The llama.cpp project is absolutely amazing. Our goal was to build with/extend the project (vs try to be an alternative)", and "Ollama was originally inspired by the 'server' example".

That server example grew into a standalone program, `llama-server`. Ollama's first releases compiled llama.cpp in; from v0.0.18 (2023-09-06) Ollama ran the server example as a separate process (inside its own process again, as a library, from January to April 2024), replaced it with engines of its own from late 2024, and in 2026 went back to running it.

The README had dropped every mention of llama.cpp two days before the Hacker News post, in the rewrite of 2023-07-18 (commit `e3cc4d5e`, first shipped in v0.0.7). The credit came back on 2024-04-17, in a new section headed "Supported backends" that names the "llama.cpp project founded by Georgi Gerganov" (commit `9755cf91`, first shipped in v0.1.33), and is still there at v0.35.0.

## Who are these people, and where did they come from? {#who-are-these-people}

Y Combinator's directory lists two founders, **Jeffrey Morgan** and **Michael Chiang**, and puts the company in its Winter 2021 batch, in San Francisco. Morgan's account, in the company's funding post of 2026-07-09: "Michael and I first met in college, where we started our first company, Kitematic, which made Docker dead-simple to run. In 2015, it was acquired by Docker. There, our work became Docker Desktop". Docker's announcement of "the acquisition of Kitematic" is dated 2015-03-12.

BetaKit, a Canadian tech news site, adds the college: the two "met at the University of Waterloo". It dates the product from its "launching in 2023", and reports that the company "was initially Toronto-based" and that "Morgan said it relocated to Silicon Valley". Y Combinator's batch and BetaKit's launch year are two different dates, and no source we read gives a founding date, or the founders' own account of the name, so this page offers neither.

Why a model server, from two people who had built a container tool? Morgan, in the same Hacker News thread: after "working on the Docker project for a number of years", "the recent rise in open source language models made us think something similar needed to exist for large language models too". The shape shows: a registry of models on its own site, a `Modelfile` that packages weights with their settings the way a Dockerfile packages an image, `pull`, `push`, `create`, `cp` and `rm`.

The repository went up on GitHub on 2023-06-26 under Morgan's own account (his Hacker News post links it as `jmorganca/ollama`); about eighteen minutes later, a commit titled "`proto` -> `ollama`" renamed the project. On 2026-09-30 GitHub counts 181,963 stars.

## What have they shipped, and what ran underneath? {#what-have-they-shipped}

The releases that mark where the project bent, each dated to GitHub's release record or, where the row says so, to the version's git tag or the announcement. The last column names the program that actually ran the model.

| Date (UTC) | Release | What changed | What ran the model |
|---|---|---|---|
| 2023-07-08 | v0.0.1 | "an early preview"; macOS; one command pulls a model from its library and runs it; "REST API to use with your application" | llama.cpp compiled into Ollama's own program through Go bindings (cgo), in the same process |
| 2023-09-06 | v0.0.18 | The model moved out of Ollama's own process (commit `42998d79`, "subprocess llama.cpp server (#401)", 2023-08-30) | llama.cpp's `server` example as a separate program, one per backend (Metal, CUDA, CPU) by v0.1.0 |
| 2023-09-23 | v0.1.0 | Linux, with NVIDIA acceleration "out-of-the-box" | The same |
| 2024-02-15 (announced) | — | Windows, in preview | llama.cpp's server code back inside Ollama's own process, as a library, from v0.1.18 (2024-01-03) until v0.1.32 (2024-04-10) moved it out again |
| 2024-07-02 | v0.2.0 | "Concurrency": several requests to one model at once, and several models loaded at once, after an experimental preview in v0.1.33 (2024-04-28) | The server example as a separate program again |
| 2024-11-05 (git tag) | v0.4.0 | Llama 3.2 Vision, announced 2024-11-06 | Ollama's own runner, written in Go over llama.cpp's code |
| 2025-03-11 | v0.6.0 | Gemma 3, in four sizes | Ollama's new engine (models written in Go over GGML, the tensor library under llama.cpp), required for Gemma 3 from this release; announced on 2025-05-15 as "Ollama's new engine for multimodal models"; llama.cpp beside it for the rest |
| 2025-07-30 (git tag) | v0.10.0 | "Parallel request processing now defaults to 1"; the new desktop app for macOS and Windows, announced the same day | The same |
| 2025-09-18 | v0.12.0 | Cloud models, in preview (announced 2025-09-19); a web search API followed in v0.12.2 (2025-09-24) | The same |
| 2026-03-27 | v0.19.0 | MLX, Apple's array library, as the engine on Apple silicon, in preview (announced 2026-03-30) | MLX, in preview, for the tags built for it on Apple silicon (the preview's post names one, a Qwen3.5-35B-A3B build); as above for the rest |
| 2026-06-01 (git tag) | v0.30.0 | GGUF models moved onto upstream llama.cpp's `llama-server`; "the vendored GGML and llama.cpp backend, CGO runner, GGML-based Go model implementations" removed (PR #16031, merged 2026-05-29); Vulkan on by default; announced 2026-06-05 | `llama-server`, started as a separate process, with "a temporary patch and compatibility layer" of Ollama's own; MLX for safetensors models, with builds on CUDA for Linux and Windows |
| 2026-09-23 | v0.34.4 | The release we read the code at; it pins llama.cpp build b11081 | The same |
| 2026-09-25 | v0.40.0-rc0 | A pre-release, a test build of a later version, published before 0.35.0: "Models run on MLX on Apple Silicon by default" | MLX by default on Apple silicon, for the architectures it supports; otherwise as at v0.34.4 |
| 2026-09-28 | v0.35.0 | Decision models on a new route, "based on TypeSafe's Jev API" (announced 2026-09-29) | As at v0.34.4; llama.cpp still at b11081 |

GitHub lists 258 releases as of 2026-09-30, 252 of them not marked pre-release, and 106 in the twelve months to that date, 100 of those not marked pre-release. The releases jumped from 0.24.0 (2026-05-14) to 0.30.0 (tagged 2026-06-01) around the engine change. A page like this one is stamped with the version it read, because a default that holds at 0.34 may not at 0.36.

## Who owns it, and who pays for it {#who-owns-it}

Ollama the project is run by Ollama the company. There is no foundation, and the code's licence file reads "MIT License" over "Copyright (c) Ollama" at v0.0.1 and at v0.35.0 alike. Under it sit llama.cpp, MIT ("Copyright (c) 2023-2026 The ggml authors"), and, on Apple silicon and the CUDA builds, Apple's MLX, MIT too, its repository created 2023-11-28.

The models carry their own licences. Most of the tags we read in its library ship the model's licence as a file beside the weights; a few ship none, [`tinyllama`](https://ollama.com/library/tinyllama:latest) and [`orca-mini`](https://ollama.com/library/orca-mini:latest) among them. The model's licence, not Ollama's, governs the weights you run, and this shelf reads it off the tag, as [the licence ledger](https://research.strata2signal.com/licences/) does.

Funding, in the company's words on 2026-07-09: "Ollama has raised $88M from Peter Fenton at Benchmark, Tomasz Tunguz at Theory Ventures, Alex Kolicich at 8VC", with Solomon Hykes, the founder of Docker, among the angels. The round itself, a $65M Series B, is named by its lead investor, Tunguz, in a post the same day: "Today we announced our investment leading Ollama's $65m Series B, alongside Benchmark, YC, & others". The earlier rounds are not itemised in any source we read, so no seed or Series A figure is here.

Two claims in the company's funding post are its own count, not ours: "serving 8.9 million developers", and "used by 85% of the Fortune 500". TNW (The Next Web) put the team at fourteen people on 2026-07-09.

What it sells is the cloud, and that business dates from August 2025. A page headed **Turbo Preview**, "Supercharge models with faster hardware", was up by 2025-08-05 at $20 a month; its questions and answers said "Ollama does not log or retain any queries made via Turbo mode" and "All hardware is located in the United States". On 2025-09-19 the same idea reappeared as **cloud models**, "now in preview", tags ending in `-cloud` that a signed-in local Ollama proxies to ollama.com, again with "does not retain your data".

On 2026-08-31, having "received feedback that GPU-time based billing was difficult to predict", the company announced "transparent per-token pricing": **Pro** at $20 a month with $60 of usage included, **Max** at $100 a month with $300, and **Team** at an introductory $500 a month with $1,000 shared. Read on 2026-09-30, the pricing page shows those three, a free tier with "Starter usage credits included", and a custom-priced Enterprise plan. The local server needs none of it: no account, no key. What it does fetch from ollama.com by default, and the two switches that turn it off, are under [who can talk to it](#who-can-talk-to-it).

## What it speaks, and where it runs {#what-it-speaks-and-where-it-runs}

**Its own API came first**: `/api/generate` and `/api/chat` for answers, `/api/embed` for embeddings, and the library routes that pull, push, create, copy, show, list and delete models, plus `/api/ps` for what is loaded. Each answer carries timings — how long the model took to load, how many prompt tokens were read and how many came from the cache, how many tokens were written and how long that took.

**The dialects arrived in this order**, each dated to its announcement:

- An OpenAI-style chat completions route, 2024-02-08, "making it possible to use more tooling and applications with Ollama locally".
- Dedicated embedding models, 2024-04-08; its own route had served embeddings since v0.0.14 (2023-08-10).
- Tool calling, 2024-07-25, on both its own route and the OpenAI-style one.
- Structured outputs, 2024-12-06, "a specific format defined by a JSON schema", passed in the `format` field.
- A switch for a model's thinking, 2025-05-30: the `think` field.
- An Anthropic-style messages route, 2026-01-16, in v0.14.0 and later.
- A decision route, 2026-09-29, in v0.35.0, that returns "choices, probabilities, and scores instead of text", for "tasks such as ticket triage, model routing, and content classification", in its release notes' words.

At v0.34.4 the OpenAI-style surface spans chat and text completions, embeddings, models, the Responses API and audio transcription, and its documentation lists the fields it leaves out, log probabilities among them.

**Getting a model** is `ollama pull` and a name from its library; `ollama run hf.co/{username}/{repository}` for any GGUF on the Hugging Face Hub, "without creating a new Modelfile" (Hugging Face's page, read 2026-09-30); or a `Modelfile` whose `FROM` names a local GGUF or a safetensors folder. Several models load and unload on demand, up to "3 * the number of GPUs" at once by default (its FAQ).

**Where it runs**, from its documentation at v0.34.4 (unchanged at v0.35.0):

| Platform | What its documentation says |
|---|---|
| Operating systems | macOS, Windows, Linux; an official Docker image |
| Apple silicon | Metal; MLX in preview since v0.19.0 (2026-03-27) and the default in the v0.40.0 pre-release (2026-09-25), by their release notes |
| NVIDIA cards | "compute capability 5.0+ and driver version 550 and newer"; 5.0 through 6.2 (the GTX 750, 900 and 10 series among them) need driver 570 or newer |
| AMD cards | ROCm ("the AMD ROCm v7 driver on Linux"); Vulkan besides |
| Intel and other graphics | Vulkan, "enabled by default when the backend is installed", on Windows and Linux |
| Processor only | Yes; v0.1.0's notes promised "CPU only, and small hobby gaming GPUs to super powerful workstation graphics cards" |
| Two or more cards | "If the model will entirely fit on any single GPU, Ollama will load the model on that GPU"; otherwise it "will be spread across all the available GPUs" |
| A model larger than the card | Loaded anyway, the rest in the computer's memory; `ollama ps` prints the split under PROCESSOR, as "48%/52% CPU/GPU" in its FAQ's example |

That last row is a choice with a cost: a model too big for the card still loads, and runs slower. This shelf watched that happen on an RTX 5080 (16 GB) in [RTX 5080 vs RTX 3080 Ti](https://research.strata2signal.com/rtx-5080-vs-rtx-3080-ti/).

## How it serves a request {#how-it-serves-a-request}

`ollama serve` listens on the loopback address at Ollama's default local port and waits. The first request for a model starts a **runner** for it. For a GGUF model, since 0.30, that runner is upstream llama.cpp's `llama-server` from Ollama's own bundle, started as a separate process on the loopback address with flags that include a context size and a slot count. For a safetensors model the runner is Ollama's MLX engine instead. When the model has been idle for the keep-alive, five minutes by default, the runner is stopped and the memory comes back.

**Every request gets a slot, and a slot is a whole window.** `OLLAMA_NUM_PARALLEL` sets the slots, one by default, and the runner is started with a context of window × slots. In its FAQ's words: "a 2K context with 4 parallel requests will result in an 8K context", and the memory needed "will scale by `OLLAMA_NUM_PARALLEL` * `OLLAMA_CONTEXT_LENGTH`". With one slot a second request waits, with no cap on how many. The queue, 512 requests by default, holds only those waiting for a model to load, and beyond it the server turns new ones away with an error (HTTP 503, "server busy, please try again"). Some model families are held at one slot whatever the setting (eleven architectures at v0.34.4), and every embedding model is.

**The window is fixed when the runner starts.** With `OLLAMA_CONTEXT_LENGTH` unset, the window is chosen from the memory of every card the server can see, added up: 262,144 tokens from 47 GiB, 32,768 tokens from 23 GiB, 4,096 tokens below that. The thresholds sit "slightly lower" than 48 and 24 GiB "to account for small differences in the exact value" (a comment in the code), so one 24 GB card gets 32,768 tokens. If a load at an automatic window fails for memory, it is retried once at a smaller one, and the log says so. Because the window is passed to the runner at start, a request for a GGUF model that asks for a different one means a new runner.

**The prompt cache is `llama-server`'s, at its defaults.** Every generation request Ollama sends its runner carries `cache_prompt: true`, so the runner can skip re-reading the part of a conversation it has already read. What the cache keeps, and which slot a new request lands in, are left to the bundled program (build b11081, documented in `llama-server`'s README): Ollama sets none of the prompt cache's flags.

**A long chat is trimmed, quietly.** On the chat route `truncate` defaults to true for GGUF models (the code turns it off for MLX models). Messages that overflow the window are dropped from the front, keeping the system messages and the newest message, and the log records it only at debug level.

## Who can talk to it {#who-can-talk-to-it}

Its documentation answers in one sentence: the local API "does not require authentication". The code agrees. At v0.34.4 the router puts two checks in front of every route. A CORS filter decides which web pages may call it from a browser: by default, only pages on the computer itself and local apps. While the server is bound to the loopback address, an allowed-hosts check rejects a request whose `Host` header is not local — a guard against a browser trick that makes a remote page look local. Neither asks who you are. The routes that pull, push, copy and delete models are as open as the ones that answer, and the server speaks plain HTTP, with no TLS (the protocol behind https) of its own. The keys Ollama does issue, sent as `Authorization: Bearer`, are for ollama.com.

That leans entirely on the default. Ollama binds the loopback address, so only the computer itself can reach it; its FAQ's answer to "How can I expose Ollama on my network?" is to change `OLLAMA_HOST`, and its recipes for a proxy, ngrok and Cloudflare Tunnel forward every route as-is. Ollama's official Docker image makes that change itself: at v0.34.4 it sets `OLLAMA_HOST` to every interface, and every `docker run` line in its Docker guide publishes the port on the host. Bound to every interface, Ollama is a model server that answers anyone who can reach the port.

Exposure has been measured from outside more than once. The security firm Wiz wrote up CVE-2024-37032 on 2024-06-24: a flaw in model downloads that let a malicious registry overwrite files on the server, and from there run code on it. Wiz called it "extremely severe in Docker installations, as the server runs with root privileges and listens on 0.0.0.0 by default" and counted "over 1,000 exposed instances". Ollama had acknowledged the report and committed a fix on 2024-05-05, the day it arrived ("in about 4 hours", in Wiz's words), and shipped v0.1.34 within days.

On 2026-01-29 SentinelOne's SentinelLABS, which had partnered with Censys "to scan and map internet-reachable Ollama deployments", reported "175,108 unique Ollama hosts across 130 countries" over 293 days of scanning. "Nearly half of observed hosts are configured with tool-calling capabilities that enable them to execute code, access APIs, and interact with external systems", it found, and exposing one "requires only a single configuration change: setting the service to bind to 0.0.0.0 or a public interface".

The maintainers' position is on record, in two comments of February 2025. On the community pull request that would add server-side keys (#9131, still open on 2026-09-30), a member of the team wrote on 2025-02-19: "In the short-term we don't plan to add an API key system to the Ollama server itself. Usually I recommend using a VPN" — he named one — "to manage access to Ollama server instances". A collaborator on the related issue, the day before: "The core team currently recommends using a proxy".

The engine underneath does take keys: `llama-server` has `--api-key`, and it serves TLS from a key file. Ollama starts it on the loopback address and passes neither flag through.

**What leaves the machine.** Its FAQ says "Ollama runs locally. We don't see your prompts or data when you run locally". The code at v0.34.4 shows what does go out unless the cloud features are off. At start the server fetches two things from ollama.com: since v0.22.1 (2026-04-30), a list of recommended models, fetched again about every four hours, sooner after a failed fetch; and since v0.23.2 (2026-05-07), the catalogue of its cloud models, so that it can describe them without a round trip. Each of those requests is signed with the key the installation generated, so calls from one installation carry the same public key. `OLLAMA_NO_CLOUD=1`, or `disable_ollama_cloud` in `~/.ollama/server.json`, stops both, and turns off the cloud models and web search with them. Its FAQ calls this "local only mode", and the log then reads `Ollama cloud disabled: true`.

## The defaults that moved {#the-defaults-that-moved}

Each row is a setting a reader of an older tutorial may have wrong today, dated to the release or commit that moved it.

| Date (UTC) | What moved | Before → after | Where it is written |
|---|---|---|---|
| 2024-04-28 | Serving more than one request at once | One at a time → `OLLAMA_NUM_PARALLEL` and `OLLAMA_MAX_LOADED_MODELS` as opt-in "experimental concurrency features" | v0.1.33 release notes |
| 2024-07-02 | The same, on by default | Opt-in → the slot count "will auto-select either 4 or 1 based on available memory" | v0.2.0 release notes; the FAQ at v0.2.0 |
| 2025-04-29 | The default window, and the automatic slot count | 2,048 → 4,096 tokens; 4 or 1 → 2 or 1 slots | Commits `44b466ee`, "config: update default context length to 4096", and `fe5b9bb2`, "lower default num parallel to 2". The FAQ said "4 or 1" until 2025-07-08 |
| 2025-07-08 | Requests at once | Chosen from free memory → one, "with the intent going forward that parallelism is explicit and will no longer be dynamically determined" | Commit `20c3266e` (#11330); first shipped in v0.10.0 |
| 2025-09-11 | How memory is planned, for models on Ollama's new engine (until 0.30, 2026-06-01) | Estimated → measured "exact" before a model runs; models still on llama.cpp kept the old estimates | PR #12252, "llm: Enable new memory estimates by default", first shipped in v0.11.11 (opt-in since v0.11.5, 2025-08-15); the "New model scheduling" post of 2025-09-23 |
| 2025-12-13 | Flash attention, which saves memory as the window grows | Opt-in (`OLLAMA_FLASH_ATTENTION=1`) → automatic where supported, model by model from 2025-08-27 | v0.13.4 release notes; v0.11.8's for gpt-oss |
| 2026-02-02 | The default window, again | 4,096 (at least 8,192 for gpt-oss and Qwen3-VL models on 20 GiB or more) → 4,096 / 32,768 / 262,144 tokens by the memory of the cards it sees | Commit `0334ffa6` (PR #13946, merged 2026-02-02; first shipped in v0.15.5, 2026-02-03); the context-length page. The FAQ still says "4096" at v0.35.0 |
| 2026-06-01 | The engine for GGUF models | Ollama's own engines → upstream `llama-server` | PR #16031; v0.30.0 |
| 2026-06-01 | Vulkan | Off → "enabled by default", for AMD and Intel cards | The 0.30 post of 2026-06-05; `OLLAMA_VULKAN` |
| 2026-08-12 | The repeat penalty, a sampler setting that pushes a model away from tokens it has just used | 1.1 → 1.0, "off", for models that set none, "matching other engines" | v0.32.10 release notes (a pre-release; the first full release carrying it was v0.32.11, 2026-08-14); our own [The free speed wasn't free](https://research.strata2signal.com/the-free-speed-wasnt-free/) |
| 2026-09-25 | The engine on Apple silicon | llama.cpp → MLX by default, in the v0.40.0 pre-release | v0.40.0-rc0 release notes |

*A footnote on numbers. This page carries no measurement of our own, on purpose: every figure on it comes from the source named beside it. The workshop's own measurements of Ollama beside vLLM, another open-source model server, on one EVGA RTX 3090 XC3 Ultra (24 GB) — one request at a time and several at once, the same model on both — are coming soon, on a separate page.*

## What to take with you {#what-to-take-with-you}

- **It began with llama.cpp compiled in, and since 0.30 (2026-06-01) it runs llama.cpp's own server again.** In between came that server as a separate process from 2023-09-06 and engines of its own from late 2024.
- **Two founders who had sold Kitematic to Docker in 2015 run it as a company, with no foundation.** It had raised $88M by 2026-07-09, most recently a $65M Series B, and has sold a cloud service since August 2025, priced per token since 2026-08-31.
- **Nothing on the local server asks who you are, and by default it calls ollama.com.** It binds the loopback address, except in its official Docker image, and takes no key; in February 2025 the maintainers recommended a VPN or a proxy. At v0.34.4 its calls go out at start and about every four hours, signed with the installation's key, and two documented switches stop them.
- **Our busiest model's 2026-08-16 speed-up was one default moving.** On 2026-08-12 the penalty on repeated words went from 1.1 to off for models that set none. Two other defaults: since v0.10.0 (2025-07-30) Ollama serves one request at a time by default, each slot a whole window of memory, and since v0.15.5 (2026-02-03) the default window is 4,096, 32,768 or 262,144 tokens by the memory of every card it sees, added up.

## How to check our work — and see it live {#how-to-check-our-work-and-see-it-live}

- **Ask the rules desk.** [RuleSage](https://rulesage-live.strata2signal.com/), our board-game rules helper, is free and needs no account. The doorman named in this page's first section, a model served by Ollama on a card in this building, reads questions there too.
- **Read the source at the tag.** Every file named here is linked at v0.34.4 under Sources, so a phone will do; at a computer, `git clone https://github.com/ollama/ollama && git -C ollama checkout v0.34.4`, then read `envconfig/config.go` (the defaults), `server/routes.go` (the window tiers and the two checks in front of every route), `server/sched.go` (the one-slot list and the queue-full error), `llm/llama_server.go` (the runner's flags and `cache_prompt: true`) and `server/model_recommendations.go` (the call to ollama.com about every four hours).
- **Read the guide the project wrote.** `docs/faq.mdx` at the tag answers the slots and the exposure recipes in the project's own words, and `docs/context-length.mdx` the window.
- **See it on your own install.** `ollama ps` prints the split under PROCESSOR when a model did not fit the card, and with `OLLAMA_DEBUG=1` the log says it is truncating "messages which exceed context length" when a chat outgrows its window.

## The rest of the seminar {#the-rest-of-the-seminar}

- [A short history of vLLM](https://research.strata2signal.com/a-short-history-of-vllm/): the other model server in this set.
- [A short history of Mistral](https://research.strata2signal.com/a-short-history-of-mistral/): the model at our doors, and the first of these histories.
- [The free speed wasn't free](https://research.strata2signal.com/the-free-speed-wasnt-free/): the sampler default in the table above, as it landed on this shelf.
- [The compressed photograph](https://research.strata2signal.com/the-compressed-photograph/): what the 4-bit tags in Ollama's library mean.
- [RTX 5080 vs RTX 3080 Ti](https://research.strata2signal.com/rtx-5080-vs-rtx-3080-ti/): Ollama's planner deciding what fits on an RTX 5080 (16 GB).
- [What 150 Watts Buys](https://research.strata2signal.com/what-150-watts-buys/): what a lower power cap cost the two models our products serve, on an RTX PRO 6000 Blackwell (96 GB), one measured through Ollama and the other through vLLM.
- [ONNX Runtime's telemetry on Linux, measured](https://research.strata2signal.com/onnx-runtime-telemetry-on-linux/): an outbound call, and the method we use for one.
- [Reading the answer instead of writing it](https://research.strata2signal.com/reading-the-answer/): decision models, the kind Ollama 0.35.0 added a route for, and the six-way test (which of six dinner guests said a line) this workshop scores models with.

## Who ran this, and thanks {#who-ran-this-and-thanks}

Thanks to **Ollama** (MIT), its team and its contributors, whose release notes, commits and documentation are dated and still up, which is why this page could be written from primary sources rather than from memory; to **Georgi Gerganov** and **the ggml authors** for **llama.cpp** (MIT), the engine at the bottom of most of this shelf's language-model numbers, and for a server README that states its defaults; to **Apple** for **MLX** (MIT). On NVIDIA cards that stack runs over **CUDA** (NVIDIA, proprietary, the one closed piece in it). The founding and funding lines rest on Ollama's own posts, Jeffrey Morgan's words on **Hacker News**, **Docker**'s 2015 announcement, **Y Combinator**'s directory, the lead investor **Tomasz Tunguz**'s own post, and reports by **TNW** and **BetaKit**; the exposure history on **Wiz**'s disclosure and the **SentinelLABS** and **Censys** research, in SentinelLABS's own report; the Hugging Face path on **Hugging Face**'s own documentation; the Turbo page on the **Internet Archive**'s Wayback Machine, which kept it after the live address moved; and the release, commit and pull-request dates on **GitHub**'s record of the repository. None of them owed us anything. A small human team asked for this page, chose what it would and would not claim, and signed off on it; a fleet of AI agents fetched and read every source listed below and drafted it under that team's rulings.

## Sources {#sources}

Every external link below was fetched and read between 2026-09-30 18:42 and 2026-10-01 06:52 (UTC); the version and default lines, and the sibling page's link, were read again on 2026-10-02 (UTC). Source files were read at the tags named; the tag-pinned URLs reproduce each read. A release is dated by GitHub's release record, except where its tag commit is more than a week later or the record predates the merge cited for it; there the tag commit's date is used, and the entry says so.

- GitHub — [ollama/ollama](https://github.com/ollama/ollama) (repository created 2023-06-26T19:39:32Z; MIT; 181,963 stars on 2026-09-30; 258 releases, 252 not marked pre-release) and its [releases](https://github.com/ollama/ollama/releases): [v0.0.1](https://github.com/ollama/ollama/releases/tag/v0.0.1) 2023-07-08 ("an early preview release"), [v0.0.7](https://github.com/ollama/ollama/releases/tag/v0.0.7) 2023-07-19 (the README rewrite), [v0.0.14](https://github.com/ollama/ollama/releases/tag/v0.0.14) 2023-08-10 (the first embeddings route), [v0.0.18](https://github.com/ollama/ollama/releases/tag/v0.0.18) 2023-09-06 (llama.cpp's server example run as a separate process), [v0.1.0](https://github.com/ollama/ollama/releases/tag/v0.1.0) 2023-09-23 (Linux, "GPU acceleration enabled out-of-the-box for Nvidia GPUs", the CPU-only line), [v0.1.18](https://github.com/ollama/ollama/releases/tag/v0.1.18) 2024-01-03 (the first to carry PR #1146), [v0.1.32](https://github.com/ollama/ollama/releases/tag/v0.1.32) 2024-04-10 (the first to carry PR #3218), [v0.1.33](https://github.com/ollama/ollama/releases/tag/v0.1.33) 2024-04-28 ("Experimental concurrency features"), [v0.1.34](https://github.com/ollama/ollama/releases/tag/v0.1.34) 2024-05-07, [v0.2.0](https://github.com/ollama/ollama/releases/tag/v0.2.0) 2024-07-02 ("Concurrency"), [v0.4.0](https://github.com/ollama/ollama/releases/tag/v0.4.0) (Llama 3.2 Vision; its tag commit `9d71bcc3` is dated 2024-11-05; the release record carries an earlier 2024-10-21 stamp, which is why the tag date is used), [v0.6.0](https://github.com/ollama/ollama/releases/tag/v0.6.0) 2025-03-11 (Gemma 3), [v0.10.0](https://github.com/ollama/ollama/releases/tag/v0.10.0) ("Parallel request processing now defaults to 1"; its tag commit `6dcc5dfb` is dated 2025-07-30; the release record carries an earlier 2025-07-18 stamp, which is why the tag date is used), [v0.11.5](https://github.com/ollama/ollama/releases/tag/v0.11.5) 2025-08-15 (the new memory estimates, opt-in), [v0.11.8](https://github.com/ollama/ollama/releases/tag/v0.11.8) 2025-08-27 (flash attention on by default for gpt-oss), [v0.11.11](https://github.com/ollama/ollama/releases/tag/v0.11.11) 2025-09-11 (on by default for the new engine, #12252), [v0.12.0](https://github.com/ollama/ollama/releases/tag/v0.12.0) 2025-09-18, [v0.12.2](https://github.com/ollama/ollama/releases/tag/v0.12.2) 2025-09-24 (the web search API), [v0.13.4](https://github.com/ollama/ollama/releases/tag/v0.13.4) 2025-12-13 ("Enable Flash Attention automatically for models by default"), [v0.14.0](https://github.com/ollama/ollama/releases/tag/v0.14.0) 2026-01-10, [v0.15.5](https://github.com/ollama/ollama/releases/tag/v0.15.5) 2026-02-03 (the three window tiers in its notes), [v0.19.0](https://github.com/ollama/ollama/releases/tag/v0.19.0) 2026-03-27, [v0.22.1](https://github.com/ollama/ollama/releases/tag/v0.22.1) (the first to fetch the recommended models at start, [PR #15868](https://github.com/ollama/ollama/pull/15868), merged 2026-04-29; its tag commit `c7c2837c` is dated 2026-04-30; the release record carries an earlier 2026-04-28 stamp, before that merge, which is why the tag date is used), [v0.23.2](https://github.com/ollama/ollama/releases/tag/v0.23.2) 2026-05-07 (the first to fetch the cloud catalogue at start, [PR #15967](https://github.com/ollama/ollama/pull/15967)), [v0.24.0](https://github.com/ollama/ollama/releases/tag/v0.24.0) 2026-05-14, [v0.30.0](https://github.com/ollama/ollama/releases/tag/v0.30.0) (its tag commit `2c71d8d7` is dated 2026-06-01; the release record carries an earlier 2026-05-13 stamp, which is why the tag date is used), [v0.32.10](https://github.com/ollama/ollama/releases/tag/v0.32.10) 2026-08-12 (pre-release; models that set no repeat penalty "now default to 1.0 (off) instead of 1.1, matching other engines"), [v0.32.11](https://github.com/ollama/ollama/releases/tag/v0.32.11) 2026-08-14 (the first full release carrying it), [v0.34.4](https://github.com/ollama/ollama/releases/tag/v0.34.4) 2026-09-23, [v0.40.0-rc0](https://github.com/ollama/ollama/releases/tag/v0.40.0-rc0) 2026-09-25 (pre-release; "Models run on MLX on Apple Silicon by default"), [v0.35.0](https://github.com/ollama/ollama/releases/tag/v0.35.0) 2026-09-28 (decision models, "based on TypeSafe's Jev API")
- GitHub — the changes that moved the defaults: [`44b466ee`](https://github.com/ollama/ollama/commit/44b466eeb2) 2025-04-29, [`fe5b9bb2`](https://github.com/ollama/ollama/commit/fe5b9bb21b) 2025-04-29 ("lower default num parallel to 2"; at v0.9.6 [`server/sched.go`](https://github.com/ollama/ollama/blob/v0.9.6/server/sched.go) tries 2 slots, then 1, while [`docs/faq.md`](https://github.com/ollama/ollama/blob/v0.9.6/docs/faq.md) still says "auto-select either 4 or 1"), [`20c3266e`](https://github.com/ollama/ollama/commit/20c3266e94) 2025-07-08 (#11330, the "parallelism is explicit" sentence; it also rewrote the FAQ's "4 or 1" line), [PR #12252](https://github.com/ollama/ollama/pull/12252), "llm: Enable new memory estimates by default", merged 2025-09-11 ("Models running on the llama engine will continue to use the original style of memory estimation"), and [`0334ffa6`](https://github.com/ollama/ollama/commit/0334ffa625) 2026-02-02 (merged as [PR #13946](https://github.com/ollama/ollama/pull/13946), opened 2026-01-28, the commit's author date; first shipped in v0.15.5, 2026-02-03, whose notes list the three tiers)
- GitHub — the commits and pull requests that moved the engine and the README's credit, and the rename: [`42998d79`](https://github.com/ollama/ollama/commit/42998d797d) 2023-08-30 ("subprocess llama.cpp server (#401)", first shipped in v0.0.18); [PR #1146](https://github.com/ollama/ollama/pull/1146), "Add cgo implementation for llama.cpp", merged 2023-12-22 (to "link directly via cgo instead of running a subprocess"; first shipped in v0.1.18, 2024-01-03); [PR #3218](https://github.com/ollama/ollama/pull/3218), "Switch back to subprocessing for llama.cpp", merged 2024-04-02 (first shipped in v0.1.32, 2024-04-10); [`e3cc4d5e`](https://github.com/ollama/ollama/commit/e3cc4d5eac) 2023-07-18 ("update `README.md` with new syntax", after which the README names llama.cpp nowhere until v0.1.33); [`9755cf91`](https://github.com/ollama/ollama/commit/9755cf9173) 2024-04-17 ("acknowledge the amazing work done by Georgi and team!"); and the rename commit ["`proto` -> `ollama`"](https://github.com/ollama/ollama/commit/df5fdd6647), 2023-06-26, about eighteen minutes after the repository went up
- GitHub — [PR #16031](https://github.com/ollama/ollama/pull/16031), "runner: Remove CGO engines, use llama-server exclusively for GGML models", merged 2026-05-29 (the sentences quoted above; GitHub's compare shows v0.30.0 eight commits ahead of its merge)
- Ollama source at tag v0.34.4, unchanged at v0.35.0 where noted: [`LICENSE`](https://github.com/ollama/ollama/blob/v0.34.4/LICENSE) (also read at v0.0.1; unchanged at v0.35.0); [`LLAMA_CPP_VERSION`](https://github.com/ollama/ollama/blob/v0.34.4/LLAMA_CPP_VERSION) (b11081 at both tags); [`envconfig/config.go`](https://github.com/ollama/ollama/blob/v0.34.4/envconfig/config.go) (the loopback default, `OLLAMA_NUM_PARALLEL` 1, `OLLAMA_MAX_QUEUE` 512, keep-alive 5m, `OLLAMA_CONTEXT_LENGTH` 0, `OLLAMA_VULKAN`, `OLLAMA_NO_CLOUD`, the default browser origins); [`Dockerfile`](https://github.com/ollama/ollama/blob/v0.34.4/Dockerfile) (the image sets `OLLAMA_HOST` to every interface; unchanged at v0.35.0); [`server/routes.go`](https://github.com/ollama/ollama/blob/v0.34.4/server/routes.go) (the two checks in front of every route, the route table, the summed memory tiers with their 47/23 GiB comment, `truncate` and its MLX exception, `truncateNativeChatMessages` and its debug line); [`server/sched.go`](https://github.com/ollama/ollama/blob/v0.34.4/server/sched.go) (the one-slot list, embedding models at one slot, the retry at a smaller window, the queue-full error and the loads it applies to, the reload on a new window and its MLX exception); [`llm/llama_server.go`](https://github.com/ollama/ollama/blob/v0.34.4/llm/llama_server.go) (the runner's flags, `cache_prompt: true` on generation requests, the uncapped wait for a slot); [`server/prompt.go`](https://github.com/ollama/ollama/blob/v0.34.4/server/prompt.go) (the trimming rule and its debug line); [`server/model_recommendations.go`](https://github.com/ollama/ollama/blob/v0.34.4/server/model_recommendations.go), [`server/model_show_cache.go`](https://github.com/ollama/ollama/blob/v0.34.4/server/model_show_cache.go) and [`server/cloud_proxy.go`](https://github.com/ollama/ollama/blob/v0.34.4/server/cloud_proxy.go) (the two outbound fetches, the four-hour interval, the signed header); [`server/images.go`](https://github.com/ollama/ollama/blob/v0.34.4/server/images.go) (safetensors → MLX); [`llama/compat/`](https://github.com/ollama/ollama/tree/v0.34.4/llama/compat) (the patch and the compatibility layer, neither of which sets a prompt-cache flag); the v0.34.4 release assets (an MLX bundle for Linux and for Windows); earlier trees: [v0.0.1 `llama/llama.go`](https://github.com/ollama/ollama/blob/v0.0.1/llama/llama.go) (llama.cpp linked in through cgo bindings, "Copyright (c) 2023 go-skynet authors"), [v0.0.14 `server/routes.go`](https://github.com/ollama/ollama/blob/v0.0.14/server/routes.go) (the `/api/embeddings` route), [v0.0.18 `llm/ggml_llama.go`](https://github.com/ollama/ollama/blob/v0.0.18/llm/ggml_llama.go) (the server example started as a separate process), [v0.1.0 `llm/llama.go`](https://github.com/ollama/ollama/blob/v0.1.0/llm/llama.go) (the per-backend `server` binaries), [v0.4.0 `llama/runner/runner.go`](https://github.com/ollama/ollama/blob/v0.4.0/llama/runner/runner.go) (the Go runner), [v0.6.0 `fs/ggml/ggml.go`](https://github.com/ollama/ollama/blob/v0.6.0/fs/ggml/ggml.go) and [`llm/server.go`](https://github.com/ollama/ollama/blob/v0.6.0/llm/server.go) (the new engine required for Gemma 3), the READMEs at [v0.0.1](https://github.com/ollama/ollama/blob/v0.0.1/README.md), [v0.0.7](https://github.com/ollama/ollama/blob/v0.0.7/README.md) (the `Modelfile`; no mention of llama.cpp), [v0.1.33](https://github.com/ollama/ollama/blob/v0.1.33/README.md) ("Supported backends") and [v0.35.0](https://github.com/ollama/ollama/blob/v0.35.0/README.md)
- Ollama documentation at tag v0.34.4, unchanged at v0.35.0: [`docs/faq.mdx`](https://github.com/ollama/ollama/blob/v0.34.4/docs/faq.mdx) (the 4,096 line, the concurrency section and its 2K × 4 = 8K sentence, the queue, keep-alive, the exposure and proxy recipes, "local only mode", the multi-card rule, flash attention, "We don't see your prompts"); [`docs/context-length.mdx`](https://github.com/ollama/ollama/blob/v0.34.4/docs/context-length.mdx) (the tiers by card memory); [`docs/gpu.mdx`](https://github.com/ollama/ollama/blob/v0.34.4/docs/gpu.mdx) (compute capability and driver lines, ROCm v7, Vulkan, Metal); [`docs/docker.mdx`](https://github.com/ollama/ollama/blob/v0.34.4/docs/docker.mdx) (each `docker run` line publishes the port on the host); [`docs/api/authentication.mdx`](https://github.com/ollama/ollama/blob/v0.34.4/docs/api/authentication.mdx) and its live copy at [docs.ollama.com](https://docs.ollama.com/api/authentication) ("does not require authentication"); [`docs/api/openai-compatibility.mdx`](https://github.com/ollama/ollama/blob/v0.34.4/docs/api/openai-compatibility.mdx) (the unticked fields); [`docs/api.md`](https://github.com/ollama/ollama/blob/v0.34.4/docs/api.md) (`format`, `think`, and the timing fields of an answer; at v0.35.0 it adds only a `typical_p` deprecation line); [`docs/windows.mdx`](https://github.com/ollama/ollama/blob/v0.34.4/docs/windows.mdx) (the MLX-on-CUDA bundle); and, at [v0.2.0, `docs/faq.md`](https://github.com/ollama/ollama/blob/v0.2.0/docs/faq.md) ("auto-select either 4 or 1 based on available memory")
- Ollama's library — the `latest` tag pages of [`tinyllama`](https://ollama.com/library/tinyllama:latest) and [`orca-mini`](https://ollama.com/library/orca-mini:latest), which list no licence among their files, and of [`mistral`](https://ollama.com/library/mistral:latest), which does; the same read in the public manifests at `registry.ollama.ai/v2/library/<model>/manifests/latest`; and, read on 2026-10-01, those manifests for the first 40 models in the library's Popular order (its default): 37 carry a licence layer, and 3 do not (`glm-ocr`, `tinyllama` and `qwen3-embedding`)
- Ollama's blog — [OpenAI compatibility](https://ollama.com/blog/openai-compatibility) 2024-02-08; [Windows preview](https://ollama.com/blog/windows-preview) 2024-02-15; [Embedding models](https://ollama.com/blog/embedding-models) 2024-04-08; [Tool support](https://ollama.com/blog/tool-support) 2024-07-25; [Llama 3.2 Vision](https://ollama.com/blog/llama3.2-vision) 2024-11-06 ("Download Ollama 0.4"); [Structured outputs](https://ollama.com/blog/structured-outputs) 2024-12-06; [Ollama's new engine for multimodal models](https://ollama.com/blog/multimodal-models) 2025-05-15 ("has so far relied on the ggml-org/llama.cpp project"); [Thinking](https://ollama.com/blog/thinking) 2025-05-30; [Ollama's new app](https://ollama.com/blog/new-app) 2025-07-30; [Cloud models](https://ollama.com/blog/cloud-models) 2025-09-19; [New model scheduling](https://ollama.com/blog/new-model-scheduling) 2025-09-23; [Web search](https://ollama.com/blog/web-search) 2025-09-24; [Claude Code with Anthropic API compatibility](https://ollama.com/blog/claude) 2026-01-16 ("Ollama v0.14.0 and later"); [Ollama is now powered by MLX on Apple Silicon in preview](https://ollama.com/blog/mlx) 2026-03-30 ("Download Ollama 0.19"); [Improved performance and model support with GGUF](https://ollama.com/blog/improved-performance-and-model-support-with-gguf) 2026-06-05 (Ollama 0.30; "Vulkan is now enabled by default"); [Ollama: all aboard open models](https://ollama.com/blog/all-aboard-open-models) 2026-07-09 (Kitematic, Docker, the $88M sentence, the investors, the two counts); [Ollama's transparent pricing](https://ollama.com/blog/transparent-pricing) 2026-08-31 (the three plans, with the Team plan's "introductory pricing"; "GPU-time based billing was difficult to predict"); [Ollama now supports Jev-style decision models](https://ollama.com/blog/ollama-now-supports-jev-style-decision-models) 2026-09-29; and the [pricing page](https://ollama.com/pricing) as read 2026-09-30
- Ollama's Turbo page as archived by the Internet Archive on 2025-08-05, [its earliest capture](https://web.archive.org/web/20250805172304/https://ollama.com/turbo): "Turbo Preview", "$20/mo", the retention and hardware-location answers. The live address now redirects to the pricing page
- Hacker News — [Show HN: Ollama – Run LLMs on your Mac](https://news.ycombinator.com/item?id=36802582), 2023-07-20 (the founder's post and comment quoted above)
- GitHub — [ollama/ollama#9131](https://github.com/ollama/ollama/pull/9131) (a community pull request, still open on 2026-09-30; [the maintainer's comment of 2025-02-19](https://github.com/ollama/ollama/pull/9131#issuecomment-2669623160)) and [#8536](https://github.com/ollama/ollama/issues/8536) ([the collaborator's comment of 2025-02-18](https://github.com/ollama/ollama/issues/8536#issuecomment-2667153335))
- Y Combinator — [Ollama](https://www.ycombinator.com/companies/ollama) (founders, Winter 2021, San Francisco)
- Docker — [Kitematic a Docker GUI joins the Docker family](https://www.docker.com/blog/kitematic-a-docker-gui-joins-the-docker-family/), 2015-03-12
- Theory Ventures — Tomasz Tunguz, [8.9 Million AI Users](https://tomtunguz.com/ollama-series-b/), 2026-07-09 ("our investment leading Ollama's $65m Series B"; "Theory led the Series B")
- TNW (The Next Web) — [Ollama raises $65M as its open-model runner hits nearly 9M developers](https://thenextweb.com/news/ollama-65m-series-b-theory-ventures-open-models), 2026-07-09 (the team of 14)
- BetaKit — [Kitematic founders nab $65 million USD to help devs run AI models with Ollama](https://betakit.com/kitematic-founders-nab-65-million-usd-to-help-devs-run-ai-models-locally-with-ollama/), 2026-07-20 (the University of Waterloo; "launching in 2023"; "initially Toronto-based"; "Morgan said it relocated to Silicon Valley")
- Wiz — [Probllama: Ollama Remote Code Execution Vulnerability (CVE-2024-37032) – Overview and Mitigations](https://www.wiz.io/blog/probllama-ollama-vulnerability-cve-2024-37032), 2024-06-24 (the timeline from 2024-05-05 and the fix "in about 4 hours", the upgrade "to version 0.1.34 or newer", the Docker sentence, the exposed-instance count)
- SentinelLABS (SentinelOne), with Censys — [Silent Brothers | Ollama Hosts Form Anonymous AI Network Beyond Platform Guardrails](https://www.sentinelone.com/labs/silent-brothers-ollama-hosts-form-anonymous-ai-network-beyond-platform-guardrails/), 2026-01-29, updated 2026-03-06 (293 days of scanning; 175,108 hosts in 130 countries; nearly half with tool calling; "a single configuration change")
- GitHub — [ggml-org/llama.cpp](https://github.com/ggml-org/llama.cpp) (created 2023-03-10; MIT, "Copyright (c) 2023-2026 The ggml authors" in its [`LICENSE` at b11081](https://github.com/ggml-org/llama.cpp/blob/b11081/LICENSE); "LLM inference in C/C++"; 129,980 stars on 2026-09-30), [PR #2398 GGUF](https://github.com/ggml-org/llama.cpp/pull/2398) (merged 2023-08-21), and [`tools/server/README.md` at b11081](https://github.com/ggml-org/llama.cpp/blob/b11081/tools/server/README.md) (the prompt cache's defaults, `--api-key`, the TLS flags and `--host`)
- GitHub — [ml-explore/mlx](https://github.com/ml-explore/mlx) (created 2023-11-28; MIT)
- Hugging Face — [Use Ollama with any GGUF Model on Hugging Face Hub](https://huggingface.co/docs/hub/en/ollama) (the `hf.co/{username}/{repository}` form; "without creating a new Modelfile")
- This workshop's own shelf, read on 2026-10-01 in the markdown copy that each of its 65 released exhibits publishes beside itself (`index.md`): 29 print a figure in tokens a second, and 22 of those name Ollama (the count behind "most" in the short version and the opening section)
- This workshop's own released pages, linked above: [A short history of Mistral](https://research.strata2signal.com/a-short-history-of-mistral/) ("taken through Ollama (MIT) over llama.cpp (MIT)"; "the rules desk uses the same seat"), [The free speed wasn't free](https://research.strata2signal.com/the-free-speed-wasnt-free/) (the upgrade from 0.32.9 to 0.32.13 and the restart of 2026-08-16), [RTX 5080 vs RTX 3080 Ti](https://research.strata2signal.com/rtx-5080-vs-rtx-3080-ti/), [the licence ledger](https://research.strata2signal.com/licences/), [how the print lab works](https://research.strata2signal.com/how-the-print-lab-works/) (its doorman, "a small language model, Mistral Small 3.2"; "Mistral Small 3.2 (Mistral AI, Apache-2.0)"; "both served by Ollama (MIT)"), [The compressed photograph](https://research.strata2signal.com/the-compressed-photograph/), [What 150 Watts Buys](https://research.strata2signal.com/what-150-watts-buys/) ("the two models our products actually serve": one measured through Ollama, the other through vLLM as "the candidate fast lane"; its hardware line names the RTX PRO 6000 Blackwell (96 GB)), [ONNX Runtime's telemetry on Linux, measured](https://research.strata2signal.com/onnx-runtime-telemetry-on-linux/), [Reading the answer instead of writing it](https://research.strata2signal.com/reading-the-answer/), [A short history of vLLM](https://research.strata2signal.com/a-short-history-of-vllm/)

Corrections and later measurements will be added below, each dated (UTC), with a window at both ends where one applies and saying in plain words what it counts.

<!-- derived 2026-10-02 (UTC) by tools/derive_md.py from the pour source.
     source html sha256: c2542c8b490b0f5b811a5bf8efd66ccfe9a67a1c5eb15c01b0ac26b10a1e4596
     derivation sha256:  b14a3b7e12acc549c999ca15740de31811896a2be3b5f0493a007b71642e6788
     the {#id} on each heading is the anchor that heading carries on the page. -->
