The notes — a model server, who wrote it, who owns it, and every default that moved, dated

A short history of Ollama

exhibit sixty-eight The notes
Published 2026-10-02 (UTC)
A small (human) team and a fleet of AI agents.

Two founders who had sold their first company to Docker in 2015, a one-command model server that began with llama.cpp compiled in and came back to llama.cpp's own server in 2026, and every default we found that moved under its users, dated.

One of four pages on two model servers: A short history of Ollama (this page) · A short history of vLLM · The bench: both servers on one RTX 3090 (coming soon) · The distilled findings and recs (coming soon)

ask about this page → assistant.strata2signal.com · in beta, still being tested

hardware on this page: RTX PRO 6000 Blackwell (96 GB) · RTX 3090 (24 GB) · RTX 5080 (16 GB) → research.strata2signal.com/hardware/ · the roster is in beta, still being tested

the short version

Most of the language-model readings this workshop has published were taken through Ollama, the program that turns an open-weight model into one command. It pulls the weights, loads them when asked, and by default answers over a local address and lets them go five minutes after the last request. This is the history of the project and the company behind it: a first release on 2023-07-08, llama.cpp underneath, then engines of its own, then llama.cpp's own server again from 2026, and $88M raised by July 2026. It is also a dated record of the defaults that moved under its users, among them a penalty on repeated words, the context window, the number of requests it serves at once, and the engine itself.

6,710 words, about 30 minutes to read.

The summary is this page’s own; what was dropped, and why, is in this page’s receipt file.

If you're new here: strata→signal is a small workshop (plus a friendly dog with a white patch) that builds on its own machines and writes up what it measures. Every date below was read on 2026-09-30 and 2026-10-01 (UTC) from the source credited beside it — among them Ollama's own blog, release notes, documentation and source code at its release tags, GitHub's record of its commits and releases, two press reports where no primary exists, the security researchers' reports, and its maintainers' words in public threads. Where we found no source, the claim is not here.

Six words this page leans on. A token is the unit a language model reads and writes: a word or a piece of one. A model's context window, the window for short, is how many tokens it can hold at once: the conversation so far and the answer it is writing, together. A slot is room for one request; Ollama serves as many requests at once as it has slots, and reserves a whole window of memory for each. A tag is a name in Ollama's library, like gemma3:4b, that pulls one particular build of a model. safetensors is a common file format for a model's weights; most of Ollama's library ships in GGUF, llama.cpp's single-file format, instead. The loopback address is the one a computer uses to talk to itself: a server that listens only there cannot be reached from another machine.

The server under most of our numbers

On 2026-08-16 the busiest model on one of this workshop's machines started writing faster than it ever had, and nobody had touched the model, its settings or the card. Ollama had just been upgraded, from 0.32.9 to 0.32.13, and a default had moved with it: the penalty on repeated words, from 1.1 to off for models that set none. The free speed wasn't free is that story.

Ollama also runs the doorman at our doors, the Mistral Small 3.2 model (Apache-2.0) that reads what strangers type into the print lab and at the rules desk. A short history of Mistral credited Ollama in one line: every reading there was "taken through Ollama (MIT) over llama.cpp (MIT)". Most of this shelf's language-model numbers could say the same. This page is the history behind them, down to every default we could find that moved.

The engine underneath: llama.cpp

On 2023-03-10 Georgi Gerganov created llama.cpp, a C and C++ program for running these models on ordinary computers, processor or graphics card. It is MIT-licensed, and on 2026-09-30 GitHub counts 129,980 stars on it. On 2023-08-21 it merged GGUF, the single-file model format most of Ollama's library ships in (PR #2398, opened 2023-07-26).

Ollama's first README, at v0.0.1 on 2023-07-08, put its debt in the first line: "Run large language models with llama.cpp", and under features, "Fast inference server written in Go, powered by llama.cpp". Jeffrey Morgan, one of its two founders, said the same on the day he showed it to Hacker News, 2023-07-20: "The llama.cpp project is absolutely amazing. Our goal was to build with/extend the project (vs try to be an alternative)", and "Ollama was originally inspired by the 'server' example".

That server example grew into a standalone program, llama-server. Ollama's first releases compiled llama.cpp in; from v0.0.18 (2023-09-06) Ollama ran the server example as a separate process (inside its own process again, as a library, from January to April 2024), replaced it with engines of its own from late 2024, and in 2026 went back to running it.

The README had dropped every mention of llama.cpp two days before the Hacker News post, in the rewrite of 2023-07-18 (commit e3cc4d5e, first shipped in v0.0.7). The credit came back on 2024-04-17, in a new section headed "Supported backends" that names the "llama.cpp project founded by Georgi Gerganov" (commit 9755cf91, first shipped in v0.1.33), and is still there at v0.35.0.

Who are these people, and where did they come from?

Y Combinator's directory lists two founders, Jeffrey Morgan and Michael Chiang, and puts the company in its Winter 2021 batch, in San Francisco. Morgan's account, in the company's funding post of 2026-07-09: "Michael and I first met in college, where we started our first company, Kitematic, which made Docker dead-simple to run. In 2015, it was acquired by Docker. There, our work became Docker Desktop". Docker's announcement of "the acquisition of Kitematic" is dated 2015-03-12.

BetaKit, a Canadian tech news site, adds the college: the two "met at the University of Waterloo". It dates the product from its "launching in 2023", and reports that the company "was initially Toronto-based" and that "Morgan said it relocated to Silicon Valley". Y Combinator's batch and BetaKit's launch year are two different dates, and no source we read gives a founding date, or the founders' own account of the name, so this page offers neither.

Why a model server, from two people who had built a container tool? Morgan, in the same Hacker News thread: after "working on the Docker project for a number of years", "the recent rise in open source language models made us think something similar needed to exist for large language models too". The shape shows: a registry of models on its own site, a Modelfile that packages weights with their settings the way a Dockerfile packages an image, pull, push, create, cp and rm.

The repository went up on GitHub on 2023-06-26 under Morgan's own account (his Hacker News post links it as jmorganca/ollama); about eighteen minutes later, a commit titled "proto -> ollama" renamed the project. On 2026-09-30 GitHub counts 181,963 stars.

What have they shipped, and what ran underneath?

The releases that mark where the project bent, each dated to GitHub's release record or, where the row says so, to the version's git tag or the announcement. The last column names the program that actually ran the model.

Date (UTC)ReleaseWhat changedWhat ran the model
2023-07-08v0.0.1"an early preview"; macOS; one command pulls a model from its library and runs it; "REST API to use with your application"llama.cpp compiled into Ollama's own program through Go bindings (cgo), in the same process
2023-09-06v0.0.18The model moved out of Ollama's own process (commit 42998d79, "subprocess llama.cpp server (#401)", 2023-08-30)llama.cpp's server example as a separate program, one per backend (Metal, CUDA, CPU) by v0.1.0
2023-09-23v0.1.0Linux, with NVIDIA acceleration "out-of-the-box"The same
2024-02-15 (announced)—Windows, in previewllama.cpp's server code back inside Ollama's own process, as a library, from v0.1.18 (2024-01-03) until v0.1.32 (2024-04-10) moved it out again
2024-07-02v0.2.0"Concurrency": several requests to one model at once, and several models loaded at once, after an experimental preview in v0.1.33 (2024-04-28)The server example as a separate program again
2024-11-05 (git tag)v0.4.0Llama 3.2 Vision, announced 2024-11-06Ollama's own runner, written in Go over llama.cpp's code
2025-03-11v0.6.0Gemma 3, in four sizesOllama's new engine (models written in Go over GGML, the tensor library under llama.cpp), required for Gemma 3 from this release; announced on 2025-05-15 as "Ollama's new engine for multimodal models"; llama.cpp beside it for the rest
2025-07-30 (git tag)v0.10.0"Parallel request processing now defaults to 1"; the new desktop app for macOS and Windows, announced the same dayThe same
2025-09-18v0.12.0Cloud models, in preview (announced 2025-09-19); a web search API followed in v0.12.2 (2025-09-24)The same
2026-03-27v0.19.0MLX, Apple's array library, as the engine on Apple silicon, in preview (announced 2026-03-30)MLX, in preview, for the tags built for it on Apple silicon (the preview's post names one, a Qwen3.5-35B-A3B build); as above for the rest
2026-06-01 (git tag)v0.30.0GGUF models moved onto upstream llama.cpp's llama-server; "the vendored GGML and llama.cpp backend, CGO runner, GGML-based Go model implementations" removed (PR #16031, merged 2026-05-29); Vulkan on by default; announced 2026-06-05llama-server, started as a separate process, with "a temporary patch and compatibility layer" of Ollama's own; MLX for safetensors models, with builds on CUDA for Linux and Windows
2026-09-23v0.34.4The release we read the code at; it pins llama.cpp build b11081The same
2026-09-25v0.40.0-rc0A pre-release, a test build of a later version, published before 0.35.0: "Models run on MLX on Apple Silicon by default"MLX by default on Apple silicon, for the architectures it supports; otherwise as at v0.34.4
2026-09-28v0.35.0Decision models on a new route, "based on TypeSafe's Jev API" (announced 2026-09-29)As at v0.34.4; llama.cpp still at b11081

GitHub lists 258 releases as of 2026-09-30, 252 of them not marked pre-release, and 106 in the twelve months to that date, 100 of those not marked pre-release. The releases jumped from 0.24.0 (2026-05-14) to 0.30.0 (tagged 2026-06-01) around the engine change. A page like this one is stamped with the version it read, because a default that holds at 0.34 may not at 0.36.

Who owns it, and who pays for it

Ollama the project is run by Ollama the company. There is no foundation, and the code's licence file reads "MIT License" over "Copyright (c) Ollama" at v0.0.1 and at v0.35.0 alike. Under it sit llama.cpp, MIT ("Copyright (c) 2023-2026 The ggml authors"), and, on Apple silicon and the CUDA builds, Apple's MLX, MIT too, its repository created 2023-11-28.

The models carry their own licences. Most of the tags we read in its library ship the model's licence as a file beside the weights; a few ship none, tinyllama and orca-mini among them. The model's licence, not Ollama's, governs the weights you run, and this shelf reads it off the tag, as the licence ledger does.

Funding, in the company's words on 2026-07-09: "Ollama has raised $88M from Peter Fenton at Benchmark, Tomasz Tunguz at Theory Ventures, Alex Kolicich at 8VC", with Solomon Hykes, the founder of Docker, among the angels. The round itself, a $65M Series B, is named by its lead investor, Tunguz, in a post the same day: "Today we announced our investment leading Ollama's $65m Series B, alongside Benchmark, YC, & others". The earlier rounds are not itemised in any source we read, so no seed or Series A figure is here.

Two claims in the company's funding post are its own count, not ours: "serving 8.9 million developers", and "used by 85% of the Fortune 500". TNW (The Next Web) put the team at fourteen people on 2026-07-09.

What it sells is the cloud, and that business dates from August 2025. A page headed Turbo Preview, "Supercharge models with faster hardware", was up by 2025-08-05 at $20 a month; its questions and answers said "Ollama does not log or retain any queries made via Turbo mode" and "All hardware is located in the United States". On 2025-09-19 the same idea reappeared as cloud models, "now in preview", tags ending in -cloud that a signed-in local Ollama proxies to ollama.com, again with "does not retain your data".

On 2026-08-31, having "received feedback that GPU-time based billing was difficult to predict", the company announced "transparent per-token pricing": Pro at $20 a month with $60 of usage included, Max at $100 a month with $300, and Team at an introductory $500 a month with $1,000 shared. Read on 2026-09-30, the pricing page shows those three, a free tier with "Starter usage credits included", and a custom-priced Enterprise plan. The local server needs none of it: no account, no key. What it does fetch from ollama.com by default, and the two switches that turn it off, are under who can talk to it.

What it speaks, and where it runs

Its own API came first: /api/generate and /api/chat for answers, /api/embed for embeddings, and the library routes that pull, push, create, copy, show, list and delete models, plus /api/ps for what is loaded. Each answer carries timings — how long the model took to load, how many prompt tokens were read and how many came from the cache, how many tokens were written and how long that took.

The dialects arrived in this order, each dated to its announcement:

  • An OpenAI-style chat completions route, 2024-02-08, "making it possible to use more tooling and applications with Ollama locally".
  • Dedicated embedding models, 2024-04-08; its own route had served embeddings since v0.0.14 (2023-08-10).
  • Tool calling, 2024-07-25, on both its own route and the OpenAI-style one.
  • Structured outputs, 2024-12-06, "a specific format defined by a JSON schema", passed in the format field.
  • A switch for a model's thinking, 2025-05-30: the think field.
  • An Anthropic-style messages route, 2026-01-16, in v0.14.0 and later.
  • A decision route, 2026-09-29, in v0.35.0, that returns "choices, probabilities, and scores instead of text", for "tasks such as ticket triage, model routing, and content classification", in its release notes' words.

At v0.34.4 the OpenAI-style surface spans chat and text completions, embeddings, models, the Responses API and audio transcription, and its documentation lists the fields it leaves out, log probabilities among them.

Getting a model is ollama pull and a name from its library; ollama run hf.co/{username}/{repository} for any GGUF on the Hugging Face Hub, "without creating a new Modelfile" (Hugging Face's page, read 2026-09-30); or a Modelfile whose FROM names a local GGUF or a safetensors folder. Several models load and unload on demand, up to "3 * the number of GPUs" at once by default (its FAQ).

Where it runs, from its documentation at v0.34.4 (unchanged at v0.35.0):

PlatformWhat its documentation says
Operating systemsmacOS, Windows, Linux; an official Docker image
Apple siliconMetal; MLX in preview since v0.19.0 (2026-03-27) and the default in the v0.40.0 pre-release (2026-09-25), by their release notes
NVIDIA cards"compute capability 5.0+ and driver version 550 and newer"; 5.0 through 6.2 (the GTX 750, 900 and 10 series among them) need driver 570 or newer
AMD cardsROCm ("the AMD ROCm v7 driver on Linux"); Vulkan besides
Intel and other graphicsVulkan, "enabled by default when the backend is installed", on Windows and Linux
Processor onlyYes; v0.1.0's notes promised "CPU only, and small hobby gaming GPUs to super powerful workstation graphics cards"
Two or more cards"If the model will entirely fit on any single GPU, Ollama will load the model on that GPU"; otherwise it "will be spread across all the available GPUs"
A model larger than the cardLoaded anyway, the rest in the computer's memory; ollama ps prints the split under PROCESSOR, as "48%/52% CPU/GPU" in its FAQ's example

That last row is a choice with a cost: a model too big for the card still loads, and runs slower. This shelf watched that happen on an RTX 5080 (16 GB) in RTX 5080 vs RTX 3080 Ti.

How it serves a request

ollama serve listens on the loopback address at Ollama's default local port and waits. The first request for a model starts a runner for it. For a GGUF model, since 0.30, that runner is upstream llama.cpp's llama-server from Ollama's own bundle, started as a separate process on the loopback address with flags that include a context size and a slot count. For a safetensors model the runner is Ollama's MLX engine instead. When the model has been idle for the keep-alive, five minutes by default, the runner is stopped and the memory comes back.

Every request gets a slot, and a slot is a whole window. OLLAMA_NUM_PARALLEL sets the slots, one by default, and the runner is started with a context of window × slots. In its FAQ's words: "a 2K context with 4 parallel requests will result in an 8K context", and the memory needed "will scale by OLLAMA_NUM_PARALLEL * OLLAMA_CONTEXT_LENGTH". With one slot a second request waits, with no cap on how many. The queue, 512 requests by default, holds only those waiting for a model to load, and beyond it the server turns new ones away with an error (HTTP 503, "server busy, please try again"). Some model families are held at one slot whatever the setting (eleven architectures at v0.34.4), and every embedding model is.

The window is fixed when the runner starts. With OLLAMA_CONTEXT_LENGTH unset, the window is chosen from the memory of every card the server can see, added up: 262,144 tokens from 47 GiB, 32,768 tokens from 23 GiB, 4,096 tokens below that. The thresholds sit "slightly lower" than 48 and 24 GiB "to account for small differences in the exact value" (a comment in the code), so one 24 GB card gets 32,768 tokens. If a load at an automatic window fails for memory, it is retried once at a smaller one, and the log says so. Because the window is passed to the runner at start, a request for a GGUF model that asks for a different one means a new runner.

The prompt cache is llama-server's, at its defaults. Every generation request Ollama sends its runner carries cache_prompt: true, so the runner can skip re-reading the part of a conversation it has already read. What the cache keeps, and which slot a new request lands in, are left to the bundled program (build b11081, documented in llama-server's README): Ollama sets none of the prompt cache's flags.

A long chat is trimmed, quietly. On the chat route truncate defaults to true for GGUF models (the code turns it off for MLX models). Messages that overflow the window are dropped from the front, keeping the system messages and the newest message, and the log records it only at debug level.

Who can talk to it

Its documentation answers in one sentence: the local API "does not require authentication". The code agrees. At v0.34.4 the router puts two checks in front of every route. A CORS filter decides which web pages may call it from a browser: by default, only pages on the computer itself and local apps. While the server is bound to the loopback address, an allowed-hosts check rejects a request whose Host header is not local — a guard against a browser trick that makes a remote page look local. Neither asks who you are. The routes that pull, push, copy and delete models are as open as the ones that answer, and the server speaks plain HTTP, with no TLS (the protocol behind https) of its own. The keys Ollama does issue, sent as Authorization: Bearer, are for ollama.com.

That leans entirely on the default. Ollama binds the loopback address, so only the computer itself can reach it; its FAQ's answer to "How can I expose Ollama on my network?" is to change OLLAMA_HOST, and its recipes for a proxy, ngrok and Cloudflare Tunnel forward every route as-is. Ollama's official Docker image makes that change itself: at v0.34.4 it sets OLLAMA_HOST to every interface, and every docker run line in its Docker guide publishes the port on the host. Bound to every interface, Ollama is a model server that answers anyone who can reach the port.

Exposure has been measured from outside more than once. The security firm Wiz wrote up CVE-2024-37032 on 2024-06-24: a flaw in model downloads that let a malicious registry overwrite files on the server, and from there run code on it. Wiz called it "extremely severe in Docker installations, as the server runs with root privileges and listens on 0.0.0.0 by default" and counted "over 1,000 exposed instances". Ollama had acknowledged the report and committed a fix on 2024-05-05, the day it arrived ("in about 4 hours", in Wiz's words), and shipped v0.1.34 within days.

On 2026-01-29 SentinelOne's SentinelLABS, which had partnered with Censys "to scan and map internet-reachable Ollama deployments", reported "175,108 unique Ollama hosts across 130 countries" over 293 days of scanning. "Nearly half of observed hosts are configured with tool-calling capabilities that enable them to execute code, access APIs, and interact with external systems", it found, and exposing one "requires only a single configuration change: setting the service to bind to 0.0.0.0 or a public interface".

The maintainers' position is on record, in two comments of February 2025. On the community pull request that would add server-side keys (#9131, still open on 2026-09-30), a member of the team wrote on 2025-02-19: "In the short-term we don't plan to add an API key system to the Ollama server itself. Usually I recommend using a VPN" — he named one — "to manage access to Ollama server instances". A collaborator on the related issue, the day before: "The core team currently recommends using a proxy".

The engine underneath does take keys: llama-server has --api-key, and it serves TLS from a key file. Ollama starts it on the loopback address and passes neither flag through.

What leaves the machine. Its FAQ says "Ollama runs locally. We don't see your prompts or data when you run locally". The code at v0.34.4 shows what does go out unless the cloud features are off. At start the server fetches two things from ollama.com: since v0.22.1 (2026-04-30), a list of recommended models, fetched again about every four hours, sooner after a failed fetch; and since v0.23.2 (2026-05-07), the catalogue of its cloud models, so that it can describe them without a round trip. Each of those requests is signed with the key the installation generated, so calls from one installation carry the same public key. OLLAMA_NO_CLOUD=1, or disable_ollama_cloud in ~/.ollama/server.json, stops both, and turns off the cloud models and web search with them. Its FAQ calls this "local only mode", and the log then reads Ollama cloud disabled: true.

The defaults that moved

Each row is a setting a reader of an older tutorial may have wrong today, dated to the release or commit that moved it.

Date (UTC)What movedBefore → afterWhere it is written
2024-04-28Serving more than one request at onceOne at a time → OLLAMA_NUM_PARALLEL and OLLAMA_MAX_LOADED_MODELS as opt-in "experimental concurrency features"v0.1.33 release notes
2024-07-02The same, on by defaultOpt-in → the slot count "will auto-select either 4 or 1 based on available memory"v0.2.0 release notes; the FAQ at v0.2.0
2025-04-29The default window, and the automatic slot count2,048 → 4,096 tokens; 4 or 1 → 2 or 1 slotsCommits 44b466ee, "config: update default context length to 4096", and fe5b9bb2, "lower default num parallel to 2". The FAQ said "4 or 1" until 2025-07-08
2025-07-08Requests at onceChosen from free memory → one, "with the intent going forward that parallelism is explicit and will no longer be dynamically determined"Commit 20c3266e (#11330); first shipped in v0.10.0
2025-09-11How memory is planned, for models on Ollama's new engine (until 0.30, 2026-06-01)Estimated → measured "exact" before a model runs; models still on llama.cpp kept the old estimatesPR #12252, "llm: Enable new memory estimates by default", first shipped in v0.11.11 (opt-in since v0.11.5, 2025-08-15); the "New model scheduling" post of 2025-09-23
2025-12-13Flash attention, which saves memory as the window growsOpt-in (OLLAMA_FLASH_ATTENTION=1) → automatic where supported, model by model from 2025-08-27v0.13.4 release notes; v0.11.8's for gpt-oss
2026-02-02The default window, again4,096 (at least 8,192 for gpt-oss and Qwen3-VL models on 20 GiB or more) → 4,096 / 32,768 / 262,144 tokens by the memory of the cards it seesCommit 0334ffa6 (PR #13946, merged 2026-02-02; first shipped in v0.15.5, 2026-02-03); the context-length page. The FAQ still says "4096" at v0.35.0
2026-06-01The engine for GGUF modelsOllama's own engines → upstream llama-serverPR #16031; v0.30.0
2026-06-01VulkanOff → "enabled by default", for AMD and Intel cardsThe 0.30 post of 2026-06-05; OLLAMA_VULKAN
2026-08-12The repeat penalty, a sampler setting that pushes a model away from tokens it has just used1.1 → 1.0, "off", for models that set none, "matching other engines"v0.32.10 release notes (a pre-release; the first full release carrying it was v0.32.11, 2026-08-14); our own The free speed wasn't free
2026-09-25The engine on Apple siliconllama.cpp → MLX by default, in the v0.40.0 pre-releasev0.40.0-rc0 release notes

A footnote on numbers. This page carries no measurement of our own, on purpose: every figure on it comes from the source named beside it. The workshop's own measurements of Ollama beside vLLM, another open-source model server, on one EVGA RTX 3090 XC3 Ultra (24 GB) — one request at a time and several at once, the same model on both — are coming soon, on a separate page.

What to take with you

  • It began with llama.cpp compiled in, and since 0.30 (2026-06-01) it runs llama.cpp's own server again. In between came that server as a separate process from 2023-09-06 and engines of its own from late 2024.
  • Two founders who had sold Kitematic to Docker in 2015 run it as a company, with no foundation. It had raised $88M by 2026-07-09, most recently a $65M Series B, and has sold a cloud service since August 2025, priced per token since 2026-08-31.
  • Nothing on the local server asks who you are, and by default it calls ollama.com. It binds the loopback address, except in its official Docker image, and takes no key; in February 2025 the maintainers recommended a VPN or a proxy. At v0.34.4 its calls go out at start and about every four hours, signed with the installation's key, and two documented switches stop them.
  • Our busiest model's 2026-08-16 speed-up was one default moving. On 2026-08-12 the penalty on repeated words went from 1.1 to off for models that set none. Two other defaults: since v0.10.0 (2025-07-30) Ollama serves one request at a time by default, each slot a whole window of memory, and since v0.15.5 (2026-02-03) the default window is 4,096, 32,768 or 262,144 tokens by the memory of every card it sees, added up.

How to check our work — and see it live

  • Ask the rules desk. RuleSage, our board-game rules helper, is free and needs no account. The doorman named in this page's first section, a model served by Ollama on a card in this building, reads questions there too.
  • Read the source at the tag. Every file named here is linked at v0.34.4 under Sources, so a phone will do; at a computer, git clone https://github.com/ollama/ollama && git -C ollama checkout v0.34.4, then read envconfig/config.go (the defaults), server/routes.go (the window tiers and the two checks in front of every route), server/sched.go (the one-slot list and the queue-full error), llm/llama_server.go (the runner's flags and cache_prompt: true) and server/model_recommendations.go (the call to ollama.com about every four hours).
  • Read the guide the project wrote. docs/faq.mdx at the tag answers the slots and the exposure recipes in the project's own words, and docs/context-length.mdx the window.
  • See it on your own install. ollama ps prints the split under PROCESSOR when a model did not fit the card, and with OLLAMA_DEBUG=1 the log says it is truncating "messages which exceed context length" when a chat outgrows its window.

The rest of the seminar

Who ran this, and thanks

Thanks to Ollama (MIT), its team and its contributors, whose release notes, commits and documentation are dated and still up, which is why this page could be written from primary sources rather than from memory; to Georgi Gerganov and the ggml authors for llama.cpp (MIT), the engine at the bottom of most of this shelf's language-model numbers, and for a server README that states its defaults; to Apple for MLX (MIT). On NVIDIA cards that stack runs over CUDA (NVIDIA, proprietary, the one closed piece in it). The founding and funding lines rest on Ollama's own posts, Jeffrey Morgan's words on Hacker News, Docker's 2015 announcement, Y Combinator's directory, the lead investor Tomasz Tunguz's own post, and reports by TNW and BetaKit; the exposure history on Wiz's disclosure and the SentinelLABS and Censys research, in SentinelLABS's own report; the Hugging Face path on Hugging Face's own documentation; the Turbo page on the Internet Archive's Wayback Machine, which kept it after the live address moved; and the release, commit and pull-request dates on GitHub's record of the repository. None of them owed us anything. A small human team asked for this page, chose what it would and would not claim, and signed off on it; a fleet of AI agents fetched and read every source listed below and drafted it under that team's rulings.

Sources

Every external link below was fetched and read between 2026-09-30 18:42 and 2026-10-01 06:52 (UTC); the version and default lines, and the sibling page's link, were read again on 2026-10-02 (UTC). Source files were read at the tags named; the tag-pinned URLs reproduce each read. A release is dated by GitHub's release record, except where its tag commit is more than a week later or the record predates the merge cited for it; there the tag commit's date is used, and the entry says so.

Corrections and later measurements will be added below, each dated (UTC), with a window at both ends where one applies and saying in plain words what it counts.

elsewhere in the workshop

a strata→signal property · hello@strata2signal.com · say hello