# A short history of Mistral

*A French lab that put its first model on a torrent, every open-weight release since with its licence, and what its 24B model does on a 3090 and a 96 GB workstation card.*

*Published 2026-09-16 (UTC) · A small (human) team and a fleet of AI agents.*

**the short version:** A 24-billion-parameter French model reads what strangers type at two of this workshop's doors, before anything else runs. This is the history of the company behind it — founded April 2023, first model out by torrent, open weights at almost every size since — and every number we have measured on it, including a bench that set a 24 GB card beside a 96 GB one.

5,259 words · about 24 minutes (at 220 words/min) · 7 tables · data kit: no

https://research.strata2signal.com/a-short-history-of-mistral/

---

*If you're new here: [strata→signal](https://strata2signal.com) is a small workshop (plus a friendly dog with a white patch) that builds on its own machines and writes up what it measures. Every date below was read on 2026-09-16 (UTC) from the source credited beside it — Mistral's announcements and model cards, four press reports where no primary exists, and our own bench files. Where we found no source, the claim is not here.*

## The question that put a model at the door {#the-question-that-put-a-model-at-the-door}

A rules question arrived at this workshop's rules desk about a game ability whose printed name also reads, out of context, as a slur. A keyword list has one behaviour available to it: refuse the phrase — which means refuse the player holding the card, every time, and tell them nothing.

So what stands at those doors is not a word list. It is a 24-billion-parameter language model that reads the whole question, lets the rules question through, and refuses the same words used the other way. It cost nothing to download, runs on a card in this building, and is French: `mistral-small3.2:24b`, Apache-2.0, from a company now raising billions on the argument that giving the weights away is the business.

## Who are these people, and where did they come from? {#who-are-these-people-and-where-did-they-come-from}

Mistral AI was founded in April 2023, in Paris — The Register's "the Paris-based startup". Its own [about page](https://mistral.ai/about) dates it in one line, "In April 2023, Mistral was born to put frontier AI in everyone's hands", and names three founders: **Arthur Mensch**, CEO; **Guillaume Lample**, Chief Science Officer; **Timothée Lacroix**, CTO.

Where they came from explains what they built first. Mensch researched at Google DeepMind; Lample and Lacroix came from Meta, and both are authors on Meta's [LLaMA paper](https://arxiv.org/abs/2302.13971), submitted 2023-02-27, weeks before the company existed. The Register [reported the seed round](https://www.theregister.com/2023/06/15/mistral_ai/) on 2023-06-15 calling the company four weeks old, a month later than April; and Mensch's own words in the Series A announcement are "Since the creation of Mistral AI in May". Three sources, three months, no primary incorporation date — and we have not picked between them.

## La frontière — what they said they were for {#la-frontiere-what-they-said-they-were-for}

**La frontière, entre les mains de tout le monde.** English only borrowed that first word: *frontier* came around 1400 from the Old French *frontiere*, "boundary-line of a country" ([Online Etymology Dictionary](https://www.etymonline.com/word/frontier)). A French lab promising the frontier in everyone's hands is using a French word for it — and that sentence is ours, not theirs.

Their mission, read 2026-09-16 (UTC): **"Our mission is to make frontier AI open to all, and together solve the world's hardest problems."** The Series D frames it as sovereignty — data, models and compute staying where the customer is.

They mostly kept it, and said so where they did not — which is why the licence column below is the one that matters. Returning to Apache-2.0 for Mistral Small 3 in January 2025: *"We're renewing our commitment to using Apache 2.0 license for our general purpose models, as we progressively move away from MRL-licensed models"* — MRL being the research licence some of the larger models shipped under.

## What have they shipped, and under what licence? {#what-have-they-shipped-and-under-what-licence}

| model | licence | announced | size | what it changed |
|---|---|---|---|---|
| Mistral 7B | Apache 2.0 | 2023-09-27 | 7.3B | the first release, and it went out by torrent — a 13.4 GB download with a few hundred seeders on the day |
| Mixtral 8x7B | Apache 2.0 | 2023-12-11 | 46.7B total, 12.9B per token | the first mixture of experts anybody could download and run |
| Mistral Large (1) | **API only** | 2024-02-26 | not published | the commercial turn; announced on Azure, with le Chat beside it |
| Mixtral 8x22B | Apache 2.0 | 2024-04-17 | 141B total, 39B active, 64K context | "the most permissive open-source licence", in their words |
| Codestral | **Mistral AI Non-Production License** | 2024-05-29 | 22B, 32k context, 80+ languages | the first code model, and the licence exception |
| Mistral NeMo | Apache 2.0 | 2024-07-18 | 12B, 128k context | built with NVIDIA; introduced the Tekken tokenizer |
| Pixtral 12B | Apache 2.0 | 2024-09-17 | 12B (400M vision encoder) | vision; the page now marks it deprecated |
| Mistral Small 3 | Apache 2.0 | 2025-01-30 | 24B | the return to Apache-2.0, and small enough to quantize onto one consumer card |
| Mistral Small 3.1 | Apache 2.0 | 2025-03-17 | 24B, 128k context | multimodal, and twenty-five languages on the card |
| Mistral Medium 3 | **API / deployment** | 2025-05-07 | not published | priced at $0.4 in, $2 out per million tokens |
| Devstral (Small 1.0) | Apache 2.0 | 2025-05-21 | 24B, 128k context | built with All Hands AI; 46.8% on SWE-Bench Verified |
| Magistral Small / Medium | Small Apache 2.0; **Medium: API preview, no licence published** | 2025-06-10 | Small 24B | their first reasoning model |
| Mistral Small 3.2 | Apache 2.0 | 2025-06 (card `2506`) | 24B, 128k context | instruction following, repetition, function calling — **this is the one we run**, served here at 32k rather than its full 128k |
| Mistral 3 (Large 3 + Ministral 3) | Apache 2.0 | 2025-12-02 | Large 3: 675B total, 41B active (MoE); Ministral 3 at 14B, 8B, 3B | a frontier-size model released open-weight — and the Ministral 3 trio is the part of it that fits on one consumer card |
| Mistral Small 4 | Apache 2.0 | 2026-03-16 | 119B total, 6B active per token (MoE, 128 experts), 256k context | the newest Apache-2.0 release, and the one the API pricing page sells most cheaply |
| Mistral Medium 3.5 | **Modified MIT** | 2026-04-28 (card `26-04`) | 128B dense, 256k context | one set of weights for instructions, reasoning and code; 77.6% on SWE-Bench Verified |

Two words in that size column are not interchangeable. A **dense** model uses every parameter on every token; a **mixture of experts** routes each token through a fraction of them, so it costs the disk of the big number at nearer the speed of the small one. `Q4_K_M`, below, is the 4-bit quantisation most people run: a quarter the size of 16-bit weights, at a small cost in quality.

This table is the licence spine, not the catalogue: a search for "mistral" on the ollama.com library returned twenty entries on 2026-09-16, eleven of them Mistral AI's own. Le Chat, announced beside Mistral Large on 2024-02-26, now carries a rename notice on its own announcement page: *"Le Chat is now Vibe—one agent and one licence across work and code, with every conversation, setting, and plan carried over."* Read 2026-09-16 (UTC).

Read that licence column top to bottom and a shape appears: the models small enough to run at home stayed Apache-2.0, and the ones that needed a data centre are where the word "open" started carrying conditions.

## What the company is worth {#what-the-company-is-worth}

Five rounds in three years, each dated to the announcement that carried it.

| round | date | raised | valuation | lead investor |
|---|---|---|---|---|
| seed | [2023-06-15](https://www.theregister.com/2023/06/15/mistral_ai/) | $113M (€105M) | $259M (€240M) — the report says "is now valued at", not post-money | Lightspeed Venture Partners, with Eric Schmidt, Xavier Niel and Bpifrance contributing |
| Series A | [2023-12-11](https://techcrunch.com/2023/12/11/mistral-ai-a-paris-based-openai-rival-closed-its-415-million-funding-round/) | $415M (€385M) | roughly $2 billion, attributed to Bloomberg | Andreessen Horowitz |
| Series B | [2024-06-11](https://techcrunch.com/2024/06/11/paris-based-ai-startup-mistral-ai-raises-640-million) | ~$640M (€600M, equity and debt) | $6 billion "following this funding round" | General Catalyst |
| Series C | [2025-09-09](https://mistral.ai/news/mistral-ai-raises-1-7-b-to-accelerate-technological-progress-with-ai/) | $2.00B (€1.7B at the ECB rate of 1.1744 on 2025-09-09) | $13.74B **post-money** (€11.7B at that rate) | ASML Holding NV |
| Series D | [2026-09-08](https://mistral.ai/news/mistral-makes-sovereign-open-weight-ai-to-frontier/) | $3.48B (€3B at the ECB rate of 1.1614 on 2026-09-08) | more than $24.39B **post-money** (€21B at that rate) | Samsung Electronics, with Scaleup Europe Fund (managed by EQT) and PSG Equity as co-leads |

Every figure here is in dollars: the first three rounds as their sources printed them, the C and the D converted from euros at the European Central Bank's reference rate for the announcement day, which each row carries. Only the last two say "post-money" in their source; the others keep their report's wording, because the difference matters and guessing is not reporting. Mistral's claim for the last row: **"the largest equity fundraising round ever completed by a European technology company, three years after the company's launch."** We have no view on the valuation.

## Do they train models in French? {#do-they-train-models-in-french}

Non — ils entraînent sur des données multilingues où le français figure au premier rang. Not "in French", then, but on multilingual data in which French is first-class — the more checkable thing: the [Mistral Small 3.1 model card](https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503) lists twenty-five languages, English first.

The tokenizer is where that claim becomes measurable. Tekken, introduced with [Mistral NeMo](https://mistral.ai/news/mistral-nemo) in July 2024, was trained on more than 100 languages and is "~30% more efficient at compressing source code, Chinese, Italian, French, German, Spanish, and Russian" than the SentencePiece tokenizer it replaced, and beat Llama 3's on about 85% of all languages. Fewer tokens for the same French sentence is less money and less waiting.

One caution, from our bench rather than theirs: those figures say Tekken beat its predecessor and Llama 3's, not that it is the densest tokenizer in the field. On one frozen English prompt, `mistral-small3.2:24b` read **1,023 tokens** where `gemma4:26b` read **536** of the same bytes ([two hours on battery](/two-hours-on-battery/)) — double the prefill, and a bigger cache on every decode step after.

Mistral publishes no proportions, so anyone quoting a percentage is guessing.

## Why is it called Mistral? {#why-is-it-called-mistral}

The wind first. Larousse gives one sense — *"Vent violent, froid, turbulent et sec, qui souffle du secteur nord, sur la France méditerranéenne, entre les méridiens de Sète et de Toulon"* — and one etymology, the joke the company has been telling ever since, printed as Larousse prints it: *"ancien provençal maestral, du bas latin magistralis, magistrat"*. The reasoning model they shipped in 2025 is [**Magistral**](https://mistral.ai/news/magistral): that root, come back round.

The rest of the family is the wind's name with the job spliced into it: [**Mixtral**](https://mistral.ai/news/mixtral-of-experts) for the mixture of experts, [**Codestral**](https://mistral.ai/news/codestral) for code, [**Pixtral**](https://mistral.ai/news/pixtral-12b) for vision, [**Devstral**](https://mistral.ai/news/devstral) for software engineering, [**Voxtral**](https://mistral.ai/news/voxtral) for speech, and [**les Ministraux**](https://mistral.ai/news/ministraux), the edge models, announced 2024-10-16 as "Un Ministral, des Ministraux" — a singular-and-plural joke landing on Mistral 7B's first anniversary.

The chat product never took the wind's name, and has given up its own: [**Le Chat**](https://mistral.ai/news/le-chat-mistral) is **Vibe** — neither French nor a wind.

## What does it do on our own machines? {#what-does-it-do-on-our-own-machines}

Nothing here was measured for this page; every row was published earlier. Seven of the nine readings are on `mistral-small3.2:24b`, the 24B dense Apache-2.0 release at Q4_K_M — two instruments, two figures: 14.14 GiB of weights in the model store, about 17 GiB of a card once it serves a 32,768-token window. That is why it fits a 24 GB card and not a 16 GB one. The last two rows name larger Mistral models.

| what was measured, and on which model | reading | power cap | when, and where to check it |
|---|---|---|---|
| writing speed on an RTX PRO 6000 Blackwell (96 GB), single warm calls, 128 tokens out | 95.0 / 94.6 / 93.5 tokens a second | 600 / 500 / 450 W | 2026-08-27 · [what 150 watts buys](/what-150-watts-buys/) |
| power and heat under its own load, same card | 594.6 W mean across its first four-minute leg, 87 °C peak across the eight minutes — the hottest arm of the serving bench it was read from | uncapped in that arm | 2026-08-25 · [what 150 watts buys](/what-150-watts-buys/) |
| what a power cap costs this model, same card | clock −12.0%, throughput −1.4% at one stream; −18.7% and −2.0% at sixteen | 600 W → 450 W | 2026-08-25 · [what 150 watts buys](/what-150-watts-buys/) |
| on processors alone, no graphics card | 7.1 tokens a second on a server, 2.7 on a laptop, 3.3 on a mini PC — and 97.7 s before the first word on the mini | not a card | 2026-08-25 · [two hours on battery](/two-hours-on-battery/), [the two-hour machine priced](/the-two-hour-machine-priced/) |
| the same processors, threads tuned | 2.7 → 5.1 (1.90×) on the laptop; 3.3 → 4.0 (1.20×) on the mini | not a card | 2026-08-25 · [two hours on battery](/two-hours-on-battery/) |
| on a used consumer card too small to hold it, half in RAM | 5.13 tokens a second split, against 2.71 on that box's processor alone — 1.89× | cap not set by that harness | 2026-09-10 · this workshop's own bench record, `gpu-3080-2026-09-10/RESULTS.md` (not yet a published page) |
| what this model's tokenizer does to one fixed prompt | 1,023 tokens where a Gemma tokenizer read 536 of the same text | not a card | 2026-08-25 · [two hours on battery](/two-hours-on-battery/) |
| `mistral-medium-3.5:128b`, `mistral-large-3` and a 23.6B q4 `mistral-small` seat sitting our grounded-judge trial — 43 sentences, 27 that must be killed as unsourced and 16 that must be left alone | medium-3.5 killed 27/27 and kept 15/16 — PASS; large-3 killed 27/27 and kept 15/16 — PASS; the small seat killed 27/27 but kept only 10/16 — FAIL, on preservation rather than on catching | not a card measurement | 2026-08 · [the seat trials](/seat-trials/) |
| `mistral-large-3` judging our work from outside | carried its audition first time and judged 5 of 5 sealed rounds | not a card measurement | 2026-08-13 · [outside judges](/outside-judges/) |

That trial is a different job from standing at a door: a grounding judge must leave true claims alone, and the small seat over-refused. A doorman answers one yes-or-no question about one short string — the task it does keep. Nothing here says a Mistral model failed a safety exam; no such exam was run.

On 2026-09-16 two cards sat in one box for the first time, and the cheap one has something to say.

## What we measured on two cards in one box {#what-we-measured-on-two-cards-in-one-box}

A bench ran in the small hours of 2026-09-16 (UTC) that set a used **RTX 3090 (24 GB)** beside an **RTX PRO 6000 Blackwell (96 GB)** in one box: one kernel, one driver, one ollama build, one model store, one frozen prompt, the same weight files the live seats hold. Nothing is box against box. It was pre-registered, including the sentence that would have refuted it — *if a 3090 reads within 15% of the workstation card on either model, that is the headline* — which did not fire: 39% behind on one model, 63% on the other.

Every row: a 3090 at 250 W against the workstation card at 420 W, one frozen 2,101-byte prompt — 1,023 tokens to Mistral's tokenizer, 536 to Gemma's — 256 tokens out.

| measured | RTX 3090 (24 GB) — a bench instance | RTX PRO 6000 Blackwell (96 GB) — the served seat |
|---|---|---|
| `mistral-small3.2:24b`, 24B **dense**, decode | 34.13 tokens a second | 92.77 tokens a second |
| `gemma4:26b`, 26B **MoE**, decode | 126.12 tokens a second | 205.85 tokens a second |
| watts while decoding, dense / MoE | 228.2 W / 206.4 W | 328.3 W / 182.5 W |
| energy per 1,000 tokens, dense / MoE | 6,680 J / 1,632 J | 3,538 J / 887 J |
| first token on a prompt it has already seen, dense / MoE | 98 ms / 167 ms | 28 ms / 162 ms |

*One file per column: `m4-3090.json`, `m4-6000-live.json`, `m5-3090.json`, `m5-6000-live.json`. The 96 GB column is a live serving seat: a second bench copy would not fit beside the services already on it — 12,965 MiB free against 16,645 MiB wanted (`m4-6000-bench.json`), as the pre-registration predicted. A 96 GB card in service is already full.*

**Read the dense row first; it explains the rest.** Writing one token from a dense model means reading every weight in it, once — at Q4_K_M, 14.14 GiB per token. So 34.13 tokens a second is about 482 GiB/s of fetching: **55% of the 936 GB/s NVIDIA publishes for a GeForce RTX 3090**, and at 350 W, cap out of the way, 50.02 tokens a second is 707 GiB/s — four fifths of it. Decode is a fetching problem, not an arithmetic one. A mixture of experts changes that: `gemma4:26b` routes each token through [about 4 billion active parameters](https://ollama.com/library/gemma4), a few gigabytes instead of fourteen, and on the same card, cap and prompt runs **3.7× faster** on a model *larger* on disk.

The same arithmetic on the workstation card lands in the same place: 92.77 tokens a second is 1,311 GiB/s against the 1,792 GB/s NVIDIA publishes for it. **Unthrottled, both cards read the dense model at about four fifths of their own memory bandwidth** — 81% for a 3090 at its board default, 79% for a 96 GB card at 420 W. What separates them is how much bandwidth they have, not how well either uses it.

Two earlier benches on other boxes say it from below — controls, not like-for-like: on a 3080 split with system memory, 34.40 tokens a second for the MoE against 5.13 for the dense Mistral; on a mini PC's processors alone, 12.66 against 3.31.

## What a power cap actually costs {#what-a-power-cap-actually-costs}

Fetching is also what costs watts, which is why a 3090's two arms behave differently under one 250 W limit. The dense model sat at 228.2 W — 91% of cap — clock pulled to **495–675 MHz** across three runs at 64–66 °C, nowhere near thermal (`m4-3090.json`). The MoE drew 206.4 W, 82.6% of the same cap, clock following the work (`m5-3090.json`). A dense model under a cap is power-bound; a sparse one is not yet. So the ladder moved the cap:

| model | power cap | tokens a second | vs 250 W | energy per 1,000 tokens |
|---|---|---|---|---|
| `mistral-small3.2:24b` (dense) | 250 W | 34.13 | — | 6,680 J |
| `mistral-small3.2:24b` (dense) | 300 W | 48.18 | +41.2% | 5,673 J |
| `mistral-small3.2:24b` (dense) | 350 W (the board's default) | 50.02 | +46.6% | 5,318 J |
| `gemma4:26b` (MoE) | 250 W | 126.12 | — | 1,632 J |
| `gemma4:26b` (MoE) | 350 W (the board's default) | 136.49 | +8.2% | 1,823 J |

*250 W rows: `m4-3090.json`, `m5-3090.json`. The rungs above: `capladder-mistral-small3.2-24b.json`, `capladder-gemma4-26b.json`.*

Almost all the money is in the first fifty watts: 250 → 300 W buys the dense model 41%, the next fifty 3.8%; the MoE, never power-bound, gains 8.2% for a hundred. The cap on this card was ruled to **300 W** minutes after that table was measured — the figure the dollars below use.

**What VRAM is worth.** `qwen3.5:122b` — a mixture of experts of 122 billion parameters, 10 billion active per token, 75.7 GiB of weights — fits on neither card, so it was split across both. Same caps: a 3090 at 250 W, the workstation card at 420 W.

| what the cards could offer | on card | tokens a second | source |
|---|---|---|---|
| the workshop's seats up | 39.5% | 38.22 | `q1-split-resident-seatsup.json` |
| the seats stopped, inside one authorised window | 89.6% | 71.48 | `q1-split-resident-seatsoff.json` |
| the 96 GB card alone, seats stopped | 63.0% | *refused — does not fit* | `q1-6000-whole.json` |

Freeing the seats' memory nearly doubled throughput — **1.87×**, same model, same box, same minute, nothing else changed: 87% more tokens a second from the same silicon. **This box holds 90% of a 122B model across two cards, and none of it whole on one.**

A 3090 takes spillover, too: the MoE at one, two and four simultaneous streams produced **127.1 / 195.8 / 261.4 tokens a second** aggregate at a 250 W cap, its fans at zero RPM until the fourth stream woke them to about half speed (`concurrency-3090.json`).

One arm is not reported, and the reason is the finding. The embedding arm **voided itself against its own pre-registered gate**: five batches spread 34.9% around their median where the gate allowed 15% (`embed-3090.json`). Half-second batches on a card that fast measure scheduling, not silicon — and a rung rejected by its own gate gets no ratio, so no number from it is here. It is owed a longer batch and a second run.

## Is it cheaper to run it yourself? {#is-it-cheaper-to-run-it-yourself}

Two of those numbers turn into money. The bench measures energy per thousand output tokens, so the arithmetic runs joules → kilowatt-hours per million → dollars at **18.34 cents a kilowatt-hour**: the US residential average for June 2026 (Energy Information Administration, *Electric Power Monthly*, table 5.6.A), a national figure and not this workshop's bill.

- `mistral-small3.2:24b` on a 3090 at its ruled 300 W cap: **5,673 J per 1,000 tokens** (`capladder-mistral-small3.2-24b.json`) → 1.576 kWh per million → **$0.29 per million**.
- `gemma4:26b` on the workstation card at 420 W: **887 J per 1,000 tokens** (`m5-6000-live.json`) → 0.246 kWh per million → **$0.045 per million**.

Per token written the 96 GB card is the more frugal — 3,538 J against 6,680 dense, 887 against 1,632 MoE, roughly 1.9× and 1.8×.

Beside Mistral's own list prices, read from [the API pricing page](https://mistral.ai/pricing/api) on 2026-09-16 (UTC):

| what you are paying for | input | output |
|---|---|---|
| Mistral Small 4, on their machines | $0.15 | $0.6 |
| Mistral Medium 3.5, on their machines | $1.5 | $7.5 |
| Mistral Large 3, on their machines | $0.5 | $1.5 |
| Ministral 3 (8B), on their machines | $0.15 | $0.15 |
| `mistral-small3.2:24b` on our own 3090 at a 300 W cap, electricity only | — | $0.29 |
| `gemma4:26b` on our own 96 GB card at 420 W, electricity only | — | $0.045 |

**Those are not the same kind of number, and the comparison is worth exactly what that caveat allows.** The electricity figure is the watts that left the card while it wrote output tokens, and nothing else: no hardware, no cooling, no idle hours, no prompt-processing energy, nobody's time — and it is the whole card's draw while it decoded, floor included, of which about 86 W on the shared 96 GB seat belongs to the services already resident there. The API figure is a published list price for a managed service, which includes all of that and a margin. The pair is good for a floor: running a 24B model yourself is cents of electricity per million tokens, so the rest of the bill is the machine, not the meter.

One oddity in their own table: **Mistral Large 3 is cheaper than Mistral Medium 3.5**, $0.5 in and $1.5 out against $1.5 and $7.5. That is what the page said on 2026-09-16 (UTC); list prices move, so it is stamped, not asserted.

**What the card cost.** The used 3090 in this box was bought for **$1,449.99 in May 2026**. A card price is a reading, not a fact, so here is what the same search found on 2026-09-13, 06:02–06:05Z —

- the same model, an EVGA XC3 Ultra, **renewed at $1,849.99** — one Amazon seller, one unit;
- the cheapest **new** listing, a ZOTAC Trinity OC at **$1,829.99**, with a "usually ships within 5 to 6 days" caveat;
- an EVGA K|NGP|N Hybrid, **new at $1,839.00**;
- eleven renewed cards of other makes, **$1,529.99 to $1,799.99**.

At what was paid, that is **$11.50 per token a second** on the MoE (126.12 tok/s) and **$30.10** on the dense Mistral at 300 W (48.18 tok/s). No price is printed for the 96 GB workstation card: this workshop holds none of record, and an invented one is worse than a blank.

## Where the doorman actually stands {#where-the-doorman-actually-stands}

In the [print lab](/how-the-print-lab-works/), Mistral Small 3.2 reads the word you type and answers one question — is this safe to paint and publish on a family board — before a picture is drawn. In the [beat lab](/how-the-beat-lab-works/), it checks that a wish is a musical request and nothing else; the rules desk uses the same seat. It has stood at those doors since 2026-08-26.

The card whose printed name started this page still goes through it, and still comes back a rules question rather than a refusal.

## What to take with you {#what-to-take-with-you}

- **They shipped the first one as a torrent.** September 2023, 13.4 GB, no gate, no waitlist.
- **"Open" is a per-release word, not a company word.** Apache-2.0 on what fits one card, conditions on what needs a data centre.
- **They do not train "in French."** French is first-class in a multilingual mix, and their tokenizer read our English prompt at nearly twice Gemma's count.
- **$113M to more than $24B in three years,** on the argument that giving the weights away is the business.
- **Sparse beats big, which is why a doorman is cheap.** A dense 24B reads all 14.14 GiB of itself per token written; a same-size mixture of experts reads a fraction and runs 3.7× faster — so the door costs 34 tokens a second on one used card, cents of electricity per million, and nothing a stranger types leaves the building.

## How to check our work {#how-to-check-our-work}

Every external link below was fetched and read on 2026-09-16 (UTC).

- Mistral AI — [About](https://mistral.ai/about) (mission, founding month, the three founders and their roles)
- Touvron et al. — [LLaMA: Open and Efficient Foundation Language Models](https://arxiv.org/abs/2302.13971), 2023-02-27 (Lacroix and Lample in the author list)
- The Register — [Mistral AI seed round](https://www.theregister.com/2023/06/15/mistral_ai/), 2023-06-15 (€105M / $113M, a $259M valuation, Lightspeed, Schmidt, Niel, Bpifrance; "the Paris-based startup"; "Europe's largest ever seed round", per Dealroom.co; the company "four weeks" old; Mensch at DeepMind, Lacroix and Lample from Meta)
- Online Etymology Dictionary — [frontier](https://www.etymonline.com/word/frontier) (the Old French *frontiere*, around 1400)
- Mistral AI — [Mistral 7B](https://mistral.ai/news/announcing-mistral-7b), 2023-09-27
- TechCrunch — [Mistral AI makes its first large language model free for everyone](https://techcrunch.com/2023/09/27/mistral-ai-makes-its-first-large-language-model-free-for-everyone/), 2023-09-27 (the 13.4 GB torrent)
- Mistral AI — [Mixtral of experts](https://mistral.ai/news/mixtral-of-experts), 2023-12-11
- Mistral AI — [Au Large](https://mistral.ai/news/mistral-large), 2024-02-26 (Mistral Large, Azure, le Chat)
- Mistral AI — [Le Chat](https://mistral.ai/news/le-chat-mistral), 2024-02-26 (and the Vibe rename notice on the same page)
- Mistral AI — [Cheaper, Better, Faster, Stronger (Mixtral 8x22B)](https://mistral.ai/news/mixtral-8x22b), 2024-04-17
- Mistral AI — [Codestral](https://mistral.ai/news/codestral), 2024-05-29 (the Non-Production License)
- Mistral AI — [Mistral NeMo](https://mistral.ai/news/mistral-nemo), 2024-07-18 (NVIDIA, Tekken, the compression figures, the SentencePiece baseline)
- Mistral AI — [Pixtral 12B](https://mistral.ai/news/pixtral-12b), 2024-09-17
- Mistral AI — [Mistral Small 3](https://mistral.ai/news/mistral-small-3), 2025-01-30 (the Apache-2.0 commitment and the MRL sentence)
- Mistral AI — [Mistral Small 3.1](https://mistral.ai/news/mistral-small-3-1), 2025-03-17
- Hugging Face — [mistralai/Mistral-Small-3.1-24B-Instruct-2503](https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503). The twenty-five languages, in the card's own order: "English, French, German, Greek, Hindi, Indonesian, Italian, Japanese, Korean, Malay, Nepali, Polish, Portuguese, Romanian, Russian, Serbian, Spanish, Swedish, Turkish, Ukrainian, Vietnamese, Arabic, Bengali, Chinese, Farsi."
- Hugging Face — [mistralai/Mistral-Small-3.2-24B-Instruct-2506](https://huggingface.co/mistralai/Mistral-Small-3.2-24B-Instruct-2506) (the model this workshop runs; licence apache-2.0)
- Mistral AI — [Medium is the new large (Mistral Medium 3)](https://mistral.ai/news/mistral-medium-3), 2025-05-07
- Mistral AI — [Devstral](https://mistral.ai/news/devstral), 2025-05-21
- Hugging Face — [mistralai/Devstral-Small-2505](https://huggingface.co/mistralai/Devstral-Small-2505) (24B, 128k, apache-2.0)
- Mistral AI — [Magistral](https://mistral.ai/news/magistral), 2025-06-10
- Mistral AI — [Introducing Mistral 3](https://mistral.ai/news/mistral-3/), 2025-12-02 (Large 3 and the Ministral 3 trio, all Apache 2.0)
- Mistral AI — [Remote agents in Vibe, powered by Mistral Medium 3.5](https://mistral.ai/news/vibe-remote-agents-mistral-medium-3-5/), 2026-05-22 (128B dense, 256k, modified MIT, self-hosting)
- Hugging Face — [mistralai/Mistral-Medium-3.5-128B](https://huggingface.co/mistralai/Mistral-Medium-3.5-128B) (Modified MIT License, 128B, 256k)
- Mistral AI — [Mistral raises €3B to make sovereign, open-weight AI the technology frontier](https://mistral.ai/news/mistral-makes-sovereign-open-weight-ai-to-frontier/), 2026-09-08 (the Series D — €3B, $3.48B at that day's ECB rate — the post-money wording, and the European-record claim)
- ollama.com — the [library search for "mistral"](https://ollama.com/search?q=mistral), read 2026-09-16 (twenty entries, eleven of them Mistral AI's own)
- Mistral AI — [Mistral Small 4](https://mistral.ai/news/mistral-small-4/), 2026-03-16 (119B total, "6B active parameters per token (8B including embedding and output layers)", Apache 2.0)
- Hugging Face — [mistralai/Mistral-Small-4-119B-2603](https://huggingface.co/mistralai/Mistral-Small-4-119B-2603) (Apache 2.0 — and it says 6.5B activated per token where the announcement says 6B; the table above follows the announcement, and the disagreement is the vendor's own)
- Mistral AI — [Mistral Medium 3.5 model card](https://docs.mistral.ai/models/model-cards/mistral-medium-3-5-26-04) (released 2026-04-28; Modified MIT)
- Mistral AI — [API pricing](https://mistral.ai/pricing/api), read 2026-09-16 (the four list prices quoted above)
- TechCrunch — [Mistral AI closed its $415 million funding round](https://techcrunch.com/2023/12/11/mistral-ai-a-paris-based-openai-rival-closed-its-415-million-funding-round/), 2023-12-11 (the Series A, a16z, "roughly $2 billion" attributed to Bloomberg, and Mensch's "Since the creation of Mistral AI in May")
- TechCrunch — [Mistral AI raises $640 million](https://techcrunch.com/2024/06/11/paris-based-ai-startup-mistral-ai-raises-640-million), 2024-06-11 (the Series B, General Catalyst, "$6 billion valuation following this funding round")
- Mistral AI — [Mistral AI raises 1.7B€](https://mistral.ai/news/mistral-ai-raises-1-7-b-to-accelerate-technological-progress-with-ai/), 2025-09-09 (the Series C, ASML, 11.7B€ post-money — $13.74B at that day's ECB rate)
- European Central Bank — [the daily US dollar euro reference rate](https://data.ecb.europa.eu/data/datasets/EXR/EXR.D.USD.EUR.SP00.A), series `EXR.D.USD.EUR.SP00.A`, read 2026-09-16: 1.1744 dollars to the euro on 2025-09-09 and 1.1614 on 2026-09-08, which are the two rates the funding table converts at
- Larousse — [mistral](https://www.larousse.fr/dictionnaires/francais/mistral/51794) (the wind, the single numbered sense, and the *magistralis* etymology exactly as printed)
- Mistral AI — [Un Ministral, des Ministraux](https://mistral.ai/news/ministraux), 2024-10-16 (the edge models and the name)
- Mistral AI — [Voxtral](https://mistral.ai/news/voxtral), 2025-07-15 (speech understanding, 24B and 3B)
- NVIDIA — [NVIDIA Ampere GA102 GPU Architecture](https://images.nvidia.com/aem-dam/en-zz/Solutions/geforce/ampere/pdf/NVIDIA-ampere-GA102-GPU-Architecture-Whitepaper-V1.pdf) (936 GB/s of peak memory bandwidth on a GeForce RTX 3090, 384-bit, 19.5 Gbps GDDR6X)
- NVIDIA — [RTX PRO 6000 Blackwell Workstation Edition](https://www.nvidia.com/en-us/products/workstations/professional-desktop-gpus/rtx-pro-6000/) (96 GB GDDR7 with ECC, 1792 GB/s)
- ollama.com — [gemma4](https://ollama.com/library/gemma4) (the 26B tag is a mixture of experts with about 4B active parameters, from Google DeepMind) and [qwen3.5](https://ollama.com/library/qwen3.5) (the 122b tag)
- Hugging Face — [Qwen/Qwen3.5-122B-A10B](https://huggingface.co/Qwen/Qwen3.5-122B-A10B) (apache-2.0, 122B total with 10B activated per token — the model behind the `qwen3.5:122b` tag benched above)
- This workshop's own dated card-price read, `docs/RECON-3090-DUO-RIG-PRICE-2026-09-13.md`, 2026-09-13 06:02–06:05Z (the four listings quoted above; not a published page)
- This workshop's own published benches, linked inline above: [what 150 watts buys](/what-150-watts-buys/), [two hours on battery](/two-hours-on-battery/), [the two-hour machine priced](/the-two-hour-machine-priced/), [seat trials](/seat-trials/), [outside judges](/outside-judges/), [how the print lab works](/how-the-print-lab-works/), [how the beat lab works](/how-the-beat-lab-works/)
- The two controls quoted in the two-card section, from an earlier bench on other boxes at other settings: the MoE split between a 3080 and system memory read 34.40 tokens a second (`gpu-3080-m5-split.json`) at 129.9 W — the watts are in that bench's own `RESULTS.md`, not in the result file — and 12.66 on a mini PC's processors alone (`cpu-mini-m5.json`); the dense Mistral read 5.13 (`gpu-3080-m4-split.json`) and 3.31 (`cpu-mini-m4.json`). Neither harness set a power cap.
- The two-card bench's own result files, named beside every figure above: `m4-3090.json`, `m4-6000-live.json`, `m4-6000-bench.json`, `m5-3090.json`, `m5-6000-live.json`, `capladder-mistral-small3.2-24b.json`, `capladder-gemma4-26b.json`, `concurrency-3090.json`, `embed-3090.json`, `q1-split-resident-seatsup.json`, `q1-split-resident-seatsoff.json`, `q1-6000-whole.json`. The run names each file for the box it ran on and then for the arm, and this shelf does not publish box names, so every file is cited here by its arm half, which is unique within the run.

## Who ran this, and thanks {#who-ran-this-and-thanks}

Thanks to Mistral AI, whose announcements and model cards are dated, detailed and still up — which is why this page could be written from primary sources rather than from memory. The model itself is **Mistral Small 3.2** (Mistral AI, Apache-2.0); every reading republished above was taken through **Ollama** (MIT) over **llama.cpp** (MIT) and **CUDA** (NVIDIA, proprietary, under the CUDA Toolkit EULA) — the one closed piece in the stack, and the one that comes with the cards — on 4-bit `Q4_K_M` weights packaged by the people who quantize and host them for everyone else. The comparison model in the tokenizer row and in the two-card bench is **Gemma 4** (Google DeepMind): the licence blob this hub read off its own `gemma4:26b` tag on 2026-08-11 is stock Apache 2.0, and it is on our [licences](/licences/) page. The 122B model is **Qwen 3.5** (Qwen, Alibaba), whose model card publishes apache-2.0 — this shelf has never read the licence packaged with that ollama tag itself, and says so rather than implying a read it did not run. The electricity price is the **U.S. Energy Information Administration's** published residential average; the wind's definition is **Larousse's** and the word's history the **Online Etymology Dictionary's**; the two memory-bandwidth figures are **NVIDIA's** own published specifications. None of them owed us anything. A small human team asked for this page, chose what it would and would not claim, and signed the numbers; a fleet of AI agents fetched and read every source listed above, pulled the measured rows out of benches this hub had already published, and drafted it under that team's rulings.

Corrections and later measurements will be added below, each dated (UTC), with a window at both ends where one applies and saying in plain words what it counts.

<!-- derived 2026-09-16 (UTC) by tools/derive_md.py from the pour source.
     source html sha256: 61e4b1e8495f311fccb5a7283e08577c658037f9fc280a17a6772bfed3c28004
     derivation sha256:  11d53e8466f57f7f471a385bc6b0c68f6580d981579be1de7fcb6f4924a27cf6
     the {#id} on each heading is the anchor that heading carries on the page. -->
