The notes — a company, its stated goals, and every open-weight model it has shipped

A short history of Mistral

exhibit forty-nine The notes
Published 2026-09-16 (UTC)
A small (human) team and a fleet of AI agents.

A French lab that put its first model on a torrent, every open-weight release since with its licence, and what its 24B model does on a 3090 and a 96 GB workstation card.

ask about this page assistant.strata2signal.com · in beta, still being tested

the short version

A 24-billion-parameter French model reads what strangers type at two of this workshop's doors, before anything else runs. This is the history of the company behind it — founded April 2023, first model out by torrent, open weights at almost every size since — and every number we have measured on it, including a bench that set a 24 GB card beside a 96 GB one.

5,259 words, about 24 minutes to read.

The summary is this page’s own; the receipt lines were drafted by a model on this workshop’s network and every figure in them is in the article, checked before this page went out — what was dropped, and why, is in this page’s receipt file.

If you're new here: strata→signal is a small workshop (plus a friendly dog with a white patch) that builds on its own machines and writes up what it measures. Every date below was read on 2026-09-16 (UTC) from the source credited beside it — Mistral's announcements and model cards, four press reports where no primary exists, and our own bench files. Where we found no source, the claim is not here.

The question that put a model at the door

A rules question arrived at this workshop's rules desk about a game ability whose printed name also reads, out of context, as a slur. A keyword list has one behaviour available to it: refuse the phrase — which means refuse the player holding the card, every time, and tell them nothing.

So what stands at those doors is not a word list. It is a 24-billion-parameter language model that reads the whole question, lets the rules question through, and refuses the same words used the other way. It cost nothing to download, runs on a card in this building, and is French: mistral-small3.2:24b, Apache-2.0, from a company now raising billions on the argument that giving the weights away is the business.

Who are these people, and where did they come from?

Mistral AI was founded in April 2023, in Paris — The Register's "the Paris-based startup". Its own about page dates it in one line, "In April 2023, Mistral was born to put frontier AI in everyone's hands", and names three founders: Arthur Mensch, CEO; Guillaume Lample, Chief Science Officer; Timothée Lacroix, CTO.

Where they came from explains what they built first. Mensch researched at Google DeepMind; Lample and Lacroix came from Meta, and both are authors on Meta's LLaMA paper, submitted 2023-02-27, weeks before the company existed. The Register reported the seed round on 2023-06-15 calling the company four weeks old, a month later than April; and Mensch's own words in the Series A announcement are "Since the creation of Mistral AI in May". Three sources, three months, no primary incorporation date — and we have not picked between them.

La frontière — what they said they were for

La frontière, entre les mains de tout le monde. English only borrowed that first word: frontier came around 1400 from the Old French frontiere, "boundary-line of a country" (Online Etymology Dictionary). A French lab promising the frontier in everyone's hands is using a French word for it — and that sentence is ours, not theirs.

Their mission, read 2026-09-16 (UTC): "Our mission is to make frontier AI open to all, and together solve the world's hardest problems." The Series D frames it as sovereignty — data, models and compute staying where the customer is.

They mostly kept it, and said so where they did not — which is why the licence column below is the one that matters. Returning to Apache-2.0 for Mistral Small 3 in January 2025: "We're renewing our commitment to using Apache 2.0 license for our general purpose models, as we progressively move away from MRL-licensed models" — MRL being the research licence some of the larger models shipped under.

What have they shipped, and under what licence?

modellicenceannouncedsizewhat it changed
Mistral 7BApache 2.02023-09-277.3Bthe first release, and it went out by torrent — a 13.4 GB download with a few hundred seeders on the day
Mixtral 8x7BApache 2.02023-12-1146.7B total, 12.9B per tokenthe first mixture of experts anybody could download and run
Mistral Large (1)API only2024-02-26not publishedthe commercial turn; announced on Azure, with le Chat beside it
Mixtral 8x22BApache 2.02024-04-17141B total, 39B active, 64K context"the most permissive open-source licence", in their words
CodestralMistral AI Non-Production License2024-05-2922B, 32k context, 80+ languagesthe first code model, and the licence exception
Mistral NeMoApache 2.02024-07-1812B, 128k contextbuilt with NVIDIA; introduced the Tekken tokenizer
Pixtral 12BApache 2.02024-09-1712B (400M vision encoder)vision; the page now marks it deprecated
Mistral Small 3Apache 2.02025-01-3024Bthe return to Apache-2.0, and small enough to quantize onto one consumer card
Mistral Small 3.1Apache 2.02025-03-1724B, 128k contextmultimodal, and twenty-five languages on the card
Mistral Medium 3API / deployment2025-05-07not publishedpriced at $0.4 in, $2 out per million tokens
Devstral (Small 1.0)Apache 2.02025-05-2124B, 128k contextbuilt with All Hands AI; 46.8% on SWE-Bench Verified
Magistral Small / MediumSmall Apache 2.0; Medium: API preview, no licence published2025-06-10Small 24Btheir first reasoning model
Mistral Small 3.2Apache 2.02025-06 (card 2506)24B, 128k contextinstruction following, repetition, function calling — this is the one we run, served here at 32k rather than its full 128k
Mistral 3 (Large 3 + Ministral 3)Apache 2.02025-12-02Large 3: 675B total, 41B active (MoE); Ministral 3 at 14B, 8B, 3Ba frontier-size model released open-weight — and the Ministral 3 trio is the part of it that fits on one consumer card
Mistral Small 4Apache 2.02026-03-16119B total, 6B active per token (MoE, 128 experts), 256k contextthe newest Apache-2.0 release, and the one the API pricing page sells most cheaply
Mistral Medium 3.5Modified MIT2026-04-28 (card 26-04)128B dense, 256k contextone set of weights for instructions, reasoning and code; 77.6% on SWE-Bench Verified

Two words in that size column are not interchangeable. A dense model uses every parameter on every token; a mixture of experts routes each token through a fraction of them, so it costs the disk of the big number at nearer the speed of the small one. Q4_K_M, below, is the 4-bit quantisation most people run: a quarter the size of 16-bit weights, at a small cost in quality.

This table is the licence spine, not the catalogue: a search for "mistral" on the ollama.com library returned twenty entries on 2026-09-16, eleven of them Mistral AI's own. Le Chat, announced beside Mistral Large on 2024-02-26, now carries a rename notice on its own announcement page: "Le Chat is now Vibe—one agent and one licence across work and code, with every conversation, setting, and plan carried over." Read 2026-09-16 (UTC).

Read that licence column top to bottom and a shape appears: the models small enough to run at home stayed Apache-2.0, and the ones that needed a data centre are where the word "open" started carrying conditions.

What the company is worth

Five rounds in three years, each dated to the announcement that carried it.

rounddateraisedvaluationlead investor
seed2023-06-15$113M (€105M)$259M (€240M) — the report says "is now valued at", not post-moneyLightspeed Venture Partners, with Eric Schmidt, Xavier Niel and Bpifrance contributing
Series A2023-12-11$415M (€385M)roughly $2 billion, attributed to BloombergAndreessen Horowitz
Series B2024-06-11~$640M (€600M, equity and debt)$6 billion "following this funding round"General Catalyst
Series C2025-09-09$2.00B (€1.7B at the ECB rate of 1.1744 on 2025-09-09)$13.74B post-money (€11.7B at that rate)ASML Holding NV
Series D2026-09-08$3.48B (€3B at the ECB rate of 1.1614 on 2026-09-08)more than $24.39B post-money (€21B at that rate)Samsung Electronics, with Scaleup Europe Fund (managed by EQT) and PSG Equity as co-leads

Every figure here is in dollars: the first three rounds as their sources printed them, the C and the D converted from euros at the European Central Bank's reference rate for the announcement day, which each row carries. Only the last two say "post-money" in their source; the others keep their report's wording, because the difference matters and guessing is not reporting. Mistral's claim for the last row: "the largest equity fundraising round ever completed by a European technology company, three years after the company's launch." We have no view on the valuation.

Do they train models in French?

Non — ils entraînent sur des données multilingues où le français figure au premier rang. Not "in French", then, but on multilingual data in which French is first-class — the more checkable thing: the Mistral Small 3.1 model card lists twenty-five languages, English first.

The tokenizer is where that claim becomes measurable. Tekken, introduced with Mistral NeMo in July 2024, was trained on more than 100 languages and is "~30% more efficient at compressing source code, Chinese, Italian, French, German, Spanish, and Russian" than the SentencePiece tokenizer it replaced, and beat Llama 3's on about 85% of all languages. Fewer tokens for the same French sentence is less money and less waiting.

One caution, from our bench rather than theirs: those figures say Tekken beat its predecessor and Llama 3's, not that it is the densest tokenizer in the field. On one frozen English prompt, mistral-small3.2:24b read 1,023 tokens where gemma4:26b read 536 of the same bytes (two hours on battery) — double the prefill, and a bigger cache on every decode step after.

Mistral publishes no proportions, so anyone quoting a percentage is guessing.

Why is it called Mistral?

The wind first. Larousse gives one sense — "Vent violent, froid, turbulent et sec, qui souffle du secteur nord, sur la France méditerranéenne, entre les méridiens de Sète et de Toulon" — and one etymology, the joke the company has been telling ever since, printed as Larousse prints it: "ancien provençal maestral, du bas latin magistralis, magistrat". The reasoning model they shipped in 2025 is Magistral: that root, come back round.

The rest of the family is the wind's name with the job spliced into it: Mixtral for the mixture of experts, Codestral for code, Pixtral for vision, Devstral for software engineering, Voxtral for speech, and les Ministraux, the edge models, announced 2024-10-16 as "Un Ministral, des Ministraux" — a singular-and-plural joke landing on Mistral 7B's first anniversary.

The chat product never took the wind's name, and has given up its own: Le Chat is Vibe — neither French nor a wind.

What does it do on our own machines?

Nothing here was measured for this page; every row was published earlier. Seven of the nine readings are on mistral-small3.2:24b, the 24B dense Apache-2.0 release at Q4_K_M — two instruments, two figures: 14.14 GiB of weights in the model store, about 17 GiB of a card once it serves a 32,768-token window. That is why it fits a 24 GB card and not a 16 GB one. The last two rows name larger Mistral models.

what was measured, and on which modelreadingpower capwhen, and where to check it
writing speed on an RTX PRO 6000 Blackwell (96 GB), single warm calls, 128 tokens out95.0 / 94.6 / 93.5 tokens a second600 / 500 / 450 W2026-08-27 · what 150 watts buys
power and heat under its own load, same card594.6 W mean across its first four-minute leg, 87 °C peak across the eight minutes — the hottest arm of the serving bench it was read fromuncapped in that arm2026-08-25 · what 150 watts buys
what a power cap costs this model, same cardclock −12.0%, throughput −1.4% at one stream; −18.7% and −2.0% at sixteen600 W → 450 W2026-08-25 · what 150 watts buys
on processors alone, no graphics card7.1 tokens a second on a server, 2.7 on a laptop, 3.3 on a mini PC — and 97.7 s before the first word on the mininot a card2026-08-25 · two hours on battery, the two-hour machine priced
the same processors, threads tuned2.7 → 5.1 (1.90×) on the laptop; 3.3 → 4.0 (1.20×) on the mininot a card2026-08-25 · two hours on battery
on a used consumer card too small to hold it, half in RAM5.13 tokens a second split, against 2.71 on that box's processor alone — 1.89×cap not set by that harness2026-09-10 · this workshop's own bench record, gpu-3080-2026-09-10/RESULTS.md (not yet a published page)
what this model's tokenizer does to one fixed prompt1,023 tokens where a Gemma tokenizer read 536 of the same textnot a card2026-08-25 · two hours on battery
mistral-medium-3.5:128b, mistral-large-3 and a 23.6B q4 mistral-small seat sitting our grounded-judge trial — 43 sentences, 27 that must be killed as unsourced and 16 that must be left alonemedium-3.5 killed 27/27 and kept 15/16 — PASS; large-3 killed 27/27 and kept 15/16 — PASS; the small seat killed 27/27 but kept only 10/16 — FAIL, on preservation rather than on catchingnot a card measurement2026-08 · the seat trials
mistral-large-3 judging our work from outsidecarried its audition first time and judged 5 of 5 sealed roundsnot a card measurement2026-08-13 · outside judges

That trial is a different job from standing at a door: a grounding judge must leave true claims alone, and the small seat over-refused. A doorman answers one yes-or-no question about one short string — the task it does keep. Nothing here says a Mistral model failed a safety exam; no such exam was run.

On 2026-09-16 two cards sat in one box for the first time, and the cheap one has something to say.

What we measured on two cards in one box

A bench ran in the small hours of 2026-09-16 (UTC) that set a used RTX 3090 (24 GB) beside an RTX PRO 6000 Blackwell (96 GB) in one box: one kernel, one driver, one ollama build, one model store, one frozen prompt, the same weight files the live seats hold. Nothing is box against box. It was pre-registered, including the sentence that would have refuted it — if a 3090 reads within 15% of the workstation card on either model, that is the headline — which did not fire: 39% behind on one model, 63% on the other.

Every row: a 3090 at 250 W against the workstation card at 420 W, one frozen 2,101-byte prompt — 1,023 tokens to Mistral's tokenizer, 536 to Gemma's — 256 tokens out.

measuredRTX 3090 (24 GB) — a bench instanceRTX PRO 6000 Blackwell (96 GB) — the served seat
mistral-small3.2:24b, 24B dense, decode34.13 tokens a second92.77 tokens a second
gemma4:26b, 26B MoE, decode126.12 tokens a second205.85 tokens a second
watts while decoding, dense / MoE228.2 W / 206.4 W328.3 W / 182.5 W
energy per 1,000 tokens, dense / MoE6,680 J / 1,632 J3,538 J / 887 J
first token on a prompt it has already seen, dense / MoE98 ms / 167 ms28 ms / 162 ms

One file per column: m4-3090.json, m4-6000-live.json, m5-3090.json, m5-6000-live.json. The 96 GB column is a live serving seat: a second bench copy would not fit beside the services already on it — 12,965 MiB free against 16,645 MiB wanted (m4-6000-bench.json), as the pre-registration predicted. A 96 GB card in service is already full.

Read the dense row first; it explains the rest. Writing one token from a dense model means reading every weight in it, once — at Q4_K_M, 14.14 GiB per token. So 34.13 tokens a second is about 482 GiB/s of fetching: 55% of the 936 GB/s NVIDIA publishes for a GeForce RTX 3090, and at 350 W, cap out of the way, 50.02 tokens a second is 707 GiB/s — four fifths of it. Decode is a fetching problem, not an arithmetic one. A mixture of experts changes that: gemma4:26b routes each token through about 4 billion active parameters, a few gigabytes instead of fourteen, and on the same card, cap and prompt runs 3.7× faster on a model larger on disk.

The same arithmetic on the workstation card lands in the same place: 92.77 tokens a second is 1,311 GiB/s against the 1,792 GB/s NVIDIA publishes for it. Unthrottled, both cards read the dense model at about four fifths of their own memory bandwidth — 81% for a 3090 at its board default, 79% for a 96 GB card at 420 W. What separates them is how much bandwidth they have, not how well either uses it.

Two earlier benches on other boxes say it from below — controls, not like-for-like: on a 3080 split with system memory, 34.40 tokens a second for the MoE against 5.13 for the dense Mistral; on a mini PC's processors alone, 12.66 against 3.31.

What a power cap actually costs

Fetching is also what costs watts, which is why a 3090's two arms behave differently under one 250 W limit. The dense model sat at 228.2 W — 91% of cap — clock pulled to 495–675 MHz across three runs at 64–66 °C, nowhere near thermal (m4-3090.json). The MoE drew 206.4 W, 82.6% of the same cap, clock following the work (m5-3090.json). A dense model under a cap is power-bound; a sparse one is not yet. So the ladder moved the cap:

modelpower captokens a secondvs 250 Wenergy per 1,000 tokens
mistral-small3.2:24b (dense)250 W34.136,680 J
mistral-small3.2:24b (dense)300 W48.18+41.2%5,673 J
mistral-small3.2:24b (dense)350 W (the board's default)50.02+46.6%5,318 J
gemma4:26b (MoE)250 W126.121,632 J
gemma4:26b (MoE)350 W (the board's default)136.49+8.2%1,823 J

250 W rows: m4-3090.json, m5-3090.json. The rungs above: capladder-mistral-small3.2-24b.json, capladder-gemma4-26b.json.

Almost all the money is in the first fifty watts: 250 → 300 W buys the dense model 41%, the next fifty 3.8%; the MoE, never power-bound, gains 8.2% for a hundred. The cap on this card was ruled to 300 W minutes after that table was measured — the figure the dollars below use.

What VRAM is worth. qwen3.5:122b — a mixture of experts of 122 billion parameters, 10 billion active per token, 75.7 GiB of weights — fits on neither card, so it was split across both. Same caps: a 3090 at 250 W, the workstation card at 420 W.

what the cards could offeron cardtokens a secondsource
the workshop's seats up39.5%38.22q1-split-resident-seatsup.json
the seats stopped, inside one authorised window89.6%71.48q1-split-resident-seatsoff.json
the 96 GB card alone, seats stopped63.0%refused — does not fitq1-6000-whole.json

Freeing the seats' memory nearly doubled throughput — 1.87×, same model, same box, same minute, nothing else changed: 87% more tokens a second from the same silicon. This box holds 90% of a 122B model across two cards, and none of it whole on one.

A 3090 takes spillover, too: the MoE at one, two and four simultaneous streams produced 127.1 / 195.8 / 261.4 tokens a second aggregate at a 250 W cap, its fans at zero RPM until the fourth stream woke them to about half speed (concurrency-3090.json).

One arm is not reported, and the reason is the finding. The embedding arm voided itself against its own pre-registered gate: five batches spread 34.9% around their median where the gate allowed 15% (embed-3090.json). Half-second batches on a card that fast measure scheduling, not silicon — and a rung rejected by its own gate gets no ratio, so no number from it is here. It is owed a longer batch and a second run.

Is it cheaper to run it yourself?

Two of those numbers turn into money. The bench measures energy per thousand output tokens, so the arithmetic runs joules → kilowatt-hours per million → dollars at 18.34 cents a kilowatt-hour: the US residential average for June 2026 (Energy Information Administration, Electric Power Monthly, table 5.6.A), a national figure and not this workshop's bill.

  • mistral-small3.2:24b on a 3090 at its ruled 300 W cap: 5,673 J per 1,000 tokens (capladder-mistral-small3.2-24b.json) → 1.576 kWh per million → $0.29 per million.
  • gemma4:26b on the workstation card at 420 W: 887 J per 1,000 tokens (m5-6000-live.json) → 0.246 kWh per million → $0.045 per million.

Per token written the 96 GB card is the more frugal — 3,538 J against 6,680 dense, 887 against 1,632 MoE, roughly 1.9× and 1.8×.

Beside Mistral's own list prices, read from the API pricing page on 2026-09-16 (UTC):

what you are paying forinputoutput
Mistral Small 4, on their machines$0.15$0.6
Mistral Medium 3.5, on their machines$1.5$7.5
Mistral Large 3, on their machines$0.5$1.5
Ministral 3 (8B), on their machines$0.15$0.15
mistral-small3.2:24b on our own 3090 at a 300 W cap, electricity only$0.29
gemma4:26b on our own 96 GB card at 420 W, electricity only$0.045

Those are not the same kind of number, and the comparison is worth exactly what that caveat allows. The electricity figure is the watts that left the card while it wrote output tokens, and nothing else: no hardware, no cooling, no idle hours, no prompt-processing energy, nobody's time — and it is the whole card's draw while it decoded, floor included, of which about 86 W on the shared 96 GB seat belongs to the services already resident there. The API figure is a published list price for a managed service, which includes all of that and a margin. The pair is good for a floor: running a 24B model yourself is cents of electricity per million tokens, so the rest of the bill is the machine, not the meter.

One oddity in their own table: Mistral Large 3 is cheaper than Mistral Medium 3.5, $0.5 in and $1.5 out against $1.5 and $7.5. That is what the page said on 2026-09-16 (UTC); list prices move, so it is stamped, not asserted.

What the card cost. The used 3090 in this box was bought for $1,449.99 in May 2026. A card price is a reading, not a fact, so here is what the same search found on 2026-09-13, 06:02–06:05Z —

  • the same model, an EVGA XC3 Ultra, renewed at $1,849.99 — one Amazon seller, one unit;
  • the cheapest new listing, a ZOTAC Trinity OC at $1,829.99, with a "usually ships within 5 to 6 days" caveat;
  • an EVGA K|NGP|N Hybrid, new at $1,839.00;
  • eleven renewed cards of other makes, $1,529.99 to $1,799.99.

At what was paid, that is $11.50 per token a second on the MoE (126.12 tok/s) and $30.10 on the dense Mistral at 300 W (48.18 tok/s). No price is printed for the 96 GB workstation card: this workshop holds none of record, and an invented one is worse than a blank.

Where the doorman actually stands

In the print lab, Mistral Small 3.2 reads the word you type and answers one question — is this safe to paint and publish on a family board — before a picture is drawn. In the beat lab, it checks that a wish is a musical request and nothing else; the rules desk uses the same seat. It has stood at those doors since 2026-08-26.

The card whose printed name started this page still goes through it, and still comes back a rules question rather than a refusal.

What to take with you

  • They shipped the first one as a torrent. September 2023, 13.4 GB, no gate, no waitlist.
  • "Open" is a per-release word, not a company word. Apache-2.0 on what fits one card, conditions on what needs a data centre.
  • They do not train "in French." French is first-class in a multilingual mix, and their tokenizer read our English prompt at nearly twice Gemma's count.
  • $113M to more than $24B in three years, on the argument that giving the weights away is the business.
  • Sparse beats big, which is why a doorman is cheap. A dense 24B reads all 14.14 GiB of itself per token written; a same-size mixture of experts reads a fraction and runs 3.7× faster — so the door costs 34 tokens a second on one used card, cents of electricity per million, and nothing a stranger types leaves the building.

How to check our work

Every external link below was fetched and read on 2026-09-16 (UTC).

Who ran this, and thanks

Thanks to Mistral AI, whose announcements and model cards are dated, detailed and still up — which is why this page could be written from primary sources rather than from memory. The model itself is Mistral Small 3.2 (Mistral AI, Apache-2.0); every reading republished above was taken through Ollama (MIT) over llama.cpp (MIT) and CUDA (NVIDIA, proprietary, under the CUDA Toolkit EULA) — the one closed piece in the stack, and the one that comes with the cards — on 4-bit Q4_K_M weights packaged by the people who quantize and host them for everyone else. The comparison model in the tokenizer row and in the two-card bench is Gemma 4 (Google DeepMind): the licence blob this hub read off its own gemma4:26b tag on 2026-08-11 is stock Apache 2.0, and it is on our licences page. The 122B model is Qwen 3.5 (Qwen, Alibaba), whose model card publishes apache-2.0 — this shelf has never read the licence packaged with that ollama tag itself, and says so rather than implying a read it did not run. The electricity price is the U.S. Energy Information Administration's published residential average; the wind's definition is Larousse's and the word's history the Online Etymology Dictionary's; the two memory-bandwidth figures are NVIDIA's own published specifications. None of them owed us anything. A small human team asked for this page, chose what it would and would not claim, and signed the numbers; a fleet of AI agents fetched and read every source listed above, pulled the measured rows out of benches this hub had already published, and drafted it under that team's rulings.

Corrections and later measurements will be added below, each dated (UTC), with a window at both ends where one applies and saying in plain words what it counts.

elsewhere in the workshop

a strata→signal property · hello@strata2signal.com · say hello