{
 "schema_version": "1.0",
 "page_kind": "article",
 "pack_kind": "generate",
 "slug": "a-short-history-of-mistral",
 "title": "A Short History of Mistral",
 "dek": "A French lab that put its first model on a torrent, every open-weight release since with its licence, and what its 24B model does on a 3090 and a 96 GB workstation card.",
 "published": "2026-09-16",
 "series": [
  "notes"
 ],
 "licence": "CC BY 4.0",
 "status": "pending-judge",
 "pack_note": "no judge is seated, so every chip is held",
 "notice": "items with published:false are held and are not this page's published words",
 "source": {
  "url": "https://research.strata2signal.com/a-short-history-of-mistral/",
  "md_url": "https://research.strata2signal.com/a-short-history-of-mistral/index.md",
  "md_sha": "12249483ca5ec3b646944fef2d4b59810deb8d7013974d2f0511cf00b38614e8",
  "html_sha": "773583945afe373a4f710c05249aeb921c1b719ae578c8b7d0031f2910ccde0d",
  "anchors_sha": "f1e3995d67fabbd57dac7d1b571d1ae9ba3d4cb0fefa6ad78f794812606a6026"
 },
 "short": {
  "paragraph": "A 24-billion-parameter French model reads what strangers type at two of this workshop's doors, before anything else runs. This is the history of the company behind it - founded April 2023, first model out by torrent, open weights at almost every size since - and every number we have measured on it, including a bench that set a 24 GB card beside a 96 GB one.",
  "from_dek": false,
  "counts": {
   "words": 5259,
   "minutes": 24,
   "tables": 7,
   "kit": false
  },
  "bullets": [
   {
    "text": "the tokenizer is about 30%",
    "figure": "30%",
    "cite": "do-they-train-models-in-french",
    "quote": "Tekken, introduced with Mistral NeMo in July 2024, was trained on more than 100 languages and is \"~30% more efficient at compressing source code, Chinese, Italian, French, German, Spanish, and Russian\" than the SentencePiece tokenizer it replaced, and beat Llama 3's on about 85% of all languages.",
    "span": [
     10540,
     10543
    ]
   },
   {
    "text": "the mixture of experts runs 3.7×",
    "figure": "3.7×",
    "cite": "what-to-take-with-you",
    "quote": "A dense 24B reads all 14.14 GiB of itself per token written; a same-size mixture of experts reads a fraction and runs 3.7× faster - so the door costs 34 tokens a second on one used card, cents of electricity per million, and nothing a stranger types leaves the building.",
    "span": [
     26417,
     26422
    ]
   }
  ]
 },
 "sections": [
  {
   "id": "the-question-that-put-a-model-at-the-door",
   "heading": "The question that put a model at the door",
   "level": 2,
   "span": [
    1245,
    2041
   ],
   "chunks": [
    [
     1245,
     2041
    ]
   ],
   "chars": 796,
   "digest": "a rules question about a game ability whose name reads as a slur led to the use of a 24-billion-parameter language model. the model, mistral-small3.2:24b, reads the whole question and refuses the same words used the other way.",
   "digest_skipped": null
  },
  {
   "id": "who-are-these-people-and-where-did-they-come-from",
   "heading": "Who are these people, and where did they come from?",
   "level": 2,
   "span": [
    2041,
    3134
   ],
   "chunks": [
    [
     2041,
     3134
    ]
   ],
   "chars": 1093,
   "digest": "Mistral AI was founded in April 2023 in Paris. the founders are Arthur Mensch, Guillaume Lample, and Timothée Lacroix. Mensch researched at Google DeepMind, while Lample and Lacroix came from Meta and are authors on Meta's LLaMA paper.",
   "digest_skipped": null
  },
  {
   "id": "la-frontiere-what-they-said-they-were-for",
   "heading": "La frontière — what they said they were for",
   "level": 2,
   "span": [
    3134,
    4251
   ],
   "chunks": [
    [
     3134,
     4251
    ]
   ],
   "chars": 1117,
   "digest": "the mission of Mistral AI is to make frontier AI open to all and solve the world's hardest problems. the Series D frames it as sovereignty regarding data, models and compute. they use the Apache-2.0 license for general purpose models.",
   "digest_skipped": null
  },
  {
   "id": "what-have-they-shipped-and-under-what-licence",
   "heading": "What have they shipped, and under what licence?",
   "level": 2,
   "span": [
    4251,
    7905
   ],
   "chunks": [
    [
     4251,
     7905
    ]
   ],
   "chars": 3654,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "what-the-company-is-worth",
   "heading": "What the company is worth",
   "level": 2,
   "span": [
    7905,
    9931
   ],
   "chunks": [
    [
     7905,
     9931
    ]
   ],
   "chars": 2026,
   "digest": "Mistral AI has completed five rounds in three years. the Series D, dated 2026-09-08, raised $3.48B (€3B at the ECB rate of 1.1614 on 2026-09-08) with a post-money valuation of more than $24.39B (€21B at that rate).",
   "digest_skipped": null
  },
  {
   "id": "do-they-train-models-in-french",
   "heading": "Do they train models in French?",
   "level": 2,
   "span": [
    9931,
    11297
   ],
   "chunks": [
    [
     9931,
     11297
    ]
   ],
   "chars": 1366,
   "digest": "models are trained on multilingual data where French is first-class. the Tekken tokenizer is ~30% more efficient at compressing certain languages than SentencePiece. on one frozen English prompt, mistral-small3.2:24b read 1,023 tokens where gemma4:26b read 536.",
   "digest_skipped": null
  },
  {
   "id": "why-is-it-called-mistral",
   "heading": "Why is it called Mistral?",
   "level": 2,
   "span": [
    11297,
    12628
   ],
   "chunks": [
    [
     11297,
     12628
    ]
   ],
   "chars": 1331,
   "digest": "the name comes from the wind. the family includes Mixtral for mixture of experts, Codestral for code, Pixtral for vision, Devstral for software engineering, Voxtral for speech, and les Ministraux for edge models. Le Chat is now Vibe.",
   "digest_skipped": null
  },
  {
   "id": "what-does-it-do-on-our-own-machines",
   "heading": "What does it do on our own machines?",
   "level": 2,
   "span": [
    12628,
    15922
   ],
   "chunks": [
    [
     12628,
     15922
    ]
   ],
   "chars": 3294,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "what-we-measured-on-two-cards-in-one-box",
   "heading": "What we measured on two cards in one box",
   "level": 2,
   "span": [
    15922,
    18960
   ],
   "chunks": [
    [
     15922,
     18960
    ]
   ],
   "chars": 3038,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "what-a-power-cap-actually-costs",
   "heading": "What a power cap actually costs",
   "level": 2,
   "span": [
    18960,
    21905
   ],
   "chunks": [
    [
     18960,
     21905
    ]
   ],
   "chars": 2945,
   "digest": "a 3090's behavior changes under a 250 W limit. for mistral-small3.2:24b, a 300 W cap provides 48.18 tokens a second, which is +41.2% compared to 250 W. the 96 GB card can hold 90% of a 122B model across two cards.",
   "digest_skipped": null
  },
  {
   "id": "is-it-cheaper-to-run-it-yourself",
   "heading": "Is it cheaper to run it yourself?",
   "level": 2,
   "span": [
    21905,
    25112
   ],
   "chunks": [
    [
     21905,
     25112
    ]
   ],
   "chars": 3207,
   "digest": "running mistral-small3.2:24b on a 3090 at a 300 W cap costs $0.29 per million tokens for electricity. Mistral Large 3 is priced at $0.5 in and $1.5 out on their machines. the used 3090 was bought for $1,449.99 in May 2026.",
   "digest_skipped": null
  },
  {
   "id": "where-the-doorman-actually-stands",
   "heading": "Where the doorman actually stands",
   "level": 2,
   "span": [
    25112,
    25703
   ],
   "chunks": [
    [
     25112,
     25703
    ]
   ],
   "chars": 591,
   "digest": "in the print lab, beat lab, and rules desk, Mistral Small 3.2 reads words and answers questions. the card whose name started the page still goes through it and returns a rules question.",
   "digest_skipped": null
  },
  {
   "id": "what-to-take-with-you",
   "heading": "What to take with you",
   "level": 2,
   "span": [
    25703,
    26572
   ],
   "chunks": [
    [
     25703,
     26572
    ]
   ],
   "chars": 869,
   "digest": "Mistral AI shipped its first model as a 13.4 GB torrent. open weights are provided for models that fit on one card. sparse models like mixture of experts run 3.7× faster than dense models.",
   "digest_skipped": null
  },
  {
   "id": "how-to-check-our-work",
   "heading": "How to check our work",
   "level": 2,
   "span": [
    26572,
    35312
   ],
   "chunks": [
    [
     26572,
     35312
    ]
   ],
   "chars": 8740,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "who-ran-this-and-thanks",
   "heading": "Who ran this, and thanks",
   "level": 2,
   "span": [
    35312,
    37170
   ],
   "chunks": [
    [
     35312,
     37170
    ]
   ],
   "chars": 1858,
   "digest": null,
   "digest_skipped": "credits"
  }
 ],
 "chips": [
  {
   "id": "c-789bcb2c",
   "q": "how does a language model handle sensitive words in a rules question?",
   "a": "A 24-billion-parameter language model reads the entire question to allow legitimate rules queries through while refusing specific words when used as slurs.",
   "cites": [
    "the-question-that-put-a-model-at-the-door"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-1a518f14",
   "q": "what is the stated mission of mistral ai?",
   "a": "The mission is to make frontier AI open to all and to work together to solve the world's hardest problems.",
   "cites": [
    "la-frontiere-what-they-said-they-were-for"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-cded6462",
   "q": "how does the company's valuation change over its funding rounds?",
   "a": "The company grew from a $259M valuation at its seed round to more than $24.39B post-money following its Series D round.",
   "cites": [
    "what-the-company-is-worth"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-1529c0af",
   "q": "how does the tekken tokenizer compare to previous versions?",
   "a": "The tokenizer is approximately 30% more efficient at compressing source code, Chinese, Italian, French, German, Spanish, and Russian than the SentencePiece tokenizer it replaced.",
   "cites": [
    "do-they-train-models-in-french"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-2da0a31a",
   "q": "what is the etymological origin of the name mistral?",
   "a": "The name comes from the ancient Provençal maestral, derived from the Latin magistralis, meaning magistrate.",
   "cites": [
    "why-is-it-called-mistral"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-392c9d49",
   "q": "how does a mixture of experts model differ from a dense model?",
   "a": "A dense model uses every parameter for every token, whereas a mixture of experts routes each token through a fraction of the parameters, allowing it to run faster.",
   "cites": [
    "what-we-measured-on-two-cards-in-one-box"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-67ed85a3",
   "q": "how does a power cap affect the performance of a dense model?",
   "a": "A dense model is power-bound; increasing the cap from 250 W to 350 W can increase throughput from 34.13 to 50.02 tokens a second.",
   "cites": [
    "what-a-power-cap-actually-costs"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-e3648745",
   "q": "what is the electricity cost for running mistral small 3.2?",
   "a": "Running the model on a 3090 at a 300 W cap costs approximately $0.29 per million tokens based on electricity alone.",
   "cites": [
    "is-it-cheaper-to-run-it-yourself"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  }
 ],
 "related": [
  {
   "slug": "three-librarians",
   "why": "shares ground with § Where the doorman actually stands · § The doorman stands off this road entirely"
  },
  {
   "slug": "home-inference-and-diffusion-rig-cost",
   "why": "shares ground with § What to take with you · § What this rig is for, in plain words"
  },
  {
   "slug": "ask-about-this-page",
   "why": "shares ground with § Where the doorman actually stands · § The room behind every article, and what it opens"
  }
 ],
 "thanks": "## Who ran this, and thanks {#who-ran-this-and-thanks}\n\nThanks to Mistral AI, whose announcements and model cards are dated, detailed and still up — which is why this page could be written from primary sources rather than from memory. The model itself is **Mistral Small 3.2** (Mistral AI, Apache-2.0); every reading republished above was taken through **Ollama** (MIT) over **llama.cpp** (MIT) and **CUDA** (NVIDIA, proprietary, under the CUDA Toolkit EULA) — the one closed piece in the stack, and the one that comes with the cards — on 4-bit `Q4_K_M` weights packaged by the people who quantize and host them for everyone else. The comparison model in the tokenizer row and in the two-card bench is **Gemma 4** (Google DeepMind): the licence blob this hub read off its own `gemma4:26b` tag on 2026-08-11 is stock Apache 2.0, and it is on our [licences](/licences/) page. The 122B model is **Qwen 3.5** (Qwen, Alibaba), whose model card publishes apache-2.0 — this shelf has never read the licence packaged with that ollama tag itself, and says so rather than implying a read it did not run. The electricity price is the **U.S. Energy Information Administration's** published residential average; the wind's definition is **Larousse's** and the word's history the **Online Etymology Dictionary's**; the two memory-bandwidth figures are **NVIDIA's** own published specifications. None of them owed us anything. A small human team asked for this page, chose what it would and would not claim, and signed the numbers; a fleet of AI agents fetched and read every source listed above, pulled the measured rows out of benches this hub had already published, and drafted it under that team's rulings.\n\nCorrections and later measurements will be added below, each dated (UTC), with a window at both ends where one applies and saying in plain words what it counts.",
 "kit": null,
 "seat_class": {
  "writer": "gemma-class",
  "judge": null,
  "writer_runtime": "vllm"
 },
 "bench": {
  "bullets_written": 3,
  "bullets_kept": 2,
  "digests_written": 12,
  "digests_kept": 10,
  "chips_written": 8,
  "chips_kept": 8,
  "chips_grounded": 0,
  "chips_published": 0,
  "dropped_by": {
   "new_noun": 0,
   "figure": 0,
   "length": 0,
   "cite": 0,
   "judge": 0,
   "quote": 0,
   "directive": 0,
   "redaction": 2,
   "profanity": 0,
   "identifier": 1
  },
  "new_noun_tokens_checked": [
   "0.29",
   "0.5",
   "023",
   "1",
   "1.1614",
   "1.5",
   "122B",
   "13.4",
   "14.14",
   "2023",
   "2026",
   "2026-09-08",
   "21B",
   "24",
   "24-billion-parameter",
   "24.39B",
   "24b",
   "250",
   "259M",
   "26b",
   "3",
   "3.2",
   "3.48B",
   "3.7×",
   "30%",
   "300",
   "3090",
   "3090's",
   "34.13",
   "350",
   "3B",
   "41.2%",
   "449.99",
   "48.18",
   "50.02",
   "536",
   "7.1",
   "90%",
   "92.77",
   "93.5",
   "94.6",
   "95.0",
   "96",
   "AI",
   "Apache-2.0",
   "April",
   "Arthur",
   "Chat",
   "Chinese",
   "Codestral",
   "D",
   "DeepMind",
   "Devstral",
   "ECB",
   "English",
   "French",
   "GB",
   "German",
   "GiB",
   "Google",
   "Guillaume",
   "Italian",
   "LLaMA",
   "Lacroix",
   "Lample",
   "Large",
   "Latin",
   "May",
   "Mensch",
   "Meta",
   "Meta's",
   "Ministraux",
   "Mistral",
   "Mixtral",
   "PRO",
   "Paris",
   "Pixtral",
   "Proven",
   "Q4KM",
   "Russian",
   "SentencePiece",
   "Series",
   "Small",
   "Spanish",
   "Tekken",
   "Timoth",
   "Vibe",
   "Voxtral",
   "W",
   "gemma4",
   "mistral-small3.2"
  ],
  "new_noun_tokens_withheld": 8,
  "figure_definition": "v3",
  "drops": [
   {
    "kind": "digest",
    "reason": "redaction",
    "item": null,
    "item_chars": 206,
    "withheld": "a string carrying a fenced literal",
    "detail": ""
   },
   {
    "kind": "digest",
    "reason": "redaction",
    "item": null,
    "item_chars": 218,
    "withheld": "a string carrying a fenced literal",
    "detail": ""
   },
   {
    "kind": "bullet",
    "reason": "identifier",
    "item": "the Series D round raised $3.48B",
    "detail": "3.48b is in the cited section only as code or in a table cell"
   }
  ]
 },
 "generated_utc": "2026-09-16T09:12:03Z"
}
