senn-techsenn-tech
AI & Development
AI & Development2026-10-07· By Franz Senn

Mistral Large 4 among the open weights: index scores, prices and the bytes behind them

Mistral has announced Mistral Large 4, a mixture-of-experts model with around a trillion parameters, and has promised the weights by the end of October. The heise coverage frames it as catching up with the frontier. We checked two things: what an independent measurement says about quality, and what the weights will weigh once they exist. For that we measured the checkpoint size of every comparison model through the Hugging Face API and benchmarked our own inference lane under load. The index score is the only quality number in this story that does not come from the vendor.

Mistral AI brand mark, stepped bars in yellow, orange and red
Mistral calls the model le Chonk in its own announcement, ML4 informally. The brand graphic comes from the Commons collection, every vendor figure in this post from the official announcement and the model card, both retrieved on 7 October 2026. (Quelle: Wikimedia Commons, Datei Mistral AI logo (2025–).svg)

What Mistral states and what was measured

The official announcement is titled Introducing Mistral Large 4 and ships with the vendor's own benchmark charts. It lists the model at roughly 1,000 billion parameters and 49 billion active, the model card at 1.05 trillion and 52 billion. Training is described as from scratch on 3,800 Grace Blackwell in Mistral's own European data centres, with a reinforcement run on 3,000 GPUs producing 33 billion tokens a day, about 16 billion of them trainable. A substantial share of the training data is called multilingual, spanning more than 160 languages including every official language of the European Union.

On coding the page names DeepSWE v1.1 at 61.7 percent, SWE-Atlas-QnA at 59.4 percent and Terminal-Bench 4 at 28.3 percent, combined into a coding agent index of 49.8 percent, said to pass DeepSeek V4 Pro of 13 August and Qwen3.8 Max. Lakera reports 93.3 percent at safety level B3, the KORA assessment 1.691 out of 2. In the Artificial Analysis cyber index Mistral wants to sit among the top five models worldwide, on 82 percent for reproducing and patching a real vulnerability in open source software. The reason given for that gap is the most remarkable sentence on the page: Claude Opus 5.5 and GPT-6 Astra landed near zero on the same test because they refuse the work.

Mistral draws the comparison with China itself. In Surge AI's blind human evaluation on real coding tasks Large 4 scored 3.74 out of 5 on average, second of five: ahead of Kimi K3 at 3.59 and GLM-5.3 at 3.60, behind Claude Opus 5 at 4.22. On AutomationBench with its 657 business processes the page names 59.9 percent and puts Kimi K3, MiMo-V2.6-Pro and DeepSeek V4 Pro behind it. On visual grounding (Dense 200) it reports 42 percent against 41 percent for GPT-6 Astra. All of it vendor figures with vendor charts, none of which we can recompute. They are quoted because they name the quantities we can measure: index, price and bytes.

The Artificial Analysis numbers, as of 7 October 2026, are another matter. The Large 4 preview scores 38 on Intelligence Index v4.3.2, rank 64 out of 225 models in its comparison class, where the class median is 26. Output speed is 116.1 tokens per second, time to first token 1.46 seconds. Running the index produced 200 million output tokens against a class median of 81 million, and Artificial Analysis turns that into 1.13 dollars per task. The context window reads 524,000 tokens measured against one million advertised. The model is listed as proprietary, because the weights are not up yet.

Two price levels exist: promotional prices of 0.68 dollars per million input and 2.09 dollars per million output tokens for a limited time, list prices of 1.36 and 4.18 dollars after that. The comparison below works with list prices.

Intelligence Index, one scaleMiMo-V2.6-Pro46 · Xiaomi, MITGLM-5.345 · Z.ai, restrictedKimi K344 · Moonshot, restrictedDeepSeek V4.1 Flash39 · MITMistral Large 4 (preview)38 · weights missingQwen3.8 27B34 · Apache 2.0Mistral Medium 3.514 · 128B, modified MITMistral Small 411 · 119B, Apache 2.0052
Artificial Analysis Intelligence Index v4.3.2, each value read from its own model page on 7 October 2026. Large 4 sits in the proprietary comparison class, whose median is 26; the small open class has a median of 8. (Quelle: Artificial Analysis)

Four Chinese models on the same scale

Every row below comes from a separate model page at Artificial Analysis, all retrieved on 7 October 2026.

ModelIndexTokens/s$/M input$/M output$ per taskContextParameters total/activeLicence
MiMo-V2.6-Pro (Xiaomi)4646.50.430.870.131M1.0T / 42BMIT
GLM-5.3 (Z.ai)4577.21.404.402.011M753B / 40Bown, restricted
Kimi K3 (Moonshot)4443.73.0015.002.001M2.8T / 104Bown, restricted
DeepSeek V4.1 Flash39227.90.301.200.271M552B / 16BMIT
Mistral Large 4 (preview)38116.11.364.181.13524k1.05T / 52Bnone yet
Qwen3.8 27B (Alibaba)3447.10.503.001.01256k27BApache 2.0
Mistral Medium 3.514170.01.507.500.50256k128Bmodified MIT
Mistral Small 411171.30.150.600.02256k119B / 6.5BApache 2.0

Two readings matter. MiMo-V2.6-Pro sits in the same size class as Large 4, carries 1.0 trillion parameters and an MIT licence, and leads the index by eight points. Per index task it costs 0.13 dollars against 1.13 dollars, roughly one ninth. And DeepSeek V4.1 Flash reaches almost the same index score as the Large 4 preview while serving at twice the output speed for a fraction of the cost per task. Mistral's own mid-sized and small models, Medium 3.5 and Small 4, are fast and cheap and sit far behind on the index.

Xiaomi brand mark
MiMo-V2.6-Pro is a Xiaomi model: 1.0 trillion parameters, MIT licence, index 46 and 0.13 dollars per task. The leader of this scale comes out of China and is free to use. (Quelle: Simple Icons, Datei xiaomi.svg)
DeepSeek logo
DeepSeek V4.1 Flash holds 39 index points at 227.9 tokens per second, 510 GB of weights and an MIT licence. That is the model we could put into our own infrastructure without a licence dispute or a vendor lock. (Quelle: RoadMaster19 / Wikimedia, CC BY 4.0)
Cost per index task in US centsMistral Small 42MiMo-V2.6-Pro13DeepSeek V4.1 Flash27Mistral Medium 3.550Qwen3.8 27B101Mistral Large 4113 · previewKimi K3200GLM-5.32010240
Weighted cost per Intelligence Index task in US cents, lower is better. Artificial Analysis sums input, cache, thinking and answer tokens from each price list across the ten sub-evaluations. (Quelle: Artificial Analysis)
LMArena Code Arena chart: cloud of models, a green efficiency frontier, Mistral Large 4 at score 1534 and 3.47 dollars
Screenshot from LMArena's Code Arena, retrieved 7 October 2026. Large 4 sits at 1534 points and a blended 3.47 dollars per million tokens at a 3-to-1 ratio, listed as proprietary. The green line is the efficiency frontier across all models, and to sit below it means another model is both cheaper and better. That is exactly where Large 4 is, while mimo-v2-6-flash and glm-5-3-flash sit on the line at the top right, at roughly a sixteenth of the price as read off the logarithmic axis. Those are the flash variants, not the flagship models in the table above. The score axis is a blind vote on real coding tasks, so it is a different yardstick from the Intelligence Index. (Quelle: Screenshot: LMArena, Code Arena)

The arena also exposes a blurring we can recompute ourselves. Artificial Analysis lists 1.36 dollars per million input and 4.18 dollars per million output tokens for the preview, which blended at 3 to 1 gives 2.07 dollars. The arena computes with 3.47 dollars. Neither chart shows which token mix each one assumes.

Open is not the same as free to use

The licence text decides whether a model may enter a customer project. DeepSeek V4.1 Flash and MiMo-V2.6-Pro are MIT, Qwen3.8 and Mistral Small 4 are Apache 2.0. For GLM-5.3 and Kimi K3, Artificial Analysis lists bespoke licences with commercial restrictions, and the same for the modified MIT of Mistral Medium 3.5. Large 4 has no licence file today, only a promise that weights are coming. Read the file whenever you hear the word open. Two of the four Chinese flagships put real intent into those clauses, and Mistral's own 128 billion model follows the same pattern.

What the weights actually weigh

Feasibility is decided in GPU memory. We measured the size of every release repository through the Hugging Face API: the sum of all file sizes in the repository as offered for download, as of 7 October 2026. That is neither a compressed image nor an estimate, it is the bytes someone would put on our disk.

Checkpoint size in GB, measuredKimi K31561 · 2.8T parametersGLM-5.3756Mistral Large 3 675B682 · 1 byte per paramMiMo-V2.6-Pro574 · RL releaseDeepSeek V4.1 Flash510Large 3 as NVFP4403 · 4-bitour four GPUs137 · 127.4 GiBSmall 4 as NVFP471 · full 242 GBQwen3.8 27B INT421 · full 56 GB01700
Our own measurement on 7 October 2026 via the Hugging Face API, summing all file sizes in each repository. All sizes in GB (10^9 bytes) as the interface reports them. The green row is our own GPU estate at 127.4 GiB, which is 137 GB. (Quelle: Hugging Face Hub API)

The green row is where our planning stops. Kimi K3 needs eleven times our GPU memory, GLM-5.3 five times, DeepSeek V4.1 Flash three and a half times. For the predecessor Large 3 we measured both variants: 682 GB as shipped and 403 GB as the NVFP4 quantisation, still three times our box. Large 4 has nothing to measure yet, so we do arithmetic: 1.05 trillion parameters at 0.5 bytes per parameter gives 525 GB, and scaling the predecessor's measured NVFP4 layout gives 627 GB. Both are conclusions from measured sizes rather than measurements, and the range sits at about four times our estate.

What runs on our machines

We operate four RTX 5090 in one host with 91 GB of system memory, 127.4 GiB of GPU memory in total, measured at 32,607 MiB per card. It serves a mixture-of-experts model with 48 layers and 512 experts per layer, a native context window of 262,144 tokens capped to 220,000 by the server, and the KV cache in NVFP4. The build is a pinned vLLM revision with local patches, because the NVFP4 path has an open defect upstream on our GPU generation. The index of that checkpoint declares 182.7 GB of weight tensors while the four cards offer 137 GB, and the reason it runs anyway is weight and KV quantisation. A second host with a single card serves as fallback: a 27 billion parameter model as AWQ-INT4, whose repository measures 21.0 GB.

Four requests with short prompts and 200 output tokens each, the box sitting at 76 to 84 percent utilisation: a median of 70.8 tokens per second in decode, range 66.7 to 75.4, and a median of 0.25 seconds to the first token. Artificial Analysis measures output speed through vendor APIs, we measured single requests on a busy box. The numbers are not directly comparable, the order of magnitude is: our lane at 70.8 tokens per second lands between GLM-5.3 (77.2) and Qwen3.8 27B (47.1). The four vendor APIs we captured for time to first token lie between 1.46 seconds for Large 4 and 5.56 seconds for MiMo-V2.6-Pro, on those same model pages. The short path without a provider queue is our real advantage rather than the compute.

From paper into our machinesWeights there?licence file and byte sumFits the GPUs?137 GB limit, 4x 5090Own test rundecode, TTFT, replay
The check we run before switching a lane. Step two is the hard cut for models of this size: anything above 137 GB is out of scope for self-hosting, whatever its quality. (Quelle: Own measurements and operating practice, October 2026)

Capital and origin

Mistral frames Large 4 as the first milestone on a product road paid for by that Series D of 3 billion euros, and calls the round the largest equity round a European technology company has ever closed. It was closed in early September 2026 at a post-money valuation above 21 billion euros. Samsung Electronics led it, alongside EQT Scaleup Europe Fund and PSG Equity, with Nvidia and ASML reported as participants by several outlets. European are the seat and part of the investor base.

Samsung wordmark
Samsung Electronics leads the round that pays for this model. Part of the picture of independent European AI is that the semiconductor suppliers in that same round are a South Korean and a US company. (Quelle: Datei via Wikimedia Commons)

The model is trained on hardware that is not on sale in Europe in that configuration, and the price competition arrives from China.

What we do with this

When the weights land at the end of October, we check in this order: open and read the licence file, measure the repository byte sum, then run a test pass on our box with tasks from real operations. Until then Large 4 appears in none of our pricing or architecture decisions, because neither the licence nor the bytes are known and the preview's index score sits behind the free MIT models. Anyone looking for an open model for a customer project today finds it in the Apache and MIT class up to 40 billion parameters, which runs on our own hardware, or in the API of an MIT model that has passed a data protection review. The open question is whether the context window stays at one million tokens after verification or at the measured 524,000.

Further reading

Questions?
Is Mistral Large 4 the best European open model at launch?+

Not by the independent Artificial Analysis index. The preview scores 38 points, MiMo-V2.6-Pro scores 46, GLM-5.3 scores 45, Kimi K3 scores 44 and DeepSeek V4.1 Flash scores 39. The Large 4 weights are not published either; the announced date is the end of October 2026. Scores retrieved on 7 October 2026.

What does one task cost with Mistral Large 4?+

Artificial Analysis calculates 1.13 US dollars per Intelligence Index task for the Large 4 preview. MiMo-V2.6-Pro is at 0.13 dollars, DeepSeek V4.1 Flash at 0.27 dollars, GLM-5.3 and Kimi K3 at about 2 dollars each. Large 4 output tokens cost 2.09 dollars per million on promotion and 4.18 dollars at list price.

Could we run Mistral Large 4 on our own hardware?+

Not according to our arithmetic. The model card names 1.05 trillion parameters. At 0.5 bytes per parameter that is 525 GB, and scaled to the measured NVFP4 layout of its predecessor it is 627 GB. Our four RTX 5090 provide 127.4 GiB in total, which is 137 GB. The model needs at least three and a half times our GPU memory. Figures as of 7 October 2026.