Mistral Large 4 among the open weights: index scores, prices and the bytes behind them
Mistral has announced Mistral Large 4, a mixture-of-experts model with around a trillion parameters, and has promised the weights by the end of October. The heise coverage frames it as catching up with the frontier. We checked two things: what an independent measurement says about quality, and what the weights will weigh once they exist. For that we measured the checkpoint size of every comparison model through the Hugging Face API and benchmarked our own inference lane under load. The index score is the only quality number in this story that does not come from the vendor.
What Mistral states and what was measured
The official announcement is titled Introducing Mistral Large 4 and ships with the vendor's own benchmark charts. It lists the model at roughly 1,000 billion parameters and 49 billion active, the model card at 1.05 trillion and 52 billion. Training is described as from scratch on 3,800 Grace Blackwell in Mistral's own European data centres, with a reinforcement run on 3,000 GPUs producing 33 billion tokens a day, about 16 billion of them trainable. A substantial share of the training data is called multilingual, spanning more than 160 languages including every official language of the European Union.
On coding the page names DeepSWE v1.1 at 61.7 percent, SWE-Atlas-QnA at 59.4 percent and Terminal-Bench 4 at 28.3 percent, combined into a coding agent index of 49.8 percent, said to pass DeepSeek V4 Pro of 13 August and Qwen3.8 Max. Lakera reports 93.3 percent at safety level B3, the KORA assessment 1.691 out of 2. In the Artificial Analysis cyber index Mistral wants to sit among the top five models worldwide, on 82 percent for reproducing and patching a real vulnerability in open source software. The reason given for that gap is the most remarkable sentence on the page: Claude Opus 5.5 and GPT-6 Astra landed near zero on the same test because they refuse the work.
Mistral draws the comparison with China itself. In Surge AI's blind human evaluation on real coding tasks Large 4 scored 3.74 out of 5 on average, second of five: ahead of Kimi K3 at 3.59 and GLM-5.3 at 3.60, behind Claude Opus 5 at 4.22. On AutomationBench with its 657 business processes the page names 59.9 percent and puts Kimi K3, MiMo-V2.6-Pro and DeepSeek V4 Pro behind it. On visual grounding (Dense 200) it reports 42 percent against 41 percent for GPT-6 Astra. All of it vendor figures with vendor charts, none of which we can recompute. They are quoted because they name the quantities we can measure: index, price and bytes.
The Artificial Analysis numbers, as of 7 October 2026, are another matter. The Large 4 preview scores 38 on Intelligence Index v4.3.2, rank 64 out of 225 models in its comparison class, where the class median is 26. Output speed is 116.1 tokens per second, time to first token 1.46 seconds. Running the index produced 200 million output tokens against a class median of 81 million, and Artificial Analysis turns that into 1.13 dollars per task. The context window reads 524,000 tokens measured against one million advertised. The model is listed as proprietary, because the weights are not up yet.
Two price levels exist: promotional prices of 0.68 dollars per million input and 2.09 dollars per million output tokens for a limited time, list prices of 1.36 and 4.18 dollars after that. The comparison below works with list prices.
Four Chinese models on the same scale
Every row below comes from a separate model page at Artificial Analysis, all retrieved on 7 October 2026.
| Model | Index | Tokens/s | $/M input | $/M output | $ per task | Context | Parameters total/active | Licence |
|---|---|---|---|---|---|---|---|---|
| MiMo-V2.6-Pro (Xiaomi) | 46 | 46.5 | 0.43 | 0.87 | 0.13 | 1M | 1.0T / 42B | MIT |
| GLM-5.3 (Z.ai) | 45 | 77.2 | 1.40 | 4.40 | 2.01 | 1M | 753B / 40B | own, restricted |
| Kimi K3 (Moonshot) | 44 | 43.7 | 3.00 | 15.00 | 2.00 | 1M | 2.8T / 104B | own, restricted |
| DeepSeek V4.1 Flash | 39 | 227.9 | 0.30 | 1.20 | 0.27 | 1M | 552B / 16B | MIT |
| Mistral Large 4 (preview) | 38 | 116.1 | 1.36 | 4.18 | 1.13 | 524k | 1.05T / 52B | none yet |
| Qwen3.8 27B (Alibaba) | 34 | 47.1 | 0.50 | 3.00 | 1.01 | 256k | 27B | Apache 2.0 |
| Mistral Medium 3.5 | 14 | 170.0 | 1.50 | 7.50 | 0.50 | 256k | 128B | modified MIT |
| Mistral Small 4 | 11 | 171.3 | 0.15 | 0.60 | 0.02 | 256k | 119B / 6.5B | Apache 2.0 |
Two readings matter. MiMo-V2.6-Pro sits in the same size class as Large 4, carries 1.0 trillion parameters and an MIT licence, and leads the index by eight points. Per index task it costs 0.13 dollars against 1.13 dollars, roughly one ninth. And DeepSeek V4.1 Flash reaches almost the same index score as the Large 4 preview while serving at twice the output speed for a fraction of the cost per task. Mistral's own mid-sized and small models, Medium 3.5 and Small 4, are fast and cheap and sit far behind on the index.


The arena also exposes a blurring we can recompute ourselves. Artificial Analysis lists 1.36 dollars per million input and 4.18 dollars per million output tokens for the preview, which blended at 3 to 1 gives 2.07 dollars. The arena computes with 3.47 dollars. Neither chart shows which token mix each one assumes.
Open is not the same as free to use
The licence text decides whether a model may enter a customer project. DeepSeek V4.1 Flash and MiMo-V2.6-Pro are MIT, Qwen3.8 and Mistral Small 4 are Apache 2.0. For GLM-5.3 and Kimi K3, Artificial Analysis lists bespoke licences with commercial restrictions, and the same for the modified MIT of Mistral Medium 3.5. Large 4 has no licence file today, only a promise that weights are coming. Read the file whenever you hear the word open. Two of the four Chinese flagships put real intent into those clauses, and Mistral's own 128 billion model follows the same pattern.
What the weights actually weigh
Feasibility is decided in GPU memory. We measured the size of every release repository through the Hugging Face API: the sum of all file sizes in the repository as offered for download, as of 7 October 2026. That is neither a compressed image nor an estimate, it is the bytes someone would put on our disk.
The green row is where our planning stops. Kimi K3 needs eleven times our GPU memory, GLM-5.3 five times, DeepSeek V4.1 Flash three and a half times. For the predecessor Large 3 we measured both variants: 682 GB as shipped and 403 GB as the NVFP4 quantisation, still three times our box. Large 4 has nothing to measure yet, so we do arithmetic: 1.05 trillion parameters at 0.5 bytes per parameter gives 525 GB, and scaling the predecessor's measured NVFP4 layout gives 627 GB. Both are conclusions from measured sizes rather than measurements, and the range sits at about four times our estate.
What runs on our machines
We operate four RTX 5090 in one host with 91 GB of system memory, 127.4 GiB of GPU memory in total, measured at 32,607 MiB per card. It serves a mixture-of-experts model with 48 layers and 512 experts per layer, a native context window of 262,144 tokens capped to 220,000 by the server, and the KV cache in NVFP4. The build is a pinned vLLM revision with local patches, because the NVFP4 path has an open defect upstream on our GPU generation. The index of that checkpoint declares 182.7 GB of weight tensors while the four cards offer 137 GB, and the reason it runs anyway is weight and KV quantisation. A second host with a single card serves as fallback: a 27 billion parameter model as AWQ-INT4, whose repository measures 21.0 GB.
Four requests with short prompts and 200 output tokens each, the box sitting at 76 to 84 percent utilisation: a median of 70.8 tokens per second in decode, range 66.7 to 75.4, and a median of 0.25 seconds to the first token. Artificial Analysis measures output speed through vendor APIs, we measured single requests on a busy box. The numbers are not directly comparable, the order of magnitude is: our lane at 70.8 tokens per second lands between GLM-5.3 (77.2) and Qwen3.8 27B (47.1). The four vendor APIs we captured for time to first token lie between 1.46 seconds for Large 4 and 5.56 seconds for MiMo-V2.6-Pro, on those same model pages. The short path without a provider queue is our real advantage rather than the compute.
Capital and origin
Mistral frames Large 4 as the first milestone on a product road paid for by that Series D of 3 billion euros, and calls the round the largest equity round a European technology company has ever closed. It was closed in early September 2026 at a post-money valuation above 21 billion euros. Samsung Electronics led it, alongside EQT Scaleup Europe Fund and PSG Equity, with Nvidia and ASML reported as participants by several outlets. European are the seat and part of the investor base.
The model is trained on hardware that is not on sale in Europe in that configuration, and the price competition arrives from China.
What we do with this
When the weights land at the end of October, we check in this order: open and read the licence file, measure the repository byte sum, then run a test pass on our box with tasks from real operations. Until then Large 4 appears in none of our pricing or architecture decisions, because neither the licence nor the bytes are known and the preview's index score sits behind the free MIT models. Anyone looking for an open model for a customer project today finds it in the Apache and MIT class up to 40 billion parameters, which runs on our own hardware, or in the API of an MIT model that has passed a data protection review. The open question is whether the context window stays at one million tokens after verification or at the measured 524,000.
Further reading
- Artificial Analysis: Mistral Large 4, index, speed and prices of the preview (7 October 2026)
- Artificial Analysis: DeepSeek V4.1 Flash · GLM-5.3 · Kimi K3 · MiMo-V2.6-Pro · Qwen3.8 27B
- Artificial Analysis: Mistral Medium 3.5 · Mistral Small 4
- Mistral: official announcement of Large 4 with the vendor's own benchmark charts, model card and preview documentation
- heise online: Mistral Large 4, the LLM boulder from the EU
- LMArena: Code Arena with the efficiency frontier for coding tasks, state of 7 October 2026
- Mistral: Series D announcement and TechCrunch: Mistral raises 3 billion euros on the funding round
- Hugging Face: DeepSeek V4.1 Flash · Z.ai GLM-5.3 · Moonshot Kimi K3 · Xiaomi MiMo V2.6 Pro RL · Mistral Large 3 675B NVFP4 · Qwen3.8 27B
- Our own AI stack with LiteLLM and vLLM, how the lanes are built
- Seven days of local AI in the field, operating without cloud models
- Two GB300 stations in a paired setup, the hardware class above our estate
Is Mistral Large 4 the best European open model at launch?+
Not by the independent Artificial Analysis index. The preview scores 38 points, MiMo-V2.6-Pro scores 46, GLM-5.3 scores 45, Kimi K3 scores 44 and DeepSeek V4.1 Flash scores 39. The Large 4 weights are not published either; the announced date is the end of October 2026. Scores retrieved on 7 October 2026.
What does one task cost with Mistral Large 4?+
Artificial Analysis calculates 1.13 US dollars per Intelligence Index task for the Large 4 preview. MiMo-V2.6-Pro is at 0.13 dollars, DeepSeek V4.1 Flash at 0.27 dollars, GLM-5.3 and Kimi K3 at about 2 dollars each. Large 4 output tokens cost 2.09 dollars per million on promotion and 4.18 dollars at list price.
Could we run Mistral Large 4 on our own hardware?+
Not according to our arithmetic. The model card names 1.05 trillion parameters. At 0.5 bytes per parameter that is 525 GB, and scaled to the measured NVFP4 layout of its predecessor it is 627 GB. Our four RTX 5090 provide 127.4 GiB in total, which is 137 GB. The model needs at least three and a half times our GPU memory. Figures as of 7 October 2026.
senn-tech