Mistral AI released Mistral Large 4, internally codenamed Le Chonk, as a public preview on 6 October 2026. The model features 1.05 trillion total parameters, with 49 billion active per token, and includes a 1.6 billion parameter vision encoder capable of handling a 1 million token context window. Training occurred from scratch across 3,800 NVIDIA Grace Blackwell GPUs within Mistral’s European datacentres.
In this article
Architecture and training
Le Chonk operates as a hybrid instruct-and-reasoning Mixture of Experts system. Approximately 4.7% of the weights activate per token, allowing the massive model to run at mid-tier pricing levels. While the full 1.05 trillion parameters must reside in memory, the activation count dictates compute requirements rather than hardware costs.
Mistral has not yet disclosed the expert count, top-k routing, or layer layout. These details will accompany the weights upon release. The training data covered more than 160 languages, including every official European Union language.
Performance results
In cybersecurity, the model scored 93% on Cybench and 82% on CyberGym-E2E. These figures place Mistral Large 4 in the global top five on the Artificial Analysis Cyber Index. Mistral notes that several closed frontier models score near zero on CyberGym-E2E because they refuse the task. Reproducing a vulnerability to prove it is real is standard defensive work, and provider-level refusals block this process.
For agentic coding, the model achieved 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, and 28.3% on Terminal-Bench 4.0. This yields a combined Artificial Analysis Coding Agent Index score of 49.8%. Mistral states these were evaluated privately ahead of the harness going public, so they are not yet independently reproducible.
A blind human evaluation run with Surge AI offers a more honest signal. Professional annotators rated Mistral Large 4 Preview at 3.74 out of 5. This placed it second among five models, ahead of GLM-5.3 at 3.60 and Kimi K3 at 3.59, but behind Claude Opus 5 at 4.22.
On safety, the model resists 93.3% of attacks on Lakera’s B3 benchmark and scores 1.691 out of a maximum 2.0 on KORABench.
Comparison with open-weight rivals
The following table compares Mistral Large 4 against DeepSeek V4 Pro, Kimi K3, and GLM-5.3.
Specification comparison
- Total parameters: Mistral Large 4 has 1.05T, compared to 1.6T for DeepSeek V4 Pro, 2.8T for Kimi K3, and an unpublished figure for GLM-5.3.
- Active per token: Mistral Large 4 and DeepSeek V4 Pro activate 49B, while Kimi K3 activates approximately 104B. GLM-5.3 figures are not officially published.
- Context window: All four models support a 1M token window.
- Native image input: Available for Mistral Large 4 and Kimi K3. DeepSeek V4 Pro and GLM-5.3 do not support this natively.
- Weights availability: Mistral Large 4 weights are not yet available, promised by end of October 2026. DeepSeek V4 Pro weights are on Hugging Face. Kimi K3 has been available since 27 July 2026. GLM-5.3 weights are available per Artificial Analysis.
- License: Mistral has not yet announced its license. DeepSeek uses MIT, Kimi uses Modified MIT, and GLM-5.3 uses the GLM-5.3 License.
- API price per 1M in/out: Mistral charges $1.36 for input and $4.18 for output. DeepSeek pricing varies by provider. Kimi charges $3.00 for input and $15.00 for output. GLM-5.3 charges $1.40 for input and $4.40 for output.
- Released: Mistral Large 4 launched on 6 October 2026. DeepSeek V4 Pro released in August 2026 (0813 build). Kimi K3 released on 16 July 2026. GLM-5.3 released on 14 August 2026.
Current capabilities
The preview API supports function calling, structured outputs, document QnA, batching, and the Agents and Conversations endpoints. Cached input is priced at $0.14 per 1M tokens. This rate materially changes the economics of long-context agent loops at a 1M window.
What it means
Developers can access the preview API immediately for function calling and document QnA tasks. However, self-hosting remains impossible until the weights ship at the end of October 2026. The pricing structure suggests a viable option for projects requiring high context windows, particularly where cached inputs are common.


