Benchmarks · OCTOBER 9, 2026
Mistral Large 4 'Le Chonk' Scores 59.9% on AutomationBench, Topping Every Chinese Open-Weight Rival
Mistral's 1T-parameter MoE enters public preview at $1.36/$4.18 per million tokens with 49B active parameters, 93% on Cybench, and a 38 on the Artificial Analysis Intelligence Index. Weights are scheduled for October 27.
Mistral launched a public preview of Large 4 ("Le Chonk") on October 6, 2026, posting a 59.9% score on AutomationBench, the 657-workflow evaluation that most closely tracks the pipeline work small B2B teams actually automate across Gmail, Sheets, Slack, and Salesforce. On that single number, the Paris lab beats every Chinese open-weight competitor it chose to benchmark against. Weights are scheduled for October 27, after a roughly three-week red-teaming window with developers, cybersecurity leaders, and state authorities.
The architecture: 1.05 trillion total parameters in a mixture-of-experts, with 49 billion active per token (roughly 4.7%), a 1.6B vision encoder, and a claimed 1M-token context. Artificial Analysis measures the context window closer to 512K in preview, clocks output at 116.1 tokens per second with a 1.46-second time to first token, and places Large 4 at 38 on its Intelligence Index against a median of 26 for comparable reasoning models.
Priced at $1.36 input and $4.18 output per million tokens on Mistral's API (with cached input at $0.14, and promo pricing of $0.68/$2.09 via OpenRouter), For a small team wiring up outreach, qualification, or CRM automation, that's the pitch.
The rest of the benchmark card is a selective tour. Large 4 scores 93% on Cybench's 40 security-competition challenges, 93.3% resistance on Lakera's B3 prompt-injection suite, 1.691 of 2.0 on KORABench, 1,393 Elo on Mistral's own AA-Briefcase long-horizon eval, and 15% on Harvey's Legal Agent Benchmark, where Vals.ai has Kimi K3 at 12.92%, MiMo V2.6 Pro at 10.83%, and GLM-5.3 at 8.33%. On DeepSWE v1.1, Mistral reports 61.7% against Reflection's Beam at 44.4% and Qwen 3.8 Max at 51%. The live DeepSWE leaderboard tells a less flattering story: GLM-5.3 and Kimi K3 reach roughly 69% under best configurations, and GPT-6 Astra, Gemini 3.8 Flash, and Claude Opus 5 cluster around 74%.
Surge AI's blind human evaluation of five models ranked Large 4 Preview second at 3.74, behind Claude Opus 5 at 4.22 and ahead of GLM-5.3 (3.60) and Kimi K3 (3.59) on the 1–5 scale.
The efficiency story is where sovereignty talk starts to matter. Mistral trained ML4 "using only 4,000 Nvidia GPUs which is two to three times less than our Chinese competitors," VP Science Pierre Stock told TechCrunch. Mistral's own post cites 3,800 Grace Blackwell GPUs over roughly two months, across 160+ languages. Large 3, for comparison, was 675B total and 41B active, trained on 3,000 H200s.
The context is a €3 billion Series D at a post-money valuation above €21 billion (roughly $24 billion), 125+ enterprise customers including Airbus, ASML, and HSBC, and a European open-weights pitch aimed squarely at buyers who read "DeepSeek" and feel a procurement headache coming. Every score above is vendor-reported. The October 27 drop is when the number that matters, replication, finally arrives.
Sources
- Introducing Mistral Large 4
- Mistral's new 1T model aims to leapfrog closed and open rivals
- Mistral Large 4 Preview, Intelligence, Performance & Price Analysis
- Mistral debuts Large 4 'Le Chonk', a 1-trillion parameter text output model
- Mistral AI Releases Mistral Large 4 (Le Chonk): A 1.05T Parameter Multimodal MoE