THE CRUNCH

Mistral AI has released a public preview of Mistral Large 4, its largest model yet: a natively multimodal model with one trillion parameters, 49 billion of them active, trained in the company's own European data centres. The API is live through Mistral Studio, and model weights are expected at the end of October. Mistral claims it is the best open-weight model from the US or Europe on aggregated benchmarks, and the company has nicknamed it "le Chonk".

Independent numbers back the leap, with caveats. On the Artificial Analysis Intelligence Index, which aggregates ten benchmarks, ML4 scores 38 points, a big jump from the 9 points of Mistral Large 3 and 14 for Mistral Medium 3.5. It also edges past Z.ai's GLM-5.2. But Claude Opus 5.5 (Max) still leads with 58 points, and ML4 trails several closed models and Chinese open-weight models overall.

Security is the main pitch. Mistral says ML4 ranks among the top five models worldwide on the Artificial Analysis Cyber Index and leads open-weight models developed outside China by a wide margin. On a test that asks a model to reproduce a real vulnerability in open-source software and then patch it, ML4 hits 82 percent, the highest score of any model. The company argues that defending software often starts with proving a vulnerability is real, work that closed models' safety filters block, and that losing provider access mid-incident is itself a risk. ML4 is also designed to run in private clouds or on-premise.

That said, the comparison is not purely about capability. According to Mistral, Claude Opus 5.5 and GPT-6 Astra score near zero on the vulnerability test because they refuse the task entirely, so the benchmark measures provider policies as much as model ability. Mistral simultaneously touts high refusal rates on malicious cyber prompts from JailbreakBench, StrongREJECT and AgentHarm, though it does not explain how the model reliably distinguishes legitimate vulnerability research from attack preparation.

Elsewhere, ML4 scores 49.8 percent on the Artificial Analysis Coding Agent Index, ahead of Deepseek V4 Pro and Qwen3.8 Max, and placed second out of five models in a blind code-quality evaluation run with Surge AI, where professional annotators rated it 3.74 out of 5 against Claude Opus 5's leading 4.22. Until the weights ship, Mistral is red-teaming the model with security firms, vetted partners and government agencies, who get a version with reduced moderation and expanded cyber capabilities.

WHAT HAPPENS NEXT

Model weights are expected at the end of October, after red-teaming with security firms, vetted partners and government agencies.