Mistral AI Unleashes Large 4 Preview and Powerful 1T Model

Mistral AI has launched Mistral Large 4, or 'Le Chonk,' a new 1-trillion parameter multimodal model aiming to carve a "third way in AI." Trained on significantly fewer GPUs, it boasts impressive benchmarks in cybersecurity, coding, and visual tasks, with plans for an open-weight release by late 2026.
Uche Emeka
Uche Emeka • AI • 1 hour ago • 3 minute read •
Mistral AI Unleashes Large 4 Preview and Powerful 1T Model

French AI lab Mistral AI has introduced Mistral Large 4 (ML4), a new large multimodal model, aiming to establish "a third way in AI" and compete with both American and Chinese rivals. Nicknamed 'Le Chonk' due to its 1 trillion parameters, this model positions Mistral AI as a distinct alternative amidst the growing divide between closed and open AI models.

ML4, which boasts 1 trillion parameters with 49 billion active parameters in its public preview, was trained entirely on Mistral's own compute infrastructure. The company utilized 3,800 NVIDIA Grace Blackwell GPUs in its European datacenters, a significantly lower number than its competitors, demonstrating optimized resource use. Its training data is comprehensive, spanning over 160 languages, including every official language in the European Union.

While initially accessible via a public guardrail endpoint, Mistral AI plans to release the model's weights by the end of October 2026, alongside architecture details, additional benchmarks, and its post-training methodology. This open-weight approach is intended to facilitate auditing and enable private-cloud and on-premise operations, allowing organizations to apply their own policies and address mounting security concerns, particularly among enterprise and institutional clients. Mistral's VP Science, Pierre Stock, emphasized that open-source weights would be used for defense, not malicious attacks.

Mistral AI hopes ML4 will be best-in-class among open-weight models globally and potentially outperform closed models in specific areas critical to its customers, where multimodal capabilities add value. Key optimized use cases include cybersecurity, finance, and chip design. These strategic areas align with Mistral's main backers: Dutch giant ASML, which led its Series C, and Samsung, which led its Series D last month at a €21 billion valuation. Mistral aims to remain a frontier lab, not merely an inference provider, with 'Le Chonk' reinforcing its position.

The model has undergone rigorous cybersecurity testing, including red-teaming with cybersecurity leaders and state authorities. Mistral Large 4 ranks among the top five models on the Artificial Analysis Cyber Index, achieving an 82 percent score on a test requiring it to reproduce and patch real vulnerabilities in open-source software, the highest score recorded. Furthermore, it solves 93 percent of Cybench's 40 security-competition exercises and resisted 93.3 percent of attacks on Lakera’s public B3 AI Security Benchmark. Internal testing indicates its utility for malware analysis, vulnerability prioritization, and detection-rule writing, with Mistral arguing that provider-level refusals can obstruct legitimate vulnerability research and incident response.

In terms of software engineering, Mistral Large 4 demonstrates strong performance, scoring 61.7 percent on DeepSWE v1.1, 59.4 percent on SWE-Atlas-QnA, and 28.3 percent on Terminal-Bench 4, resulting in a combined Coding Agent Index score of 49.8 percent. A blind human evaluation by Surge AI, where professional annotators rated coding outputs, placed ML4 Preview second among five models with a score of 3.74, behind Claude Opus 5.

Beyond coding, ML4 achieved a 59.9 percent score on AutomationBench, which covers 657 business workflows across applications like Gmail, Google Sheets, Slack, and Salesforce. The company also showcased its multimodal capabilities through visual tasks involving technical drawings, PDFs, and satellite imagery, reporting a Dense 200 visual-grounding result of 42 percent, slightly outperforming GPT-6-Astra in Mistral’s internal testing.

Mistral Large 4 utilizes the same training, customization, and reinforcement-learning (RL) environment, Mistral Forge, offered to its customers. The model's RL library integrates chat, scientific problem-solving, safety alignment, factuality, and long-running tool-use tasks within a shared interface. These training environments leverage resources such as code sandboxes, web search, and external APIs, with verification conducted through reward models, unit tests, LLM judges, and static checks. A training run at a reported scale of 3,000 GPUs produces approximately 33 billion tokens per day, including about 16 billion trainable completion tokens. The reinforcement-learning run behind the public preview remains in progress, with the full details and open weights eagerly anticipated.

Loading...