Signal SentryUpdated Oct 6, 16:30 UTCPM Drop
Models · scorecard

Mistral Large 4

What the vendor claims, next to what independent boards measured. The two are never mixed in one column.

At a glance

Maker
Mistral
Price per 1M tokens
$0.68 in · $2.09 out
Context window
524K
Input → output
text, image → text
Added to OpenRouter
Oct 6, 2026

Source: OpenRouter, as of Oct 6, 15:00 UTC. Model page on OpenRouter

What the vendor claims

Quoted word for word from the vendor's own launch post or model card. Claims, not measurements.

Scores the vendor claims, each with its source and the quoted line
BenchmarkClaimedSource
DeepSWE v1.1The source says: ML4 excels across software engineering, repository understanding, and complex terminal workflows, scoring 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, and 28.3% on Terminal-Bench 4.61.7%Mistral · Oct 6From Drop #008 · Tue, Oct 6 · Evening
SWE-Atlas-QnAThe source says: ML4 excels across software engineering, repository understanding, and complex terminal workflows, scoring 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, and 28.3% on Terminal-Bench 4.59.4%Mistral · Oct 6From Drop #008 · Tue, Oct 6 · Evening
Terminal-Bench 4The source says: ML4 excels across software engineering, repository understanding, and complex terminal workflows, scoring 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, and 28.3% on Terminal-Bench 4.28.3%Mistral · Oct 6From Drop #008 · Tue, Oct 6 · Evening
AutomationBenchThe source says: On AutomationBench — 657 business workflows across apps like Gmail, Google Sheets, Slack, and Salesforce — it scores 59.9%, ahead of Kimi K3, MiMo-V2.6-Pro, and DeepSeek V4 Pro.59.9%Mistral · Oct 6From Drop #008 · Tue, Oct 6 · Evening
CybenchThe source says: It also solves 93% of the challenges in Cybench, a set of 40 exercises drawn from security competitions, one of the highest scores reported for an open-weight model.93%Mistral · Oct 6From Drop #008 · Tue, Oct 6 · Evening

Independent results

Measured by Epoch AI or LMArena, not by the vendor. "Listed as" is the source's own name for the model, so you can check our match.

No independent results yet. They appear here as Epoch AI and LMArena add this model.

In the Drops

Models

Mistral previews Large 4, 'Le Chonk', a 1T-parameter model with 49B active

Mistral says ML4 is a natively multimodal 1T-parameter model with 49B active, in preview API now, with weights by end of month. OpenRouter lists it at $0.68 in and $2.09 out per 1M tokens.

Why it matters: Mistral claims it is the strongest open-weight model built outside China and pitches self-hosted cyber defense as an answer to closed-model refusals and lost access.

From Drop #008 · Tue, Oct 6 · Evening