
What is Mistral Large 4?
Mistral Large 4 is Mistral's new flagship, launched on October 6, 2026 as a public preview and nicknamed "Le Chonk". Per its model card, it is a mixture-of-experts LLM with 1.05T total parameters, 52B active per token, a 1.6B vision encoder and a 1M-token context window. It is a hybrid model: you can ask it to reason first with a reasoning_effort setting, or answer straight away.

A quick note on where I am coming from. I write and run SEO on eesel's blog, and the people who search "Mistral Large 4 review" are rarely curious about parameter counts. They have a narrower question: is it good enough to switch to, and for what? That is the question this review answers. For the full spec sheet, see my Mistral Large 4 overview, and for every pricing tier, my Mistral Large 4 pricing breakdown.
I have not run the model myself, so my takes lean on the people who have: Artificial Analysis's independent tests, Mistral's own docs, and the hands-on reports on Hacker News, Reddit and X in the first three days.
Mistral Large 4 review at a glance
Here is the short version, before the detail.
| What I checked | Result | Source |
|---|---|---|
| Overall intelligence | 38.4 on AA Intelligence Index (median for its price tier: 26) | Artificial Analysis |
| Jump over Large 3 | 9.3 to 38.4, about 4.1x | AA: Large 3 |
| Rank among close rivals | Behind Kimi K3 (43.6) and Qwen 3.8 Max (45.4) | AA |
| Price (preview sale) | $0.68 in / $2.09 out per 1M tokens | Mistral pricing |
| Cost per AA task | $1.13 at list, about $0.57 at sale | Artificial Analysis |
| Hallucination rate | 41.9% on AA-Omniscience | Artificial Analysis |
| Output speed | 116.1 tokens/s; ~18.7 s to first answer token | Artificial Analysis |
| Verbosity | 200M tokens to run the index vs 81M median | Artificial Analysis |
| Open weights | Promised end of October 2026, not out yet | Mistral |
| Best at | Cyber tasks without refusals, legal agent work, visual grounding | Mistral launch charts |
What did Mistral actually fix?
The honest headline is that Mistral is back in the race. Last December, Mistral Large 3 scored 9.3 on the same index, which put it nowhere near the frontier. Large 4 scores 38.4. That is the difference between a model you used for cheap summaries and one you could run an AI agent on.
Plotly's team ran it through their data-analytics benchmark on launch day and saw the same jump from a different angle:
"It's 10x cheaper than Mistral Medium 3.5 from April and goes from 58% to 74% correct. Definitely a generational shift."
The catch is where that jump came from. Large 4 is a reasoning model and Large 3 was not, so a lot of the gain was bought with thinking tokens.

The score went up about 4.1x. The tokens it writes per task went up 12.5x, and the cost to run the full index at list price went up 34.8x, from $46 to $1,602, per Artificial Analysis. That is not a knock on the model; every lab paid the same toll to get reasoning. But it does mean the per-token price is a poor guide to what Large 4 will cost you. Artificial Analysis calls it "very verbose": 200M tokens to complete the index against a median of 81M.
How Mistral Large 4 compares on independent benchmarks
This is where the review gets less flattering. Mistral's launch post is full of charts, and most of them show Large 4 winning. The independent picture is more mixed.

On Artificial Analysis's v4.3.2 index, which blends 10 evaluations, here is where Large 4 lands among the models people are comparing it to:
| Model | AA Index | Price in / out per 1M (list) | Cost per AA task | Output speed |
|---|---|---|---|---|
| Claude Sonnet 5.5 (Xhigh) | 51.9 | $2 / $10 | $2.01 | 100.7 t/s |
| Qwen 3.8 Max | 45.4 | $2 / $6 | $5.41 | 37.6 t/s |
| Kimi K3 | 43.6 | $3 / $15 | $2.00 | 41.7 t/s |
| Gemini 3.8 Flash | 40.9 | $0.75 / $3.75 | $1.24 | 125.2 t/s |
| DeepSeek V4.1 Flash | 39.5 | $0.30 / $1.20 | $0.27 | 216.1 t/s |
| Mistral Large 4 Preview | 38.4 | $1.36 / $4.18 | $1.13 | 116.1 t/s |
| GPT-5.6 Luna | 37.3 | $0.20 / $1.20 | $0.18 | 117.0 t/s |
All figures from Artificial Analysis model pages, fetched October 8, 2026. Large 4's cost per task roughly halves at the current sale price.
Two things jump out. First, Large 4 is cheaper per task than every model above it and streams faster than Kimi K3 and Qwen 3.8 Max. Second, DeepSeek V4.1 Flash and GPT-5.6 Luna sit within about one point of it and cost a fraction as much per task, even after Mistral's 50% sale. If your job is "decent answers, lots of them", those two are the harder competition.
Mistral's own Terminal-Bench 4.0 chart is a good example of how framing matters. It shows Large 4 at 28, ahead of Qwen 3.8 Max and Kimi K3, but well behind GLM-5.3 at 40.

Some Hacker News readers were blunter about the chart choices. One commenter checked a multimodal chart against Artificial Analysis's public numbers and found a gap:
"I don't know this benchmark but AA's GDP.pdf listing has Sol 5.6 Max at a 27, even non-reasoning beats the 12.8. It's extremely weird to cherry pick Sol 5.6 and then also lie about the published third party benchmark score."
I would not go as far as "lie". Benchmark versions and settings differ, and Mistral's charts may use a different run. But it is a good reason to read vendor charts as the best case and treat the Artificial Analysis page as the baseline.
Where Mistral Large 4 really is strong
Being mid-pack overall does not mean it is mid-pack everywhere. Three areas hold up under scrutiny.
Security work without refusals. On Artificial Analysis's Cyber Index, as charted by Mistral, Large 4 scores 50, ahead of Kimi K3 and DeepSeek V4.1 Flash at 41. The more interesting part of the chart is the grey bars: other frontier models had a large share of their attempts stopped by safety blocks.

Mistral co-founder Guillaume Lample pitched exactly this angle on launch day, saying open models "do not refuse to help". If you run a security team and your current model keeps declining to look at exploit code in your own stack, that is a real reason to test Large 4.
Legal and document work. Mistral's chart from Vals.ai has Large 4 leading Harvey's legal agent benchmark, and Artificial Analysis gives it 81.3% on AA-LCR, its long-context reasoning test. Long contracts are a good fit for a 1M-token window.
Speed once it starts talking. At 116.1 tokens per second on Mistral's API, it streams faster than Kimi K3 and Qwen 3.8 Max by a wide margin. One early tester on Hacker News liked it as a daily driver for that reason:
"It is very fast via openrouter (significantly better than Kimi K3) on webui. Very verbose and starts to forget instructions after awhile it seems, but it gave me quite a lot of good info during a half an hour chat on C and embedded programming."
Note the middle of that quote. "Starts to forget instructions" is the kind of thing that matters a lot more in production than in a chat window.
Where Mistral Large 4 falls short
The weak spots are just as specific.
Factual recall and hallucination. This is the number I would flag in red if I were evaluating it for anything customer-facing. On AA-Omniscience, which asks knowledge questions and rewards a model for saying "I don't know" instead of guessing, Large 4 got 25.8% right and had a 41.9% hallucination rate, for an overall score of -5.3, per Artificial Analysis. One Hacker News tester saw the same thing in a trivia game:
"Mistral Large 4 is not very good. It can sometimes solve a game with ~40 guesses whereas the top models like Gemini 3.8 Flash or Grok 4.7 can one shot most puzzles. My benchmark here aligns with AA Omniscience."
Time to first answer. It is a reasoning model, so it thinks before it speaks. Artificial Analysis measured about 18.7 seconds to the first answer token on a long prompt and 28.4 seconds at 100k tokens. That is fine for a back-office agent and painful for a live chat widget.
Agentic coding. It trails GLM-5.3 on Terminal-Bench and scored 26.8% on Artificial Analysis's own run. Some hands-on reports were harsher still:
"Played around with it a bit, for code analysis, image classifications, and chess capability. Feels quite dated and far behind what I usually utilize."
To be fair to Mistral, this is a preview. Lample said on X that the reinforcement learning run "is still in flight", and a final version is due with the weights. Scores may move before the end of October.
Is "open weights" a reason to pick it today?
Not yet. The model card carries an "Open" badge, but the license field is empty, and as of October 9 the Mistral Hugging Face page has no Large 4 repo. Artificial Analysis lists it as proprietary until the weights land.
Even when they ship, a 1.05T-parameter model is not something most teams will run on a spare GPU. The most-liked reply to Mistral's launch tweet said it best:
"me with 12 GB of VRAM reading "open weights, 1.05T parameters""
The practical value of open weights here is for large companies and governments that want to host a capable model on their own hardware in Europe. For everyone else, it is the API, and the API is a preview with a one-month deprecation notice, not the six months a general-availability model gets. Building a production workflow on a preview means accepting that it can change under you with four weeks' warning.
Who should use Mistral Large 4?
After all of the above, my answer comes down to two questions: does your data need to stay in Europe (or on your own servers), and does raw accuracy matter most for your use case?

The sovereignty point is a real one, and plenty of commenters said so plainly:
"In most corporate workflows, it is probably good enough. And you don't have to sell your soul to Xi or Trump. As a european, I had feared much worse."
Try the fit check below to see where your use case lands.
What matters most for your workload?
Pick one to see my call.
What this means if you run a support team
A lot of the people reading model reviews right now are support leads wondering whether a cheaper, European model could power their bot. I get it. But the hallucination number is the one I would sit with.
I have watched what happens when a support bot answers from the model's own memory instead of your docs. Earlier this year, a few paying eesel customers had a bot that, when the knowledge base search came back empty, filled the gap from its training data instead of handing off. One sent made-up subscription details to real customers. Another answered a support question with "Oxygen", the element from the periodic table. That was a fallback problem, not a model problem, and a bigger model would not have fixed it. It is why every eesel rollout now gets simulated against past tickets before it touches a customer.
A CX lead at a supplements brand handling about 7,000 tickets put the real requirement better than any benchmark:
"I need an AI who is only handling the tickets that it's confident to handle and all the other ones, leave them alone."
That is a routing and confidence score problem first and a model problem second. A model with a 41.9% hallucination rate makes that layer more important, not less. If you do build your own on Large 4, my guides on RAG for support, stopping hallucinations and the best LLM for support are a good place to start. And if your reason for Mistral is EU data rules, check what SOC 2 and GDPR actually require of the vendor first; the model's home country is only one part of it.
Mistral Large 4 alternatives worth testing
If the review has you hesitating, here is where I would look instead, by what you need:
| You want | Look at | Why |
|---|---|---|
| More accuracy, still affordable | Sonnet 5.5 or Kimi K3 | 51.9 and 43.6 on the AA index |
| Same score, lower cost | DeepSeek V4 Flash family, GPT-5.6 Luna | 39.5 and 37.3 at $0.18 to $0.27 per task |
| Open weights you can run today | DeepSeek V4.1 Flash, GLM-5.3 | Both listed as open-weight next to Large 4 in this AA comparison |
| A cheaper Mistral | Large 3 or Small 4 (see Mistral AI pricing) | $0.50/$1.50 and $0.15/$0.60 per 1M |
For a longer list, my Mistral alternatives roundup and the head-to-heads on Claude vs Mistral and ChatGPT vs Mistral go deeper.
The verdict
Mistral Large 4 is a real comeback from a lab that looked like it had stalled. The score jump is real, the price is fair while the sale lasts, and the speed is good. If you need a European provider or a model that will actually help with security work, it is the one I would test first.
For everyone else, the numbers do not yet justify switching. It sits about level with models that cost a quarter as much per task, it hallucinates more than I would accept for customer-facing work, and the version you test today is a preview that can change on a month's notice. I would check back after the final release and weights at the end of October, then decide.
Try eesel
If you came here because you are picking a model for a support bot, here is the shortcut: the model is infrastructure, and eesel is the employee. eesel's AI helpdesk teammate plugs into Zendesk, Freshdesk or Gorgias, learns from your past tickets and help center, and runs a simulation over hundreds of your real tickets so you see its answers before a customer does. Model choice, grounding and the "I don't know, passing to a human" fallback are handled for you, and you pay per ticket handled, so a verbose model never shows up on your bill. Plans start free on the pricing page.

Try eesel and replay your last month of tickets before you decide anything.
Frequently Asked Questions
Is Mistral Large 4 any good?
How much does Mistral Large 4 cost?
Is Mistral Large 4 open source?
Is Mistral Large 4 better than Kimi K3 or Qwen 3.8 Max?
Can I use Mistral Large 4 for a customer support chatbot?
Why would I pick Mistral Large 4 over a cheaper model?
Is Mistral Large 4 fast?
What are the best Mistral Large 4 alternatives?

Article by
Kurnia Kharisma
Kurnia is a software engineer and writer at eesel AI with two years of SEO experience, writing about AI tools, helpdesk software, and customer support. He pairs a developer's understanding of how these products are built with search-driven research into what actually ranks and resonates with the people searching for them.








