Mistral Large 4 review: a real jump that still lands mid-pack

Kurnia Kharisma
Written by

Kurnia Kharisma

Katelin Teen
Reviewed by

Katelin Teen

Last edited October 9, 2026

Expert Verified
Hand-drawn illustration of a reviewer weighing a new AI model on a scale, with the Mistral logo

What is Mistral Large 4?

Mistral Large 4 is Mistral's new flagship, launched on October 6, 2026 as a public preview and nicknamed "Le Chonk". Per its model card, it is a mixture-of-experts LLM with 1.05T total parameters, 52B active per token, a 1.6B vision encoder and a 1M-token context window. It is a hybrid model: you can ask it to reason first with a reasoning_effort setting, or answer straight away.

Mistral Large 4 model card showing the public preview badge, the open tag, 1M context, and the sale price next to the struck-through list price, as taken from Mistral docs
Mistral Large 4 model card showing the public preview badge, the open tag, 1M context, and the sale price next to the struck-through list price, as taken from Mistral docs

A quick note on where I am coming from. I write and run SEO on eesel's blog, and the people who search "Mistral Large 4 review" are rarely curious about parameter counts. They have a narrower question: is it good enough to switch to, and for what? That is the question this review answers. For the full spec sheet, see my Mistral Large 4 overview, and for every pricing tier, my Mistral Large 4 pricing breakdown.

I have not run the model myself, so my takes lean on the people who have: Artificial Analysis's independent tests, Mistral's own docs, and the hands-on reports on Hacker News, Reddit and X in the first three days.

Mistral Large 4 review at a glance

Here is the short version, before the detail.

What I checkedResultSource
Overall intelligence38.4 on AA Intelligence Index (median for its price tier: 26)Artificial Analysis
Jump over Large 39.3 to 38.4, about 4.1xAA: Large 3
Rank among close rivalsBehind Kimi K3 (43.6) and Qwen 3.8 Max (45.4)AA
Price (preview sale)$0.68 in / $2.09 out per 1M tokensMistral pricing
Cost per AA task$1.13 at list, about $0.57 at saleArtificial Analysis
Hallucination rate41.9% on AA-OmniscienceArtificial Analysis
Output speed116.1 tokens/s; ~18.7 s to first answer tokenArtificial Analysis
Verbosity200M tokens to run the index vs 81M medianArtificial Analysis
Open weightsPromised end of October 2026, not out yetMistral
Best atCyber tasks without refusals, legal agent work, visual groundingMistral launch charts

What did Mistral actually fix?

The honest headline is that Mistral is back in the race. Last December, Mistral Large 3 scored 9.3 on the same index, which put it nowhere near the frontier. Large 4 scores 38.4. That is the difference between a model you used for cheap summaries and one you could run an AI agent on.

Plotly's team ran it through their data-analytics benchmark on launch day and saw the same jump from a different angle:

Hacker News

"It's 10x cheaper than Mistral Medium 3.5 from April and goes from 58% to 74% correct. Definitely a generational shift."

The catch is where that jump came from. Large 4 is a reasoning model and Large 3 was not, so a lot of the gain was bought with thinking tokens.

Hand-drawn bar pairs comparing Mistral Large 3 with Large 4: Intelligence Index 9.3 to 38.4 (4.1x), output tokens per task 5,566 to 69,330 (12.5x), cost to run the index $46 to $1,602 (34.8x)
Hand-drawn bar pairs comparing Mistral Large 3 with Large 4: Intelligence Index 9.3 to 38.4 (4.1x), output tokens per task 5,566 to 69,330 (12.5x), cost to run the index $46 to $1,602 (34.8x)

The score went up about 4.1x. The tokens it writes per task went up 12.5x, and the cost to run the full index at list price went up 34.8x, from $46 to $1,602, per Artificial Analysis. That is not a knock on the model; every lab paid the same toll to get reasoning. But it does mean the per-token price is a poor guide to what Large 4 will cost you. Artificial Analysis calls it "very verbose": 200M tokens to complete the index against a median of 81M.

How Mistral Large 4 compares on independent benchmarks

This is where the review gets less flattering. Mistral's launch post is full of charts, and most of them show Large 4 winning. The independent picture is more mixed.

Hand-drawn two-panel comparison: Mistral's launch charts show Cyber Index 50 in first place, a lead on the legal agent test and on visual grounding; independent tests show an Intelligence Index of 38.4, a 41.9% hallucination rate and 26.8% on Terminal-Bench
Hand-drawn two-panel comparison: Mistral's launch charts show Cyber Index 50 in first place, a lead on the legal agent test and on visual grounding; independent tests show an Intelligence Index of 38.4, a 41.9% hallucination rate and 26.8% on Terminal-Bench

On Artificial Analysis's v4.3.2 index, which blends 10 evaluations, here is where Large 4 lands among the models people are comparing it to:

ModelAA IndexPrice in / out per 1M (list)Cost per AA taskOutput speed
Claude Sonnet 5.5 (Xhigh)51.9$2 / $10$2.01100.7 t/s
Qwen 3.8 Max45.4$2 / $6$5.4137.6 t/s
Kimi K343.6$3 / $15$2.0041.7 t/s
Gemini 3.8 Flash40.9$0.75 / $3.75$1.24125.2 t/s
DeepSeek V4.1 Flash39.5$0.30 / $1.20$0.27216.1 t/s
Mistral Large 4 Preview38.4$1.36 / $4.18$1.13116.1 t/s
GPT-5.6 Luna37.3$0.20 / $1.20$0.18117.0 t/s

All figures from Artificial Analysis model pages, fetched October 8, 2026. Large 4's cost per task roughly halves at the current sale price.

Two things jump out. First, Large 4 is cheaper per task than every model above it and streams faster than Kimi K3 and Qwen 3.8 Max. Second, DeepSeek V4.1 Flash and GPT-5.6 Luna sit within about one point of it and cost a fraction as much per task, even after Mistral's 50% sale. If your job is "decent answers, lots of them", those two are the harder competition.

Mistral's own Terminal-Bench 4.0 chart is a good example of how framing matters. It shows Large 4 at 28, ahead of Qwen 3.8 Max and Kimi K3, but well behind GLM-5.3 at 40.

Mistral's Terminal-Bench 4.0 chart showing Mistral Large 4 Preview at 28, DeepSeek V4 Pro at 10, Qwen 3.8 Max at 17, Kimi K3 at 21 and GLM-5.3 at 40, as taken from Mistral
Mistral's Terminal-Bench 4.0 chart showing Mistral Large 4 Preview at 28, DeepSeek V4 Pro at 10, Qwen 3.8 Max at 17, Kimi K3 at 21 and GLM-5.3 at 40, as taken from Mistral

Some Hacker News readers were blunter about the chart choices. One commenter checked a multimodal chart against Artificial Analysis's public numbers and found a gap:

Hacker News

"I don't know this benchmark but AA's GDP.pdf listing has Sol 5.6 Max at a 27, even non-reasoning beats the 12.8. It's extremely weird to cherry pick Sol 5.6 and then also lie about the published third party benchmark score."

I would not go as far as "lie". Benchmark versions and settings differ, and Mistral's charts may use a different run. But it is a good reason to read vendor charts as the best case and treat the Artificial Analysis page as the baseline.

Where Mistral Large 4 really is strong

Being mid-pack overall does not mean it is mid-pack everywhere. Three areas hold up under scrutiny.

Security work without refusals. On Artificial Analysis's Cyber Index, as charted by Mistral, Large 4 scores 50, ahead of Kimi K3 and DeepSeek V4.1 Flash at 41. The more interesting part of the chart is the grey bars: other frontier models had a large share of their attempts stopped by safety blocks.

Mistral's chart of the Artificial Analysis Cyber Index: Mistral Large 4 Preview leads with 50 successes, while Qwen 3.8, Opus 5.5 and GPT-6 Astra show large safety-block bars, as taken from Mistral
Mistral's chart of the Artificial Analysis Cyber Index: Mistral Large 4 Preview leads with 50 successes, while Qwen 3.8, Opus 5.5 and GPT-6 Astra show large safety-block bars, as taken from Mistral

Mistral co-founder Guillaume Lample pitched exactly this angle on launch day, saying open models "do not refuse to help". If you run a security team and your current model keeps declining to look at exploit code in your own stack, that is a real reason to test Large 4.

Legal and document work. Mistral's chart from Vals.ai has Large 4 leading Harvey's legal agent benchmark, and Artificial Analysis gives it 81.3% on AA-LCR, its long-context reasoning test. Long contracts are a good fit for a 1M-token window.

Speed once it starts talking. At 116.1 tokens per second on Mistral's API, it streams faster than Kimi K3 and Qwen 3.8 Max by a wide margin. One early tester on Hacker News liked it as a daily driver for that reason:

Hacker News

"It is very fast via openrouter (significantly better than Kimi K3) on webui. Very verbose and starts to forget instructions after awhile it seems, but it gave me quite a lot of good info during a half an hour chat on C and embedded programming."

Note the middle of that quote. "Starts to forget instructions" is the kind of thing that matters a lot more in production than in a chat window.

Where Mistral Large 4 falls short

The weak spots are just as specific.

Factual recall and hallucination. This is the number I would flag in red if I were evaluating it for anything customer-facing. On AA-Omniscience, which asks knowledge questions and rewards a model for saying "I don't know" instead of guessing, Large 4 got 25.8% right and had a 41.9% hallucination rate, for an overall score of -5.3, per Artificial Analysis. One Hacker News tester saw the same thing in a trivia game:

Hacker News

"Mistral Large 4 is not very good. It can sometimes solve a game with ~40 guesses whereas the top models like Gemini 3.8 Flash or Grok 4.7 can one shot most puzzles. My benchmark here aligns with AA Omniscience."

Time to first answer. It is a reasoning model, so it thinks before it speaks. Artificial Analysis measured about 18.7 seconds to the first answer token on a long prompt and 28.4 seconds at 100k tokens. That is fine for a back-office agent and painful for a live chat widget.

Agentic coding. It trails GLM-5.3 on Terminal-Bench and scored 26.8% on Artificial Analysis's own run. Some hands-on reports were harsher still:

Reddit

"Played around with it a bit, for code analysis, image classifications, and chess capability. Feels quite dated and far behind what I usually utilize."

To be fair to Mistral, this is a preview. Lample said on X that the reinforcement learning run "is still in flight", and a final version is due with the weights. Scores may move before the end of October.

Is "open weights" a reason to pick it today?

Not yet. The model card carries an "Open" badge, but the license field is empty, and as of October 9 the Mistral Hugging Face page has no Large 4 repo. Artificial Analysis lists it as proprietary until the weights land.

Even when they ship, a 1.05T-parameter model is not something most teams will run on a spare GPU. The most-liked reply to Mistral's launch tweet said it best:

"me with 12 GB of VRAM reading "open weights, 1.05T parameters""

The practical value of open weights here is for large companies and governments that want to host a capable model on their own hardware in Europe. For everyone else, it is the API, and the API is a preview with a one-month deprecation notice, not the six months a general-availability model gets. Building a production workflow on a preview means accepting that it can change under you with four weeks' warning.

Who should use Mistral Large 4?

After all of the above, my answer comes down to two questions: does your data need to stay in Europe (or on your own servers), and does raw accuracy matter most for your use case?

Hand-drawn 2x2 quadrant: horizontal axis from raw accuracy matters less to raw accuracy matters most, vertical axis from any cloud is fine to must stay in Europe or self-host; the top-left quadrant, labelled Mistral Large 4 fits, is highlighted with examples of EU data rules and security research
Hand-drawn 2x2 quadrant: horizontal axis from raw accuracy matters less to raw accuracy matters most, vertical axis from any cloud is fine to must stay in Europe or self-host; the top-left quadrant, labelled Mistral Large 4 fits, is highlighted with examples of EU data rules and security research

The sovereignty point is a real one, and plenty of commenters said so plainly:

Reddit

"In most corporate workflows, it is probably good enough. And you don't have to sell your soul to Xi or Trump. As a european, I had feared much worse."

Try the fit check below to see where your use case lands.

What matters most for your workload?

Pick one to see my call.

Strong fit. Large 4 is the most capable model from a European lab, served from Mistral's own EU infrastructure, with open weights promised by the end of October 2026. Regional inference costs 1.1x list. Wait for the weights and license before planning a self-hosted rollout.
Strong fit, worth a test. It leads the Artificial Analysis Cyber Index on Mistral's chart (50 vs 41 for Kimi K3) and was not held back by safety blocks the way several frontier models were. Run it on your own tasks, since Terminal-Bench (26.8%) shows its agentic coding still trails.
Probably not. DeepSeek V4.1 Flash scores 39.5 at $0.27 per Artificial Analysis task and GPT-5.6 Luna scores 37.3 at $0.18. Large 4 scores 38.4 at about $0.57 even on its half-price sale.
No. At 38.4 on the Intelligence Index it sits well behind Claude Sonnet 5.5 (51.9 at Xhigh), Qwen 3.8 Max (45.4) and Kimi K3 (43.6). Pay for one of those, or wait for the final Large 4 release and recheck.
Only with strong guardrails. A 41.9% hallucination rate and ~18.7 seconds to the first answer token are the two numbers that matter for support. Ground every answer in your help center, route low-confidence tickets to a human, and replay past tickets before going live. Or use a support teammate that handles that layer for you.

What this means if you run a support team

A lot of the people reading model reviews right now are support leads wondering whether a cheaper, European model could power their bot. I get it. But the hallucination number is the one I would sit with.

I have watched what happens when a support bot answers from the model's own memory instead of your docs. Earlier this year, a few paying eesel customers had a bot that, when the knowledge base search came back empty, filled the gap from its training data instead of handing off. One sent made-up subscription details to real customers. Another answered a support question with "Oxygen", the element from the periodic table. That was a fallback problem, not a model problem, and a bigger model would not have fixed it. It is why every eesel rollout now gets simulated against past tickets before it touches a customer.

A CX lead at a supplements brand handling about 7,000 tickets put the real requirement better than any benchmark:

"I need an AI who is only handling the tickets that it's confident to handle and all the other ones, leave them alone."

That is a routing and confidence score problem first and a model problem second. A model with a 41.9% hallucination rate makes that layer more important, not less. If you do build your own on Large 4, my guides on RAG for support, stopping hallucinations and the best LLM for support are a good place to start. And if your reason for Mistral is EU data rules, check what SOC 2 and GDPR actually require of the vendor first; the model's home country is only one part of it.

Mistral Large 4 alternatives worth testing

If the review has you hesitating, here is where I would look instead, by what you need:

You wantLook atWhy
More accuracy, still affordableSonnet 5.5 or Kimi K351.9 and 43.6 on the AA index
Same score, lower costDeepSeek V4 Flash family, GPT-5.6 Luna39.5 and 37.3 at $0.18 to $0.27 per task
Open weights you can run todayDeepSeek V4.1 Flash, GLM-5.3Both listed as open-weight next to Large 4 in this AA comparison
A cheaper MistralLarge 3 or Small 4 (see Mistral AI pricing)$0.50/$1.50 and $0.15/$0.60 per 1M

For a longer list, my Mistral alternatives roundup and the head-to-heads on Claude vs Mistral and ChatGPT vs Mistral go deeper.

The verdict

Mistral Large 4 is a real comeback from a lab that looked like it had stalled. The score jump is real, the price is fair while the sale lasts, and the speed is good. If you need a European provider or a model that will actually help with security work, it is the one I would test first.

For everyone else, the numbers do not yet justify switching. It sits about level with models that cost a quarter as much per task, it hallucinates more than I would accept for customer-facing work, and the version you test today is a preview that can change on a month's notice. I would check back after the final release and weights at the end of October, then decide.

Try eesel

If you came here because you are picking a model for a support bot, here is the shortcut: the model is infrastructure, and eesel is the employee. eesel's AI helpdesk teammate plugs into Zendesk, Freshdesk or Gorgias, learns from your past tickets and help center, and runs a simulation over hundreds of your real tickets so you see its answers before a customer does. Model choice, grounding and the "I don't know, passing to a human" fallback are handled for you, and you pay per ticket handled, so a verbose model never shows up on your bill. Plans start free on the pricing page.

eesel's Skills list showing Simulation, which runs the agent against multiple past tickets and returns a scored performance report, alongside Support Analytics and Self Review
eesel's Skills list showing Simulation, which runs the agent against multiple past tickets and returns a scored performance report, alongside Support Analytics and Self Review

Try eesel and replay your last month of tickets before you decide anything.

Frequently Asked Questions

Is Mistral Large 4 any good?
It is a real step up: Artificial Analysis scores it 38.4 on its Intelligence Index, against 9.3 for Large 3. But that puts it mid-pack, behind Kimi K3 and Qwen 3.8 Max, and level with cheaper models. My full Mistral Large 4 overview covers the specs.
How much does Mistral Large 4 cost?
During the preview it is on sale at $0.68 per 1M input tokens and $2.09 per 1M output tokens, half the $1.36 / $4.18 list price, with no end date. I break down batch, regional and Priority tiers in my Mistral Large 4 pricing guide.
Is Mistral Large 4 open source?
Not yet. Mistral calls it open-weight and promises the weights by the end of October 2026, but as of October 9 there is no Hugging Face repo and no license named. Until then it is API-only, like most of the new Mistral models at launch.
Is Mistral Large 4 better than Kimi K3 or Qwen 3.8 Max?
On independent tests, no. Artificial Analysis has Kimi K3 at 43.6 and Qwen 3.8 Max at 45.4, against 38.4 for Large 4. Large 4 is cheaper per task than both and faster than both on Mistral's own API.
Can I use Mistral Large 4 for a customer support chatbot?
You can, but its 41.9% hallucination rate on AA-Omniscience means it needs strict grounding in your knowledge base and a fallback when retrieval comes back empty. My guide on AI hallucinations in support covers the guardrails.
Why would I pick Mistral Large 4 over a cheaper model?
Two reasons hold up: you need a European provider (or self-hosted weights once they ship), or you do security work where other frontier models block requests. If neither applies, DeepSeek and GPT-5.6 Luna score about the same for far less per task.
Is Mistral Large 4 fast?
It streams at 116.1 tokens per second on Mistral's API, well above the 77 tokens per second median for its price tier. But it thinks first, so the first answer token takes around 18.7 seconds on a long prompt, which is slow for live AI customer support chat.
What are the best Mistral Large 4 alternatives?
For more accuracy, Sonnet 5.5 or Kimi K3. For the same score at a lower cost per task, DeepSeek V4.1 Flash or GPT-5.6 Luna. My Mistral alternatives roundup compares them by use case.

Share this article

Kurnia Kharisma

Article by

Kurnia Kharisma

Kurnia is a software engineer and writer at eesel AI with two years of SEO experience, writing about AI tools, helpdesk software, and customer support. He pairs a developer's understanding of how these products are built with search-driven research into what actually ranks and resonates with the people searching for them.

Related Posts

All posts →
Hand-drawn illustration of two people looking at a price tag and a speed gauge next to the Mistral logo
Trending

Mistral Large 4 pricing: API rates, service tiers, and the real cost per task

Mistral Large 4 pricing is $0.68 in and $2.09 out per 1M tokens on sale. Regional and Priority tiers are priced off the $1.36 / $4.18 list price, not the sale.

Rama AdiRama AdiOct 8, 2026
Xiaomi MiMo V2.6 review illustration
Trending

Xiaomi MiMo V2.6 review: is the cheap open frontier worth it?

A hands-on review of Xiaomi MiMo V2.6: what the benchmarks really say, the value math that makes it interesting, who should run it, and where it still trails.

Kurnia KharismaKurnia KharismaSep 24, 2026
Xiaomi MiMo V2.6 pricing and API cost illustration
Trending

Xiaomi MiMo V2.6 pricing: every plan, model, and API cost in 2026

Xiaomi MiMo V2.6 pricing, broken down: free open weights, and API costs from $0.14 per million tokens. Here is what each variant actually costs to run.

Kurnia KharismaKurnia KharismaSep 23, 2026
Xiaomi MiMo V2.6 open-weight model illustration
Trending

Xiaomi MiMo V2.6: specs, benchmarks, and how to run the open model

Xiaomi's MiMo V2.6 is an open-weight, omnimodal model family with a 1M-token context and MIT license. Here are the real specs, benchmarks, and API pricing.

Rama AdiRama AdiSep 23, 2026
Illustration of a team reviewing Gemini 3.8 Flash, with a speed gauge, a rocket, and a verdict checkmark
Trending

Gemini 3.8 Flash review: fast, verbose, and not the upgrade the number implies

A hands-on Gemini 3.8 Flash review: what it's good at, where it falls down, the 13-second catch nobody quoted, and whether to switch from 3.7 Flash.

Kurnia KharismaKurnia KharismaSep 8, 2026
Illustration of a fast-moving robot coding on a laptop while a person watches, representing Gemini 3.8 Flash
Trending

Gemini 3.8 Flash: what it is, honest benchmarks, and my review

Google shipped Gemini 3.8 Flash on September 2, 2026, three weeks after 3.7. Same price, better scores, and one line of fine print that changes the answer.

KiraKiraSep 3, 2026
Illustration of a developer and a colleague working with a fast AI coding agent
Trending

Gemini 3.7 Flash review: a great model that stopped being cheap

I put Google's Gemini 3.7 Flash against its own benchmarks and its own price list. It is fast and sharp, but it is no longer the cheap high-volume workhorse.

Rama AdiRama AdiAug 14, 2026
Hand-drawn illustration of three people around a table discussing a stack of server blocks topped with the Mistral logo
Trending

Mistral Large 4: specs, pricing, benchmarks, and who should use it

Mistral Large 4 is a 1T-parameter open-weight preview at $0.68/$2.09 per 1M tokens. A huge jump for Mistral, a mid-pack score next to rivals, and a few real niches.

KiraKiraOct 8, 2026
Hand-drawn illustration of a person sprinting with a tall stack of small task cards while another person watches, for a Claude Haiku 5.5 review
Trending

Claude Haiku 5.5 review: a real bargain, but only under 100k tokens

My Claude Haiku 5.5 review: Luna-level prices, much better scores, and real speed, as long as your prompts stay under 100k tokens and effort stays low.

Riellvriany IndriawanRiellvriany IndriawanOct 8, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free