
What is Mistral Large 4?
I build AI agents for a living at eesel, so every frontier launch lands on my desk and comes with the same question, which is whether eesel's agents should run on this. With Mistral Large 4 the honest answer is "not yet, but watch it", and the reasons for that are more interesting than what the benchmark chart suggests.
Mistral calls it its "largest and most capable model to date". Per the launch post, it was trained on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European data centers, and it is also the first milestone on the roadmap that Mistral's €3 billion Series D is funding. If you need a refresher on the company itself, my Mistral AI overview covers the history, and the Mistral AI reviews roundup covers what users think of the older models.
Here are the specs that matter, straight from the docs model card:
| Spec | Mistral Large 4 |
|---|---|
| Release | October 6, 2026, public preview (v26.10) |
| Architecture | Granular mixture-of-experts, hybrid instruct and reasoning |
| Parameters | 1.05T total, 52B active, plus a 1.6B vision encoder |
| Context window | 1M tokens |
| Input | Text and images |
| API model ID | mistral-large-4 |
| Features | Function calling, structured outputs, document QnA, batching, Agents API, built-in tools |
| Weights | Promised by end of October 2026, license not yet named |
| Training languages | 160+, including every official EU language |
There is one small wrinkle in terms of the numbers: the docs card says 52B active parameters, while Mistral's own launch post on X says 49B. Artificial Analysis also lists the context window at 524,288 tokens, half the 1M on the card, so I would plan around the smaller number until Mistral clarifies this.

How Mistral Large 4 works under the hood
The headline number, 1.05T parameters, is a little bit misleading if you look at it alone. Mistral Large 4 is a mixture-of-experts model: a router picks a small slice of the network for each token, so only about 5% of the weights (52B) do any work on a given token. That is why the model can be huge and also reasonably fast. Artificial Analysis measured 116.1 output tokens per second on Mistral's API, well above the 77.0 median for reasoning models at a similar price.

It is also a hybrid model, so you can use it as a plain instruct model or turn on reasoning with the reasoning_effort parameter shown in the docs example. Reasoning is where the cost and latency hide (every thinking step is billed as output tokens), and I'll come back to that later, because it is the most important thing to understand before putting it in front of customers.
The other design choice that I think is worth knowing is that Mistral trained it on its own infrastructure and serves the preview from the same place, including "a European deployment that Mistral operates end-to-end". For teams that work under EU data rules, that is a bigger deal than any benchmark is. It also matters for SOC 2 and GDPR reviews, where one of the first questions is usually exactly where the model runs.
Mistral Large 4 pricing
Mistral Large 4 is priced per token through the API, and right now there is a launch sale on it, at exactly half of the list price. The API pricing page shows both:
| Model (per 1M tokens) | Input | Cached input | Output |
|---|---|---|---|
| Mistral Large 4 (launch sale) | $0.68 | $0.07 | $2.09 |
| Mistral Large 4 (list price) | $1.36 | $0.14 | $4.18 |
| Mistral Large 3 | $0.50 | $0.05 | $1.50 |
| Mistral Medium 3.5 | $1.50 | $0.15 | $7.50 |
| Mistral Small 4 | $0.15 | $0.015 | $0.60 |
| GLM 5.3 (hosted by Mistral) | $1.40 | $0.14 | $4.40 |

Three things stand out to me:
- The sale has no end date. At list price, output is 2.8x what Large 3 cost. Budget on $1.36 / $4.18, and then treat the sale as a bonus on top.
- It undercuts Mistral's own Medium 3.5. On sale, Large 4 output is $2.09 against $7.50 for Medium 3.5, so there is little reason to stay on Medium for any new work.
- Batch is cheaper again. Mistral's pricing FAQ says batch requests cost 50% less and cached input up to 90% less.

If you would rather use it as a chat app than through an API, it sits inside Mistral's consumer product, which is now called Vibe (formerly Le Chat): Free at $0, Pro at $14.99/month, Team at $24.99/user/month, and Enterprise on request. The pricing page doesn't say which Vibe plans get Large 4, so it is worth checking that before you buy a seat for it. My full Mistral AI pricing guide breaks down the plans.
What a support reply costs on Large 4
Token prices are easier to judge when there is a real workload attached to them. Say a support reply reads 8,000 input tokens (the ticket thread plus a few help center articles pulled in by RAG) and writes 3,000 output tokens, reasoning included. On the sale price that comes to about $0.012 per reply, while at list price it is about $0.023.
That is cheap, and it's why I'd tell most support teams that token price is rarely the line item that decides an AI support rollout. Wrong answers and the hours spent on babysitting the bot cost far more, which is also the theme of my AI support agent cost breakdown.
Mistral Large 4 benchmarks: Mistral's charts vs independent tests
Mistral published a lot of charts for this launch, and they are worth reading carefully, since they are honest in a slightly unusual way. Mistral's own charts show rivals beating Large 4 on four of the six benchmarks I could read.

On coding, it scores 61.7% on DeepSWE v1.1, then 28.3% on Terminal-Bench 4. Kimi K3 is higher on DeepSWE (68), and GLM-5.3 is higher on Terminal-Bench (40), which both come from Mistral's own charts.

The independent picture from Artificial Analysis lines up with that. Here is where Large 4 sits on their Intelligence Index, with the cost to complete one benchmark task:
| Model | AA Intelligence Index | Price in / out per 1M | Cost per task |
|---|---|---|---|
| Claude Sonnet 5.5 (Max) | 56.0 | $2 / $10 | $5.46 |
| Qwen 3.8 Max | 45.4 | $2 / $6 | $5.41 |
| Kimi K3 (Max) | 43.6 | $3 / $15 | $2.00 |
| Claude Haiku 5.5 (Max) | 43.4 | $0.10 / $0.50 | $0.21 |
| Gemini 3.8 Flash (High) | 40.9 | $0.75 / $3.75 | $1.24 |
| DeepSeek V4.1 Flash (Max) | 39.5 | $0.30 / $1.20 | $0.27 |
| Mistral Large 4 Preview | 38.4 | $1.36 / $4.18 | $1.13 |
| GPT-5.6 Luna (Max) | 37.3 | $0.20 / $1.20 | $0.18 |
| DeepSeek V4 Pro (Max) | 36.0 | $1.32 / $3.96 | $0.67 |
| Mistral Large 3 | 9.3 | $0.50 / $1.50 | $0.03 |
There are two readings of the same table, and both of them are true. Against Mistral's past, Large 4 is a generational leap: 4.1x the score of Large 3. Against today's field, it trails Kimi K3 and Qwen 3.8 Max, and lands just below DeepSeek V4.1 Flash and just above GPT-5.6 Luna, both of which cost a fraction as much per task.
The reason for this is token appetite. Artificial Analysis counted 69,330 output tokens per task for Large 4, of which 52,011 were reasoning. Large 3 used 5,566. A reasoning model that thinks this hard will turn a moderate per-token price into a high per-task price.

Artificial Analysis prices its runs at list rates, and because the sale halves every rate, including cached input, the sale price works out to roughly $0.57 per task. That's still about 2x DeepSeek V4.1 Flash for a slightly lower score. If you're comparing open models on cost, my DeepSeek V4 Flash pricing and Kimi K3 pricing posts have the full rate cards.
Human raters tell a kinder story than the index does. In Mistral's human evaluations, run by a third party, Large 4 won 69.2% of head-to-heads against both GLM-5.3 and Kimi K3, and 38.5% against Claude Opus 5.

Where Mistral Large 4 is strongest
Strip away the overall index and a clear shape appears, where Large 4 is not the best all-rounder but still leads in a few specific places.
Cyber defense. This is the headline claim, and the numbers do back it up. Large 4 scored 50 on the Artificial Analysis Cyber Index, ahead of Kimi K3 at 41, GPT-6 Astra at 33 and Claude Opus 5.5 at 29, per Mistral's launch post. It hit 82% on a reproduce-then-patch test and 93% of Cybench. On that patch test, Mistral notes the closed frontier models "score near zero on the same test because they refuse", so part of the lead there is about willingness, not skill.

Mistral co-founder Guillaume Lample made the case directly:
"Open models are the best defense for enterprises today: capable, auditable, self-hosted and governed under enterprise policies -- and more importantly, they do not refuse to help."
Legal agent work. On the Vals.ai Harvey Legal Agent benchmark it scored 15.8, ahead of Kimi K3 at 12.9 and GPT-6 Astra at 5.4. The absolute numbers are low, but the lead is clear.
Long documents and charts. It scored 81.3% on Artificial Analysis's long-context reasoning test (AA-LCR) and 42% on Dense 200 visual grounding, edging GPT-6 Astra's 41%. Plotly's data team saw the same thing in practice (more below).
Safety under attack. It resisted 93.3% of attacks on the Lakera B3 agent security benchmark, the highest score Mistral reported among competitors.
What the community says about Mistral Large 4
The Hacker News launch thread hit 1,990 points and 1,144 comments, and the mood is split in roughly the same way the benchmark table suggests, with people impressed by the leap but unconvinced it beats the field.
The most useful hands-on test came from someone at Plotly who ran it on their data analytics benchmark:
"It's 10x cheaper than Mistral Medium 3.5 from April and goes from 58% to 74% correct. Definitely a generational shift. It's not on the Pareto curve yet, but it's good enough for data analytics"
Price came up a lot in the replies. One commenter put it next to the cheaper "flash" tier models:
"So, I think it would have to be significantly better than GLM 5.3-flash to be worth it. GLM 5.3-flash is already very good."
For the head-to-head with closed models, my Claude vs Mistral and ChatGPT vs Mistral comparisons are a good next read.
Reddit was harsher. On r/LocalLLaMA, a user who runs their own chess benchmark wrote:
"Played around with it a bit, for code analysis, image classifications, and chess capability. Feels quite dated and far behind what I usually utilize."
And yet the sovereignty argument keeps coming back, which is the one I'd take most seriously for business buyers:
"In most corporate workflows, it is probably good enough. And you don't have to sell your soul to Xi or Trump. As a european, I had feared much worse."
The most-liked reply on Mistral's launch post on X was a joke that says a lot about the "open weights" framing: "me with 12 GB of VRAM reading 'open weights, 1.05T parameters'" (An Xuan on X). Which brings me to the part that is, for now, still only a promise.
Open weights: what "open" means today
Mistral tags Large 4 as open-weight, but as of October 8, 2026 you can't download it. There's no Mistral Large 4 repo on Mistral's Hugging Face page, and no license has been named. Lample says the RL run "is still in flight" and that a final version ships with the weights before the end of the month.

That has two practical consequences for you:
- Benchmarks may move. The model you test today is a preview, and Mistral's lifecycle policy gives Public Preview models only one month of notice before changes, against six months for general availability. Re-run your evals when the final version lands.
- The license decides everything. Large 3 shipped under Apache 2.0, while Medium 3.5 uses a modified MIT license. Until Mistral names one for Large 4, "open" is a direction, not a contract.
Even with weights, most teams won't run it themselves, since at 1.05T parameters you are in multi-GPU server territory, so the realistic paths are Mistral's API, a European deployment through Mistral, or a hosted provider once those appear. If you want something you can run on your own hardware today, my list of Mistral alternatives includes smaller open models.
Should you use Mistral Large 4 for customer support?
This is the question I get asked the most, so here is my straight answer: Large 4 can write a good support reply, but it should never answer from its own memory.
On Artificial Analysis's knowledge test (AA-Omniscience), it answered 25.8% of questions correctly and had a 41.9% hallucination rate, meaning that when it didn't know the answer, it often just guessed. In a helpdesk, a confident guess about your refund policy or about your product's compatibility is the worst answer possible.
I've seen this failure mode up close myself. On one eesel sales call, a Danish B2B vehicle-telematics team running Zendesk at about 200 tickets a month, worried their bot would confirm unsupported car models, because their knowledge base said "we support all models". A better model would not have fixed that, because the model was reading the docs correctly and the docs were wrong. That is why every rollout I work on is tested against historical tickets before it goes live, and why the guides on how to stop hallucinations start with the knowledge base, not the model.
A CX lead at a DTC supplements brand, handling around 7,000 tickets, put the real requirement better than any benchmark can:
"I need an AI who is only handling the tickets that it's confident to handle and all the other ones, leave them alone."
Latency is the second thing that needs a check. With reasoning on, Artificial Analysis measured about 18.7 seconds before the first answer token on a long prompt. That is fine for email and ticket replies but slow for live chat, where you'd want a lower reasoning_effort or a faster model. My comparison of the best AI model for tickets goes deeper on this trade-off.
Here's how I'd sum it up for a support team:
| Use case | Large 4 fit | Why |
|---|---|---|
| Drafting replies for agents to review | Good | Strong writing, cheap per reply, a human in the loop checks the output |
| Autonomous email or ticket replies | Only with guardrails | Needs grounding in your docs, a confidence score and a human fallback |
| Live chat | Weak with reasoning on | ~18.7 s to first answer token |
| EU data residency requirements | Strong | Built and served in Europe by Mistral |
| Reading screenshots and PDFs customers attach | Strong | Native image input, good document and chart scores |
If you're choosing between models for a support stack, my guide to the best LLM for support compares the main options, and RAG vs fine-tuning explains why grounding beats training for help centers.
The model is infrastructure, eesel is the employee
Here is the shift that I'd ask you to make. A model like Mistral Large 4 is infrastructure: an engine that predicts text. What a support team actually needs is a teammate that knows the product, works inside the helpdesk, knows when to hand off, and can be checked before it talks to customers. Picking the engine is the easier half of the job.
Karel at GENERAL BYTES said it plainly when explaining why his team didn't build their own:
"We could try to write our own LLM application but we didn't want to invest our time into that. We wanted something that we would not have to maintain."
That's the job of eesel, an AI teammate platform where you hire ready-to-work teammates for specific jobs.
For support, that's the AI helpdesk teammate: it plugs into Zendesk, Freshdesk or Gorgias, learns from your past tickets and help center, and runs a simulation over hundreds of your past tickets so you can see how it would have answered before switching it on. For content teams, the AI blog writer does the same for long-form posts.
If you're weighing it against your helpdesk's built-in AI, I compared eesel vs Zendesk AI separately.
To be clear about the limits here: eesel runs on models from OpenAI, Anthropic and Google, per its security page, so Mistral Large 4 isn't an option you can pick inside eesel today. If your requirement is specifically "a Mistral model, hosted in Europe", calling Mistral's API directly is the more direct route.
If you came to Large 4 because you want API-level control, the eesel CLI gives you that over the teammate itself. npx @eesel/cli init chat-bubble --site <your-site> spins up a chat bubble trained on your site with no account, every command returns JSON so Claude Code, Cursor or Codex can drive the setup, and eesel mcp token turns your workspace into an MCP server. It's the same agent as the dashboard, so anything you set up from the terminal shows up there too. My post on AI agent CLIs covers why that matters.
Try eesel
If you're evaluating Mistral Large 4 because you want an AI that answers support tickets, I'd start one step higher. eesel's helpdesk teammate joins your Zendesk, Freshdesk or Gorgias queue in minutes, learns from tickets your team already answered, and shows you a scored simulation on those past tickets before it replies to anyone. You see the wrong answers in a report, not in a customer's inbox.

The free plan includes 100 credits with no card, and paid plans start at $299/month for 500 tickets or chats, per eesel pricing. Try eesel on a slice of last month's tickets and see how it does.
Frequently Asked Questions
What is Mistral Large 4?
How much does Mistral Large 4 cost?
Is Mistral Large 4 open source?
Is Mistral Large 4 better than Kimi K3 or Qwen 3.8 Max?
Is Mistral Large 4 good for customer support?
What is the Mistral Large 4 API model name?
mistral-large-4, and it accepts a reasoning_effort parameter. It supports function calling, structured outputs, batching and the Agents API. Preview models get one month of notice before changes, so pin your tests to it before you build an AI agent on top.Can I run Mistral Large 4 locally?

Article by
Kira
Kira is a writer at eesel AI with a Computer Science background and over a year of hands-on experience evaluating AI-powered customer service tools. She focuses on breaking down how helpdesk platforms and AI agents actually work so that support teams can make better buying decisions.








