
Mistral Large 4 pricing at a glance
I build the API integrations at eesel, and the helpdesk teammate runs on models from OpenAI, Anthropic and Google, so a big part of my week is reading rate cards and then looking past the headline number. Mistral's rate card for Large 4 is short, but it hides more than it shows. Here is everything on Mistral's API pricing page for the flagship family, Standard tier, as of October 8, 2026.

| Model | Input (per 1M) | Cached input (per 1M) | Output (per 1M) | Notes |
|---|---|---|---|---|
| Mistral Large 4 (sale) | $0.68 | $0.07 | $2.09 | Public preview, 1M context |
| Mistral Large 4 (list) | $1.36 | $0.14 | $4.18 | Struck through on the page |
| Mistral Large 3 | $0.50 | $0.05 | $1.50 | Previous flagship, still listed |
| Mistral Medium 3.5 | $1.50 | $0.15 | $7.50 | Previous reasoning model |
| Mistral Small 4 | $0.15 | $0.015 | $0.60 | Budget tier |
| Ministral 3 14B | $0.20 | $0.02 | $0.20 | Small dense model |
| Z.ai GLM 5.3 (hosted by Mistral) | $1.40 | $0.14 | $4.40 | Third-party model |
There are a few things worth knowing before you put these numbers into a spreadsheet:
- The sale is exactly 50% off, across all three meters. The Large 4 model card shows the same struck-through pattern and gives no end date.
- Mistral's own launch post quotes list price. The Mistral Large 4 announcement lists $1.36 / $4.18, so the sale is a promotion layered on top, not the "real" price.
- Cached input is a 90% discount. $0.07 against $0.68 on sale. It matters more than usual for this model, and I explain why further down.
- Large 4 costs more per token than Large 3. List price went from $0.50 / $1.50 to $1.36 / $4.18, roughly 2.7x on input and 2.8x on output. Even on sale it is 36% more on input and 39% more on output.
For the history of how Mistral got here, my roundup of new Mistral models covers the Large 3 and Medium 3.5 launches.
What "sale price" actually means for a preview model
Large 4 is a Public Preview on Mistral's API, and the label changes how you should read the price. Mistral's model lifecycle policy says three things about preview models that matter to a buyer:
- They are priced at the same rate as General Availability models. So the sale is not a "preview discount" that disappears at GA by design. It is a separate promotion.
- They have no guaranteed path to General Availability, and may be retired before reaching it.
- The deprecation notice for a preview model is 1 month, against 6 months for a GA model. Once retired, requests return a
404.
Taken together, the safe planning number is the list price, plus a migration plan that you could execute in four weeks. Mistral co-founder Guillaume Lample on X has already said the final version is coming:
"The RL run behind this preview is still in flight and shows no sign of saturation -- we will release a final version before the end of the month along with the weights of the model."
The lifecycle doc also notes that preview models can't use a -latest alias, and that aliases "may expose you to silent updates in model behavior and pricing." My advice is the same one I give for any model in production: pin the exact version identifier and re-price when you move.
One token, four prices: batch, standard, regional and Priority
The pricing page has Standard, Batch and Priority tabs and also a regional inference toggle, while the multipliers sit on separate docs pages. Most coverage of Mistral Large 4 pricing skips this part, and here the bill can more than triple.

| Tier | How it's priced | Input / cached / output per 1M | Latency | Uptime SLA |
|---|---|---|---|---|
| Batch | 50% discount | $0.34 / $0.035 / $1.045 if it stacks on the sale; $0.68 / $0.07 / $2.09 off list | Queued, up to 24 hours | None |
| Standard (sale) | Launch promotion | $0.68 / $0.07 / $2.09 | Seconds to minutes | None |
| Standard (list) | Rate card | $1.36 / $0.14 / $4.18 | Seconds to minutes | None |
| Regional (EU or US) | 1.1x standard list | $1.496 / $0.154 / $4.598 | Seconds to minutes | None |
| Priority Tier | 1.75x standard list | $2.38 / $0.245 / $7.315 | Seconds | 99.5% |
The sources for each row:
- Batch. Mistral's batch processing docs promise "a 50% discount" and up to 1 million requests per batch. The docs don't say whether that discount applies to the sale price or the list price. Notice that 50% off list lands exactly on the current sale price.
- Regional inference. The regional inference page says it is "billed at 1.1× standard list pricing (a 10% upcharge) for input tokens, output tokens, cached reads, and cache writes." Regions today are the EU and the US.
- Priority Tier. The Priority Tier page says the multiplier "is 1.75x standard list pricing." It requires account setup through sales, falls back to Standard when you exceed your limits, and is the only tier with an uptime SLA.

This matters more for Mistral than for most vendors. The strongest reason people give for choosing Mistral is that it is European. Mistral's own launch post on X leads with it:
"Forged in Europe end-to-end and is deployable from Europe via our own Mistral Cloud infrastructure."
But the endpoint that guarantees inference runs inside the EU is the regional one, and its 1.1x multiplier is stated against list price. Both docs pages are written in general terms and say nothing about Large 4's sale, so Mistral may apply the multiplier to the discounted rate in practice. I would not budget on that. If residency is the reason you're buying, budget at roughly 2.2x the headline price until your invoice says otherwise.
The same logic goes for a support team that needs an SLA. Priority Tier is the only option with one, and at 1.75x list it is $7.32 per 1M output tokens and $2.38 input. That input rate is higher than the $2 that Claude Sonnet 5.5 pricing starts at, and the output rate is much closer to Sonnet's $10 than to the $2.09 headline.
Why the per-token price undersells the bill
Rate cards compare models per token, but your invoice compares them per job. For reasoning models the two rankings disagree. A reasoning model writes a long internal draft before it answers, then you pay output rates for every word of it.
Artificial Analysis ran the full Intelligence Index on Mistral Large 4 Preview and published a cost breakdown. At list price the whole run cost $1,601.97, and here is where the money went:

- Reasoning output: $742.45. Large 4 averaged 69,330 output tokens per task, and 52,011 of them, 75%, were reasoning.
- Input: $766.79, of which $479.04 was cache reads. Long agentic tasks re-send their context every turn, which is why the 90% cache discount is worth engineering for.
- The answer you actually read: $92.73. That is under 6% of the bill.
The useful unit is cost per task, then. Artificial Analysis puts Large 4 at $1.13 per Index task at list price, which works out to roughly $0.57 at the sale rate if every meter halves. Next to it, the table compares models that land near it on the same index.

| Model | Intelligence Index | Price in / out per 1M | Cost per Index task | Output tokens per task |
|---|---|---|---|---|
| GPT-5.6 Luna (Max) | 37.3 | $0.20 / $1.20 | $0.178 | 41,235 |
| DeepSeek V4.1 Flash (Max) | 39.5 | $0.30 / $1.20 | $0.265 | 88,574 |
| Mistral Large 4 Preview (sale, derived) | 38.4 | $0.68 / $2.09 | ~$0.57 | 69,330 |
| Mistral Large 4 Preview (list) | 38.4 | $1.36 / $4.18 | $1.132 | 69,330 |
| Gemini 3.8 Flash (High) | 40.9 | $0.75 / $3.75 | $1.243 | 71,003 |
| Kimi K3 (Max) | 43.6 | $3 / $15 | $2.000 | 48,455 |
| Claude Sonnet 5.5 (Xhigh) | 51.9 | $2 / $10 | $2.012 | 74,810 |
All figures come from each model's Artificial Analysis page, fetched October 8, 2026; the sale-rate row is my arithmetic. My read is that even on sale, Large 4 costs about double DeepSeek V4.1 Flash per task for a score within about a point (the older model's rates are in my DeepSeek V4 Flash pricing guide). It is cheaper per task than Kimi K3, but Kimi scores 5 points higher. The same goes for Gemini 3.8 Flash, whose rates are in my Gemini 3.8 Flash pricing guide.
Hacker News readers noticed the same tension on launch day. One commenter liked the rate card but was not sold:
"The pricing ($.68 in/$.07 cached/$2.09 out) makes it much cheaper than Kimi K3, GLM 5.3, and Meta Muse Spark 1.3. That's great! But also much more expensive than GLM 5.3-flash and Spark 1.3 Contributor (the Meta-takes-your-data pricing of Spark 1.3)."
Another commenter pointed to a different benchmark, where the per-test cost went the other way even against a rival that is pricier per token:
"Otherwise, behind on the broader Pareto frontier, but not by much (Vals Index: 48.05% vs. GLM-5.3's 53.51%; $13.78 vs. $7.25 per test). Many companies will prefer it over Chinease models."
That second quote is the whole lesson in one line: GLM 5.3 costs more per token on Mistral's own price list ($1.40 / $4.40 vs list $1.36 / $4.18) and still came out cheaper per test in that run, because it used fewer tokens to get there.
Mistral Large 4 cost calculator
Put in your own traffic. The calculator uses the published per-1M rates above, applies the regional and Priority multipliers to list price as Mistral's docs describe, and treats batch as 50% off the sale rate. Output tokens should include reasoning, which for Large 4 has averaged about 3 reasoning tokens for every answer token on Artificial Analysis.
Worked examples: what three real workloads would cost
Abstract numbers are easy to nod along to. These are workloads I would actually expect someone to price, run through the published rates.
| Workload | Assumptions | Sale | List | Regional | Priority |
|---|---|---|---|---|---|
| Support chatbot | 20,000 conversations/month, 8,000 input tokens each (70% cached), 1,200 output tokens | $90.64 | $181.28 | $199.41 | $317.24 |
| Same bot, no caching | As above, 0% cached | $158.96 | $317.92 | $349.71 | $556.36 |
| Agentic tasks | 1,000 tasks/month sized like the Artificial Analysis index | ~$566 | $1,132 | ~$1,245 | ~$1,981 |
| Batch document processing | 1M documents, 3,000 input + 500 output tokens each | $3,085 | $6,170 | n/a | n/a |
Three takeaways from that table:
- Caching is worth more than the sale. On the support bot, turning on caching at list price ($181) almost matches the uncached sale price ($159). If you structure prompts so the system prompt and knowledge base chunks sit at the front and stay stable, you keep most of that saving after the sale ends.
- Tier choice moves the bill more than model choice. The same bot spans $91 to $317 a month depending only on which Mistral tier serves it.
- The batch doc job is where Large 4's sale shines. If the batch discount stacks on the sale price, that 1M-document job drops to about $1,542. If it doesn't, it lands at $3,085, which is still the sale rate.
Vibe plans: do you need the API at all?
If you just want to use Mistral's models for your own work rather than build on them, the consumer side is Vibe (the product formerly known as Le Chat). Prices from Mistral's pricing page:

| Plan | Price | What you get |
|---|---|---|
| Free | $0 | Limited messages, web searches and coding sessions, 100+ connectors, Mistral Studio access |
| Pro | $14.99/month | More usage, all-day coding in CLI, IDE or web, $25.50/month in API credits |
| Pro (students) | $5.99/month | Verified students; $15/month in API credits |
| Team | $24.99/user/month | Higher limits, up to 30GB storage per user, domain verification, data export; the page's calculator starts at 2 users ($50/month) |
| Enterprise | Contact sales | Custom models and agents, audit logs, SAML SSO, white label |
One detail worth flagging: the Pro plan includes $25.50 a month of API credits. At the sale rate that's about 12M output tokens of Large 4, which is plenty to prototype with before you commit to the API. What the page does not say is which Vibe plans run Large 4 in the chat app itself, so don't assume Free gets the flagship. My Mistral AI reviews roundup and ChatGPT vs Mistral comparison cover Vibe as a workspace in more depth.
Self-hosting: the price after the weights ship
Mistral says Large 4 is open-weight, with weights due "end of October." Today there is no license named and no public weights, so self-hosted pricing is still a forecast. Two facts shape that forecast.
First, size. The model card lists 1.05T total parameters and 52B active, plus a 1.6B vision encoder. Running a 1T-parameter model needs a multi-GPU server, not a laptop. The most-liked reply on Mistral's launch tweet made the point better than I can:
"me with 12 GB of VRAM reading "open weights, 1.05T parameters""
Second, licensing. Mistral's pricing FAQ says its open-weight models are "Apache 2.0 licensed for research/individual use; while commercial deployments require a Mistral license with separate terms." That line is written about older models like Mistral 7B, and Large 4's license hasn't been published, but it is a reason not to assume "open weights" means "free for commercial use." If self-hosting is your plan, wait for the license text before you price the GPUs. In the meantime, open-source AI agents built on already-released models are the safer bet.
Is Mistral Large 4 worth the price?
The answer depends on which of Mistral's strengths you are paying for, so I would be specific.
Worth it if:
- You need EU-hosted inference from a European vendor. That's the clearest case, and even at 1.1x list you're paying $4.60 per 1M output tokens, which is still well under most US frontier models. A Reddit commenter in r/singularity put the trade plainly:
"In most corporate workflows, it is probably good enough. And you don't have to sell your soul to Xi or Trump. As a european, I had feared much worse."
- Your work matches its benchmark strengths. Mistral's charts show it leading on cyber defense and Harvey's legal agent benchmark, and Lample says it scores 82% on vulnerability reproduction and patching. For those jobs, $0.57 a task is good value.
- You're running big batch jobs during the sale. Long-context document work at $0.68 input is cheap, and Artificial Analysis scores its long-context reasoning at 81.3%.
Skip it if:
- You want the cheapest model at this intelligence level. DeepSeek V4.1 Flash and GPT-5.6 Luna get within about a point for a third to a half of the per-task cost. My DeepSeek V4 Flash review covers the DeepSeek side. For OpenAI's tiers, see GPT-5.6 pricing.
- Factual recall matters. Artificial Analysis measured a 41.9% hallucination rate on its AA-Omniscience knowledge test. For anything customer-facing, that means grounding every answer in your own docs with RAG.
- You need stable pricing for a year. A preview model with a 1-month deprecation notice and an undated sale is a moving target.
For a broader view of how Large 4 stacks up, see my Mistral Large 4 overview or the list of Mistral alternatives.
If you're choosing between vendors rather than models, Claude vs Mistral compares the full products. So does my Gemini vs Mistral breakdown.
What the token price leaves out for support teams
Many people pricing Large 4 this week are pricing it for one job, a support bot. If that is you, the table above already shows the uncomfortable part. Even at Priority rates, the model bill for 20,000 conversations is a few hundred dollars. The expensive part is everything around the model, meaning retrieval, fallbacks, testing, the helpdesk connection, and also the person who maintains it all.
That's the trade an eesel customer, Karel at GENERAL BYTES, described when they chose not to build:
"We could try to write our own LLM application but we didn't want to invest our time into that. We wanted something that we would not have to maintain."
The other thing a cheaper model won't fix is a bot that answers when it shouldn't. With a 41.9% hallucination rate on knowledge questions, the controls matter more than the rate card: answer only from your own docs, hand off when unsure with a human in the loop, and test on real past tickets before going live. My write-up on AI hallucinations in support covers the failure modes.
If you're still at the model-picking stage, my guides to the best LLM for support and how to train an AI support agent go deeper on both halves of the job.
The way I see it, the model is infrastructure, and what a support team wants to hire is the employee who uses it. For a wider look at that category, see the best AI helpdesk software roundup.
Try eesel
If you came here pricing Mistral Large 4 for a support bot, there is a shortcut. eesel's AI helpdesk teammate joins your existing Zendesk, Freshdesk or Gorgias queue, learns from your past tickets and help center, and runs a simulation against hundreds of your past tickets before it ever replies to a customer. It's billed per ticket handled, starting at $299 for 500 credits on the pricing page, so whether the model underneath thinks for 1,000 tokens or 50,000 is eesel's problem, not your invoice. Two honest notes: eesel runs on models from OpenAI, Anthropic and Google, not Mistral, and EU data residency is a $500/month add-on.

Developers can set it up from a terminal too: the eesel CLI prints JSON for every command, so Claude Code or Cursor can connect your helpdesk, edit instructions and run test chats for you. The free plan includes 100 credits with no card. Try eesel.
Frequently Asked Questions
How much does Mistral Large 4 cost?
When does the Mistral Large 4 sale price end?
Is Mistral Large 4 cheaper than Kimi K3 or Qwen 3.8 Max?
Does Mistral Large 4 have batch pricing?
How much does EU data residency cost on Mistral Large 4?
Is Mistral Large 4 free to use?
What does Mistral Large 4 pricing mean for a customer support bot?

Article by
Rama Adi
Rama is a software engineer at eesel AI with two years of experience writing about B2B SaaS, AI tools, and customer support technology. Based in Bali, Indonesia, he brings a developer's perspective to product comparisons — cutting through marketing copy to what the integrations and APIs actually do.








