Mistral Large 4 pricing: API rates, service tiers, and the real cost per task

Rama Adi
Written by

Rama Adi

Katelin Teen
Reviewed by

Katelin Teen

Last edited October 8, 2026

Expert Verified
Hand-drawn illustration of two people looking at a price tag and a speed gauge next to the Mistral logo

Mistral Large 4 pricing at a glance

I build the API integrations at eesel, and the helpdesk teammate runs on models from OpenAI, Anthropic and Google, so a big part of my week is reading rate cards and then looking past the headline number. Mistral's rate card for Large 4 is short, but it hides more than it shows. Here is everything on Mistral's API pricing page for the flagship family, Standard tier, as of October 8, 2026.

Mistral API pricing page showing Mistral Large 4 at a sale price of $0.68 input, $0.07 cached and $2.09 output next to the struck-through list prices, as taken from Mistral docs
Mistral API pricing page showing Mistral Large 4 at a sale price of $0.68 input, $0.07 cached and $2.09 output next to the struck-through list prices, as taken from Mistral docs
ModelInput (per 1M)Cached input (per 1M)Output (per 1M)Notes
Mistral Large 4 (sale)$0.68$0.07$2.09Public preview, 1M context
Mistral Large 4 (list)$1.36$0.14$4.18Struck through on the page
Mistral Large 3$0.50$0.05$1.50Previous flagship, still listed
Mistral Medium 3.5$1.50$0.15$7.50Previous reasoning model
Mistral Small 4$0.15$0.015$0.60Budget tier
Ministral 3 14B$0.20$0.02$0.20Small dense model
Z.ai GLM 5.3 (hosted by Mistral)$1.40$0.14$4.40Third-party model

There are a few things worth knowing before you put these numbers into a spreadsheet:

  • The sale is exactly 50% off, across all three meters. The Large 4 model card shows the same struck-through pattern and gives no end date.
  • Mistral's own launch post quotes list price. The Mistral Large 4 announcement lists $1.36 / $4.18, so the sale is a promotion layered on top, not the "real" price.
  • Cached input is a 90% discount. $0.07 against $0.68 on sale. It matters more than usual for this model, and I explain why further down.
  • Large 4 costs more per token than Large 3. List price went from $0.50 / $1.50 to $1.36 / $4.18, roughly 2.7x on input and 2.8x on output. Even on sale it is 36% more on input and 39% more on output.

For the history of how Mistral got here, my roundup of new Mistral models covers the Large 3 and Medium 3.5 launches.

What "sale price" actually means for a preview model

Large 4 is a Public Preview on Mistral's API, and the label changes how you should read the price. Mistral's model lifecycle policy says three things about preview models that matter to a buyer:

  1. They are priced at the same rate as General Availability models. So the sale is not a "preview discount" that disappears at GA by design. It is a separate promotion.
  2. They have no guaranteed path to General Availability, and may be retired before reaching it.
  3. The deprecation notice for a preview model is 1 month, against 6 months for a GA model. Once retired, requests return a 404.

Taken together, the safe planning number is the list price, plus a migration plan that you could execute in four weeks. Mistral co-founder Guillaume Lample on X has already said the final version is coming:

"The RL run behind this preview is still in flight and shows no sign of saturation -- we will release a final version before the end of the month along with the weights of the model."

The lifecycle doc also notes that preview models can't use a -latest alias, and that aliases "may expose you to silent updates in model behavior and pricing." My advice is the same one I give for any model in production: pin the exact version identifier and re-price when you move.

One token, four prices: batch, standard, regional and Priority

The pricing page has Standard, Batch and Priority tabs and also a regional inference toggle, while the multipliers sit on separate docs pages. Most coverage of Mistral Large 4 pricing skips this part, and here the bill can more than triple.

Hand-drawn staircase chart titled same output token, four prices: launch sale $2.09, list price $4.18, EU or US regional $4.60, Priority Tier $7.32, per 1M output tokens
Hand-drawn staircase chart titled same output token, four prices: launch sale $2.09, list price $4.18, EU or US regional $4.60, Priority Tier $7.32, per 1M output tokens
TierHow it's pricedInput / cached / output per 1MLatencyUptime SLA
Batch50% discount$0.34 / $0.035 / $1.045 if it stacks on the sale; $0.68 / $0.07 / $2.09 off listQueued, up to 24 hoursNone
Standard (sale)Launch promotion$0.68 / $0.07 / $2.09Seconds to minutesNone
Standard (list)Rate card$1.36 / $0.14 / $4.18Seconds to minutesNone
Regional (EU or US)1.1x standard list$1.496 / $0.154 / $4.598Seconds to minutesNone
Priority Tier1.75x standard list$2.38 / $0.245 / $7.315Seconds99.5%

The sources for each row:

  • Batch. Mistral's batch processing docs promise "a 50% discount" and up to 1 million requests per batch. The docs don't say whether that discount applies to the sale price or the list price. Notice that 50% off list lands exactly on the current sale price.
  • Regional inference. The regional inference page says it is "billed at 1.1× standard list pricing (a 10% upcharge) for input tokens, output tokens, cached reads, and cache writes." Regions today are the EU and the US.
  • Priority Tier. The Priority Tier page says the multiplier "is 1.75x standard list pricing." It requires account setup through sales, falls back to Standard when you exceed your limits, and is the only tier with an uptime SLA.
Mistral docs page explaining Priority Tier, including the requirement to contact an account executive and the service tier comparison table, as taken from Mistral docs
Mistral docs page explaining Priority Tier, including the requirement to contact an account executive and the service tier comparison table, as taken from Mistral docs

This matters more for Mistral than for most vendors. The strongest reason people give for choosing Mistral is that it is European. Mistral's own launch post on X leads with it:

"Forged in Europe end-to-end and is deployable from Europe via our own Mistral Cloud infrastructure."

But the endpoint that guarantees inference runs inside the EU is the regional one, and its 1.1x multiplier is stated against list price. Both docs pages are written in general terms and say nothing about Large 4's sale, so Mistral may apply the multiplier to the discounted rate in practice. I would not budget on that. If residency is the reason you're buying, budget at roughly 2.2x the headline price until your invoice says otherwise.

The same logic goes for a support team that needs an SLA. Priority Tier is the only option with one, and at 1.75x list it is $7.32 per 1M output tokens and $2.38 input. That input rate is higher than the $2 that Claude Sonnet 5.5 pricing starts at, and the output rate is much closer to Sonnet's $10 than to the $2.09 headline.

Why the per-token price undersells the bill

Rate cards compare models per token, but your invoice compares them per job. For reasoning models the two rankings disagree. A reasoning model writes a long internal draft before it answers, then you pay output rates for every word of it.

Artificial Analysis ran the full Intelligence Index on Mistral Large 4 Preview and published a cost breakdown. At list price the whole run cost $1,601.97, and here is where the money went:

Hand-drawn stacked bar titled where $1,602 went: input $767, thinking $742, answer $93, with a note that the part you read is 6%
Hand-drawn stacked bar titled where $1,602 went: input $767, thinking $742, answer $93, with a note that the part you read is 6%
  • Reasoning output: $742.45. Large 4 averaged 69,330 output tokens per task, and 52,011 of them, 75%, were reasoning.
  • Input: $766.79, of which $479.04 was cache reads. Long agentic tasks re-send their context every turn, which is why the 90% cache discount is worth engineering for.
  • The answer you actually read: $92.73. That is under 6% of the bill.

The useful unit is cost per task, then. Artificial Analysis puts Large 4 at $1.13 per Index task at list price, which works out to roughly $0.57 at the sale rate if every meter halves. Next to it, the table compares models that land near it on the same index.

Hand-drawn bar chart of cost per benchmark task: GPT-5.6 Luna $0.18 score 37, DeepSeek V4.1 Flash $0.27 score 40, Mistral Large 4 sale $0.57 score 38, Mistral Large 4 list $1.13 score 38, Gemini 3.8 Flash $1.24 score 41, Kimi K3 $2.00 score 44
Hand-drawn bar chart of cost per benchmark task: GPT-5.6 Luna $0.18 score 37, DeepSeek V4.1 Flash $0.27 score 40, Mistral Large 4 sale $0.57 score 38, Mistral Large 4 list $1.13 score 38, Gemini 3.8 Flash $1.24 score 41, Kimi K3 $2.00 score 44
ModelIntelligence IndexPrice in / out per 1MCost per Index taskOutput tokens per task
GPT-5.6 Luna (Max)37.3$0.20 / $1.20$0.17841,235
DeepSeek V4.1 Flash (Max)39.5$0.30 / $1.20$0.26588,574
Mistral Large 4 Preview (sale, derived)38.4$0.68 / $2.09~$0.5769,330
Mistral Large 4 Preview (list)38.4$1.36 / $4.18$1.13269,330
Gemini 3.8 Flash (High)40.9$0.75 / $3.75$1.24371,003
Kimi K3 (Max)43.6$3 / $15$2.00048,455
Claude Sonnet 5.5 (Xhigh)51.9$2 / $10$2.01274,810

All figures come from each model's Artificial Analysis page, fetched October 8, 2026; the sale-rate row is my arithmetic. My read is that even on sale, Large 4 costs about double DeepSeek V4.1 Flash per task for a score within about a point (the older model's rates are in my DeepSeek V4 Flash pricing guide). It is cheaper per task than Kimi K3, but Kimi scores 5 points higher. The same goes for Gemini 3.8 Flash, whose rates are in my Gemini 3.8 Flash pricing guide.

Hacker News readers noticed the same tension on launch day. One commenter liked the rate card but was not sold:

Hacker News

"The pricing ($.68 in/$.07 cached/$2.09 out) makes it much cheaper than Kimi K3, GLM 5.3, and Meta Muse Spark 1.3. That's great! But also much more expensive than GLM 5.3-flash and Spark 1.3 Contributor (the Meta-takes-your-data pricing of Spark 1.3)."

Another commenter pointed to a different benchmark, where the per-test cost went the other way even against a rival that is pricier per token:

Hacker News

"Otherwise, behind on the broader Pareto frontier, but not by much (Vals Index: 48.05% vs. GLM-5.3's 53.51%; $13.78 vs. $7.25 per test). Many companies will prefer it over Chinease models."

That second quote is the whole lesson in one line: GLM 5.3 costs more per token on Mistral's own price list ($1.40 / $4.40 vs list $1.36 / $4.18) and still came out cheaper per test in that run, because it used fewer tokens to get there.

Mistral Large 4 cost calculator

Put in your own traffic. The calculator uses the published per-1M rates above, applies the regional and Priority multipliers to list price as Mistral's docs describe, and treats batch as 50% off the sale rate. Output tokens should include reasoning, which for Large 4 has averaged about 3 reasoning tokens for every answer token on Artificial Analysis.

Worked examples: what three real workloads would cost

Abstract numbers are easy to nod along to. These are workloads I would actually expect someone to price, run through the published rates.

WorkloadAssumptionsSaleListRegionalPriority
Support chatbot20,000 conversations/month, 8,000 input tokens each (70% cached), 1,200 output tokens$90.64$181.28$199.41$317.24
Same bot, no cachingAs above, 0% cached$158.96$317.92$349.71$556.36
Agentic tasks1,000 tasks/month sized like the Artificial Analysis index~$566$1,132~$1,245~$1,981
Batch document processing1M documents, 3,000 input + 500 output tokens each$3,085$6,170n/an/a

Three takeaways from that table:

  1. Caching is worth more than the sale. On the support bot, turning on caching at list price ($181) almost matches the uncached sale price ($159). If you structure prompts so the system prompt and knowledge base chunks sit at the front and stay stable, you keep most of that saving after the sale ends.
  2. Tier choice moves the bill more than model choice. The same bot spans $91 to $317 a month depending only on which Mistral tier serves it.
  3. The batch doc job is where Large 4's sale shines. If the batch discount stacks on the sale price, that 1M-document job drops to about $1,542. If it doesn't, it lands at $3,085, which is still the sale rate.

Vibe plans: do you need the API at all?

If you just want to use Mistral's models for your own work rather than build on them, the consumer side is Vibe (the product formerly known as Le Chat). Prices from Mistral's pricing page:

Mistral Vibe pricing page showing Free, Pro at $14.99 per month, Team at $24.99 per user per month and Enterprise, as taken from Mistral
Mistral Vibe pricing page showing Free, Pro at $14.99 per month, Team at $24.99 per user per month and Enterprise, as taken from Mistral
PlanPriceWhat you get
Free$0Limited messages, web searches and coding sessions, 100+ connectors, Mistral Studio access
Pro$14.99/monthMore usage, all-day coding in CLI, IDE or web, $25.50/month in API credits
Pro (students)$5.99/monthVerified students; $15/month in API credits
Team$24.99/user/monthHigher limits, up to 30GB storage per user, domain verification, data export; the page's calculator starts at 2 users ($50/month)
EnterpriseContact salesCustom models and agents, audit logs, SAML SSO, white label

One detail worth flagging: the Pro plan includes $25.50 a month of API credits. At the sale rate that's about 12M output tokens of Large 4, which is plenty to prototype with before you commit to the API. What the page does not say is which Vibe plans run Large 4 in the chat app itself, so don't assume Free gets the flagship. My Mistral AI reviews roundup and ChatGPT vs Mistral comparison cover Vibe as a workspace in more depth.

Self-hosting: the price after the weights ship

Mistral says Large 4 is open-weight, with weights due "end of October." Today there is no license named and no public weights, so self-hosted pricing is still a forecast. Two facts shape that forecast.

First, size. The model card lists 1.05T total parameters and 52B active, plus a 1.6B vision encoder. Running a 1T-parameter model needs a multi-GPU server, not a laptop. The most-liked reply on Mistral's launch tweet made the point better than I can:

"me with 12 GB of VRAM reading "open weights, 1.05T parameters""

Second, licensing. Mistral's pricing FAQ says its open-weight models are "Apache 2.0 licensed for research/individual use; while commercial deployments require a Mistral license with separate terms." That line is written about older models like Mistral 7B, and Large 4's license hasn't been published, but it is a reason not to assume "open weights" means "free for commercial use." If self-hosting is your plan, wait for the license text before you price the GPUs. In the meantime, open-source AI agents built on already-released models are the safer bet.

Is Mistral Large 4 worth the price?

The answer depends on which of Mistral's strengths you are paying for, so I would be specific.

Worth it if:

  • You need EU-hosted inference from a European vendor. That's the clearest case, and even at 1.1x list you're paying $4.60 per 1M output tokens, which is still well under most US frontier models. A Reddit commenter in r/singularity put the trade plainly:
Reddit

"In most corporate workflows, it is probably good enough. And you don't have to sell your soul to Xi or Trump. As a european, I had feared much worse."

  • Your work matches its benchmark strengths. Mistral's charts show it leading on cyber defense and Harvey's legal agent benchmark, and Lample says it scores 82% on vulnerability reproduction and patching. For those jobs, $0.57 a task is good value.
  • You're running big batch jobs during the sale. Long-context document work at $0.68 input is cheap, and Artificial Analysis scores its long-context reasoning at 81.3%.

Skip it if:

  • You want the cheapest model at this intelligence level. DeepSeek V4.1 Flash and GPT-5.6 Luna get within about a point for a third to a half of the per-task cost. My DeepSeek V4 Flash review covers the DeepSeek side. For OpenAI's tiers, see GPT-5.6 pricing.
  • Factual recall matters. Artificial Analysis measured a 41.9% hallucination rate on its AA-Omniscience knowledge test. For anything customer-facing, that means grounding every answer in your own docs with RAG.
  • You need stable pricing for a year. A preview model with a 1-month deprecation notice and an undated sale is a moving target.

For a broader view of how Large 4 stacks up, see my Mistral Large 4 overview or the list of Mistral alternatives.

If you're choosing between vendors rather than models, Claude vs Mistral compares the full products. So does my Gemini vs Mistral breakdown.

What the token price leaves out for support teams

Many people pricing Large 4 this week are pricing it for one job, a support bot. If that is you, the table above already shows the uncomfortable part. Even at Priority rates, the model bill for 20,000 conversations is a few hundred dollars. The expensive part is everything around the model, meaning retrieval, fallbacks, testing, the helpdesk connection, and also the person who maintains it all.

That's the trade an eesel customer, Karel at GENERAL BYTES, described when they chose not to build:

"We could try to write our own LLM application but we didn't want to invest our time into that. We wanted something that we would not have to maintain."

The other thing a cheaper model won't fix is a bot that answers when it shouldn't. With a 41.9% hallucination rate on knowledge questions, the controls matter more than the rate card: answer only from your own docs, hand off when unsure with a human in the loop, and test on real past tickets before going live. My write-up on AI hallucinations in support covers the failure modes.

If you're still at the model-picking stage, my guides to the best LLM for support and how to train an AI support agent go deeper on both halves of the job.

The way I see it, the model is infrastructure, and what a support team wants to hire is the employee who uses it. For a wider look at that category, see the best AI helpdesk software roundup.

Try eesel

If you came here pricing Mistral Large 4 for a support bot, there is a shortcut. eesel's AI helpdesk teammate joins your existing Zendesk, Freshdesk or Gorgias queue, learns from your past tickets and help center, and runs a simulation against hundreds of your past tickets before it ever replies to a customer. It's billed per ticket handled, starting at $299 for 500 credits on the pricing page, so whether the model underneath thinks for 1,000 tokens or 50,000 is eesel's problem, not your invoice. Two honest notes: eesel runs on models from OpenAI, Anthropic and Google, not Mistral, and EU data residency is a $500/month add-on.

eesel helpdesk teammate Skills list showing Simulation, Support Analytics, Self Review and Blog Writer skills
eesel helpdesk teammate Skills list showing Simulation, Support Analytics, Self Review and Blog Writer skills

Developers can set it up from a terminal too: the eesel CLI prints JSON for every command, so Claude Code or Cursor can connect your helpdesk, edit instructions and run test chats for you. The free plan includes 100 credits with no card. Try eesel.

Frequently Asked Questions

How much does Mistral Large 4 cost?
Mistral Large 4 pricing is $0.68 per 1M input tokens, $0.07 per 1M cached input tokens and $2.09 per 1M output tokens on the launch sale. The list price is exactly double: $1.36, $0.14 and $4.18. My Mistral AI pricing guide covers the rest of the lineup.
When does the Mistral Large 4 sale price end?
Mistral has not published an end date. The pricing page and model card show the list price struck through with a sale label and nothing else, so budget against the $1.36 / $4.18 list price if you are planning past October 2026. The Mistral Large 4 overview tracks what else is still in preview.
Is Mistral Large 4 cheaper than Kimi K3 or Qwen 3.8 Max?
Per token, yes: on sale it undercuts Kimi K3 pricing ($3 / $15) and Qwen 3.8 Max pricing ($2 / $6). Per Artificial Analysis task it is also cheaper, but both rivals score higher on the same index.
Does Mistral Large 4 have batch pricing?
Yes. Mistral's batch API runs asynchronous jobs at a 50% discount and supports up to 1 million requests per batch. Mistral does not say whether the batch discount stacks on top of the launch sale, so check your first invoice. Batch fits back-office jobs, not live AI customer support.
How much does EU data residency cost on Mistral Large 4?
Regional inference in the EU or US is billed at 1.1x standard list pricing, so Mistral Large 4 output works out to about $4.60 per 1M tokens instead of $2.09 on sale. If you are weighing residency for a support bot, my notes on SOC 2 and GDPR cover the compliance side.
Is Mistral Large 4 free to use?
Not on the API. The free Vibe plan includes limited chat and Mistral Studio access, but Mistral's pricing page does not say which Vibe plans run Large 4. For production tokens, you pay the API rate card.
What does Mistral Large 4 pricing mean for a customer support bot?
A bot handling 20,000 conversations a month at 8,000 input and 1,200 output tokens each costs about $91 a month on sale with good caching, and about $317 on Priority Tier. The token line is usually the small part; see AI support agent cost for the rest of the bill.

Share this article

Rama Adi

Article by

Rama Adi

Rama is a software engineer at eesel AI with two years of experience writing about B2B SaaS, AI tools, and customer support technology. Based in Bali, Indonesia, he brings a developer's perspective to product comparisons — cutting through marketing copy to what the integrations and APIs actually do.

Related Posts

All posts →
Hand-drawn illustration of three people around a table discussing a stack of server blocks topped with the Mistral logo
Trending

Mistral Large 4: specs, pricing, benchmarks, and who should use it

Mistral Large 4 is a 1T-parameter open-weight preview at $0.68/$2.09 per 1M tokens. A huge jump for Mistral, a mid-pack score next to rivals, and a few real niches.

KiraKiraOct 8, 2026
Two people reviewing token meters, per-million-token price cards and a long printed bill
Trending

Anthropic API pricing in 2026: rates and workflow costs

Compare Anthropic token rates, caching and batch costs, then evaluate a ready-made support workflow through eesel CLI with separate billing and controls.

Kurnia KharismaKurnia KharismaAug 13, 2026
Illustration of a multimodal AI model turning inputs into tokens that funnel down to a dollar sign, for a GLM-5.3-Flash pricing breakdown
Trending

GLM-5.3-Flash pricing: every rate, the promo cliff, and the real cost

GLM-5.3-Flash pricing in full: the $0.075/$0.25 promo rates, the September cliff, the coding plan, and the throughput gap that changes your real cost.

Rama AdiRama AdiAug 29, 2026
Illustration of two people reviewing a Cohere Embed 5 pricing dashboard with the Cohere logo
Trending

Cohere Embed 5 pricing: what Pro and Fast really cost in 2026

Cohere Embed 5 pricing is $0.12 per 1M tokens for Pro and $0.08 for Fast, with images at $0.40. Here is the full table, Model Vault math, and the storage bill nobody prices in.

Rama AdiRama AdiOct 1, 2026
Hand-drawn illustration of two people looking at a small price tag and a much larger one beside a big blue Gemini sparkle, for a guide to Gemini 4 Argon pricing
Trending

Gemini 4 Argon pricing: the $2/$10 promo, the $4/$20 bill, and what a task really costs

Gemini 4 Argon pricing is $2/$10 per million tokens for now, then $4/$20. Here's what that means per task, per ticket, and against GPT-6.1 Sol and Claude Opus 5.5.

KiraKiraOct 1, 2026
Illustration of a very long cat stretching beside two people reviewing a scorecard, with the LongCat logo
Trending

LongCat 2.0 review: a real workhorse with one hard blocker

I graded LongCat 2.0 on seven things a buyer actually cares about, using Meituan's own files and the people who ran billions of tokens through it. It scores well on six.

KiraKiraAug 4, 2026
Illustration weighing Alibaba's Qwen 3.8 Max against DeepSeek V4 Flash
Trending

Qwen 3.8 Max vs DeepSeek V4 Flash: price, specs, real verdict

One model costs 21x more per output token than the other. That is the least interesting thing about this comparison, and here is what the specs actually decide.

KiraKiraAug 3, 2026
Hand-drawn illustration of two people working with a voice waveform that branches into audio, storage, and voice profile options, representing Eleven v4 pricing
Trending

Eleven v4 pricing in 2026: API, app credits, and agent minutes explained

Eleven v4 pricing explained: $0.08 per 1K characters at list, $0.04 for Turbo, a launch discount ending October 12, and how credits and agent minutes work.

Kurnia KharismaKurnia KharismaOct 7, 2026
Illustration of three stacked subscription cards with a lightning streak and a usage gauge, representing the ChatGPT Pro 500 tier
Trending

ChatGPT Pro 500: what $500 a month actually buys in 2026

ChatGPT Pro 500 is OpenAI's new $500/month tier with 25x Plus usage and Ultrafast. Here's the per-unit math, the 8x Ultrafast burn, and who it's really for.

Rama AdiRama AdiOct 1, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free