Grok 4.6 pricing 2026: every rate, and what teams actually pay

Kurnia Kharisma Agung Samiadjie
Written by

Kurnia Kharisma Agung Samiadjie

Katelin Teen
Reviewed by

Katelin Teen

Last edited August 12, 2026

Expert Verified
Grok 4.6 pricing 2026: every rate, and what teams actually pay

The official Grok 4.6 price card

Here is the live text-model table from xAI's own page, every model, both context bands, no rounding. The page's own footer dates it to July 3, 2026.

ModelContextInputCachedOutputLong: in / cached / out
grok-4.6500k$2.00$0.50$6.00$4.00 / $1.00 / $12.00
grok-4.5500k$2.00$0.30$6.00$4.00 / $0.60 / $12.00
grok-build-0.1256k$1.00$0.20$2.00$2.00 / $0.40 / $4.00
grok-4.31M$1.25$0.20$2.50$2.50 / $0.40 / $5.00
grok-4.20-multi-agent-03091M$1.25$0.20$2.50$2.50 / $0.40 / $5.00
grok-4.20-0309-reasoning1M$1.25$0.20$2.50$2.50 / $0.40 / $5.00
grok-4.20-0309-non-reasoning1M$1.25$0.20$2.50$2.50 / $0.40 / $5.00

Three things in that table are worth stopping on.

The 200k threshold is a whole-request flip. xAI's wording is unambiguous: models with long-context pricing bill the long rates "for all tokens in a request" once the prompt reaches the threshold. A 201,000-token prompt does not pay 2x on the last thousand. It pays 2x on all 201,000. For an agentic loop that accumulates tool output turn after turn, drifting past that line on turn nine reprices turn nine from its first token.

The flagship is not the widest context window xAI sells. grok-4.3 and the whole 4.20 family carry 1M; 4.6 carries 500k at twice the output rate. If the job is one enormous document in a single call, the older model is both roomier and cheaper.

Cached input is the only column where 4.6 is dearer than 4.5, at $0.50 against $0.30. Worth being precise about direction: xAI did not raise a price mid-flight. Grok 4.5's cache rate came down to $0.30 after its launch, and 4.6 shipped at the older figure. The effect on today's table is the same either way, and it lands hardest on workloads that resend a fixed prompt. Our Grok 4.5 pricing post has the older ladder if you are comparing generations, and the xAI pricing guide covers the full lineup including Grok Imagine and the voice models.

What people actually pay

This is the part the launch coverage never has, because it needs traffic rather than a press release. OpenRouter publishes both the listed price and the price its users are really charged, aggregated across live requests.

Bar chart comparing Grok 4.6's listed input price of $2.00 against a measured effective price of $0.74, and a listed output price of $6.00 against a measured $6.12
Bar chart comparing Grok 4.6's listed input price of $2.00 against a measured effective price of $0.74, and a listed output price of $6.00 against a measured $6.12
EndpointEffective inEffective outListedCache hit rateToken share
SpaceXAI$0.7249$6.125$2 / $690.3%64.5%
SpaceXAI (ZDR)$0.7812$6.116$2 / $688.0%35.5%
Weighted average$0.7448$6.122$2 / $6

Read the two columns separately, because they are telling you different things.

On input, the real price is roughly a third of the sticker, and the entire reason is caching. Nine in ten input tokens across live Grok 4.6 traffic are served from cache at $0.50 rather than computed at $2.00. That is not a discount xAI advertises; it is an emergent property of how people actually use the model, in long conversations with a stable prefix.

On output, the effective rate is above list. $6.122 against $6.00 means a slice of real traffic is crossing 200,000 tokens and paying the $12.00 long-context rate. Small in aggregate, but it is direct evidence that the cliff is not theoretical.

The corollary matters more than either number: your cache hit rate is the single biggest lever on your Grok 4.6 bill, worth more than any model swap. And it is not automatic in the way people assume. Caching happens on its own, but xAI's prompt caching docs recommend setting the x-grok-conv-id header specifically "to maximize your cache hit rate", because it routes a conversation's requests to the same server. Skip that header and you can be paying $2.00 where the fleet average pays $0.72.

Practitioners comparing models have noticed the same asymmetry from the other side:

Reddit

"About the cost, Grok doesn't really cost less outright. Output token is cheaper, but their cache is not that cheap. Grok use less token than Sonnet tho."

That comment is about Grok 4.5, whose cache rate is the cheaper $0.30. On 4.6 the point sharpens rather than softens.

The meters nobody budgets for

Token rates are one row of a much longer bill. Here is everything else xAI can charge you for on the same page.

Seven meter cards showing Grok 4.6 billing lines: tokens, web search, file attachments, collections RAG, file storage, priority processing, and blocked requests
Seven meter cards showing Grok 4.6 billing lines: tokens, web search, file attachments, collections RAG, file storage, priority processing, and blocked requests

Server-side tools bill per invocation, on top of tokens.

ToolRate
Web search, X search, code execution$5 / 1k calls
File attachment search$10 / 1k calls
Collections search (RAG)$2.50 / 1k calls
Image understanding, X video, remote MCPToken-based, no invocation fee
Image generationBilled at Imagine API rates

The one to watch for any retrieval build is collections search at $2.50 per 1,000. That is the meter that fires when your agent queries your own uploaded documents, which is the whole mechanism behind RAG versus a raw LLM. Every grounded answer is a tool call, and every tool call has a price that has nothing to do with the model.

Storage is a separate meter again. File storage is $0.025 per GiB per day, collection storage $0.10 per GiB per day, and downloads $0.20 per GiB transferred. A 50 GiB knowledge collection sitting idle costs $5 a day whether or not anyone asks it a question.

Priority Processing is a flat 2x on everything. Input, output, cached and reasoning tokens all double. Two fair details xAI publishes and deserves credit for: caching discounts apply before the multiplier, and you are only billed at the priority rate when the response confirms "service_tier": "priority". If the request is served at the default tier, you pay standard rates.

The batch discount excludes the flagship. The 20% Batch API discount covers grok-4.3 and the three grok-4.20 variants. It does not cover 4.6, 4.5 or grok-build-0.1. So for overnight, batchable work like ticket classification or backlog triage, the year-old grok-4.3 lands at $1.00 input and $2.00 output after the discount, half the flagship's rate. The main pricing table gives you no hint of that.

A blocked request still costs money. Violations caught before generation in the Responses API carry a $0.05 fee per request. You can pay five cents for output you never received.

Assemble your own bill

Token math alone will understate what you spend. Switch the tools on and watch the ratio move.

Which door you buy through changes the price

Same model, five purchase paths, and two of them do not sell it yet.

Five doors representing Grok 4.6 purchase paths, with the xAI API, OpenRouter and SuperGrok open and Azure AI Foundry and AWS Bedrock locked at Grok 4.3
Five doors representing Grok 4.6 purchase paths, with the xAI API, OpenRouter and SuperGrok open and Azure AI Foundry and AWS Bedrock locked at Grok 4.3

The xAI API is the reference price: $2.00 / $0.50 / $6.00, pay as you go, no published free allowance.

OpenRouter lists the identical $2.00 / $6.00 / $0.50 cache-read rate, and says plainly that it passes through provider pricing "without any markup". The cost sits elsewhere: a 5.5% fee when you buy credits, with an $0.80 minimum, 5% on crypto, and unused credits expiring after a year. Bring your own xAI key and usage above the plan allowance carries a 5% fee of what the same request would have cost. It also runs two endpoints, standard and zero-data-retention, at the same rate but different measured speed: 74 versus 80 tokens per second at the median, 1.02s versus 0.81s latency, both at 100% uptime over 30 days.

SuperGrok at $30/month is the cheapest consumer tier that names Grok 4.6 as an entitlement, per xAI's pricing page. SuperGrok Plus is $100/month. Free, SuperGrok Lite, SuperGrok Heavy, Business and Enterprise all appear in the plan comparison, but only Free, SuperGrok and SuperGrok Plus carry a published figure; the rest are contact-sales or unlisted.

Azure AI Foundry and AWS Bedrock are still a generation behind. Azure's Grok pricing page tops out at Grok-4.3 Global, at $1.25 input and $2.50 output. That is cheaper per token, and it is two releases old. If procurement mandates that inference runs through one cloud, Grok 4.6 is not purchasable for you yet at any price. It is the same pattern we saw with Grok Bot and the Grok voice agent builder: xAI ships to its own surfaces first and lets resellers catch up.

Grok 4.6 against the other flagships

ModelContextInput /1MCached /1MOutput /1M
Grok 4.6500k$2.00$0.50$6.00
Claude Sonnet 51M$2.00$0.20$10.00
Claude Opus 51M$5.00$0.50$25.00
GPT-5.6 Sol1.05M$5.00$0.50$30.00
DeepSeek V4 Flash1M$0.14$0.0028$0.28

On output, Grok 4.6 is the cheapest frontier-tier option here that is not open-weight. $6.00 against $10, $15, $25 and $30 is not a close call, and it explains why the model shows up in coding agents first: those workloads are output-heavy and context-fluid.

Flip to the cached column and the ranking inverts. Sonnet 5 caches at $0.20 against Grok 4.6's $0.50, has a 1M window with no long-context surcharge, and carries a batch discount where Grok 4.6 has none. For a workload that resends a large fixed prompt and returns three sentences, which is exactly what a support queue is, the cheaper output rate barely registers.

One fairness note on the comparison itself: cached input is not the same product across vendors. Anthropic and OpenAI charge separately to write a cache entry, Google bills hourly cache storage, and xAI publishes a read rate with no separate write fee on the page. Comparing the cached column alone flatters the vendors that split the charge. If you want the wider field, Grok 4.5 alternatives covers the near neighbours and Qwen3.8 Max covers the open-weight end.

What this actually costs on a support queue

Run the numbers at the widget's defaults and Grok 4.6 costs roughly 1.5 cents of tokens per conversation, plus another cent once retrieval and search are switched on. At 20,000 conversations a month, that is a few hundred dollars.

Now price the rest of the build: the helpdesk integration, the triage rules, the confidence gate, the escalation paths, the review loop, the on-call when something goes sideways. The tokens are rounding error against all of that.

Which is why I keep pointing buyers at cost per resolution rather than cost per million tokens, and why measuring support ROI is a more useful exercise than optimising a rate card.

Two failure modes I see repeatedly on pricing calls, both of which a token table will not warn you about.

The first is volume, not rate. A buyer testing our own product hit 200 interactions in a single evaluation day and immediately started projecting 9,000 a month. Nothing about the price had changed. What changed was their estimate of how often the thing gets used, which is always higher than the spreadsheet said.

The second is the shape of the pricing model itself. One trial user ran twelve test chats, watched the agent perform well, opened the billing page and requested cancellation on the spot. The product worked. The bill's structure was the problem. This is why I would rather quote a support buyer a number per resolved ticket than a number per million tokens: nobody can forecast their own token consumption, and everybody can count their tickets.

And the thing a cheaper model never fixes: grounding in your own content. Grok 4.6's knowledge cutoff means it knows nothing about your refund policy, your SKUs or last week's outage. That comes from your knowledge base and past tickets, and it is a build cost regardless of which model sits underneath.

The same goes for adversarial testing and for containment measurement. Real line items, and no rate card contains them.

eesel prices the outcome, not the tokens

I have spent the last few years shipping the layer between a model and a helpdesk, and the thing I would tell anyone reading a token table with a support project in mind is that you are pricing the cheapest 5% of the work.

eesel is the other 95%. It connects to Zendesk, Freshdesk and the rest in a few minutes, learns from your help centre, macros and resolved tickets so it already sounds like your team, and simulates against your own historical tickets before it replies to a single live customer. You see the accuracy on your data first, fill the gaps it surfaces, re-run, then go live.

The eesel reports dashboard showing task volume, trigger events by type, and tool approval rates for a connected helpdesk
The eesel reports dashboard showing task volume, trigger events by type, and tool approval rates for a connected helpdesk

Billing works on outcomes rather than meters: 40 cents per ticket or chat handled, per our pricing page, charged on the conversation and not per reply, with no seat fees and nothing charged for tickets your humans take. There is $50 of free usage without a card, you can route a slice of volume first, and you can cap monthly spend. No cache-key header to forget, no 200k cliff, no separate line for the search that produced the answer. Our security and enterprise pages cover the compliance side.

If you are picking a model for a support queue rather than a coding agent, the best AI agent for that job is a better starting point than any rate card, and AI customer service software covers the platform layer above it.

The numbers I would hold a rollout to are automated ticket resolution and resolution rate, not dollars per million tokens. Whether it holds up on real queues is what our customers page is for.

Try eesel free, run the simulation, and price the outcome instead of the tokens.

Share this article

Kurnia Kharisma Agung Samiadjie

Article by

Kurnia Kharisma Agung Samiadjie

Kurnia is a software engineer and writer at eesel AI with two years of SEO experience, writing about AI tools, helpdesk software, and customer support. He pairs a developer's understanding of how these products are built with search-driven research into what actually ranks and resonates with the people searching for them.

Related Posts

All posts →
Illustration of two people reviewing a stack of cost layers, seats on top and token burn at the bottom, with the Grok mark at left
Trending

Grok Bot pricing 2026: the $200 plan and the uncapped meter

Grok Bot sells at $200 a month or $120 a seat, but the sticker is only the entry fee. The docs say there is no spend cap yet, and the meter runs weekly.

Kurnia Kharisma Agung SamiadjieKurnia Kharisma Agung SamiadjieAug 13, 2026
Hand-drawn illustration with the Grok logomark and a review scorecard showing mixed star ratings across categories
Trending

Grok 4.6 review: what the eval table says once you read the losing rows

Grok 4.6 ties GPT-5.6 Sol at a third of the price, and loses two benchmarks badly. I read xAI's own eval table row by row, then checked the number the launch post left out.

Alicia Kirana UtomoAlicia Kirana UtomoAug 13, 2026
Hand-drawn illustration of two people reviewing a pricing breakdown chart with dollar signs
Trending

Grok 4.5 pricing: API rates, SuperGrok cost, and hidden fees

Grok 4.5 costs $2/$6 per 1M tokens on the API and $30-$300/month as a consumer plan. Here's the full breakdown, the hidden fees, and what it means to budget with.

Kurnia Kharisma Agung SamiadjieKurnia Kharisma Agung SamiadjieJul 9, 2026
Two people reading a large usage meter and adjusting a stack of billing dials, in Meta's blue brand colour
Trending

Meta Muse Spark 1.1 pricing: the bill has four meters

Muse Spark 1.1 lists at $1.25 in and $4.25 out per million tokens. Four separate meters decide your real bill, and the sticker is the smallest of them.

Rama Adi NugrahaRama Adi NugrahaAug 5, 2026
Illustration of a very long cat stretching beside two people reviewing a scorecard, with the LongCat logo
Trending

LongCat 2.0 review: a real workhorse with one hard blocker

I graded LongCat 2.0 on seven things a buyer actually cares about, using Meituan's own files and the people who ran billions of tokens through it. It scores well on six.

Alicia Kirana UtomoAlicia Kirana UtomoAug 4, 2026
DeepSeek V4 Flash pricing: what you'll actually be billed
Trending

DeepSeek V4 Flash pricing: what you'll actually be billed

DeepSeek V4 Flash lists at $0.14 in and $0.28 out per million tokens. Real users have posted blended rates under a cent. Here is what decides which one you get.

Alicia Kirana UtomoAlicia Kirana UtomoAug 4, 2026
DeepSeek V4 Flash: specs, pricing, and what it's really for
Trending

DeepSeek V4 Flash: specs, pricing, and what it's really for

DeepSeek V4 Flash costs $0.14 in and $0.28 out per million tokens, and it outscores DeepSeek's own expensive tier. Here's what the price card doesn't tell you.

Rama Adi NugrahaRama Adi NugrahaAug 4, 2026
Illustration of a very long cat stretched across a desk beside a server rack, with the LongCat logo
Trending

LongCat 2.0: inside Meituan's 1.6T open-weight model

LongCat 2.0 is Meituan's MIT-licensed 1.6T MoE model, priced at $0.30 per million input tokens. I read every primary source to see what actually ships.

Rama Adi NugrahaRama Adi NugrahaAug 4, 2026
Illustration comparing DeepSeek V4 Flash and Moonshot AI's Kimi K3
Trending

DeepSeek V4 Flash vs Kimi K3: which one should you run?

One model costs 29 times more per task than the other. I went through every published number on both, and the interesting part is the option in the middle that nobody should buy.

Alicia Kirana UtomoAlicia Kirana UtomoAug 4, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free