
The official Grok 4.6 price card
Here is the live text-model table from xAI's own page, every model, both context bands, no rounding. The page's own footer dates it to July 3, 2026.
| Model | Context | Input | Cached | Output | Long: in / cached / out |
|---|---|---|---|---|---|
grok-4.6 | 500k | $2.00 | $0.50 | $6.00 | $4.00 / $1.00 / $12.00 |
grok-4.5 | 500k | $2.00 | $0.30 | $6.00 | $4.00 / $0.60 / $12.00 |
grok-build-0.1 | 256k | $1.00 | $0.20 | $2.00 | $2.00 / $0.40 / $4.00 |
grok-4.3 | 1M | $1.25 | $0.20 | $2.50 | $2.50 / $0.40 / $5.00 |
grok-4.20-multi-agent-0309 | 1M | $1.25 | $0.20 | $2.50 | $2.50 / $0.40 / $5.00 |
grok-4.20-0309-reasoning | 1M | $1.25 | $0.20 | $2.50 | $2.50 / $0.40 / $5.00 |
grok-4.20-0309-non-reasoning | 1M | $1.25 | $0.20 | $2.50 | $2.50 / $0.40 / $5.00 |
Three things in that table are worth stopping on.
The 200k threshold is a whole-request flip. xAI's wording is unambiguous: models with long-context pricing bill the long rates "for all tokens in a request" once the prompt reaches the threshold. A 201,000-token prompt does not pay 2x on the last thousand. It pays 2x on all 201,000. For an agentic loop that accumulates tool output turn after turn, drifting past that line on turn nine reprices turn nine from its first token.
The flagship is not the widest context window xAI sells. grok-4.3 and the whole 4.20 family carry 1M; 4.6 carries 500k at twice the output rate. If the job is one enormous document in a single call, the older model is both roomier and cheaper.
Cached input is the only column where 4.6 is dearer than 4.5, at $0.50 against $0.30. Worth being precise about direction: xAI did not raise a price mid-flight. Grok 4.5's cache rate came down to $0.30 after its launch, and 4.6 shipped at the older figure. The effect on today's table is the same either way, and it lands hardest on workloads that resend a fixed prompt. Our Grok 4.5 pricing post has the older ladder if you are comparing generations, and the xAI pricing guide covers the full lineup including Grok Imagine and the voice models.
What people actually pay
This is the part the launch coverage never has, because it needs traffic rather than a press release. OpenRouter publishes both the listed price and the price its users are really charged, aggregated across live requests.

| Endpoint | Effective in | Effective out | Listed | Cache hit rate | Token share |
|---|---|---|---|---|---|
| SpaceXAI | $0.7249 | $6.125 | $2 / $6 | 90.3% | 64.5% |
| SpaceXAI (ZDR) | $0.7812 | $6.116 | $2 / $6 | 88.0% | 35.5% |
| Weighted average | $0.7448 | $6.122 | $2 / $6 |
Read the two columns separately, because they are telling you different things.
On input, the real price is roughly a third of the sticker, and the entire reason is caching. Nine in ten input tokens across live Grok 4.6 traffic are served from cache at $0.50 rather than computed at $2.00. That is not a discount xAI advertises; it is an emergent property of how people actually use the model, in long conversations with a stable prefix.
On output, the effective rate is above list. $6.122 against $6.00 means a slice of real traffic is crossing 200,000 tokens and paying the $12.00 long-context rate. Small in aggregate, but it is direct evidence that the cliff is not theoretical.
The corollary matters more than either number: your cache hit rate is the single biggest lever on your Grok 4.6 bill, worth more than any model swap. And it is not automatic in the way people assume. Caching happens on its own, but xAI's prompt caching docs recommend setting the x-grok-conv-id header specifically "to maximize your cache hit rate", because it routes a conversation's requests to the same server. Skip that header and you can be paying $2.00 where the fleet average pays $0.72.
Practitioners comparing models have noticed the same asymmetry from the other side:
"About the cost, Grok doesn't really cost less outright. Output token is cheaper, but their cache is not that cheap. Grok use less token than Sonnet tho."
That comment is about Grok 4.5, whose cache rate is the cheaper $0.30. On 4.6 the point sharpens rather than softens.
The meters nobody budgets for
Token rates are one row of a much longer bill. Here is everything else xAI can charge you for on the same page.

Server-side tools bill per invocation, on top of tokens.
| Tool | Rate |
|---|---|
| Web search, X search, code execution | $5 / 1k calls |
| File attachment search | $10 / 1k calls |
| Collections search (RAG) | $2.50 / 1k calls |
| Image understanding, X video, remote MCP | Token-based, no invocation fee |
| Image generation | Billed at Imagine API rates |
The one to watch for any retrieval build is collections search at $2.50 per 1,000. That is the meter that fires when your agent queries your own uploaded documents, which is the whole mechanism behind RAG versus a raw LLM. Every grounded answer is a tool call, and every tool call has a price that has nothing to do with the model.
Storage is a separate meter again. File storage is $0.025 per GiB per day, collection storage $0.10 per GiB per day, and downloads $0.20 per GiB transferred. A 50 GiB knowledge collection sitting idle costs $5 a day whether or not anyone asks it a question.
Priority Processing is a flat 2x on everything. Input, output, cached and reasoning tokens all double. Two fair details xAI publishes and deserves credit for: caching discounts apply before the multiplier, and you are only billed at the priority rate when the response confirms "service_tier": "priority". If the request is served at the default tier, you pay standard rates.
The batch discount excludes the flagship. The 20% Batch API discount covers grok-4.3 and the three grok-4.20 variants. It does not cover 4.6, 4.5 or grok-build-0.1. So for overnight, batchable work like ticket classification or backlog triage, the year-old grok-4.3 lands at $1.00 input and $2.00 output after the discount, half the flagship's rate. The main pricing table gives you no hint of that.
A blocked request still costs money. Violations caught before generation in the Responses API carry a $0.05 fee per request. You can pay five cents for output you never received.
Assemble your own bill
Token math alone will understate what you spend. Switch the tools on and watch the ratio move.
Which door you buy through changes the price
Same model, five purchase paths, and two of them do not sell it yet.

The xAI API is the reference price: $2.00 / $0.50 / $6.00, pay as you go, no published free allowance.
OpenRouter lists the identical $2.00 / $6.00 / $0.50 cache-read rate, and says plainly that it passes through provider pricing "without any markup". The cost sits elsewhere: a 5.5% fee when you buy credits, with an $0.80 minimum, 5% on crypto, and unused credits expiring after a year. Bring your own xAI key and usage above the plan allowance carries a 5% fee of what the same request would have cost. It also runs two endpoints, standard and zero-data-retention, at the same rate but different measured speed: 74 versus 80 tokens per second at the median, 1.02s versus 0.81s latency, both at 100% uptime over 30 days.
SuperGrok at $30/month is the cheapest consumer tier that names Grok 4.6 as an entitlement, per xAI's pricing page. SuperGrok Plus is $100/month. Free, SuperGrok Lite, SuperGrok Heavy, Business and Enterprise all appear in the plan comparison, but only Free, SuperGrok and SuperGrok Plus carry a published figure; the rest are contact-sales or unlisted.
Azure AI Foundry and AWS Bedrock are still a generation behind. Azure's Grok pricing page tops out at Grok-4.3 Global, at $1.25 input and $2.50 output. That is cheaper per token, and it is two releases old. If procurement mandates that inference runs through one cloud, Grok 4.6 is not purchasable for you yet at any price. It is the same pattern we saw with Grok Bot and the Grok voice agent builder: xAI ships to its own surfaces first and lets resellers catch up.
Grok 4.6 against the other flagships
| Model | Context | Input /1M | Cached /1M | Output /1M |
|---|---|---|---|---|
| Grok 4.6 | 500k | $2.00 | $0.50 | $6.00 |
| Claude Sonnet 5 | 1M | $2.00 | $0.20 | $10.00 |
| Claude Opus 5 | 1M | $5.00 | $0.50 | $25.00 |
| GPT-5.6 Sol | 1.05M | $5.00 | $0.50 | $30.00 |
| DeepSeek V4 Flash | 1M | $0.14 | $0.0028 | $0.28 |
On output, Grok 4.6 is the cheapest frontier-tier option here that is not open-weight. $6.00 against $10, $15, $25 and $30 is not a close call, and it explains why the model shows up in coding agents first: those workloads are output-heavy and context-fluid.
Flip to the cached column and the ranking inverts. Sonnet 5 caches at $0.20 against Grok 4.6's $0.50, has a 1M window with no long-context surcharge, and carries a batch discount where Grok 4.6 has none. For a workload that resends a large fixed prompt and returns three sentences, which is exactly what a support queue is, the cheaper output rate barely registers.
One fairness note on the comparison itself: cached input is not the same product across vendors. Anthropic and OpenAI charge separately to write a cache entry, Google bills hourly cache storage, and xAI publishes a read rate with no separate write fee on the page. Comparing the cached column alone flatters the vendors that split the charge. If you want the wider field, Grok 4.5 alternatives covers the near neighbours and Qwen3.8 Max covers the open-weight end.
What this actually costs on a support queue
Run the numbers at the widget's defaults and Grok 4.6 costs roughly 1.5 cents of tokens per conversation, plus another cent once retrieval and search are switched on. At 20,000 conversations a month, that is a few hundred dollars.
Now price the rest of the build: the helpdesk integration, the triage rules, the confidence gate, the escalation paths, the review loop, the on-call when something goes sideways. The tokens are rounding error against all of that.
Which is why I keep pointing buyers at cost per resolution rather than cost per million tokens, and why measuring support ROI is a more useful exercise than optimising a rate card.
Two failure modes I see repeatedly on pricing calls, both of which a token table will not warn you about.
The first is volume, not rate. A buyer testing our own product hit 200 interactions in a single evaluation day and immediately started projecting 9,000 a month. Nothing about the price had changed. What changed was their estimate of how often the thing gets used, which is always higher than the spreadsheet said.
The second is the shape of the pricing model itself. One trial user ran twelve test chats, watched the agent perform well, opened the billing page and requested cancellation on the spot. The product worked. The bill's structure was the problem. This is why I would rather quote a support buyer a number per resolved ticket than a number per million tokens: nobody can forecast their own token consumption, and everybody can count their tickets.
And the thing a cheaper model never fixes: grounding in your own content. Grok 4.6's knowledge cutoff means it knows nothing about your refund policy, your SKUs or last week's outage. That comes from your knowledge base and past tickets, and it is a build cost regardless of which model sits underneath.
The same goes for adversarial testing and for containment measurement. Real line items, and no rate card contains them.
eesel prices the outcome, not the tokens
I have spent the last few years shipping the layer between a model and a helpdesk, and the thing I would tell anyone reading a token table with a support project in mind is that you are pricing the cheapest 5% of the work.
eesel is the other 95%. It connects to Zendesk, Freshdesk and the rest in a few minutes, learns from your help centre, macros and resolved tickets so it already sounds like your team, and simulates against your own historical tickets before it replies to a single live customer. You see the accuracy on your data first, fill the gaps it surfaces, re-run, then go live.

Billing works on outcomes rather than meters: 40 cents per ticket or chat handled, per our pricing page, charged on the conversation and not per reply, with no seat fees and nothing charged for tickets your humans take. There is $50 of free usage without a card, you can route a slice of volume first, and you can cap monthly spend. No cache-key header to forget, no 200k cliff, no separate line for the search that produced the answer. Our security and enterprise pages cover the compliance side.
If you are picking a model for a support queue rather than a coding agent, the best AI agent for that job is a better starting point than any rate card, and AI customer service software covers the platform layer above it.
The numbers I would hold a rollout to are automated ticket resolution and resolution rate, not dollars per million tokens. Whether it holds up on real queues is what our customers page is for.
Try eesel free, run the simulation, and price the outcome instead of the tokens.

Article by
Kurnia Kharisma Agung Samiadjie
Kurnia is a software engineer and writer at eesel AI with two years of SEO experience, writing about AI tools, helpdesk software, and customer support. He pairs a developer's understanding of how these products are built with search-driven research into what actually ranks and resonates with the people searching for them.








