
Claude Sonnet 5.5 pricing at a glance
Claude Sonnet 5.5 launched on September 28, 2026, and Anthropic kept the price flat. This is the full API rate card from the Claude pricing docs, per million tokens, next to its two nearest siblings:
| Price per 1M tokens | Claude Sonnet 5.5 | Claude Opus 5.5 | Claude Haiku 4.5 |
|---|---|---|---|
| Input | $2 | $4 | $1 |
| Output (thinking included) | $10 | $20 | $5 |
| 5-minute cache write | $2.50 | $5 | $1.25 |
| 1-hour cache write | $4 | $8 | $2 |
| Cache read | $0.20 | $0.20 | $0.10 |
| Batch input / output | $1 / $5 | $2 / $10 | $0.50 / $2.50 |
| Fast mode | Not offered | $8 / $40 | Not offered |
| Context window | 1M | 1M | 200K |
| Max output | 128K (300K on Batch beta) | 128K (300K on Batch) | 64K |
| Long-context surcharge | None | None | n/a |

A few details on that table matter more than they first look:
- Thinking bills as output. Sonnet 5.5 uses adaptive thinking, so every reasoning token costs $10 per million, the same as the answer.
- The API default effort is
high, per the model overview. The Claude apps and Claude Code default to Medium. That one setting changes your bill more than anything else on this page. - The minimum cacheable prompt is 512 tokens, so even a short support system prompt can be cached.
- The tokenizer is unchanged from Sonnet 5, so the same text produces the same token counts. Any saving comes from the model writing less, not from friendlier counting.
If you are coming from the previous generation, nothing on the sticker moved from Sonnet 5 pricing, and it is still below the $3 and $15 of Claude Sonnet 4.6. The thing Anthropic is pitching is that the bill still drops, because the model finishes the same work in fewer tokens.
Half the price of Opus on every line but one
Here is the part most pricing summaries skip. Sonnet 5.5 is exactly half of Claude Opus 5.5 on input, output, cache writes and Batch. On cache reads, it is not cheaper at all.

That happened because Opus 5.5 got an unusually cheap cache. Its cache read multiplier is 0.05x of input, where Sonnet uses the standard 0.1x, per the prompt caching table. Opus input is $4, so 0.05x lands on $0.20. Sonnet input is $2, so 0.1x also lands on $0.20. Same number, two routes.
For a chatbot that sends one prompt and gets one answer, this barely matters, but for an agent it matters quite a lot. Every turn of an agent loop re-sends the whole conversation so far, and with caching on, most of that re-send is billed as cache reads. The longer the session runs, the bigger that line gets, and it is the one line where picking the smaller model saves you nothing.
The community spotted it on launch day:
"Cache read is the same as Opus as well where most agentic workflow cost comes from. Not quite sure where this fits well."
What a real run actually costs
A rate card tells you the price of a token. It does not tell you how many tokens a model spends to finish a job, and that is where Sonnet 5.5 gets interesting. Artificial Analysis ran its Intelligence Index at every effort level and published the measured cost per task.

The page is blunt about the cost. At max effort, Sonnet 5.5 generated 410M output tokens across the index against a class median of 88M, and AA's own summary calls it "very verbose". Here is the full ladder from AA's Sonnet 5.5 page, next to the matching rows from its Opus 5.5 page:
| Effort | Sonnet 5.5 score | Sonnet 5.5 cost per task | Output tokens per task | Opus 5.5 score | Opus 5.5 cost per task |
|---|---|---|---|---|---|
| Low | 36 | $0.41 | 13.9k | 42 | $0.55 |
| Medium | 41 | $0.59 | 19.5k | 51 | $1.34 |
| High | 47 | $1.08 | 34.5k | 54 | $1.82 |
| Xhigh | 52 | $2.74 | 73.7k | 56 | $3.46 |
| Max | 56 | $7.60 | 192.8k | 58 | $5.98 |
Read down the Sonnet column and the cost roughly doubles at each step above Medium, then nearly triples going from Xhigh to Max. Read across and the pattern flips around High. Sonnet 5.5 at max effort costs 27% more per task than Opus 5.5 at max, while scoring two points lower. It used 193k output tokens per task there, against Opus's 119k.

The pair I would stare at is Sonnet Xhigh against Opus Medium. They score 52 and 51, and Opus costs $1.34 to Sonnet's $2.74. If you find yourself turning Sonnet 5.5 up past High to get the quality you need, you end up paying more than you would for the bigger model at its default setting.
This is also the honest reading of Anthropic's own "up to 30% less per task" claim. The launch post says Sonnet 5.5 "complements Opus 5.5 best when running at lower effort settings, where it costs less per task. At higher settings, it can perform comparably at a similar cost." The AA numbers back that up exactly.
Where the money goes on an agent bill
AA also splits each run's cost by token type, and this is where the cache-read line from earlier stops being theoretical. For the full max-effort Intelligence Index run, the Sonnet 5.5 bill came to $8,977:

| Line item | Sonnet 5.5 at max | Share | Sonnet 5.5 at medium | Share |
|---|---|---|---|---|
| Cache reads | $3,605 | 40% | $180 | 26% |
| Cache writes | $1,195 | 13% | $155 | 22% |
| Fresh input | $71 | 1% | $70 | 10% |
| Output and thinking | $4,106 | 46% | $295 | 42% |
| Total | $8,977 | $701 |
Two things jump out from this. Fresh input, the line most people picture when they think "input tokens", is close to nothing, because almost every input token in an agent run is a cache read or a cache write. And cache reads, the line with no Sonnet discount, grow from a quarter of the bill at Medium to 40% at max, because longer thinking means more turns and more re-reads.
Put the two models side by side on the same run and the gap closes, which is the same story the older Opus 5 vs Sonnet 5 comparison told. At Medium, the Sonnet 5.5 index run cost $701 against Opus 5.5's $1,627, so Sonnet was 43% of the Opus bill. At max, Sonnet cost $8,977 against Opus's $8,708. Half the sticker price, 3% more on the invoice.
Worked examples at three workload shapes
Benchmarks are someone else's workload, so here is the rate-card math for three common shapes, using only the published prices. Treat the token counts as illustrative, and plug your own into the calculator below.
1. A support reply bot. Each reply sends 8,000 input tokens, 6,000 of them a cached system prompt and help articles, and gets 1,500 tokens back including thinking.
- Sonnet 5.5: $0.004 fresh input + $0.0012 cache reads + $0.015 output = about 2 cents per reply, or $202 per 10,000 tickets
- Opus 5.5: $0.008 + $0.0012 + $0.03 = about 3.9 cents, or $392 per 10,000
- Haiku 4.5: $0.002 + $0.0006 + $0.0075 = about 1 cent, or $101 per 10,000
Output is the thing that dominates here, so Sonnet's half price on output carries through almost fully. This is the shape where Sonnet 5.5 is plainly the better buy than Opus. For the wider picture of what AI support costs beyond the model, see the AI customer service cost guide.
2. A long agent session. Picture a coding or ops agent that runs 30 turns, re-reads a 60,000-token context from cache each turn, adds 3,000 fresh tokens and writes 2,000 per turn.
- Sonnet 5.5: $0.36 cache reads + $0.18 fresh input + $0.60 output = $1.14
- Opus 5.5: $0.36 + $0.36 + $1.20 = $1.92
Now suppose Sonnet needs 40 turns where Opus needs 25, which is the direction the AA token counts point at higher effort. Sonnet becomes $0.48 + $0.24 + $0.80 = $1.52, and Opus becomes $0.30 + $0.30 + $1.00 = $1.60. The "half price" model is now within 8 cents of the expensive one.
3. An overnight Batch job. You extract fields from 10,000 documents at 20,000 input tokens and 1,000 output tokens each, and nothing needs an answer in real time.
- Sonnet 5.5 on Batch: $200 input + $50 output = $250
- Opus 5.5 on Batch: $400 + $100 = $500
- Haiku 4.5 on Batch: $100 + $25 = $125
There are no cache re-reads and no agent loop, so the 2x ratio holds pretty cleanly. Batch also stacks with prompt caching, per the pricing docs.
Estimate your own monthly bill
Set the per-task token counts for your workload and see what each model costs per month. Cache writes are left out for simplicity, since on a warm cache they are a small share of a steady workload.
If the agent preset surprises you, that is the cache-read line doing its work. With 1.8M cached tokens per session, Sonnet and Opus pay the same $0.36 for the re-reads, and only the output line keeps them apart.
Effort is the price knob, and the API default is high
Anthropic lets you set effort per request, and it is the single biggest lever on a Sonnet 5.5 bill. Simon Willison ran the same SVG prompt at every level and posted the costs in the launch thread:
| Effort | Cost for one prompt | Time |
|---|---|---|
| Low | 1.6 cents | 10s |
| Medium | 1.8 cents | 11s |
| High | 2.3 cents | 17s |
| Xhigh | 5.7 cents | 41s |
| Max | $1.28 | 15m 40s, and it failed |
At max, the model "burned through 128,000 thinking tokens (taking 15 minutes to do that) and ran out before it had produced the final SVG", in his words. That is an 80x cost jump from Low to Max for a worse result.
My practical advice, as someone who wires these calls into production code:
- Set effort explicitly on every request. The API default is
high, not the Medium you see in the apps, so a naive integration starts one notch above the cheap zone. - Start at Low for routine work like classification, tagging and short replies, then step up only where your evals show a real gap.
- Treat Xhigh and Max as a signal to switch models. Past High, Opus 5.5 at Medium tends to beat Sonnet on both score and price.
- Re-run your effort tests if you are migrating from Sonnet 5. Anthropic recalibrated the levels in the what's new page, so old settings do not carry over.
The upside is large at the bottom of the dial. AA measured Sonnet 5.5 at Medium scoring 41 for $0.59 per task, against Claude Sonnet 5 at max scoring 38 for $5.09. That is a better result for about a ninth of the cost, which is the upgrade story Anthropic is selling and it holds up.
Discounts, add-ons and the smaller line items
Beyond effort, four modifiers move the per-token price, all from the Claude pricing docs:
| Modifier | Effect on Sonnet 5.5 | Notes |
|---|---|---|
| Prompt caching | Reads at $0.20 (0.1x), writes at $2.50 (5 min) or $4 (1 hour) | Pays off after one read on the 5-minute cache |
| Batch API | 50% off: $1 input, $5 output | Stacks with caching; 300K output with a beta header |
| US-only inference | 1.1x on every token category | Set with inference_geo: "us" |
| Fast mode | Not available | Opus 5.5 only, at $8 / $40 |
The tool fees sit on top of tokens. Web search is $10 per 1,000 searches, web fetch has no extra charge, and code execution is free when used alongside web search or fetch, otherwise 1,550 free container-hours a month then $0.05 an hour. Sonnet 5.5 adds 286 system-prompt tokens whenever tools are present. Claude Managed Agents sessions add $0.08 per running session-hour on top of standard token rates.
One cost that is easy to miss: forced tool use (tool_choice set to any or a named tool) now returns a 400 on Sonnet 5.5, per the migration notes. It is not a price, but if your support bot depended on it, the rework is a real cost of switching. The hub post walks through all five breaking changes.
Claude Free, Pro, Max, Team and Enterprise
If you use Sonnet 5.5 in the Claude apps rather than the API, you pay per seat, not per token. From the Claude pricing page:
| Plan | Price | Sonnet 5.5 access | Notes |
|---|---|---|---|
| Free | $0 | Yes | Anthropic says "anyone can chat with Claude using Sonnet 5.5" |
| Pro | $17/mo billed annually, $20 monthly | Yes | Adds Opus and usage credits |
| Max 5x / 20x | From $100/mo | Yes | 5x or 20x Pro usage, priority access |
| Team standard | $20/seat/mo annually, $25 monthly | Yes | For 2 to 150 people |
| Team premium | $100/seat/mo annually, $125 monthly | Yes | 5x the standard seat's usage |
| Enterprise | $20/seat/mo plus usage at API rates | Yes | Usage bills at the token rates above |
Two things shift the math here. On Free and Pro, Sonnet is the model most people will actually use, and free users now get Sonnet 5.5. On Enterprise, the seat fee is only the start, because usage bills at the same API rates as this whole post. The Claude pricing guide and the Claude Pro breakdown cover the usage limits, and the Claude Code pricing guide covers the developer side, where Medium effort is the default.
A developer on launch day put the subscription-versus-API split well:
"Sonnet 5.5 might not be useful or necessary in a Claude Code session but it could be good value in the API when you pay per token."
Bedrock, Google Cloud and Foundry
Sonnet 5.5 shipped on every major cloud on day one, as anthropic.claude-sonnet-5-5 on Amazon Bedrock and claude-sonnet-5-5 on Google Cloud and Microsoft Foundry, per the model overview. How you get billed depends on which door you use:
- Amazon Bedrock and Google Cloud are partner-operated. They invoice you directly at their own published rates, so check Bedrock pricing before you assume parity. Google publishes its own Vertex AI rates too.
- Claude Platform on AWS and Claude in Microsoft Foundry are Anthropic-operated. They rate your usage at the standard Anthropic prices, apply any negotiated discount, and convert to Claude Consumption Units at $0.01 each, billed hourly in arrears through the marketplace.
- US-only inference carries the same 1.1x multiplier on the first-party API, Claude Platform on AWS, and Foundry's US Data Zone deployments.
If you already run Claude through a cloud, the Bedrock guide and the Vertex AI guide cover the setup.
How Sonnet 5.5 prices against GPT-6 Sol and Gemini 3.8 Flash
The most striking comparison is not inside Anthropic's lineup. GPT-6 Sol has an identical headline rate card: $2 input, $10 output, $0.20 cached input, $2.50 cache write, per the GPT-6 Sol pricing breakdown.
| Per 1M tokens | Claude Sonnet 5.5 | GPT-6 Sol | Gemini 3.8 Flash |
|---|---|---|---|
| Input | $2 | $2 | $0.75 (rises to $1.50 on Jan 1, 2027) |
| Output | $10 | $10 | $3.75 (rises to $7.50) |
| Cache read | $0.20 | $0.20 | $0.075 |
| Long-context premium | None to 1M | Higher rate above 272K input | None |
| AA cost per task (setting) | $1.08 (high) | $1.05 (max) | Not comparable |
| AA Intelligence Index | 47 (high) | 48 (max) | Not comparable |
The AA rows are the tell. Sonnet 5.5 at High and GPT-6 Sol at max land at nearly the same score for nearly the same cost per task. The difference sits more in the edges: Sonnet has no long-context surcharge, so a support agent that loads a big knowledge base into context does not cross a price cliff, and Sol does above 272K. Gemini's AA figures came from an earlier version of the index, so I left them out of the head-to-head, and Gemini 3.8 Flash pricing doubles in January. The three-way API comparison goes further if you are choosing a provider rather than a model.
For support tickets specifically, the ranking is less about price than about reliability on your own data, which is why the support-ticket model ranking puts accuracy first.
What developers are saying about the price
The launch-day reaction on Hacker News was mostly about where it sits against Opus. The skeptical view:
"I feel like sonnet is priced too close to opus right now. If Sonnet 5.5 were half its current price it would make sense to use."
The measured view, which matches the AA ladder above:
"Per the charts, there is largely no point to using Sonnet 5.5 at high+ as opus low generally will give similar performance at similar or lower cost."
And the enthusiastic one, from a user comparing it to OpenAI's flagship at a lower setting:
"Sonnet 5.5 medium scores better than gpt 5.6 Sol high and is less than 1/5 the cost Absolutely bonkers"
The pattern people kept landing on was a split setup: Sonnet 5.5 at Low or Medium doing the high-volume work as a sub-agent, with Opus 5.5 orchestrating or finishing the hard parts. If neither fits, the Claude alternatives roundup covers other routes. That matches the pricing, since Sonnet's discount is real on output-heavy, short-context calls and fades on long, cache-heavy loops.
Is Claude Sonnet 5.5 worth $2 and $10?
Yes, at Low and Medium effort, and I would not hesitate there. It beats Sonnet 5's best score for a fraction of the cost, it is half the price of Opus on the output line that dominates chat-style workloads, and it has no long-context surcharge. For a support reply bot, a classifier, a drafting step or an overnight Batch job, it is the default I would reach for.
It is a weaker buy in two cases. If your workload is a long agent loop, cache reads eat into the discount because they cost the same on both models. And if you need Xhigh or Max to get the quality you want, Opus 5.5 at Medium is usually cheaper for the same score. Test Sonnet at Low first, measure, then decide whether the gap is worth $1.34 a task.
If the job is a support queue, one more cost belongs in the math: planning for spend that moves with effort and volume. One buyer in eesel's sales calls hit 200 interactions in a single test day and immediately worried about per-interaction costs at scale. Token bills make that worry harder to answer, because the same ticket can cost 1.6 cents or $1.28 depending on a setting.
Try eesel for a flat per-ticket bill
Sonnet 5.5 is the engine. If what you want is tickets answered, eesel is the employee: an AI helpdesk teammate that plugs into Zendesk, Freshdesk, Gorgias and your other tools, learns from your past tickets and help center, and lets you simulate it against historical tickets before it replies to a real customer. You never pick an effort level or watch cache reads. eesel pricing is a fixed monthly credit plan where a ticket or a chat handled is one credit, starting at $299 for 500 credits, with a free tier of 100 credits to try it.
If you work from a terminal, the eesel CLI drives the same teammate. You can connect integrations and edit its standing instructions, and also approve or deny actions and read its activity log as JSON. Coding agents like Claude Code can run it too, through the workspace's MCP server.

Try eesel and see what your queue costs per ticket instead of per token.
Frequently Asked Questions
How much does Claude Sonnet 5.5 cost per million tokens?
Is Claude Sonnet 5.5 more expensive than Claude Sonnet 5?
Is Claude Sonnet 5.5 half the price of Claude Opus 5.5?
What is the cheapest way to run Claude Sonnet 5.5?
Can I use Claude Sonnet 5.5 for free?
Does Claude Sonnet 5.5 charge extra for long context?
How much does Claude Sonnet 5.5 cost on Amazon Bedrock or Google Cloud?
How much would Claude Sonnet 5.5 cost for a customer support bot?

Article by
Rama Adi
Rama is a software engineer at eesel AI with two years of experience writing about B2B SaaS, AI tools, and customer support technology. Based in Bali, Indonesia, he brings a developer's perspective to product comparisons — cutting through marketing copy to what the integrations and APIs actually do.








