Claude Sonnet 5.5: same price, fewer tokens, and 5 breaking changes

Riellvriany Indriawan
Written by

Riellvriany Indriawan

Katelin Teen
Reviewed by

Katelin Teen

Last edited September 29, 2026

Expert Verified
Hand-drawn illustration of a support agent at a laptop quickly clearing a stack of tickets, for a post on Claude Sonnet 5.5

What Claude Sonnet 5.5 actually is

Claude Sonnet 5.5 is the mid-tier model from Anthropic, and it is the second release in the Claude 5.5 family, coming six days after Opus 5.5. The way Anthropic pitches it is as "a faster, lower-cost complement" to Opus, with Opus kept for complex work that needs careful judgment and Sonnet taking the well-scoped everyday tasks. There is also a Haiku 5.5 promised "in the coming weeks".

The Claude Platform docs page for Claude Sonnet 5.5, showing the 1M context window, 128K max output, $2 and $10 per million token pricing, and a comparison table against Fable 5.1, Opus 5.5 and Haiku 4.5, as taken from the Claude Platform docs
The Claude Platform docs page for Claude Sonnet 5.5, showing the 1M context window, 128K max output, $2 and $10 per million token pricing, and a comparison table against Fable 5.1, Opus 5.5 and Haiku 4.5, as taken from the Claude Platform docs

Here are the specs, as listed on the model overview page:

  • The model ID is claude-sonnet-5-5 on the Claude API, Google Cloud, Microsoft Foundry and Claude Platform on AWS, and anthropic.claude-sonnet-5-5 on Amazon Bedrock.
  • Context is 1M tokens in and 128K out, or 300K out on the Batch API with a beta header.
  • Reliable knowledge cutoff is June 2026, the same as Opus 5.5 and Fable 5.1.
  • Thinking is adaptive, and the default effort is high on the API but Medium in the Claude apps and Claude Code.
  • Retirement is "not sooner than September 28, 2027", which gives you at least one year before you need to move again.

The tokenizer is the same as Claude Sonnet 5, per the what's new page, meaning the same text gives you the same token counts. This detail matters more than it sounds like it does. It means every saving Anthropic is claiming comes from the model writing less, and not from some friendlier way of counting.

The real story: same price tag, smaller bill

On the sticker price, nothing moved. Sonnet 5.5 costs exactly what Sonnet 5 costs: $2 input, $10 output, $0.20 cache reads. The part that changed is how many tokens it spends to get a job finished.

The early-access numbers in the launch post are specific in a way you don't usually see:

  • Balyasny Asset Management ran 2,441 finance tasks and saw Sonnet 5.5 use about 121k tokens per answer where Sonnet 5 used 497k, while scoring higher.
  • Slack saw better results on "almost all" of its offline Slackbot evals "with about 14% fewer output tokens", and that was without changing any prompts.
  • Base44 had it finishing app builds in 3.6 iterations on average, where Opus 5 needed 7.7.
  • Lovable measured "a third fewer tool calls and roughly half the shell runs to finish a task".

Anthropic's own side-by-side makes the same point, just visually. When both were asked to build a 400-starling murmuration in one HTML file, Sonnet 5 was still writing code at 4,520 tokens:

Sonnet 5 still writing JavaScript for a starling murmuration at 4,520 output tokens, as taken from Anthropic's Sonnet 5.5 announcement
Sonnet 5 still writing JavaScript for a starling murmuration at 4,520 output tokens, as taken from Anthropic's Sonnet 5.5 announcement

Sonnet 5.5 was already done at 4,158 tokens, and by that point its flock had been flying for 12.5 seconds:

Sonnet 5.5's finished starling murmuration running after 4,158 output tokens, with the animation 12.5 seconds in, from the same Anthropic comparison
Sonnet 5.5's finished starling murmuration running after 4,158 output tokens, with the animation 12.5 seconds in, from the same Anthropic comparison

The spread between the customers is where the honest reading sits. Balyasny's 4x token drop and Slack's 14% cut are both real, which tells you the saving depends a lot on how bloated your Sonnet 5 runs were in the first place. Anthropic's headline says "up to 30% less per task", and the "up to" is doing a lot of the work in that sentence. Sonnet 5 had a known habit of running long: in Artificial Analysis testing it took 183 turns per task at max effort, the worst on the chart. If your workload hit that, you can expect the big number, and if your prompts already kept answers short then expect something closer to Slack's.

How close it gets to Opus 5.5

For anyone paying Opus prices, this is the part where the release gets interesting. On Anthropic's launch benchmark table, Sonnet 5.5 closes most of the gap to a model that costs twice as much per token.

Hand-drawn grouped bar chart comparing Sonnet 5, Sonnet 5.5 and Opus 5.5 on Terminal-Bench 4.0, CursorBench 4.0, OSWorld 2.1 and Chartography, with Sonnet 5.5 far above Sonnet 5 and close to Opus 5.5 on each
Hand-drawn grouped bar chart comparing Sonnet 5, Sonnet 5.5 and Opus 5.5 on Terminal-Bench 4.0, CursorBench 4.0, OSWorld 2.1 and Chartography, with Sonnet 5.5 far above Sonnet 5 and close to Opus 5.5 on each
BenchmarkSonnet 5.5Sonnet 5Opus 5.5GPT-6 Sol
Terminal-Bench 4.0 (agentic coding)70.6%10.3%66.4%not reported
FrontierCode 1.1 (merge-ready code)52.1% (Xhigh)42.4%54.4%49.3%
CursorBench 4.055.5%34.1%57.8%not reported
GDPval-AA v2.1 (knowledge work, Elo)1844144918461487
AA-Briefcase v1.1 (long-horizon work, Elo)1811135918221483
Humanity's Last Exam (with tools)64.5%54.9%67.7%not reported
OSWorld 2.1 (computer use)80.1%57.0%81.8%not reported
Chartography (chart reading)61.6%15.6%64.4%53.6%

A few things jump out. First, the Terminal-Bench jump from 10.3% to 70.6% is enormous, and on that one Sonnet 5.5 actually beats Opus 5.5. GDPval-AA, a test of real work across 44 occupations, has it two Elo points behind Opus 5.5 and about 400 ahead of Sonnet 5. Then against GPT-6 Sol, which is the rival Anthropic chose to chart, Sonnet 5.5 at its best effort setting leads on every row where OpenAI's model has a score, and at High effort it matches GPT-6 Sol's best FrontierCode result for about a fifth of the cost per task.

I appreciate that Anthropic is candid about the limit here. The launch post says that "in our own testing, and in that of external testers, Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment". Benchmarks tend to reward the tasks that have a clear right answer. Open-ended calls, like deciding whether an angry enterprise customer needs a refund or an apology, or maybe an escalation, are exactly where the extra Opus spend still earns its keep. The Opus 5.5 vs GPT-6 Sol comparison goes into that top tier in more depth.

One more footnote that is worth reading: FrontierCode scores drop at Max effort (46.2%) versus Xhigh (52.1%). Anthropic's explanation is that at Max, Sonnet 5.5 more often split the code review across subagents, and this caused timeouts or edits going beyond the task's scope. More thinking is not always better, which is a pattern the Opus 5 effort guide also saw.

What Sonnet 5.5 means for support teams

If I had to put one quote from the launch post in front of any support lead, it would be this one. Zendesk's Director of AI tested the model on real tickets:

"We fed Claude Sonnet 5.5 hundreds of real support use cases across replies and escalation requests. It made fewer wrong decisions and resolved tickets faster than the Claude models we use in production today. Tickets were processed 20% faster, getting our customers the help they need without the wait."

Abhinay Kathuria, Director of AI at Zendesk, in Anthropic's announcement

"Fewer wrong decisions" is the part that matters here. Speed is nice, but on a support queue a fast wrong answer costs more than a slow right one: it creates a second ticket and an angry follow-up, and sometimes a refund on top. Zendesk AI already runs Claude models in production, so when it claims better escalation decisions against its own current setup, that carries real weight. Atlassian made a related point about scale, saying teams will run their Rovo agents "up to 30% faster than they could with Sonnet 5", and Slack AI is already testing it for Slackbot.

Looking at it from the queue side, this is how I would map the Claude lineup to support jobs:

Hand-drawn staircase of four Claude models mapped to support jobs: Haiku 4.5 at $1/$5 for tagging and routing tickets, Sonnet 5.5 highlighted at $2/$10 for everyday replies and drafts, Opus 5.5 at $4/$20 for tricky escalations, and Fable 5.1 at $10/$50 for the hardest problems
Hand-drawn staircase of four Claude models mapped to support jobs: Haiku 4.5 at $1/$5 for tagging and routing tickets, Sonnet 5.5 highlighted at $2/$10 for everyday replies and drafts, Opus 5.5 at $4/$20 for tricky escalations, and Fable 5.1 at $10/$50 for the hardest problems

Sonnet 5.5 is the natural default for the bulk of a queue, meaning order questions, how-to answers, policy lookups and first drafts for an agent to approve. Haiku stays the cheap choice for ticket triage and tagging. Opus 5.5 is where I would send the tickets that need judgment, which are the ones your senior agents would take anyway. Where the 30%+ speed gain matters most is live chat, since there a customer is sitting and watching the typing indicator.

The honest caveat: a better model does not fix a missing answer. I have watched a support bot confidently tell a customer that a product was supported because the help center said "we support all models", and I have seen another one invent answers when retrieval came back empty. Neither of those was a model-quality problem. Both were grounding problems, and a smarter model can actually make a confident wrong answer sound even more convincing. That is why eesel simulates every rollout against historical tickets before an AI replies to anyone, whichever model is underneath.

Before you migrate: five things that now return errors

Most launch coverage will skip this section, and it is the one that will actually eat your afternoon. Changing claude-sonnet-5 to claude-sonnet-5-5 is one line. Behind it, the what's new page lists five breaking changes, and there is one more rule that belongs on the same checklist.

Hand-drawn checklist titled Before you swap the model ID listing thinking disabled now errors, forced tool_choice now errors, keep history append-only, old computer-use tool rejected, and custom temperature now errors, next to a terminal showing 400 invalid_request_error
Hand-drawn checklist titled Before you swap the model ID listing thinking disabled now errors, forced tool_choice now errors, keep history append-only, old computer-use tool rejected, and custom temperature now errors, next to a terminal showing 400 invalid_request_error
  1. thinking: disabled returns a 400. Send thinking: {"type": "between_tools"} instead, which turns off up-front thinking. It only works at high effort or below; at xhigh or max it errors too.
  2. Forced tool use is gone. tool_choice set to any or a named tool returns an error. The fix is auto plus strict tool use.
  3. Thinking blocks are tied to the model and the account. If you edit earlier messages, the system prompt or the tools and then replay a Sonnet 5.5 thinking block, newer accounts get a 400. So keep conversations append-only.
  4. The old computer_20251124 tool is rejected on the Claude API and Google Cloud. Bedrock still accepts it.
  5. Some advisor pairings are rejected. Opus 4.8, Opus 4.7 and Sonnet 5 can no longer advise a Sonnet 5.5 executor.

About that extra rule, the docs note that setting temperature, top_p or top_k to a non-default value returns a 400 on this model.

The support trap is number two. A lot of helpdesk bots force a tool call to guarantee structured output: "always call classify_ticket", "always call draft_reply". On Sonnet 5.5, that request just fails outright. If you built your own integration on the Anthropic API, grep for tool_choice first, before you touch the model ID.

There is also a quieter change, one that does not error but will still confuse users. Text the model writes between tool calls now comes back inside thinking blocks, and at the default display: "omitted" that text is empty. If your chat widget streams "Checking your order status..." while the agent works, it will go silent until you set a display value. Anthropic has also recalibrated the effort levels, so re-run your effort tests rather than carrying the Sonnet 5 settings over. The AI agent handoff guide is worth a read if a model swap changes when your bot escalates.

Last thing is safety. Sonnet 5.5 is the first Sonnet to launch with cyber safeguards, because its cyber capabilities are comparable to Opus 5's. Higher-risk cybersecurity requests "visibly fall back to Sonnet 5", and a declined request returns HTTP 200 with stop_reason: "refusal". For a support bot this will rarely fire. But if you serve security or IT customers, it is worth handling the refusal case rather than showing a blank reply.

Claude Sonnet 5.5 pricing and access

This is the full rate card from Anthropic's pricing docs, per million tokens:

Price per 1M tokensClaude Sonnet 5.5Claude Sonnet 5Claude Opus 5.5Claude Haiku 4.5
Input$2$2$4$1
Output$10$10$20$5
5-minute cache write$2.50$2.50$5$1.25
1-hour cache write$4$4$8$2
Cache read$0.20$0.20$0.20$0.10
Batch input / output$1 / $5$1 / $5$2 / $10$0.50 / $2.50
Context window1M1M1M200K

A couple of notes on that table. Thinking is billed as output tokens, so a high-effort run pays $10 per million for its reasoning too. And the minimum cacheable prompt is now 512 tokens, per the model overview, and that makes it easier to cache a stable support system prompt.

To put these rates in support terms, here is some illustrative math that uses only the published prices. Say a ticket reply sends 6,000 input tokens (the ticket plus retrieved help articles) and gets back 800 output tokens. That is $0.012 of input and $0.008 of output, or about 2 cents per ticket. If 5,000 of those input tokens are a cached system prompt, the input side drops to about $0.003 and the whole reply lands near 1.1 cents. At 10,000 tickets a month, that is roughly $200 uncached versus $110 cached, before thinking tokens. Your own numbers will vary depending on effort and retrieval size, but it does show why the model line on the invoice is rarely the big cost in an AI support agent.

Here is where you can use it:

  • The Claude API, as claude-sonnet-5-5, with prompt caching and the Batch API. Zero data retention is available.
  • Claude apps. Anthropic says "anyone can chat with Claude using Sonnet 5.5 on Claude.ai" on web, iOS and Android. The paid tiers are covered in the Claude Pro pricing guide.
  • Claude Code, where the default effort is Medium. The Claude Code pricing guide covers the subscription-versus-API maths, and pinning a model takes one setting.
  • The clouds: Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS, all on day one.

If what you are weighing is providers rather than tiers, the three-way API comparison and the GPT-6 Sol pricing breakdown are the fastest sanity checks.

What developers are saying

The launch thread on Hacker News passed 550 comments in a day, and the reaction splits pretty neatly. People like the speed and they respect the benchmarks, while the pushback is mostly about where it fits next to Opus 5.5.

Hacker News

"Playing around with it for a few minutes, Sonnet 5.5 feels very fast, much quicker than Opus 5.5. Can't tell yet if it's a lot worse but the speed is definitely welcome."

The sharpest criticism has to do with the high effort settings. Simon Willison measured his usual SVG test at every effort level and found low at 1.6 cents, medium at 1.8 cents and high at 2.3 cents, but max burned through its whole thinking budget:

Hacker News

"Sonnet 5.5 has the same problem as Opus 5.5: on "max" thinking effort it burned through 128,000 thinking tokens (taking 15 minutes to do that) and ran out before it had produced the final SVG."

A Reddit user pulled the same pattern from Artificial Analysis, reporting that "sonnet 5.5 max has a cost per task of $7.60 while opus 5.5 max has $5.98", with Sonnet using 193k tokens per task to Opus's 119k (u/ex-procrastinator, Reddit). Another commenter summed up what the charts are implying:

Hacker News

"Per the charts, there is largely no point to using Sonnet 5.5 at high+ as opus low generally will give similar performance at similar or lower cost."

Anthropic says much the same in its own launch post: Sonnet 5.5 "complements Opus 5.5 best when running at lower effort settings", and "at higher settings, it can perform comparably at a similar cost". So the practical takeaway, and the thing I would tell any team that is switching over, is this. Run Sonnet 5.5 at Low or Medium effort, where the savings actually are. The API defaults to high, so the cheap version of this model is one you have to ask for. Support replies are exactly the kind of well-scoped work Medium handles well.

There is one regression worth knowing about. A developer who runs an adversarial benchmark found Sonnet 5.5 scored lower than Sonnet 5 "mainly because it is more reluctant to keep going to get an answer, instead it returns to ask the user questions" (dom96, Hacker News). For a coding agent that is annoying, but for a support bot, stopping to check instead of guessing is often exactly what you want.

A faster model still needs a teammate around it

Every model launch prompts the same conversation with technical customers: this is cheap and good enough now, so why not build our own support bot on the raw API? It is a fair question, and Sonnet 5.5 makes the model half of it easier than it has ever been.

What the rate card leaves out is everything around the model. Sonnet 5.5 is a superb engine. It is not a support agent until someone builds the retrieval over your help center and past tickets, plus the escalation rules and refusal handling, the effort tuning, and now also the migration work for five breaking changes. The model is the infrastructure, and the employee is what you build on top of it.

eesel sells that part already built. Its AI helpdesk teammate joins your existing queue in Zendesk, Freshdesk, Gorgias and more. It learns from your past tickets and help center, then drafts or sends replies with a frontier model underneath. When Anthropic ships a model like this, the switch and the breaking changes are eesel's problem, not yours.

The eesel AI helpdesk teammate's activity view in Zendesk, listing recent web conversations with their pending and resolved status and linked ticket numbers
The eesel AI helpdesk teammate's activity view in Zendesk, listing recent web conversations with their pending and resolved status and linked ticket numbers

If you do want programmatic control, eesel has a CLI that operates the same teammate and workspace the dashboard does. You can drive it from a terminal or script it, and you can also let a coding agent like Claude Code run it, so you get the agentic workflow people build on raw models without owning the grounding and retry stack. My write-ups on the AI agent CLI and MCP servers show how that fits together.

Pricing works on fixed monthly credits: a ticket or chat handled is one credit, the free plan includes 100 credits with no card, and paid plans start at $299 for 500. You can simulate the teammate against your own historical tickets before it replies to a single customer. Try eesel on your Zendesk queue free, and see the Sonnet-class speed-up on tickets you already know the answers to.

Is Claude Sonnet 5.5 worth it?

If you run Sonnet 5 today, yes, and it is one of the easier upgrades Anthropic has shipped. You pay the same rate for a model that scores far higher and finishes in fewer tokens, and it runs faster too. The only real cost is the migration work, and that is a known list of five items, not a mystery.

If you run Opus 5.5 for everyday work, test Sonnet 5.5 on your own traffic at Medium effort before renewing that habit. For well-scoped tasks, the benchmark gap is a couple of points and the per-token price gap is 2x. Once you go past high effort that advantage disappears, so if a task needs max reasoning, send it to Opus instead. Keep Opus for the open-ended judgment calls where Anthropic itself says it is still clearly stronger. One Hacker News commenter described a pattern, "80% Sonnet 5.5, Opus 5.5 to finish the last 20%", and it is a sensible way to split a queue too. For the wider field, my Sonnet 5 review and Opus 5.5 alternatives roundup are the next reads.

Frequently Asked Questions

What is Claude Sonnet 5.5?
Claude Sonnet 5.5 is Anthropic's mid-tier model, released on September 28, 2026 as the second model in the Claude 5.5 family. It sits below Claude Opus 5.5 and above Haiku 4.5, with a 1M-token context window and a June 2026 knowledge cutoff. For the wider lineup, see the Claude overview.
How much does Claude Sonnet 5.5 cost?
Claude Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens, the same as Sonnet 5. Cache reads are $0.20 and the Batch API halves both rates. The Anthropic API pricing guide covers the full rate card.
Is Claude Sonnet 5.5 cheaper than Sonnet 5?
Per token, no, the rates are identical. Per task, Anthropic says Sonnet 5.5 costs up to 30% less because it needs far fewer tokens for the same work. How much you save depends on how verbose your Sonnet 5 runs were, so meter it the way you would any AI customer service cost.
Is Claude Sonnet 5.5 better than Opus 5.5?
No, but it gets close. Sonnet 5.5 beats Opus 5.5 on Terminal-Bench 4.0 and lands within about two points on CursorBench and GDPval-AA, while Anthropic says Opus 5.5 remains clearly stronger at open-ended work that needs sustained judgment. For most everyday tasks, Sonnet 5.5 at half the price is the better buy.
Is Claude Sonnet 5.5 good for customer support?
It looks like a strong fit for everyday ticket replies. Zendesk tested it on hundreds of real support use cases and reported fewer wrong decisions and tickets processed 20% faster. The model still needs grounding in your own content, which is the job of an AI helpdesk agent built on top of it.
What breaks when I switch from Sonnet 5 to Sonnet 5.5?
Five things return errors: thinking set to disabled, forced tool choice, edited conversation history that replays thinking blocks, the old computer_20251124 tool on the Claude API, and some advisor pairings. If you run a support bot that forces a classification tool, that is the one to fix first. The AI helpdesk API guide covers how those integrations are usually wired.
Can I use Claude Sonnet 5.5 in Claude Code?
Yes. Sonnet 5.5 is available in Claude Code, where the default effort is Medium, and you can pin it with the steps in the model configuration guide. On the Claude Platform API the default effort is High.
What are the best alternatives to Claude Sonnet 5.5?
Within Anthropic's lineup, Opus 5.5 is the step up and Haiku 4.5 is the step down for routing and tagging. Outside it, GPT-6 Sol is the main rival on the launch benchmark table. The Sonnet 5 alternatives roundup covers the wider field.

Share this article

Riellvriany Indriawan

Article by

Riellvriany Indriawan

Riell is a designer and writer at eesel AI with about two years of experience researching CX platforms, AI chatbots, and helpdesk software. She combines her design background with a sharp eye for how these tools actually look and feel in practice — making her comparisons unusually visual and user-focused.

Related Posts

All posts →
Claude Sonnet 5 illustration with the Anthropic mark and a support workflow
Guides

Claude Sonnet 5: what it means for customer support

Claude Sonnet 5 brings near-Opus coding and agentic quality at mid-tier prices. Here is what the model actually changes for support teams, and what it does not.

Rama AdiRama AdiJul 1, 2026
AI pretraining
Guides

AI pretraining

Ever heard that AI is "trained on the whole internet"? That's AI pretraining, the foundational step for models like GPT. But for customer support, this general knowledge isn't enough. This guide breaks down what pretraining really is and explains why specializing an AI on your company's knowledge is the key to unlocking its true potential.

Kenneth PanganKenneth PanganOct 23, 2025
A practical guide to intents and sentiments in customer support
Guides

A practical guide to intents and sentiments in customer support

Understanding customer intents and sentiments is no longer optional. This guide breaks down what they are, why they matter, and how to use them to elevate your support.

Kenneth PanganKenneth PanganOct 27, 2025
A practical guide to the new Claude create files feature
Guides

A practical guide to the new Claude create files feature

Anthropic’s Claude now creates files like Excel sheets and PowerPoints. Useful for quick tasks, but risky for business-critical automation. Here’s the full breakdown.

Kenneth PanganKenneth PanganSep 9, 2025
A complete guide to Shift4Shop pricing in 2025
Guides

A complete guide to Shift4Shop pricing in 2026

Thinking about using Shift4Shop? Before you commit, it's crucial to understand the full picture. Our guide breaks down the official Shift4Shop pricing tiers, transaction fees, and the often-overlooked operational costs like customer support that can impact your bottom line. Discover how to build a realistic budget for your e-commerce store in 2025.

Kurnia KharismaKurnia KharismaSep 14, 2025
Image alt text
Guides

An overview of Claude Opus 4.6 pricing and capabilities

Explore our deep dive into Claude Opus 4.6 pricing. We break down the costs, new features, and practical use cases for Anthropic's latest AI model.

Katelin TeenKatelin TeenFeb 6, 2026
A complete guide to Worknet AI pricing in 2025
Guides

A complete guide to Worknet AI pricing in 2026

Searching for clear Worknet AI pricing? We analyzed their costs across multiple sources to give you the full picture, from their $75/user fee to their performance-based model, and explore a more transparent alternative.

Stevia PutriStevia PutriSep 9, 2025
Illustration of a unified support inbox pulling email, live chat, WhatsApp, social, and phone into one AI-assisted view
Guides

Multichannel customer support: what it is and how to do it well

A practical guide to multichannel customer support: the channels that matter, the multichannel vs omnichannel trap, and how AI keeps answers consistent.

Riellvriany IndriawanRiellvriany IndriawanJul 5, 2026
A complete guide to Slite pricing in 2026
Guides

A complete guide to Slite pricing in 2026

Thinking about Slite for your team's knowledge base? This guide dives deep into Slite's pricing for 2026, covering the Basic and Pro plans, the Slite Agent, and per-user costs. We'll also explore why a traditional wiki might not be enough, and how AI can help.

Kenneth PanganKenneth PanganSep 11, 2025

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free