
What is Claude Haiku 5.5?
Claude Haiku 5.5 is the small model in Anthropic's current lineup. You reach for it when a job runs thousands of times a day and each single run is short. Anthropic's launch post calls it "the cheapest, fastest, and most capable small model we've ever released" and names the jobs it was built for: summaries, compactions, database queries, classification, live customer support and browser use.

I build AI agents at eesel, and honestly most of what those agents do all day is not hard reasoning at all. They read a ticket and decide where it goes, they pull an order number out of a message, sometimes they write a three-line summary so the next person doesn't have to scroll. For years that work sat in an awkward spot: a good model cost too much for it, and a cheap model kept fumbling it. Haiku 5.5 is aimed right at that spot.
There was a hole in Anthropic's range too, and this fills it. Before this launch the cheapest current Claude was Sonnet 5.5 at $2/$10, and Haiku 4.5 at $1/$5 had fallen well behind cheaper rivals like GPT-6 Luna. As one developer put it on X a week before launch, the small, cheap model space was "mostly chinese labs" with Luna as the one exception.
Claude Haiku 5.5 specs at a glance
This is the spec sheet as it stands, taken from the Haiku 5.5 overview page in Anthropic's docs.

| Spec | Claude Haiku 5.5 |
|---|---|
| API model ID | claude-haiku-5-5 (no date suffix, no separate alias) |
| Amazon Bedrock ID | anthropic.claude-haiku-5-5 |
| Released | October 7, 2026 |
| Context window | 1M tokens |
| Max output | 128K tokens (300K on the Batch API, in beta) |
| Thinking | Adaptive, on by default |
| Default effort | medium |
| Input / output | Text and images in, text out |
| Knowledge cutoff | June 2026 |
| Platforms | Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS |
| Retirement | Not sooner than October 7, 2027 |
A couple of details in that table slip past people. First, the model ID has no dated snapshot, so claude-haiku-5-5 is the fixed name and not an alias that moves later. Second, with a beta header the Batch API can return up to 300K output tokens, which is handy if you use Haiku to write long structured files overnight.
Where Haiku 5.5 sits in the Claude lineup
Anthropic now runs four current models. Haiku 5.5 sits at the bottom of the ladder, which here is the fast and cheap end.

| Model | Best for | Latency | Price in / out per 1M | Default effort |
|---|---|---|---|---|
| Claude Fable 5.1 | Demanding reasoning, long-horizon agents | Slower | $10 / $50 | high |
| Claude Opus 5.5 | Long-running agentic coding and knowledge work | Moderate | $4 / $20 | medium |
| Claude Sonnet 5.5 | Best mix of speed and intelligence | Fast | $2 / $10 | high |
| Claude Haiku 5.5 | Classification, extraction, routing, subagents | Fastest | From $0.10 / from $0.50 | medium |
All four share the same 1M context window, 128K output cap and June 2026 knowledge cutoff, per the models overview. The window stopped being a reason to pick one model over another. What's left to decide on is price and speed, plus how hard the job really is.
Anthropic doesn't pretend Haiku is a cheap Sonnet, either. Its launch post says Sonnet 5.5 and Opus 5.5 "remain better choices for complex agentic coding tasks", and that Haiku 5.5 suits narrowly scoped work like compaction, summarization and subagent jobs. For me that line is worth more than anything else in the announcement, and I keep it in mind for every decision below.
What's new compared with Haiku 4.5
Haiku 4.5 came out a year before this one. Going by the version number you would expect a small step, but the jump is a lot bigger than that.

Anthropic's what's new page lists 13 changes: 4 new capabilities, 5 breaking changes and 4 behavior changes. Below I only go through the ones that change how you actually use the model.

Adaptive thinking and an effort setting
Haiku 5.5 is the first Haiku-class model with an adjustable effort setting. Thinking is adaptive and on by default. In practice the model decides by itself when to think and for how long, and your lever is effort instead of a fixed thinking budget. On lower effort, a simple request may get no thinking at all. What each level does, I cover further down.
A 1M window and 128K outputs
Context grew from 200k to 1M tokens and max output from 64k to 128K. Watch out here: thinking tokens count toward max_tokens. If you tuned that limit tightly for Haiku 4.5, the response can now end right after the thinking block, before any text comes out.
Browser and computer use
Haiku 5.5 supports the browser use tool on the Claude API and Google Cloud, and Anthropic shipped a browser and computer use SDK for Python and TypeScript on the same day. According to Anthropic, Haiku 5.5 is "especially well-suited" to these tasks because of its speed and price. The OSWorld 2.1 score of 72.4% (against 15.7% for Haiku 4.5) does back that claim.

Worth knowing before you get excited: the SDK ships without a browser, without a desktop, and without any URL policy. The driver is yours to bring, and the starter examples in the docs carry the label "not production code". If writing one yourself doesn't appeal, Browser Use, Browserbase, Daytona and E2B already publish ready-made integrations.
A new tokenizer
Haiku 5.5 uses the same newer tokenizer as Sonnet 5.5 and Opus 5.5, and the migration guide says the same text produces about 30% more tokens than on Haiku 4.5. This explains why Anthropic's own headline number says "around 75% less to run" and not 90%. Per token the price did fall 90%, it's just that each job eats more tokens now.
How the effort setting changes Haiku 5.5
If I had to spend my time on only one part of Haiku 5.5, it would be this one, since nothing else shifts how the model behaves as much. The effort docs describe five levels, and Haiku defaults to medium where most Claude models default to high.

Artificial Analysis ran Haiku 5.5 at every level, and the trade-off it found is not subtle. Its Intelligence Index score climbs from 29.4 at Low to 43.4 at Max, but the time to the first chunk of output, which includes thinking, goes from about 14 seconds to 415 seconds.

| Effort | AA Intelligence Index | Output tokens to run the index | Output speed | Time to first chunk |
|---|---|---|---|---|
| Low | 29.4 | 32M | 181 tokens/s | 13.6 s |
| Medium (default) | 34.5 | 54M | 137 tokens/s | 14.2 s |
| High | 37.8 | 97M | 173 tokens/s | 26.2 s |
| Xhigh | 41.2 | 180M | 188 tokens/s | 86.8 s |
| Max | 43.4 | 440M | 243 tokens/s | 415.3 s |
I read that table more like a price list in disguise. Moving from Medium to Max buys 9 points, and you pay for them with roughly 8x the output tokens. Artificial Analysis also noted in its launch post on X that moving from Xhigh to Max "adds 2 points for ~1.8x the tokens".
On a tiny job, Simon Willison's pelican test traces out the same curve:
"Low messes up the bicycle frame, but medium/high/xhigh/max all get the bicycle frame right. The max one took 5 minutes 9 seconds and cost 3.3826 cents. The cheapest one (low) cost 0.0936 cents and took 7 seconds."
The rule of thumb I landed on from building agents goes like this. Anything with a fixed answer set, tags or routing for example, starts at Low. Anything that writes prose a customer will read starts at Medium. And when a job only comes out right at High or above, that usually tells me it was a Sonnet job all along.
How good is Claude Haiku 5.5?
For its size it is very good, though you can still tell it's a small model. Below are Anthropic's published benchmarks at their headline settings.
| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| GDPval-AA v2.1 (knowledge work, Elo) | 1620 | 735 | 1437 | 1840 |
| AA-Briefcase v1.1 (knowledge work, Elo) | 1578 | 614 | 1336 | 1824 |
| OSWorld 2.1 (computer use) | 72.4% | 15.7% | 48.9% | 83.9% |
| Humanity's Last Exam (no tools) | 45.9% | 10.2% | n/a | 56.9% |
| Terminal-Bench 4.0 (agentic coding) | 39.2% | 0.0% | 16.4% | 70.6% |
| FrontierCode 1.1 (agentic coding) | 46.4% | n/a | 42.4% | 52.1% |
| Chartography (visual reasoning, no tools) | 46.4% | 6.4% | 29.1% | 61.6% |
The pattern holds across the board. Everywhere it was tested, Haiku 5.5 beats GPT-6 Luna, and it lands close to Sonnet 5.5 on knowledge work and computer use. Long agentic coding is where it drops well behind Sonnet.
On the launch page, the early customer reports tell more or less the same story, just in plainer words. HubSpot said it got "the best score we've seen on this suite yet, at 92.8%" on its CRM tasks. AlphaSense measured 0.84 against 0.76 for Haiku 4.5 on a document Q&A workload that runs about 8M calls a week. Box reported 11 points higher than Haiku 4.5 at about half the latency.
There is one caution from Artificial Analysis I'd keep in mind. Haiku 5.5 knows fewer facts than the bigger models, scoring 36% accuracy on its AA-Omniscience test, yet it says "I don't know" more readily, so its hallucination rate is 40%, against 77% for GPT-6 Luna. When a customer is going to read the output, I care about that number more than the index score. The practical takeaway: give Haiku your help center, don't rely on what it remembers.
Claude Haiku 5.5 pricing
Of the current Claude models, only Haiku 5.5 has two price tiers. Which one you pay depends on prompt length, and the pricing page lists every line.
| Price per 1M tokens | Prompts up to 100,000 tokens | Prompts over 100,000 tokens | Haiku 4.5 (flat) | Sonnet 5.5 |
|---|---|---|---|---|
| Input | $0.10 | $0.50 | $1.00 | $2.00 |
| Output | $0.50 | $2.50 | $5.00 | $10.00 |
| Cache read | $0.01 | $0.05 | $0.10 | $0.10 |
| 5-minute cache write | $0.125 | $0.625 | $1.25 | $2.50 |
| 1-hour cache write | $0.20 | $1.00 | $2.00 | $4.00 |
| Batch API | 50% off input and output | 50% off | 50% off | 50% off |
Anthropic set the line at 100,000 tokens because, per its launch post, around 90% of requests to Haiku 4.5 fell under it. With a workload shaped like that, Haiku 5.5 matches GPT-6 Luna's $0.10/$0.50 sticker. With a different shape, well, the HN thread did the math fast:
"100k tokens is an absurdly low cutoff and it is only applicable to Haiku and not Sonnet or Opus. It's a low enough cutoff that it will be quickly exceeded if you are doing anything with Agents; for typical generation or Jev-like classifiers, it's a good value and as noted in this article, that is apparently the vast majority of Haiku use."
A quieter gap sits between Claude tokens and OpenAI tokens as well. Another commenter pointed out that "modern Claude's 100K tokens are about ~60-65K modern GPT tokens", so the same document hits Haiku's line sooner than it would hit Luna's. I go deeper on the cost side in my Claude Haiku 5.5 review and in GPT-6 Luna pricing.
Put your own numbers in below to see what a month of calls costs you and where the 100k line starts to bite. Use token counts the way Haiku 5.5 counts them.
As an example, try 50,000 requests at 6,000 prompt tokens and 800 output tokens. On Haiku 5.5 that comes to $50 a month, while Sonnet 5.5 is $1,000. Then push the prompt to 120,000 tokens and you'll see Haiku jump 5x.
How to access and use Claude Haiku 5.5
Haiku 5.5 is live on every platform Anthropic sells through: the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. You can use it in the claude.ai apps too, and there Anthropic publishes its system prompt.
Here's a minimal API call. Effort is set explicitly, and there are no sampling parameters:
import anthropic
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-haiku-5-5",
max_tokens=4096,
output_config={"effort": "low"},
messages=[
{"role": "user", "content": "Tag this ticket as billing, shipping, bug or other: ..."}
],
)
for block in response.content:
if block.type == "text":
print(block.text)
A few habits that saved me debugging time later:
- Select content blocks by
type. A response can open with a thinking block even if you never asked for thinking, so don't assume the first block is the answer. - Leave out
temperature,top_pandtop_k. On Haiku 5.5, any non-default value gets you a 400 error. - Handle
stop_reason: "refusal". Haiku 5.5 runs safety classifiers that sometimes decline a request, and per the migration guide no server-side fallback model catches it.
People on a Claude plan now have a way to try it more or less for free. Starting launch week, Anthropic gives monthly API credits to Max and Team subscribers: $100 a month on Max 5x, $200 on Max 20x, and up to $500 pooled on Team. You can spend the credits on any model. What they don't cover is interactive Claude Code sessions.
Inside Claude Code itself, Haiku is a common pick for subagents, and my Claude Code model selection guide covers when to switch.
What breaks when you switch from Haiku 4.5
Calling Haiku 5.5 a drop-in replacement is only true if your code was already written for the newer Claude models. The migration guide has a 10-step checklist, and these five return errors:
| What your code does today | What happens on Haiku 5.5 | The fix |
|---|---|---|
Sends thinking with budget_tokens | 400 error | Use {"type": "adaptive"} and set effort |
Sets temperature, top_p or top_k | 400 error | Remove them |
Ends messages with an assistant turn (prefill) | 400 error | End with a user turn, use structured outputs |
Uses computer_20250124 on the Claude API or Google Cloud | Rejected | Move to computer_toolset_20260801 |
| Edits earlier turns and sends thinking blocks back | Thinking blocks invalidated | Keep the conversation append-only |
Of these, the temperature change is the one I'd plan around. A lot of classification pipelines pin temperature to 0 to get repeatable output, and on Haiku 5.5 that request just fails. If you want consistency now, structured outputs with a fixed schema is the better route. One more thing: Priority Tier isn't offered on Haiku 5.5, so a team holding a Priority commitment on 4.5 has to plan capacity some other way.
What is Claude Haiku 5.5 good for?
When I judge a small model, I ask about the job first. Is the prompt short and is there a clear right answer, and on top of that, does it run at high volume? Haiku 5.5 is built for jobs where the answer to all of it is yes.
| Job | Fit | Why |
|---|---|---|
| Ticket classification and tagging | Strong | Short prompts, fixed labels, Low effort works |
| Ticket summaries and handoff notes | Strong | Fast at Medium, cheap enough to run on every ticket |
| Routing and prioritization | Strong | Speed matters more than depth |
| Help-center answers (RAG) | Strong, if prompts stay under 100k | Lower hallucination rate than Luna, 1M window if needed |
| Subagents under a Sonnet or Opus lead | Strong | Anthropic and Cognition both point to this use |
| Browser and computer use | Good | 72.4% on OSWorld 2.1, fastest Claude |
| Long agent loops and hard coding | Weak | 39.2% vs 70.6% for Sonnet 5.5 on Terminal-Bench 4.0 |
People underrate speed. In one eesel evaluation, a buyer put the eesel support bot through a careful 67-test review and found the knowledge answers solid. They walked anyway, because the chat widget felt slow and got stuck. Good answers couldn't rescue a slow experience. Haiku 5.5 is Anthropic's fastest model at standard speed, and an OpenRouter reading shared on HN puts it at roughly twice Luna's throughput, so it speaks directly to that kind of lost deal.
On Hacker News, the developers who liked it most were the ones doing this exact kind of short, grounded work:
"Tested on my RAG system containing all of cloudflare docs (more than 3000 A4 sized highly technical documents): best
quality*speed/priceratio of any other model. And I have tested more than a 100 different models."
Another tester gave it a job too big for it, and it seemed to know its own limits:
"One interesting thing is it took a look at the job at hand, and immediately delegated it to Opus 5.5. It at least knows what it isn't good at. Very fast though, and likely best used for small subagent tasks / tightly scoped work."
What developers are saying
The Hacker News launch thread was mostly positive about price and speed. The real complaints came down to two: the 100k line, and teams whose own evals didn't line up with Anthropic's charts.
"It's absolutely better than Luna. It feels closer to a "sonnet 5.2" if that makes sense. Of course it's not as big, and hence falls-off quicker."
"At work we use haiku 4.5 for a handful of latency sensitive tasks that are fairly simple. It performs well. Just started testing 5.5 as I've been anticipating a nice improvement since it was teased. Results so far are trash. Prompt leakage even. And it's slower."
I'd take that second comment seriously. Haiku 4.5 ran without thinking by default, so a prompt tuned for it may act differently on a model that thinks at Medium effort straight out of the box. Anthropic published a Haiku 5.5 prompting guide for exactly this. If a migrated prompt gets slower, my first move would be dropping effort to Low.
The wider small-model field is covered in my posts on Claude Haiku 5.5 alternatives, Gemini 3.8 Flash and GLM-5.3 Flash.
Use Haiku-fast answers on your support queue with eesel
If a support queue is why you're reading about Haiku 5.5, then the model is the easy part. The hard part is everything around it: which tickets get the fast model, how prompts stay under the 100k line, what effort level each job runs at, and how wrong answers get caught before a customer ever sees one.
That work is what eesel's AI helpdesk teammate takes off your plate. It joins the helpdesk you already use, such as Zendesk, Freshdesk or Gorgias, learns from your past tickets and help center, and triages, tags, summarizes and replies. Every rollout gets simulated against your historical tickets before it talks to a single customer, so first you get to see how it would have handled last month's queue.

If you would rather drive it from a terminal, the eesel CLI runs the same teammate and workspace as the dashboard, so a script or a coding agent like Claude Code can set it up and check its work. It's basically the lead-and-sidekick pattern Anthropic recommends for Haiku, just pointed at your support setup. My posts on the AI agent CLI and MCP servers go deeper.
Plans start at $299 for 500 credits, though you can run it free on a slice of real tickets before paying anything. Try eesel and see whether faster answers actually show up for your customers.
Frequently Asked Questions
What is Claude Haiku 5.5?
claude-haiku-5-5. It has a 1M-token context window, up to 128K output tokens, adaptive thinking with an effort setting, and is built for high-volume work like classification, routing, extraction, summaries and subagent tasks.How much does Claude Haiku 5.5 cost?
Is Claude Haiku 5.5 better than Haiku 4.5?
What is the difference between Claude Haiku 5.5 and Sonnet 5.5?
How do I use Claude Haiku 5.5?
model: "claude-haiku-5-5", or use anthropic.claude-haiku-5-5 on Amazon Bedrock. Leave out temperature, top_p and top_k, and set output_config.effort to low, medium, high, xhigh or max. It is also on Google Cloud, Microsoft Foundry and Claude Platform on AWS. For support teams, eesel's AI helpdesk teammate runs fast models inside your helpdesk without any code.Is Claude Haiku 5.5 good for customer support?
What breaks when I switch to Claude Haiku 5.5?
budget_tokens thinking, non-default temperature, top_p or top_k, assistant message prefill, the old computer_20250124 tool on the Claude API and Google Cloud, and edits to earlier turns when you send thinking blocks back. The new tokenizer also counts about 30% more tokens for the same text. The Claude Haiku 5.5 alternatives post covers where to go if those changes do not suit you.Is Claude Haiku 5.5 free to try?

Article by
Kira
Kira is a writer at eesel AI with a Computer Science background and over a year of hands-on experience evaluating AI-powered customer service tools. She focuses on breaking down how helpdesk platforms and AI agents actually work so that support teams can make better buying decisions.








