
My verdict at a glance
I work the support queue at eesel every day, so whenever a model launches I read it with one question in mind, which is whether a customer waits less or gets a better answer. Haiku 5.5 is the first small Claude where the answer is a clear yes for a big slice of tickets, and the full scorecard is below.
| What I checked | Verdict | The number behind it |
|---|---|---|
| Price under 100k tokens | Excellent | $0.10 in / $0.50 out, 90% below Haiku 4.5 |
| Price over 100k tokens | Mediocre | $0.50 / $2.50, Luna stays cheaper here |
| Intelligence | Best in its class | 43 on AA, vs 38 for GPT-6 Luna |
| Speed | Excellent | 137-243 tokens/s on AA, by effort level |
| Agentic coding | Weak | 39.2% Terminal-Bench 4.0 vs 70.6% for Sonnet 5.5 |
| Token efficiency | Poor at high effort | ~162k output tokens per AA task at Max, ~3x Luna |
| Hallucination | Better than rivals, still real | 40% on AA-Omniscience vs 77% for Luna |
| Migration effort | Medium | 6 breaking changes, including temperature and prefill |
What Claude Haiku 5.5 is
Claude Haiku 5.5 is Anthropic's small, fast model, released on October 7, 2026, under the model ID claude-haiku-5-5. Anthropic pitches it for "high-volume, cost-sensitive tasks" like summaries, compaction, classification and database queries, and as a subagent for Opus 5.5 and Sonnet 5.5 on coding work.

The spec sheet is a big step up from Haiku 4.5. In terms of specs, the context window grows from 200k to 1M tokens, max output goes from 64k to 128k, and the knowledge cutoff moves to June 2026. It is also the first Haiku with adjustable effort and adaptive thinking, and that thinking is on by default at Medium.

The models overview describes it as the model "for high-volume, latency-sensitive tasks such as classification, extraction, and routing." That one line is a fair summary, I think, of where it shines and where it does not.
Where Claude Haiku 5.5 is strong
There is a lot to like here, and the gains over Haiku 4.5 are not small, so here is where it earns its spot.
It is close to Sonnet 5.5 on most work
On Anthropic's own launch benchmarks, Haiku 5.5 reaches 1620 on GDPval-AA against 1840 for Sonnet 5.5, and 72.4% on OSWorld 2.1 against 83.9%. For a model that costs a twentieth of Sonnet's input price, that is a striking result.
| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| GDPval-AA v2.1 (knowledge work) | 1620 | 735 | 1437 | 1840 |
| AA-Briefcase v1.1 | 1578 | 614 | 1336 | 1824 |
| OSWorld 2.1 (computer use) | 72.4% | 15.7% | 48.9% | 83.9% |
| Humanity's Last Exam, no tools | 45.9% | 10.2% | n/a | 56.9% |
| Terminal-Bench 4.0 | 39.2% | 0.0% | 16.4% | 70.6% |
| FrontierCode 1.1 | 46.4% | n/a | 42.4% | 52.1% (Xhigh) |
| Chartography, no tools | 46.4% | 6.4% | 29.1% | 61.6% |
The pattern gets clearer when you divide each Haiku score by Sonnet's, which is what the chart below does.

On knowledge work and computer use, Haiku 5.5 gets you 86-88% of Sonnet 5.5. On agentic coding it gets you 56%. That gap tells you exactly which jobs to hand it.
It beats GPT-6 Luna on intelligence and hallucination
Artificial Analysis puts Haiku 5.5 at 43 on its Intelligence Index at Max effort, ahead of GLM-5.3 Flash (42), Gemini 3.8 Flash (41) and GPT-6 Luna (38). It also found Haiku more willing to say it doesn't know, with a 40% hallucination rate on AA-Omniscience against 77% for Luna.
"Haiku 5.5 is noticeably smarter than GPT-6 Luna, so I can see their pricing strategy here. For a while Anthropic has lacked a cost effective "cheap" LLM for summarisation, compacting, RAG helpers, etc."
It is fast
Anthropic calls it "our fastest model to date" at standard speed. Artificial Analysis measured 137 to 243 output tokens per second depending on effort, against 128 for Luna at Max and 90 for Haiku 4.5. Customer reports line up: Asana saw "over a 30% reduction in latency for task completions" in its AI Teammates evals, and Box measured 11 points higher than Haiku 4.5 "at about half the latency."
Speed is not a vanity metric in support. One buyer I spoke with ran a methodical 67-test evaluation of an AI chat setup, found the answers solid, and still walked away because the chat widget was, in their words, "slow and often gets stuck." A great answer that shows up late just feels like a broken product.
It is strong at short, grounded work
The best customer quotes from the launch are all about narrow, repeated jobs. HubSpot reported 92.8% on its CRM eval suite, "the best score we've seen on this suite yet," with the lowest false positive rate on a stale-record audit. AlphaSense saw a statistically significant jump on its 8M-calls-a-week document Q&A, from 0.76 to 0.84.
The community found the same thing on retrieval-heavy work.
"Tested on my RAG system containing all of cloudflare docs (more than 3000 A4 sized highly technical documents): best
quality*speed/priceratio of any other model. And I have tested more than a 100 different models."
That is the shape of most support tickets, where you read a few help articles and the ticket and then answer the question. It is also why I think Haiku 5.5 matters more for AI ticket classification and ticket summarization than for coding.
It works as a sidekick model
Anthropic's own pitch is Haiku as the subagent for a bigger model, and the early reports back that up. Rogo uses a Haiku 5.5 subagent to pull a revenue line from a 10-K while a bigger model builds the deck. Cognition says its Devin Fusion setup, with Opus 5.5 as the lead and Haiku 5.5 as the sidekick, "holds a top-tier FrontierCode score of 66.2 while cutting cost and latency."
One Hacker News tester even watched it delegate a complex UI build to Opus on its own, noting that it "at least knows what it isn't good at." If you already run Claude Code subagents, the Haiku in Claude Code setup is the obvious place to try it first.
Where Claude Haiku 5.5 falls short
None of these are dealbreakers on their own, but taken together they decide whether Haiku 5.5 saves you money, or quietly costs more than you planned.
The 100k-token price cliff
This is the big one. Haiku 5.5 is the only current Claude model without flat pricing across its context window. Per the pricing docs, a prompt over 100,000 tokens pays the higher rate on every line, input, output and cache alike.

Anthropic says prompts under 100k make up around 90% of requests to the previous Haiku, so most classic Haiku jobs land in the cheap tier. Agent loops are a different story though, since conversation history piles up fast.
"100k tokens is an absurdly low cutoff and it is only applicable to Haiku and not Sonnet or Opus. It's a low enough cutoff that it will be quickly exceeded if you are doing anything with Agents; for typical generation or Jev-like classifiers, it's a good value"
Simon Willison's launch write-up frames the trade-off well: under 100k, Haiku matches Luna's price with higher scores, but "Above 100,000 tokens, Luna looks like a much better deal." Luna's own price step comes at 272k tokens and only rises to $0.20/$0.75.
A new tokenizer quietly adds tokens
The migration guide says the same text produces "approximately 30% more tokens than on Haiku 4.5." Simon Willison measured about 1.25x on a long prompt and called it "a hidden price increase." It also means you reach the 100k line sooner than the raw character count suggests.
One tester put some real numbers on this using a classification job.
| Model | Input tokens | Output tokens | Batch cost | Sync cost |
|---|---|---|---|---|
| Haiku 5.5 | 11,893,643 | 240,451 | $0.21 | $0.42 |
| GPT-6 Luna | 7,903,468 | 177,602 | $0.16 | $0.31 |
Those figures come from Farmadupe on Hacker News. Same per-token price, but Luna came out about 26% cheaper on the same job because it needed fewer tokens to read and write the same thing.
Max effort is a token furnace
Effort settings are new to Haiku, and the top end is expensive. Artificial Analysis found Haiku 5.5 at Max uses around 162k output tokens per Intelligence Index task, about 3x Luna at Max. Its time to first token at Max was 415 seconds.
| Effort | AA Intelligence Index | Cost to run the index | Output tokens | Time to first token |
|---|---|---|---|---|
| Low | 29.4 | $34.31 | 32M | 13.6s |
| Medium (default) | 34.5 | $56.36 | 54M | 14.2s |
| High | 37.8 | $94.07 | 97M | 26.2s |
| Xhigh | 41.2 | $157.58 | 180M | 86.8s |
| Max | 43.4 | $330.25 | 440M | 415.3s |
Going from Medium to Max adds 9 points for nearly 6x the cost. Artificial Analysis also notes these costs don't yet include the 100k price step, so long tasks will run higher than the table says. Simon Willison's pelican test makes the same point in miniature, where Low took 7 seconds and cost 0.0936 cents and Max took 5 minutes 9 seconds.
For live chat, the Medium-effort time to first token on that benchmark (about 14 seconds) is worth testing against your own tickets, since adaptive thinking runs before the answer starts. The docs say lower effort lets the model skip thinking on simple requests, and thinking can still be switched off at High effort or below.
It still gets facts wrong when it has nothing to read
Anthropic's system card is candid here. On the closed-book AA-Omniscience test, Haiku 5.5 answered 44% correctly, 32% incorrectly and abstained on 24%. That is better than Haiku 4.5 at the net level, but a roughly one-in-three wrong rate on questions it can't look up is the very reason a support model needs your knowledge base in front of it.
The system card also flags that it over-refused more than any other model in the automated behavior audit, and Artificial Analysis says its AutomationBench score of 35% is "likely understated" because of a pre-release over-refusal issue Anthropic is fixing.
Your own evals might disagree
Not everyone's tests went well on day one, and I think that is worth hearing about before you flip a production switch.
"At work we use haiku 4.5 for a handful of latency sensitive tasks that are fairly simple. It performs well. Just started testing 5.5 as I've been anticipating a nice improvement since it was teased. Results so far are trash. Prompt leakage even. And it's slower."
A second commenter replied that their evals and human pairwise tests also put Haiku 4.5 first, and they were reworking prompts as a result. Haiku 4.5 is still listed as active, with retirement not before October 15, 2026, so there is a window to run both side by side.
What breaks when you migrate
The what's new page lists several breaking changes. For support bots and classifiers, these are the ones that bite.
| Change | What happens now | The fix |
|---|---|---|
| Manual thinking budgets | budget_tokens returns a 400 | Use adaptive thinking plus effort |
| Sampling parameters | Any top_k, or temperature other than 1, returns a 400 | Remove them; steer with the prompt |
| Assistant prefill | Returns a 400, even with thinking off | End messages with a user turn |
| Append-only history | Editing system, tools or earlier turns, then replaying thinking, returns a 400 | Keep history append-only |
| Computer use tool | computer_20250124 is rejected on the API and Google Cloud | Move to computer_toolset_20260801 |
| Tokenizer | ~30% more tokens for the same text | Recount max_tokens and budgets |
The temperature change is the sneaky one for support. A lot of ticket triage pipelines run at temperature 0 for consistent labels, and that request now fails outright. Prefill is the other common trick (often used to force a JSON opening brace), and that is gone as well.
Forced tool_choice still works, which is good news after Sonnet 5.5 dropped it. The catch is that a forced tool call comes back with no thinking block, per the migration guide. If you want the model to reason before it calls your refund or routing tool, use auto and say so in the prompt.
Claude Haiku 5.5 pricing in one table
Here is the full rate card from the Claude pricing docs, with Haiku 4.5 and Sonnet 5.5 for reference. All prices are per million tokens.
| Rate | Haiku 5.5, up to 100k | Haiku 5.5, over 100k | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|---|---|
| Input | $0.10 | $0.50 | $1.00 | $2.00 |
| Output | $0.50 | $2.50 | $5.00 | $10.00 |
| 5-minute cache write | $0.125 | $0.625 | $1.25 | $2.50 |
| 1-hour cache write | $0.20 | $1.00 | $2.00 | $4.00 |
| Cache read | $0.01 | $0.05 | $0.10 | $0.10 |
| Batch input / output | $0.05 / $0.25 | $0.25 / $1.25 | $0.50 / $2.50 | $1.00 / $5.00 |

To make that concrete, take a support reply with an 8,000-token prompt (instructions, a few help articles and the ticket) and 600 output tokens, thinking included. On Haiku 5.5 that is about $0.0011 per ticket, or roughly $11 for 10,000 tickets. If the same conversation grows to a 150,000-token prompt, the input alone costs $0.075, which is five times what it would be under the line. There is no Fast mode for Haiku 5.5, and Anthropic also halved Sonnet 5.5's cache reads to $0.10 on the same day, which narrows the gap for cache-heavy agents. My Sonnet 5.5 pricing guide covers that side.
How I would use it on a support queue
The way I would run Haiku 5.5 in support follows pretty directly from the strengths and the cliffs above. Let it do the fast, short, high-volume work on every ticket and hand the hard ones up.

In practice that means:
- Triage and tagging at Low effort. Short prompts, a fixed label list, no long history. This is the job the 100k tier was built for, and it suits ticket prioritization too.
- Summaries and handoff notes at Medium. Fast enough to run on every ticket without anyone noticing the cost.
- Routine replies, grounded in your help center. Order status, password resets, return windows. Keep retrieved context trimmed so the prompt stays under 100k.
- Escalate the rest. Refund disputes, multi-step troubleshooting and anything with a long thread go to Sonnet 5.5 or a person.
This is also roughly what a real trial looks like when it goes well. In one e-commerce trial on live Zendesk traffic, triage was the strongest job for AI, at 93% accuracy and 100% spam detection, while full drafts needed more help from the team's own past replies. A small fast model fits that first job well, I think.
What developers are saying
The Hacker News launch thread passed 770 points and 380 comments, and the reaction split along one line: people with short, well-scoped jobs loved it, and people running long agent sessions did the math on the 100k cliff.
"Ran our DataAnalyticsBench benchmark on it: 9x cheaper than Haiku 4.5 and 2 letter grades better. It's also now the fastest model (using the default speeds, not trying any of the other models "Fast" mode) to complete the exam."
The same tester added that it handled every straightforward analytics question but fell short on the ones needing deeper statistical digging, and that GPT-6 Luna did a bit better, at about 30% of the cost.
"With that said, the real reason to use Haiku is that it's faster than all of these models. OpenRouter is showing an average so far of 93 tokens/sec, and AA got at least 137 in each of their benchmarks. So it might be valuable for speed at lower thinking levels."
"noticeably smarter remains to be seen in practice. For now, Haiku is a bit more expensive than Luna on < 100k token, but I just don't have any agentic work below 100k, so this is going to be 5x more expensive than shown on these charts."
The Artificial Analysis team summed it up as "notably fast, however very verbose", which matches everything above.
Who should switch, and who should wait
Here is how I would decide it, going by the job rather than by the benchmark chart.
| You are running | My call | Why |
|---|---|---|
| Classification, tagging, routing | Switch, after a short eval | Cheapest tier, fastest model, small prompts |
| RAG answers over a help center | Switch, keep prompts under 100k | Strong grounded results, 1M window if you need it |
| Subagents under a Sonnet or Opus lead | Switch | The use case Anthropic and early customers point to |
| Long agent loops past 100k tokens | Wait or pick Luna | 5x price step plus the heavier tokenizer |
| Complex agentic coding | Stay on Sonnet 5.5 | 39.2% vs 70.6% on Terminal-Bench 4.0 |
| Pipelines relying on temperature 0 or prefill | Budget migration time | Both now return a 400 error |
If you are weighing it against outside options, my posts on GPT-6 Luna alternatives, Gemini 3.8 Flash and Kimi K3 cover the rest of the small and mid-size field.
For support specifically, the guide on which LLM fits support goes wider.
A faster model does not fix a wrong answer
One thing a model review can't show you is that most support bot failures I see have nothing to do with the model. A B2B vehicle telematics team found their bot telling customers it supported car brands that were not in their database. The model was not making anything up, to be fair. Their knowledge base said they supported "all models," and the bot believed it. Haiku 5.5 would have given the same wrong answer, just faster and cheaper.
That is why every eesel rollout is simulated against historical tickets before the teammate talks to a customer. That step is where you catch the knowledge gaps that no model upgrade can fix, and with a model that gets 32% of closed-book facts wrong it matters even more.
Try eesel with fast Claude models on your queue
If Haiku 5.5's speed is what you want on your tickets, eesel is the shortcut. Its AI helpdesk teammate joins your existing helpdesk, such as Zendesk, Freshdesk or Gorgias, learns from your past tickets and help center, and triages, tags and replies with a frontier model underneath. The 100k cliff, the effort tuning, the breaking changes above, all of that becomes eesel's problem and not yours.

If you live in a terminal, the eesel CLI operates the same teammate and workspace as the dashboard. You can script it, or let a coding agent like Claude Code drive it, which is the same lead-and-sidekick pattern from this review applied to your support setup.
My posts on the AI agent CLI and MCP servers go deeper.
Paid plans start at $299 for 500 credits, and you can try it free on a slice of your real tickets first. Try eesel and see whether a faster model actually shows up as faster answers for your customers.
Is Claude Haiku 5.5 worth it?
Yes, if your work is short and repetitive. Under 100k prompt tokens and at Low or Medium effort, Claude Haiku 5.5 is the best small model I have seen from Anthropic: smarter than GPT-6 Luna at the same sticker price and faster than anything else Anthropic sells, and it sits close enough to Sonnet 5.5 on knowledge work to take a big share of a support queue.
Past either line the math flips: long agent sessions pay 5x, Max effort burns tokens, and the new tokenizer makes both of those arrive sooner. Run your own eval against Haiku 4.5 while it is still available, keep prompts trimmed, and let the bigger models handle what Haiku 5.5 hands up.
Frequently Asked Questions
Is Claude Haiku 5.5 good?
How much does Claude Haiku 5.5 cost?
Is Claude Haiku 5.5 better than Haiku 4.5?
Is Claude Haiku 5.5 better than GPT-6 Luna?
What effort level should I use with Claude Haiku 5.5?
Is Claude Haiku 5.5 good for customer support?
What breaks when migrating to Claude Haiku 5.5?
What are the alternatives to Claude Haiku 5.5?

Article by
Riellvriany Indriawan
Riell is a designer and writer at eesel AI with about two years of experience researching CX platforms, AI chatbots, and helpdesk software. She combines her design background with a sharp eye for how these tools actually look and feel in practice — making her comparisons unusually visual and user-focused.








