
Claude Haiku 5.5 pricing at a glance
Here's the thing I noticed when I looked into what people search after a model launch. The query is "Claude Haiku 5.5 pricing", but the snippet Google shows already answers $0.10 and $0.50, so what the searcher is really asking is how much a month of this costs for them, which for Haiku 5.5 depends on one number more than any other: how long your prompts are.
Claude Haiku 5.5 launched on October 7, 2026, with the model ID claude-haiku-5-5. Below is the full API rate card from the Anthropic's pricing docs, per million tokens:
| Line item | Prompt up to 100,000 tokens | Prompt over 100,000 tokens |
|---|---|---|
| Input | $0.10 | $0.50 |
| Output (including thinking) | $0.50 | $2.50 |
| Cache read (hit or refresh) | $0.01 | $0.05 |
| 5-minute cache write | $0.125 | $0.625 |
| 1-hour cache write | $0.20 | $1.00 |
| Batch input | $0.05 | $0.25 |
| Batch output | $0.25 | $1.25 |
And here is where it sit next to the models you would actually compare it with:
| Model | Input | Output | Cache read | Long prompts |
|---|---|---|---|---|
| Claude Haiku 5.5 | $0.10 | $0.50 | $0.01 | 5x on every line past 100k |
| Claude Haiku 4.5 | $1.00 | $5.00 | $0.10 | Flat, 200k context cap |
| Claude Sonnet 5.5 | $2.00 | $10.00 | $0.10 | Flat to 1M |
| Claude Opus 5.5 | $4.00 | $20.00 | $0.20 | Flat to 1M |
| Claude Fable 5.1 | $10.00 | $50.00 | $0.25 | Flat to 1M |

Some things don't show up in the table, which is worth to flag. There is no Fast mode for Haiku 5.5 (it covers Opus models only), and the migration guide says Priority Tier isn't supported either. US-only inference adds 1.1x on every token category, same like on other Claude 4.6 and later models. For the wider picture across every Anthropic model, my Anthropic API pricing guide has it.
The 100k line is per request, and caching doesn't get you under it
Every other current Claude model now include the full 1M context window at one price. The pricing page says so quite directly: Claude 4.6 and later models "(except Claude Haiku 5.5)" bill a 900k-token request at the same per-token rate as a 9k one. Haiku 5.5 is the exception here, and the docs spell out three rules which decide on your bill.
- The prompt length counts every input token. That means cache reads and also cache writes, not only the fresh text.
- Each request is priced on its own. One long request doesn't push your other requests into higher tier, and the earlier requests keep the price they were billed at.
- A request over the line pays the higher prices on everything. Input and output move to the 5x rate, and the cache moves too, even when most of the prompt is a cache hit.
That third rule is the one people tend to miss. Caching does not hide tokens from meter. If your agent drags 120,000 tokens of history into a request, it pays $0.05 per million on the cached part instead of $0.01, then $2.50 per million on its output instead of $0.50.
Here's what that does to one single request. A 90k-token prompt with 1,000 output tokens costs $0.0095. Add 20,000 more tokens of prompt and the same request is costing $0.0575.

So 20k extra tokens make one request about 6x more expensive. Not 22% more, which is what the token count alone would make you expect. The Hacker News launch thread was picking this up within minutes:
"100k tokens is an absurdly low cutoff and it is only applicable to Haiku and not Sonnet or Opus. It's a low enough cutoff that it will be quickly exceeded if you are doing anything with Agents; for typical generation or Jev-like classifiers, it's a good value and as noted in this article, that is apparently the vast majority of Haiku use."
The "vast majority" part comes from Anthropic itself. Its launch post says prompts up to 100,000 tokens made up around 90% of requests to the Haiku 4.5. If your traffic look like that, the line rarely bites, but if you're running agents it's the first thing to measure.
Your tokens got about 30% bigger
The second hidden cost is in the tokenizer. Anthropic's what's new page says the same input text produce approximately 30% more tokens on Haiku 5.5 than on Haiku 4.5. Simon Willison measured a long prompt on his Claude token counter and he got around 1.25x, which he called "a hidden price increase".
That changes the headline discount quite a bit. Looking per word of text, not per token:
| Comparison with Haiku 4.5 | List price cut | Cut per word, at 1.3x tokens |
|---|---|---|
| Prompts up to 100k tokens | 90% | about 87% |
| Prompts over 100k tokens | 50% | about 35% |
It also moves the line itself. By my arithmetic, 100,000 Haiku 5.5 tokens hold roughly the text which was 77,000 tokens on Haiku 4.5. Any Haiku 4.5 prompt above about 77k tokens today will land over the line after migration, and the migration guide tell you to recount prompts with model set to claude-haiku-5-5 for exactly this reason.
The gap with OpenAI is even bigger. One commenter estimated that Claude's 100K tokens are about "~60-65K modern GPT tokens", so the same document reach Haiku's line well before it reaches the 272k line on GPT-6 Luna. Anthropic's own net figure, which already accounts for the tokenizer, is that Haiku 5.5 costs "around 75% less to run" than Haiku 4.5 on average. That's the honest number to plan with, not the 90%. Haiku 4.5 is still listed as active on the deprecations page, with retirement not sooner than October 15, 2026, so there is time to measure before you move. If you use it inside Claude Code, my Claude Code Haiku guide covers the switch.
Effort is the other price dial
Haiku 5.5 is the first Haiku that has an effort setting, and the API default is medium. Thinking tokens bill as output, also you can't fully turn thinking off above high effort, so effort is what decides how many of those $0.50 output tokens you buy for each answer.
Artificial Analysis ran its Intelligence Index on every level. The cost of the whole run is the most clean way to see the spread:
| Effort | Intelligence Index | Cost to run the Index | Output tokens used |
|---|---|---|---|
| Low | 29.4 | $34.31 | 32M |
| Medium (API default) | 34.5 | $56.36 | 54M |
| High | 37.8 | $94.07 | 97M |
| Xhigh | 41.2 | $157.58 | 180M |
| Max | 43.4 | $330.25 | 440M |

Max costs 9.6x what Low costs for the same set of tasks, for a 14-point gain on the index. AA measured about 162k output tokens per task at Max and called the model "very verbose". At list price that come out to $0.21 per task at Max, and those figures are using the under-100k rate, since AA notes it doesn't reflect the tiered pricing yet.
A smaller test is showing the same shape. Simon Willison drew his usual pelican at each level: Low cost 0.0936 cents and took 7 seconds, while Max took 5 minutes 9 seconds and cost 3.3826 cents. Medium and everything above got the bicycle frame right.
My take is simple. For ticket tagging and routing, or short summaries, start at Low and only move up when a test set shows you're losing accuracy. For a reply that need reasoning over a policy, Medium is the default for a reason. I'd rarely pay for the Max on Haiku.
One HN commenter read AA's numbers as putting Haiku at Max level on par in cost per task with GPT-6.1 Sol at medium, which is a much stronger model. At that point you are better to step up to Sonnet 5.5 at a lower effort. My model selection guide walks through when to step up.
What a support team actually pays
Support triage is the workload which Anthropic's launch post names directly, with classification and routing plus "live customer support". It's also where I've seen the cheapest model doing the most useful work. In one eesel trial on real Zendesk traffic from an e-commerce inbox of about 1,000 tickets a month, 22% of incoming tickets turned out to be spam, and the triage pass hit 93% accuracy with 100% spam detection and zero false positives. That's a job a small model should own, and it don't need long prompts. My roundup of the best AI for ticket triage covers the tools built for it, and the AI customer service cost guide puts the model bill next to rest of the budget.
So here are two workloads at 20,000 tickets a month, each one priced from the rate card. The token counts are the assumptions, so swap in your own in the calculator below. Cache writes are left out, since a warm cache spend little on them.
Ticket tagging and routing (under the line)
Each request carries a 2,500-token system prompt with your tag list plus a 1,500-token ticket, then it returns about 400 tokens with the label and short reason.
| Setup | Per ticket | Per month (20,000 tickets) |
|---|---|---|
| Haiku 5.5, no caching | $0.0006 | $12.00 |
| Haiku 5.5, system prompt cached | $0.000375 | $7.50 |
| Haiku 5.5, cached plus Batch | about $0.00019 | about $3.75 |
| Haiku 4.5, same text (fewer tokens) | $0.0046 | $92 |
| Sonnet 5.5, no caching | $0.012 | $240 |
At this size the model bill is basically a rounding error. What costs money is everything around it, like wiring it into the helpdesk and writing the tag rules, then also checking the labels it gets wrong. The guides on AI ticket classification, ticket routing and connecting Claude to your helpdesk cover that side.
Drafting replies with knowledge base context (over the line)
Now a reply bot. Each request loads 80,000 tokens of help center articles and macros that stays the same between tickets, plus 40,000 tokens of ticket history and order data, and it writes a 2,000-token reply with its thinking. That's 120k tokens of prompt, so over the line.
| Setup | Per ticket | Per month (20,000 tickets) |
|---|---|---|
| Haiku 5.5, 120k prompt, no caching | $0.065 | $1,300 |
| Haiku 5.5, 120k prompt, 80k cached | $0.029 | $580 |
| Haiku 5.5, trimmed to 95k, no caching | $0.0105 | $210 |
| Haiku 5.5, trimmed to 95k, 80k cached | $0.0033 | $66 |
| Sonnet 5.5, 120k prompt, 80k cached | $0.108 | $2,160 |
| GPT-6 Luna, 120k prompt, 80k cached | $0.0058 | $116 |

Caching alone cuts the bill by more than half. Trimming 25k tokens of prompt cuts it by another 9x, from $580 to $66. The trim is usually just retrieval doing its job: pull the five most relevant help articles instead of whole category, and summarize old ticket history rather than pasting it in. Haiku itself is good at that summarization step, at the cheap rate.
The Luna row is using the same token counts, which flatters Claude, since OpenAI counts fewer tokens for the same text. Luna's price doesn't step up until 272k tokens, so on this shape it stay well under Haiku. More about that in the comparison below, and in my GPT-6 Luna pricing breakdown.
Estimate your own Haiku 5.5 bill
Pick a preset or put in your own numbers. The bar shows where your prompt sits against the 100k line, then the rows compare the same workload on Haiku 4.5 and Sonnet 5.5, with GPT-6 Luna also in there.
Two things worth to try. Set the Reply preset's prompt to 99,000 and then 101,000, and watch how the monthly figure jumps. Then switch on Batch, which halves every line but only fit work that can wait up to 24 hours, for example overnight tagging of a backlog or weekly ticket summaries.
Batch, caching and the smaller line items
These are the levers which stack on top of the base price, according to the pricing docs:
| Lever | What it does on Haiku 5.5 | Watch out for |
|---|---|---|
| Batch API | 50% off input and output, both tiers | Asynchronous, so not for live chat |
| 5-minute cache write | 1.25x input, pays off after one cache read | Cache expires if traffic is sparse |
| 1-hour cache write | 2x input, pays off after two cache reads | Worth it for steady daytime traffic |
| Cache read | 0.1x input ($0.01 or $0.05) | Still counts toward the 100k line |
| US-only inference | 1.1x on every token category | Global routing is the default |
| Tool use system prompt | 286 extra input tokens (auto, none) | Down from 496 on Haiku 4.5 |
| Web search | $10 per 1,000 searches, plus tokens | Results add input tokens |
| Browser use toolset | About 6,600 input tokens per request | Screenshots bill as image input |
The modifiers multiply, so a cached read in a batch job costs half of the cache price. On a model this cheap, the tool overhead matter more than it looks. A browser use agent starts each request about 6,600 tokens closer to the line before it even read a single page.
Haiku 5.5 vs Haiku 4.5, Sonnet 5.5 and the cheap-tier rivals
Price per token is only half of the comparison. Tokenizers are different, and so is how much each model writes per task. This table puts the list prices next to AA's cost per Intelligence Index task, which folds the verbosity in:
| Model | Input / output per 1M | Price change point | AA cost per task |
|---|---|---|---|
| Claude Haiku 5.5 | $0.10 / $0.50 | 5x past 100k prompt | $0.21 (Max) |
| GPT-6 Luna | $0.10 / $0.50 | 2x input, 1.5x output past 272k | $0.07 (Max) |
| Gemini 3.5 Flash-Lite | $0.30 / $2.50 | None listed | $0.19 |
| GLM-5.3 Flash | $0.15 / $0.50 | None listed | $0.25 |
| DeepSeek V4.1 Flash | $0.15 to $0.30 / $0.60 to $1.20 | Peak hours cost 2x | Not measured |
| Claude Haiku 4.5 | $1 / $5 | 200k context cap | Not measured |
| Claude Sonnet 5.5 | $2 / $10 | Flat to 1M | Not measured |
Prices are coming from each vendor's own pricing page, and the full OpenAI card is in my OpenAI API pricing guide.
If you are tempted by Gemini 3.8 Flash, keep in mind that Google's pricing page doubles its price on January 1, 2027.
The cleanest real-world comparison I could find was a developer who run the same classification job on both of them:
"I compared the costs between luna and haiku for some classification work. Luna comes out a fair amount cheaper due to more efficient tokenization and less outputs."
Their numbers: Haiku 5.5 used 11,893,643 input tokens and cost $0.42 synchronously, while GPT-6 Luna used 7,903,468 input tokens for the same work and cost $0.31. Same list price, yet Haiku's bill came out 35% higher. The Plotly's data analytics benchmark told a similar story, with Luna doing "a bit better" at about 30% of Haiku's cost.
So why pay more for Haiku? Speed, and in some tests also quality. AA measured 137 to 243 output tokens per second across effort levels, and one commenter reported OpenRouter at roughly twice of Luna's throughput. Anthropic calls it its fastest model at standard speed, which is why it names live customer support as a fit. A slow bot costs you more in human agent time than any token line. My Haiku 5.5 alternatives post goes deeper on when to switch, and the three-way API comparison covers how to pick a provider.
Against Anthropic's own models, the step up matters more. Sonnet 5.5 is 20x the input price of short-prompt Haiku, but only 4x the price of Haiku over the line, and it has no line at all. If most of your requests sit between 100k and 400k tokens, run both of them through the calculator above before assuming Haiku wins. The Sonnet 5.5 pricing breakdown (and the older Sonnet 5 pricing) also covers the same-day cut on cache read, from $0.20 to $0.10, which Anthropic says makes Sonnet 5.5 about 20% cheaper on most agentic work.
Using Haiku 5.5 on a Claude plan or with API credits
If you only want to chat with Haiku, you don't pay per token at all, and my Claude pricing guide compares the plans in full. The Claude plans page lists Haiku on all of Free, Pro, Max and Team. Pro is $17 a month billed annually ($20 monthly) while Max starts at $100, and the Team seats are $20 Standard or $100 Premium per seat on annual billing.
The newer angle is the API credits. Together with Haiku 5.5, Anthropic started to give Max and Team subscribers a monthly credit for the Claude Platform:
| Plan | Monthly API credit |
|---|---|
| Max 5x | $100 |
| Max 20x | $200 |
| Team, Standard seat | $20 per seat |
| Team, Premium seat | $100 per seat |
| Team pool cap | $500 per month |

The help article has the rules. Free, Pro and Enterprise aren't eligible, and you need to be on the plan for seven days. Credits don't roll over, which is a thing to remember, and they cover the API plus Batches, Managed Agents and the Agent SDK, also claude -p with an API key. They don't cover interactive Claude Code or the usage on Bedrock, Vertex AI or Foundry.
On Haiku 5.5 that credit goes a long way. By my arithmetic, $100 buys about 1 billion short-prompt input tokens or 200 million output tokens a month. The cached tagging workload above would run for about a year on one month of Max 5x credit. A side project or internal tool can live completely inside the credit, as long as it stay under the line.
What developers are saying about the price
The launch thread was split along one question: does your work fit inside 100k tokens?
"noticeably smarter remains to be seen in practice. For now, Haiku is a bit more expensive than Luna on < 100k token, but I just don't have any agentic work below 100k, so this is going to be 5x more expensive than shown on these charts."
The other side see the line as a fair trade for subagent work:
"I would mostly use Haiku in task or explorer subagents. I'm not saying I stay under that on every task, but I do have quite a few sessions that cap out well below that, so that price difference would be very meaningful."
And one commenter read the whole structure as more of a pricing move than a technical limit:
"I'd consider the 100k a "promotional price" to match Luna's token pricing while delivering noticeably more intelligence."
That last one is speculation, and Anthropic hasn't said it will change. Still, it's a fair reason to not build a cost model that only works if Haiku stays at $0.10 for prompts of every sizes.
Is Claude Haiku 5.5 worth it at this price?
For short, high-volume jobs, yes, clearly. Classification and tagging, routing, extraction, short summaries: these all fit under 100k tokens, they cost cents per thousand calls, and on top they run fast. Coming from Haiku 4.5, it's the most easy cost cut in the Claude lineup, about 87% per word of text. If you run Claude for Zendesk or a homegrown triage script, this is the model to move first. The Claude for customer support guide has the setup.
For long-context agents, it's more of a close call. Over the line, Haiku costs a quarter of Sonnet 5.5 per token with no ceiling on how often you cross it, and GPT-6 Luna stays cheaper all the way to 272k. Measure your prompt sizes before you migrate, then design the retrieval to stay under the line if you can.
The full quality picture is in my Haiku 5.5 review, where the HN thread also has developers who report regressions against Haiku 4.5 on their own evals. A cheaper model that need prompt rework isn't cheaper until the rework is done.
Hand eesel the triage queue instead
If the reason you're pricing Haiku 5.5 is your support queue, the token bill is the smallest part of the job, a point my AI support cost savings breakdown makes with real numbers. You still need the helpdesk integration and the tag rules, then a way to test on past tickets, and also someone who watches what it gets wrong. That's the work eesel takes off from your plate.
eesel's AI helpdesk teammate joins the helpdesk you already have, such as Zendesk, Freshdesk or Gorgias, and it learns from your past tickets, then it tags and routes tickets, drafts replies or answers them.
Before it touches a live customer, I'd run it against your historical tickets, and that simulation comes built in. You can track what it does inside the reports view:

The pricing is the opposite of a token meter. A ticket or chat is one credit no matter how long it runs, so a 120k-token thread costs same as a two-line question. Plans start at $299 a month for 500 credits and go up to $1,749 for 5,000, with a free trial of 100 credits and no card needed. Predictability matters more than it sound. One eesel trial user, on Gorgias, watched their agent handle 12 test chats well, then asked to cancel in the moment they reached the billing page. Compare that with metered tools like Zendesk AI pricing, where the bill moves with every resolution.

To be straight about the fit: if you only need a tagging script and you have engineers, the raw Haiku 5.5 API at $12 a month is cheaper, and you should use it. eesel makes sense when you want the whole teammate, from triage to replies, without building the pipeline yourself. If you'd rather automate support from the command line, the eesel CLI drives the same teammate, so you or a coding agent can script its instructions and approvals, plus the activity log. Try eesel free on a slice of your real tickets and see how the triage looks before you write a line of code.
Frequently Asked Questions
How much does Claude Haiku 5.5 cost per million tokens?
Why does Claude Haiku 5.5 cost more above 100,000 tokens?
Is Claude Haiku 5.5 cheaper than Claude Haiku 4.5?
Does prompt caching keep Claude Haiku 5.5 under the 100k price tier?
Is Claude Haiku 5.5 cheaper than GPT-6 Luna?
Can I use Claude Haiku 5.5 for free?
What is the cheapest way to run Claude Haiku 5.5 for customer support?

Article by
Kurnia Kharisma
Kurnia is a software engineer and writer at eesel AI with two years of SEO experience, writing about AI tools, helpdesk software, and customer support. He pairs a developer's understanding of how these products are built with search-driven research into what actually ranks and resonates with the people searching for them.








