Claude Opus 5: what it is and how to actually run it

Kurnia Kharisma Agung Samiadjie
Written by

Kurnia Kharisma Agung Samiadjie

Katelin Teen
Reviewed by

Katelin Teen

Last edited August 4, 2026

Expert Verified
An illustration of a person selecting one of five reasoning effort dials on a Claude Opus 5 style control strip

What Claude Opus 5 actually is

In about two months Anthropic has shipped four Claude 5-generation releases, and Opus 5 is the fourth of them. It slots in directly above Opus 4.8, at a price that has not moved. The launch framing says it "comes close to the frontier intelligence of Claude Fable 5 at half the price." Unusually for launch framing, the independent numbers mostly back that up.

Here are the specs, taken from the models overview:

SpecClaude Opus 5
Claude API model IDclaude-opus-5
AWS Bedrock IDanthropic.claude-opus-5
Google Cloud IDclaude-opus-5
Context window1M tokens, no beta header, no long-context premium
Max output128k synchronous, 300k on the Batch API
Price$5 / $25 per MTok
Adaptive thinkingYes, on by default
Effort levelslow, medium, high, xhigh, max
Knowledge cutoffMay 2026
Web fetchNot supported
Priority TierNot supported

Two rows in there tend to surprise people. First, the 1M context window is served at standard rates, so a 900k-token request bills per token at the rate a 9k one does, which the pricing docs spell out. Second, no Priority Tier at all. Opus 4.8 keeps one, so anybody sitting on a capacity commitment has to plan that part separately.

If your last look at this family was a couple of releases back, the pace is worth registering. My Opus 4.5 review is only a few months old, and it already reads like a different tier of tool, with Opus 4.6 sitting somewhere in between.

"Claude Opus 5" is five models wearing one name

Everything else in the post gets reframed by this one part, so I want to spend real time on it here.

The effort parameter is not some quality slider you nudge for taste. In the whole API it is the single biggest lever on cost and latency, and on output quality too, and the distance between its two ends is wider than what separates most competing models.

A diagram showing the single claude-opus-5 model ID fanning out into five effort settings, from low at 1223 Elo to max at 1720 Elo and $17.79 per task
A diagram showing the single claude-opus-5 model ID fanning out into five effort settings, from low at 1223 Elo to max at 1720 Elo and $17.79 per task

Artificial Analysis ran every setting through AA-Briefcase. That benchmark covers long-horizon knowledge work, built on multi-week projects carrying thousands of source files. The spread is the story here:

Effort settingAA-Briefcase EloCost per taskMinutes per taskTurns per task
max1720$17.7936.2103
xhigh1693$14.2634.391
high1606$10.4125.776
medium1470not published
low1223not published
Claude Fable 5, for scale1574$22.30

Read the Fable 5 row next to the high row and what falls out is the best value data point in the whole dataset: at high effort, Opus 5 beats Fable 5 by 32 Elo, and it does that at $10.41 a task, under half of Fable's price. Now read the same row against low. 1223, which is below GLM-5.2. Same model ID, same week, and both of them true.

GDPval-AA v2 shows the same shape, where the effort levels span 407 Elo points and output token usage ranges around 8x going from low up to max.

Chart plotting Elo score against benchmark cost on GDPval-AA v2, with Opus 5, Fable 5, Opus 4.8 and GPT-5.6 Sol each drawn as a five-point effort ladder, as taken from Anthropic
Chart plotting Elo score against benchmark cost on GDPval-AA v2, with Opus 5, Fable 5, Opus 4.8 and GPT-5.6 Sol each drawn as a five-point effort ladder, as taken from Anthropic

Worth noticing what the chart admits: down at the cheap end of the ladder, GPT-5.6 Sol sits above Opus 5 and not below it. Only once you are paying real money per task does Anthropic's line pull ahead. A single headline rank hides exactly that sort of detail.

One more wrinkle, and this is the strangest number of the whole launch. On Cognition's FrontierCode 1.1 the best Opus 5 score lands at medium effort. Not max. Anthropic explains why straight out in the system card: at higher effort the model "make[s] more changes than the task requires," and out-of-scope edits get penalised by the grader. Adding a short stay-in-scope instruction "recovered performance on most of these tasks." More thinking, then, is not monotonically better, and on one real coding benchmark at least, it is worse.

Pick your effort level

Instead of making you scroll back up to that table, here is the same decision laid out as a picker. Pick whichever setting you were planning to run, then see what it buys you.

What each Claude Opus 5 effort level actually buys

Elo, cost, wall-clock and turn count are Artificial Analysis AA-Briefcase measurements. Pick a setting.

AA-Briefcase Elo1223
Cost per tasknot published
Output tokens~1/8 of max
FrontierBench25%
Use it for: classification, extraction, tagging, summarising a known document. At this setting Opus 5 scores below GLM-5.2 on long-horizon work, so it is the wrong tool for anything open-ended. On FrontierBench it earns 25% for 64% fewer output tokens than the top setting.
AA-Briefcase Elo1470
AutomationBench24% at $0.89
FrontierCode 1.1best of all settings
Turnsfewer than high
Use it for: most day-to-day coding. This is the setting Anthropic's own FrontierCode numbers peak at, and where AutomationBench hits 24% for $0.89 a task, beating both Opus 4.8 and Fable 5 at under half their cost. If you only remember one row from this widget, make it this one.
AA-Briefcase Elo1606
Cost per task$10.41
Minutes per task25.7
Turns per task76
Use it for: the default, and Anthropic says so explicitly. It beats Fable 5 by 32 Elo at less than half Fable's cost per task, and it is the setting that produced the headline 30.2% ARC-AGI-3 score. Start here, then move down before you move up.
AA-Briefcase Elo1693
Cost per task$14.26
Minutes per task34.3
FrontierBench44.4%, the peak
Use it for: hard agentic engineering where you have already tried high and it was not enough. This is the top FrontierBench score, and on GDPval it still beats every other model while burning 25% fewer output tokens than max. Note that thinking: disabled is rejected above high, so this setting always thinks.
AA-Briefcase Elo1720
Cost per task$17.79
Minutes per task36.2
Turns per task103
Use it for: almost nothing on a schedule. It is the top score on every leaderboard Anthropic published, and it costs 71% more per task than high for 114 more Elo, takes 36 minutes, and runs 103 turns against Opus 4.8's 55. Batch work and research runs only.

Sources: Artificial Analysis AA-Briefcase and Intelligence Index runs, plus Anthropic's Opus 5 system card. Cost figures are per benchmark task, not per API call.

Why nobody online agrees on whether it is cheap

Once the effort dial is sitting in your head, launch-week chaos stops looking confusing at all.

Inside the same 48 hours, Reddit ran two top threads that contradicted each other flat out. One camp posted that token usage is amazing, and an r/ClaudeCode poster there reported eight hours of work for 5% of a weekly quota. The other camp posted about burning the 5-hour limit 10x faster, describing a limit they used to reach in six hours now arriving in under two. Then a thread titled, literally, so which is it pulled 150-plus comments from people asking about the whiplash.

Hacker News got to the answer faster than the benchmark coverage did:

Hacker News

"The chart shows max effort, used mostly by price-insensitive enterprise users. At medium effort it drops to almost half K3's cost, and is probably sufficient for 95% of coding tasks."

Hacker News

"It overthinks quite a bit above medium effort, try using that."

Underneath that first cause sits a second one, and this one catches people who never touched the effort parameter at all. Thinking is on by default on Opus 5 and it was not on Opus 4.8, so a request you never changed now spends tokens on reasoning that it used to skip. Someone caught it in the migration notes within hours of launch:

Hacker News

"The breaking changes vs. Opus 4.8 are interesting: 1. Thinking on by default: On Claude Opus 4.8, requests without a thinking field run without thinking; on Claude Opus 5, the same requests run with adaptive thinking. 2. Disabling thinking is capped at high effort"

Both camps, then, were telling the truth. Just about different configurations of the same model. That is the whole fight.

The two things that break when you migrate

Anthropic calls Opus 5 "a drop-in upgrade for Claude Opus 4.8 at the same pricing," and the migration guide then lists exactly two breaking changes.

Thinking is on by default. On 4.8, a request with no thinking field ran thinking-free; on Opus 5 that same request runs with adaptive thinking. And since max_tokens caps total output, reasoning included, any workload that used to run lean wants its ceiling revisited before it starts truncating on you.

Disabling thinking is capped at high. Sending thinking: {type: "disabled"} still works, but combine it with xhigh or max and you get a 400 back. That check runs per request, which means raising effort mid-conversation while thinking is off gets rejected even when earlier turns went through fine. Anthropic warns too that with thinking off, the model "can occasionally emit tool calls as plain text or include internal XML tags in its visible output."

The rest of the upgrade is upside, and a few pieces of it are worth switching on deliberately:

  • The prompt-cache floor dropped to 512 tokens, coming down from 1,024 on Opus 4.8 and 4,096 on Opus 4.6. Short system prompts that were never cacheable before now are, and no code change is needed for it.
  • Mid-conversation tool changes, sitting behind the mid-conversation-tool-changes-2026-07-01 beta header, let you add tools or remove them between turns, and cache hits on the earlier turns survive it.
  • A fallbacks parameter sends cyber-classifier refusals over to Opus 4.8 rather than returning an error, so requests "always route to the best available model by default rather than being blocked." Availability is the catch: not on the Batch API, not on Bedrock, and not on Google Cloud or Microsoft Foundry.
  • Opus 5 has its own rate-limit bucket. The rate limits page is explicit about it not being part of the combined Opus 4.x pool, and it lists the bucket at 1,000 RPM with 2,000,000 input tokens per minute.

Two prompting deltas are easy to miss here. Anthropic says to strip out any carried-over verification instructions, because "Claude Opus 5 verifies its own work without being told to," and it says to cap subagent spawning as well, the model delegating more readily than its predecessors did. Claude Code took that second one seriously enough to ship a hardcoded instruction telling Opus 5 not to use subagents. Reaction on HN split, predictably.

Where it leads, and where it does not

The headline is real enough. On its Intelligence Index, Artificial Analysis puts Opus 5 at max effort at 61, rank #1 against a class median of 32. Before getting excited, read the runners-up: Fable 5 at 60, GPT-5.6 Sol at 59, Kimi K3 at 57, then Opus 4.8 at 56. Artificial Analysis's own word for Opus 5 is "narrowly the most intelligent model," and the whole frontier fits inside a five-point band.

Where the wide, unambiguous wins live is all one place, agentic work. Zapier's AutomationBench drops an agent into a simulated company holding dozens of REST endpoints spread over 47 apps, and there Opus 5 scores 26.0% against 17.0 for Opus 4.8, 17.4 for Fable 5 and 18.1 for GPT-5.6 Sol.

Chart of AutomationBench pass rate against cost per task, with Opus 5's effort ladder sitting well above Fable 5, Opus 4.8 and GPT-5.6 Sol, as taken from Anthropic
Chart of AutomationBench pass rate against cost per task, with Opus 5's effort ladder sitting well above Fable 5, Opus 4.8 and GPT-5.6 Sol, as taken from Anthropic

The other real outlier is ARC-AGI-3, at 30.16% Relative Human Action Efficiency against GPT-5.6 Sol's 7.78%. And worth flagging for this post's argument: that score came out of high effort, not max.

Chart of ARC-AGI-3 score against total evaluation cost, showing Opus 5 at high effort near 30% and GPT-5.6 Sol under 10%, as taken from Anthropic
Chart of ARC-AGI-3 score against total evaluation cost, showing Opus 5 at high effort near 30% and GPT-5.6 Sol under 10%, as taken from Anthropic

Now the honest other column. Artificial Analysis measured 53% on Humanity's Last Exam, "in line with Fable 5," so a tie and not a win. Over on CritPt frontier physics it matches Fable 5, sitting behind three OpenAI models. Terminal-Bench v2.1 gives it 89%, which is "roughly in line with" the leader. Anthropic itself says Opus 5 stays behind Mythos 5 on biology research, and on offensive cybersecurity too.

Then there is the leaderboard that finally has data on it. At launch, LMArena carried no entry for Opus 5, so nobody could say what humans made of it. As of early August it is listed. The answer is not the one the Intelligence Index would predict.

A two-column comparison showing Claude Opus 5 ranked number one on the Artificial Analysis Intelligence Index and number seven on LMArena blind human preference
A two-column comparison showing Claude Opus 5 ranked number one on the Artificial Analysis Intelligence Index and number seven on LMArena blind human preference

On LMArena's text leaderboard, claude-opus-5-high sits at rank 7 with 1492 Elo, and claude-opus-5-max at rank 8 with 1490. Above the pair of them: Fable 5 at 1509, plus three older Claude models. Benchmarks measuring task completion, and humans voting on the answer they liked better, are measuring two different things, and Opus 5 demonstrates that gap cleanly. The community put it in plainer language:

Hacker News

"I compared the writing style of Opus 5 vs Fable 5, and Opus 5 continues many of the 'Claude-isms' of its 4.8 predecessor in a way that Fable broke away from."

How to actually get it

Six routes to it, and they do not all hand you the same thing.

RouteWhat you getNotes
Claude Free, $0No Opus at allSonnet and Haiku only, 200k context
Claude Pro, $20/moOpus 5 as the strongest model available$17/mo billed annually. Fable via usage credits
Claude Max 5x, $100/moOpus 5 as the default modelFable capped at 50% of weekly limits
Claude Max 20x, $200/moSame, 20x Pro usage per sessionSession-based limits, not message counts
Claude Team, $25/seat/moAll four model families$20/seat annually, teams of 2 to 150
Claude APIclaude-opus-5 at $5 / $25Batch halves it, cache hits cost $0.50

Claude Code comes with every paid plan and it draws from the same pool. Should your picker still be serving 4.8, /model claude-opus-5 works directly, and on older CLI versions /model claude-opus-5[1m] gets you the 1M-context variant. Worth knowing as well: Claude Code at first exposed only a 200K window for it, which is the kind of thing my notes on Claude Code's context window exist for.

Over on the cloud platforms, Bedrock and Google Cloud are partner-operated, they invoice you directly, and regional endpoints carry a 10% premium over the global ones. Claude Platform on AWS bills in Claude Consumption Units at $0.01 per CCU, and so does Microsoft Foundry. On top of tokens, Managed Agents add $0.08 per session-hour of active runtime.

Fast mode is the one route carrying real strings. Up to 2.5x higher output tokens per second, at exactly double the price, $10 and $50, applied across the full context window. The conditions on it: still a research preview, first-party Claude API only, gated behind either a waitlist or an account manager, and it does not work with Batch. Per token those rates land in the same place as standard Fable 5, so what you are buying there is latency and not capability. My Claude Opus 5 pricing breakdown works through what that does to a monthly bill, while Claude Code pricing covers the CLI side.

The wall between Opus 5 and a live customer

Everything above has been about picking a setting. This section is about the place where no setting helps.

A timeline showing three latency zones, with Claude Opus 5 at max effort sitting in the minutes-to-half-an-hour zone behind a marked latency wall
A timeline showing three latency zones, with Claude Opus 5 at max effort sitting in the minutes-to-half-an-hour zone behind a marked latency wall

Right now Artificial Analysis measures Opus 5 at 55.7 output tokens per second, which is rank #111 of 184, and time to first token sits at 63.43 seconds. Dropping the effort down buys back very little speed. All three of the top settings on AA-Briefcase average more than 25 minutes per task. The one-liner Artificial Analysis wrote for it: Opus 5 is "amongst the leading models in intelligence, but particularly expensive when comparing to other models of similar price. It's also notably slow and very verbose."

The second number is the one I care about more, and it points opposite to the marketing. Over Opus 4.8, Opus 5 gained 7 points of accuracy on AA-Omniscience, and at the same time its hallucination rate rose 14 points to 50%, the reason being that it answers more often when it is uncertain. Anthropic's own system card says it plainly: "The model hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall."

For a coding agent that reruns its own tests, that trade is fine. Now put the same trade on a support reply going out to a paying customer with your logo attached. A confident wrong answer there is the failure mode that costs you the account, and that is why preventing hallucinations in support ends up being a product-design problem instead of a model-selection one.

I have watched this play out enough times now to be blunt about it. Something like 10% of a working support deployment is the model. The other 90% is knowledge retrieval that stays current, plus a confidence threshold deciding what the AI is even allowed to touch, a simulation pass over historical tickets so the answer rate is known to you before any customer sees it, and then a handoff that does not drop context on the floor. One eesel customer, running a crypto-ATM network, put the build-versus-buy side of it better than I can:

"We could try to write our own LLM application but we didn't want to invest our time into that. We wanted something that we would not have to maintain."

Karel, GENERAL BYTES

A CX lead at a DTC supplements brand then described the control problem in one sentence that has stuck with me for years: the AI will never answer 100% of questions, so what she needed was an AI handling only the tickets it is confident about and leaving the rest alone. No effort parameter hands you that. It is a routing decision, and it gets made outside the model.

Try eesel

If you landed here because a model is being picked for a support queue, the honest advice is this: Opus 5 is a fine engine and the engine was never your bottleneck. eesel is the part that wraps around it. It plugs into the helpdesk you already run, it trains on your past tickets and your help center instead of on a pasted prompt, and it simulates against your own ticket history, so the resolution rate is in front of you before any customer sees the agent at all. Setup takes minutes rather than a quarter, and the pricing counts resolved tickets and not tokens, so a verbose model cannot quietly inflate the invoice.

The eesel reports dashboard showing task volume over 30 days, trigger events by type, and approval versus rejection usage per tool
The eesel reports dashboard showing task volume over 30 days, trigger events by type, and approval versus rejection usage per tool

That reporting view is the piece no raw API key gives you: which tickets the agent actually took, what triggered each run, and the points where a human stepped in. To see what your own numbers look like you can try eesel free, or read how the same decision plays out across this year's support tools and what an AI support agent costs.

For the rest of the Opus 5 picture: the full review carries the verdict on whether to upgrade at all, Opus 5 versus Fable 5 covers the tier above and Opus 5 versus Sonnet 5 the tier below, and Sonnet 5 pricing matters right now, since its introductory $2 and $10 rates expire on August 31, 2026.

Frequently Asked Questions

What is Claude Opus 5?
Claude Opus 5 is Anthropic's workhorse model, released on July 24, 2026. The model ID is claude-opus-5, it holds a 1M token context window with 128k max output, and it costs $5 per million input tokens and $25 per million output. Anthropic's docs tell you to start here for agentic coding and enterprise work, and move up to Fable 5 only when you need the ceiling. My Opus 5 review has the full verdict.
How much does Claude Opus 5 cost?
$5 per million input tokens and $25 per million output on the Claude API, which is exactly what Opus 4.8 cost. Batch halves it to $2.50 and $12.50, a cache hit is $0.50, and fast mode doubles it to $10 and $50. The catch is that Claude Opus 5 pricing per token tells you almost nothing about your bill, because the effort level decides how many tokens get spent. Full tables live in my Opus 5 pricing breakdown.
Is Claude Opus 5 better than Claude Fable 5?
On Anthropic's published benchmarks it beats Fable 5 on most rows at half the token price, and Anthropic still says it is not more capable overall. Both things can be true, since Fable is the larger model and Opus 5 is faster to a finished task. I worked through the whole thing in Opus 5 versus Fable 5.
What is the difference between the Claude Opus 5 effort levels?
There are five: low, medium, high, xhigh, max. On Artificial Analysis's AA-Briefcase they span 497 Elo points, and output token usage runs about 8x from low to max on GDPval-AA v2. The API default is high. Treat effort as your main cost control before you look at cheaper models like Sonnet 5.
Do I need to change my code to upgrade from Opus 4.8?
Two things break. Thinking is on by default now, so a request that previously ran thinking-free will spend part of your max_tokens on reasoning. And passing thinking: disabled together with xhigh or max effort returns a 400. Everything else is a drop-in swap at the same price. If you run it through the CLI, my Claude Code model selection notes cover the picker side.
Where can I use Claude Opus 5 for free?
You cannot. The Claude free tier is Sonnet and Haiku only. Opus 5 starts on Pro at $20 a month, is the default model on Max, and is available on Team and Enterprise seats. Claude Pro pricing has the plan detail, and Claude alternatives covers what to run when none of those fit.
Can Claude Opus 5 handle customer support tickets?
It can write the answer. It should not be the thing your customer waits on: 63 seconds to first token at max effort, and a hallucination rate that went up rather than down. A support deployment needs retrieval, confidence routing, and a clean handoff to a human around the model. That layer is what helpdesk AI products exist to supply.

Share this article

Kurnia Kharisma Agung Samiadjie

Article by

Kurnia Kharisma Agung Samiadjie

Kurnia is a software engineer and writer at eesel AI with two years of SEO experience, writing about AI tools, helpdesk software, and customer support. He pairs a developer's understanding of how these products are built with search-driven research into what actually ranks and resonates with the people searching for them.

Related Posts

All posts →
Illustration of a person weighing a small low-cost AI model against a larger caped flagship model on pedestals
Trending

Claude Opus 5 vs Fable 5: which should you actually run?

Fable 5 costs exactly double Opus 5. I went through both system cards, the docs and the independent benchmarks to work out when that second dollar buys anything.

Rama Adi NugrahaRama Adi NugrahaJul 27, 2026
Illustration comparing a heavyweight reasoning model against a fast balanced model on cost and capability
Trending

Claude Opus 5 vs Sonnet 5: which one should you use?

Claude Opus 5 costs 1.7x Sonnet 5 per token and still finishes some jobs cheaper. Here is the head-to-head on price, benchmarks and real cost per task.

Rama Adi NugrahaRama Adi NugrahaJul 27, 2026
Illustration of a developer at a laptop watching an agentic coding loop run through code, checks and a bot
Trending

Claude Opus 5 review: near-frontier coding at half the price

A hands-on Claude Opus 5 review: what the benchmarks actually say, the hallucination rate that went up, and whether it belongs on a live support queue.

Alicia Kirana UtomoAlicia Kirana UtomoJul 27, 2026
Illustration comparing Alibaba's Qwen3.8-Max preview with Anthropic's Claude Fable 5
Trending

Qwen3.8-Max vs Claude Fable 5: the comparison nobody can run

Alibaba called Qwen3.8-Max "second only to Fable 5". I lined both models up on price, context, availability and published evidence to see whether that holds.

Kurnia Kharisma Agung SamiadjieKurnia Kharisma Agung SamiadjieAug 3, 2026
Editorial illustration of Claude Opus 4.8, Anthropic's flagship AI model
Guides

What is Claude Opus 4.8? A clear-eyed look at Anthropic's flagship model

Claude Opus 4.8 is Anthropic's latest flagship model. Here's what changed, what it costs, and what a smarter model actually means for AI customer support.

Riellvriany IndriawanRiellvriany IndriawanJun 17, 2026
Image alt text
Guides

An overview of Claude Opus 4.6 pricing and capabilities

Explore our deep dive into Claude Opus 4.6 pricing. We break down the costs, new features, and practical use cases for Anthropic's latest AI model.

Katelin TeenKatelin TeenFeb 6, 2026
Editorial illustration for a guide to what Claude Fable 5 can do, Anthropic's most powerful AI model
Guides

What can Claude Fable 5 do? A capability-by-capability guide

What can Claude Fable 5 do? Run for days unattended, write and ship code, read 1M-token documents, and check its own work. Here's what that means in practice.

Riellvriany IndriawanRiellvriany IndriawanJun 17, 2026
Image alt text
Guides

Gemini 3 Pro vs Claude Opus 4.6: A practical comparison

This guide provides a straightforward, practical look at Gemini 3 Pro and Claude Opus 4.6. We’ll cut through the hype and focus on the real-world differences that matter when you’re putting these tools to work.

Stevia PutriStevia PutriFeb 6, 2026
GPT-5.6 versus Claude comparison hero illustration, two AI model families balanced against each other
Trending

GPT-5.6 vs Claude: which AI model wins in 2026?

A hands-on GPT-5.6 vs Claude comparison: the Sol/Terra/Luna tiers against Opus 4.8 and Sonnet 5, on pricing, benchmarks, context, and AI support agents.

Kurnia Kharisma Agung SamiadjieKurnia Kharisma Agung SamiadjieJul 10, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free