Bedrock Managed Agents pricing 2026: what an OpenAI agent on AWS actually costs

Kurnia Kharisma
Written by

Kurnia Kharisma

Katelin Teen
Reviewed by

Katelin Teen

Last edited October 1, 2026

Expert Verified
Hand-drawn illustration of a developer at a laptop beside a long receipt listing input tokens, output tokens and runtime, with a cloud and gateway diagram and the AWS logo on an orange circle

What does Bedrock Managed Agents cost?

AWS's pricing line for the preview is one sentence: "During preview, there is no additional charge for BMA beyond the underlying AWS resources your agents consume. Pricing is subject to change at general availability" (AWS What's New).

Two years of doing SEO taught me one thing about pricing queries: the word "free" sends more people to the wrong conclusion than any other. Somebody who searches "bedrock managed agents pricing" wants the number their finance team is going to see, and that number is never $0. The developer docs spell out what "underlying resources" covers: "You incur charges for model inference and the AWS resources that your application uses," and the AgentCore example "can continue to incur charges when no BMA turn is running" (AWS docs).

If you're new to the service itself, my Bedrock Managed Agents guide covers how sessions, the exec server and execution environments fit together. This one is only about the bill.

Here's every line that can show up, in one table:

Cost linePreview priceBillable unitWhen it applies
BMA service fee$0n/aAlways (subject to change at GA)
OpenAI model tokensIn-Region rate, e.g. $0.22 in / $1.32 out for GPT-5.6 LunaPer 1M input, cached and output tokensEvery turn
AgentCore Runtime v2 CPU$0.1276 per vCPU-hourPer second, active CPU onlyIf AgentCore is your execution environment
AgentCore Runtime v2 memory$0.0169 per GB-hourPer second, idle memory reclaimed after 120 secondsIf AgentCore is your execution environment
NAT gateway$0.045 per hour + $0.045 per GBPer hour (partial hours bill as full)AWS's AgentCore sample stack
S3, S3 Files, CloudWatchStandard AWS ratesStorage and requestsSkills, outputs and logs
Self-hosted computeWhatever your host costsYour existing billIf you run codex exec-server yourself

That's the whole shape of it. The lines don't all matter equally, though, so the rest of this post is about the ones that do.

Line 1: model tokens, and the 10% in-Region premium

Every OpenAI model on Bedrock lists its prices on its own model card. The detail that matters for BMA is in the preview limitations: cross-Region inference isn't supported. Which means every BMA request is paying the in-Region rate, and AWS says it plainly: "Commercial In-Region prices include a 10% fee over OpenAI rates" (GPT-5.6 Luna card).

The GPT-5.6 Luna model card in the Amazon Bedrock user guide, showing In-Region, Geo and Global inference rates for short and long context, as taken from the AWS docs

Here are the Standard-tier, short-context rates (272K input tokens or fewer) per million tokens for the OpenAI models I'd consider for an agent, all from their Bedrock cards on AWS's OpenAI models page:

ModelInput (BMA rate)Cache writeCache readOutput (BMA rate)OpenAI's own rate (in / out)
GPT-6 Luna$0.11$0.1375$0.011$0.55$0.10 / $0.50
GPT-5.6 Luna$0.22$0.275$0.022$1.32$0.20 / $1.20
GPT-6.1 Sol$2.20$2.75$0.11$11.00$2.00 / $10.00
GPT-5.6 Terra$2.20$2.75$0.22$13.20$2.00 / $12.00
GPT-5.6 Sol$4.40$5.50$0.44$22.00$4.00 / $20.00
GPT-6 Astra$11.00$13.75$1.10$55.00$10.00 / $50.00

Four details on those cards will change your estimate more than the table suggests:

  1. Long context doubles input. Once a request goes over 272,000 input tokens, the long-context rate applies "to the full request," not just the overflow (GPT-6.1 Sol card). On GPT-5.6 Sol that's $8.80 input and $33.00 output. An agent dragging a big context window through every turn will hit this quickly.
  2. No Priority or Flex discounts. Every card above says Priority and Flex tiers aren't supported, so there's no cheaper batch-style tier to drop into for overnight jobs. Standard is all you get.
  3. Caching isn't uniform. The GPT-6.1 Sol card lists cache prices but says "Explicit prompt caching is not supported for this Bedrock model." Check the card for the exact model before you count on the 90% cache-read discount.
  4. The newest models live in one Region. On the bedrock-mantle endpoint BMA uses, GPT-6 Luna and GPT-6.1 Sol are available only in us-east-1. If you deploy in Oregon or Ohio, you're on the GPT-5.6 family.

The newest generation is also the cheapest per token here: GPT-6 Luna is half the price of GPT-5.6 Luna on both input and output. If a Luna was going to be your pick anyway, test GPT-6 Luna first. My GPT-6 Luna pricing and GPT-6.1 Sol pricing posts compare them against the rest of OpenAI's lineup.

Line 2: AgentCore Runtime hours

BMA doesn't run your tools itself; they run on compute that you provide: either your own host, or Amazon Bedrock AgentCore Runtime, which is the default in AWS's example. If you use AgentCore, you pay its runtime rates (AgentCore pricing):

AgentCore Runtime resourceConsumption rateCommitted baseline
v2 microVM CPU$0.1276 per vCPU-hour$0.0997 per vCPU-hour (launching by October 2026)
v2 microVM memory$0.0169 per GB-hour$0.0132 per GB-hour (launching by October 2026)
v1 microVM CPU$0.0895 per vCPU-hourn/a
v1 microVM memory$0.00945 per GB-hourn/a
Runtime instances (EC2)EC2 On-Demand rateplus a 12% management fee
The Amazon Bedrock AgentCore pricing page, scrolling through its pay-for-what-you-use feature list, as taken from AWS

This billing model is kinder to agents than it first looks. AWS says CPU "scales to zero during I/O wait (waiting for LLM responses, tool / API calls, or database queries)," and on v2 "idle memory is reclaimed automatically" after 120 seconds. Since an agent spends most of its life waiting on the model, most of its wall-clock time never bills for CPU at all.

AWS's own pricing page has a worked example that's close to a support agent: 1 million sessions a month, 10 minutes each, 90% I/O wait, 1 vCPU and up to 2.5 GB. Its total is $0.006703 per session, or $6,703 a month for the million. And that's runtime only, before you've paid for a single token.

One thing to watch: the BMA AgentCore example sets the runtime's idle timeout and maximum lifetime to 28,800 seconds, which is eight hours. Idle memory gets reclaimed, sure, but a session you forget about is still a session.

Line 3: the NAT gateway that bills while your agent sleeps

AWS's AgentCore example stack creates a VPC with private subnets, an S3 gateway endpoint and one NAT gateway. At Amazon VPC pricing in US East, a NAT gateway costs $0.045 per hour plus $0.045 per GB processed, and partial hours bill as full hours.

Across a 730-hour month that's $32.85 before your agent does anything. For a production team that's small money. On a dev account, where someone tried the sample on a Friday and then forgot about it, it's a line that keeps showing up every month until somebody runs the cleanup steps. That's the reason AWS's own docs suggest running the example in a development account.

The other storage lines (S3 buckets for skills and outputs, S3 Files mounts, CloudWatch logs) are normally cents at test volume. Since they scale with how much your agent writes, it's worth checking them on your AWS bill once you're in production.

What one agent task actually costs

This is the worked example I use. One task, where the agent reads some code and docs and then writes a result. That's 200,000 input tokens and 20,000 output tokens, with no caching, 10 minutes on a 2 vCPU / 4 GB AgentCore session where the CPU is busy 20% of the time.

  • Runtime: 120 busy seconds x 2 vCPU x $0.1276 per hour = $0.0085, plus 600 seconds x 4 GB x $0.0169 per hour = $0.0113. Call it $0.02, and less if v2 reclaims idle memory.
  • Tokens on GPT-5.6 Luna: 0.2M x $0.22 + 0.02M x $1.32 = $0.070.
  • Tokens on GPT-5.6 Sol: 0.2M x $4.40 + 0.02M x $22.00 = $1.32.
Hand-drawn bar chart of one agent task's cost: AgentCore runtime $0.02, GPT-6 Luna $0.03, GPT-5.6 Luna $0.07, GPT-6.1 Sol $0.66, GPT-5.6 Sol $1.32 and GPT-6 Astra $3.30, for 200K input and 20K output tokens
Hand-drawn bar chart of one agent task's cost: AgentCore runtime $0.02, GPT-6 Luna $0.03, GPT-5.6 Luna $0.07, GPT-6.1 Sol $0.66, GPT-5.6 Sol $1.32 and GPT-6 Astra $3.30, for 200K input and 20K output tokens

If you take one thing from this post, take that chart. The runtime is the smallest bar on the chart. Going from GPT-6 Luna to GPT-6 Astra on the same task is a 100x jump in token cost, while the AgentCore runtime stays at about two cents. If you're trying to cut a BMA bill, the model choice is the lever. Tuning vCPUs barely moves anything.

Here's how that scales across a month on GPT-5.6 Luna, with the NAT gateway left on:

Tasks per monthTokensAgentCore runtimeNAT gatewayMonthly totalPer task
1,000$70$20$33$123$0.123
10,000$704$198$33$935$0.093
100,000$7,040$1,977$33$9,050$0.091
Three hand-drawn stacked columns of monthly cost on GPT-5.6 Luna: $123 for 1,000 tasks with a large NAT gateway share, $935 for 10,000 tasks and $9,050 for 100,000 tasks where the NAT gateway is a hairline and tokens dominate
Three hand-drawn stacked columns of monthly cost on GPT-5.6 Luna: $123 for 1,000 tasks with a large NAT gateway share, $935 for 10,000 tasks and $9,050 for 100,000 tasks where the NAT gateway is a hairline and tokens dominate

At low volume, the fixed NAT gateway makes up over a quarter of the bill. By 100,000 tasks it's noise, and tokens are almost 80%. So the advice flips depending on scale: small teams should self-host or tear down the sample stack between tests, and big teams should spend their effort on model choice and prompt caching (making the repeated context cacheable).

You can plug your own numbers into the calculator below, which uses the same rates as the tables above:

Is the 10% AWS premium worth paying?

I'd expect this is the question most buyers are really asking. The answer is more favorable to AWS than the headline suggests. Here's the same 10-minute GPT-5.6 Luna task, priced four ways:

SetupTokensRuntimePer task
OpenAI Agents API, self-hosted sandbox$0.064 (OpenAI rate)$0$0.064
Bedrock Managed Agents, self-hosted$0.070 (in-Region rate)$0$0.070
Bedrock Managed Agents on AgentCore$0.070about $0.020$0.090
OpenAI Agents API, hosted 4 GB container$0.064$0.06$0.124
Four hand-drawn receipt strips pricing the same 10-minute GPT-5.6 Luna task: OpenAI Agents API self-hosted $0.064, Bedrock Managed Agents self-hosted $0.070, Bedrock Managed Agents on AgentCore $0.090 and OpenAI Agents API hosted container $0.124, with a bracket marking the 10% AWS premium of $0.006
Four hand-drawn receipt strips pricing the same 10-minute GPT-5.6 Luna task: OpenAI Agents API self-hosted $0.064, Bedrock Managed Agents self-hosted $0.070, Bedrock Managed Agents on AgentCore $0.090 and OpenAI Agents API hosted container $0.124, with a bracket marking the 10% AWS premium of $0.006

The hosted-container line uses OpenAI's published rate of $0.12 per 20-minute session for a 4 GB container, billed by the minute with a 5-minute minimum (OpenAI pricing). My Agents API pricing post goes through every container size.

So the AWS premium on this task is about $0.006. That's real money at a million tasks a month ($6,400 or so), but it's smaller than the gap between OpenAI's hosted container and AgentCore's I/O-aware billing. Say you'd otherwise run on OpenAI's hosted sandbox: then BMA on AgentCore can come out cheaper per task, not more expensive. If you'd self-host either way, BMA costs you exactly the 10%.

What the premium buys is the part that matters to a security review: the agent runtime, model inference and your tools all stay inside your AWS account, under AWS's contract. The procurement logic was put well by one Hacker News commenter:

Hacker News

"A lot of companies already have data processing agreements and compliance sign-off for using AWS. Many are hesitant to send their data to AI startups with an incentive to train their models and a history of being.... loose with how they intake training data. Even when they do give assurances otherwise. AWS is more trusted in this aspect. If this ends up similar to Claude on Bedrock, it's the same price."

It didn't turn out to be quite the same price, but it's close, 10% more for the in-Region guarantee. If your company already has AWS paperwork signed and would need months to approve a new AI vendor, that 10% is probably the cheapest compliance work you'll buy this year.

How BMA pricing compares to other managed agent runtimes

Every major lab sells a managed agent loop now, and each charges for the runtime in its own way. Tokens are always extra, at whatever that vendor's model rates are:

RuntimeRuntime feeBillable unitModels
Bedrock Managed Agents$0 in preview, plus your computen/aOpenAI on Bedrock
Claude Managed Agents$0.08 per session-hourTime in running statusClaude only
OpenAI Agents API$0.03 to $1.92 per 20-minute containerPer minute, 5-minute minimumOpenAI only
Gemini Managed AgentsCompute not billed in previewn/aGemini only
AgentCore harnessNo harness fee, plus AgentCore RuntimePer second CPU and memoryAny Bedrock, OpenAI, Gemini or LiteLLM-compatible model

Right now two of these are free at the runtime layer, and both of them are previews. That's the pattern to notice: launch pricing in this category is generous, and none of the labs has said what its preview will cost once it's generally available.

The AgentCore harness is the one AWS-native option that's already generally available, model-agnostic and has no separate harness fee. Unless you're committed to OpenAI's harness specifically, it gets you the same runtime billing and more model choice. Claude Managed Agents adds a runtime charge but ships more features today: memory stores, multiagent and vendor sandboxes. My Anthropic API pricing post has the token side of that comparison, and the OpenAI Agents API alternatives roundup covers the wider field.

There's also a strategic angle that one LinkedIn post put better than I could:

LinkedIn

"Near-zero switching costs between frontier models on the same bill sounds like a buyer's market. It is, for now. When you can swap Claude for GPT-5.5 with a one-line code change, models start looking interchangeable, and the platform hosting them all owns the customer relationship and the pricing power."

That's a reason to keep your agent code portable, not a reason to avoid BMA.

Hidden costs to budget for

The rate cards themselves are clear enough. What catches teams out is the stuff that isn't on any rate card:

  • The GA fee nobody has announced. AWS hasn't published a post-preview price or date. If you want a placeholder, Claude's $0.08 per session-hour would add about $0.013 to the 10-minute task above. Treat that as a budgeting scenario, not a prediction.
  • Context growth across turns. Each turn resends the conversation so far, so a 20-turn task can cost far more than 20 copies of the first turn. Caching helps on the models where it's supported, and the 272K long-context cliff hurts on the ones where it isn't.
  • The things BMA doesn't include yet. The preview has no built-in long-term memory, so AWS says to "provision and authorize any application-specific datastore separately" (preview limitations). If you add AgentCore Memory, it's $0.25 per 1,000 new events plus $0.75 per 1,000 stored long-term records a month (AgentCore pricing).
  • Human review. For actions with external effects, the security docs tell you to enforce authorization and any human-in-the-loop review yourself. That's engineering time, not an AWS line item.
  • Engineering time, full stop. It's the biggest line of all, and it never shows up on the AWS bill.

On that last point, it's worth hearing from AgentCore's own critics:

Hacker News

"There's not really a good solution, as AgentCore runtime sucks and is expensive. You basically have to build this yourself because nobody is solving for self-hosted managed infra for agents, and we don't really have the time to build this sort of system on top of building our actual product."

On the numbers above, I'd disagree that the runtime is expensive. The second half of that comment, though, is the real cost of any managed runtime: building the product is still on you.

Is Bedrock Managed Agents worth it for a support agent?

Here's where I see the math go wrong most often. Someone prices a support agent on BMA, sees about nine cents per conversation and decides it'll be dirt cheap. The token math is right, too. But it leaves out the helpdesk integration, retrieval over your knowledge base, escalation rules, testing against real tickets and the on-call engineer who maintains it.

I've watched this play out in both directions at eesel for years. One churned mid-market customer, who left after a broken integration and slow support, told the eesel team "long term we will just build our own, which is so possible now with AI." And the other direction: an engineering lead at a hardware company with a 300+ article knowledge base explained why they chose to buy:

"We could try to write our own LLM application but we didn't want to invest our time into that. We wanted something that we would not have to maintain."

Both of those are reasonable calls. If you have engineers who want to own an AI agent and you need OpenAI models inside AWS, BMA's pricing won't be the thing that stops you. If you're comparing BMA against hiring a ready-made agent, compare the full cost of building vs buying instead of the token bill alone. My guide to building support agents lists every piece you'd be signing up for.

eesel: the support agent with a fixed price

If what you're actually pricing is an AI helpdesk teammate, eesel takes the token meter out of the budget. It plugs into Zendesk, Freshdesk, Gorgias and the rest of your helpdesk and learns from your past tickets and help center. Before it answers a live customer, it gets run against hundreds of your historical tickets in simulation.

eesel's pricing page, showing the free and paid teammate plans, add-ons like priority support, SSO and data residency, and customer quotes

The pricing is a credit plan rather than a token bill: a free plan with 100 credits, then paid plans from $299 a month for 500 credits, where one ticket or chat handled is one credit, with every feature and unlimited seats. Finance can put that number in a spreadsheet without guessing at context growth, and on the customer support side, Gridwise saw 73% of its tier-1 requests resolved in the first month.

And if the reason you were looking at BMA is that you want to drive agents from a terminal or a script, the eesel CLI does that for the same teammate you see in the dashboard. You can connect integrations, edit its instructions, approve or deny pending actions and read its activity, with JSON output and --dry-run on writes. It can be driven by coding agents like Claude Code, Codex and Cursor too, which I covered in my AI agent CLI post.

Try eesel free and see what a support agent costs when the plumbing is already done.

Frequently Asked Questions

How much does Bedrock Managed Agents cost?
During the public preview, Bedrock Managed Agents itself costs $0. You pay for the OpenAI model tokens your agent uses at Bedrock's in-Region rates, plus any AWS resources it runs on, such as AgentCore Runtime hours and the NAT gateway in AWS's sample stack. My full Bedrock Managed Agents guide covers how the service works, and AWS says the pricing is subject to change at general availability.
Is Bedrock Managed Agents free during the preview?
The service fee is free, but running an agent isn't. Every turn calls an OpenAI model on Bedrock, which bills per token, and the AgentCore example stack keeps charging for its NAT gateway (about $33 a month) even when no agent is running. Self-hosting the exec server on a machine you already pay for removes the runtime and NAT lines.
Why does Bedrock Managed Agents pricing cost more than OpenAI's API?
The preview doesn't support cross-Region inference, so every request uses Bedrock's in-Region rate, which AWS says includes a 10% fee over OpenAI's own prices. On GPT-5.6 Luna that's $0.22 vs $0.20 per million input tokens. Compare it with my OpenAI API pricing breakdown to see the base rates.
What is the cheapest way to run Bedrock Managed Agents?
Pick the smallest model that does the job and self-host the exec server. GPT-6 Luna costs $0.11 per million input tokens and $0.55 per million output on the bedrock-mantle endpoint in us-east-1, and a self-hosted environment has no AgentCore or NAT gateway charges. Run the cleanup steps when you finish testing so idle resources stop billing.
How does Bedrock Managed Agents pricing compare to Claude Managed Agents?
Bedrock Managed Agents has no runtime fee during preview; Claude Managed Agents charges $0.08 per session-hour while a session is running, plus Claude token rates. The bigger difference is usually the model: per-task token cost swings far more between models than any runtime fee does.
Will Bedrock Managed Agents pricing change at general availability?
Possibly. AWS's preview announcement says pricing is subject to change at general availability and doesn't publish a GA fee or date. If you're budgeting now, model a runtime fee as a scenario, and note that AgentCore's committed-baseline discounts are launching separately.
Should I build a support agent on Bedrock Managed Agents to save money?
Only if you have engineers to build and maintain it. The token bill is often small; the cost is the helpdesk integration, retrieval, escalation and testing you write yourself. A ready-made AI helpdesk agent like eesel comes on a fixed monthly plan, from $299 for 500 tickets or chats, with those parts already built.

Share this article

Kurnia Kharisma

Article by

Kurnia Kharisma

Kurnia is a software engineer and writer at eesel AI with two years of SEO experience, writing about AI tools, helpdesk software, and customer support. He pairs a developer's understanding of how these products are built with search-driven research into what actually ranks and resonates with the people searching for them.

Related Posts

All posts →
Hand-drawn illustration of two developers at a laptop below a cloud holding a friendly robot, connected to an identity shield, a locked database and a chip, with the AWS logo on an orange circle
Trending

Amazon Bedrock Managed Agents explained: OpenAI's agent harness inside your AWS account

Bedrock Managed Agents runs OpenAI's agent harness on AWS while your tools stay on your own compute. How it works, what the preview leaves out, and what it costs.

Rama AdiRama AdiOct 1, 2026
Hand-drawn illustration of a storefront of business tool icons with a buyer handing over a coin, representing OpenAI Marketplace
Trending

OpenAI Marketplace: how it works, who's on it, and the fine print

OpenAI Marketplace lets enterprises spend part of their OpenAI commitment on 32 partner tools. Here's how the money flows, who's listed, and what's unpublished.

Kurnia KharismaKurnia KharismaOct 1, 2026
Hand-drawn illustration of a person thinking beside a smiling round agent sitting on a stack of price tags next to a usage gauge, on a violet background shape
Trending

OpenAI Dots pricing: what a dot really costs, plan by plan (2026)

OpenAI Dots pricing starts at $100/month on Pro, but a two-seat Business workspace can be the cheaper route. Every plan, the usage rules, and what's priced later.

Kurnia KharismaKurnia KharismaOct 1, 2026
Illustration of a stopwatch lifting away to reveal open runway, representing a lifted usage limit
Trending

OpenAI removed Codex's 5-hour limit: what actually changed

OpenAI temporarily removed the 5-hour usage limit on Codex and ChatGPT Work. Here is what changed on July 12, what stayed, and what it means for you.

Rama AdiRama AdiJul 20, 2026
Hand-drawn illustration of a person holding a green ChatGPT key card in front of a row of open doors, with a weekly usage meter in the corner
Trending

Sign in with ChatGPT: how it works, which apps support it, and the limits

Sign in with ChatGPT lets Plus and Pro users spend their plan inside 16 partner apps. How the two permissions work, the weekly caps, and what partners still charge.

Rama AdiRama AdiOct 1, 2026
Illustration of three stacked subscription cards with a lightning streak and a usage gauge, representing the ChatGPT Pro 500 tier
Trending

ChatGPT Pro 500: what $500 a month actually buys in 2026

ChatGPT Pro 500 is OpenAI's new $500/month tier with 25x Plus usage and Ultrafast. Here's the per-unit math, the 8x Ultrafast burn, and who it's really for.

Rama AdiRama AdiOct 1, 2026
Illustration of one dominant frontier AI model surrounded by a lineup of smaller alternative models
Trending

The 8 best GPT-6 Astra alternatives in 2026

GPT-6 Astra is a brilliant agent engine at 2.5x the price for a flat intelligence bump. Here are 8 GPT-6 Astra alternatives worth testing first.

Rama AdiRama AdiSep 9, 2026
Illustration announcing Claude Fable 5.1, Anthropic's newest frontier AI model
Trending

Claude Fable 5.1: pricing, capabilities, and what it means for your team

Claude Fable 5.1 is Anthropic's most capable model yet. Here's the real pricing, what changed from Fable 5, and where it fits for support and content teams.

KiraKiraSep 2, 2026
A developer at a terminal watching coins pour out of the screen into a basket marked with an infinity symbol
Trending

Meta Muse Code pricing: what a free coding agent really costs

Meta Muse Code has no price, no plan, and no spend cap. Here is the real rate card, the three defaults that set your bill, and how to bring it down.

Rama AdiRama AdiAug 18, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free