
What does Bedrock Managed Agents cost?
AWS's pricing line for the preview is one sentence: "During preview, there is no additional charge for BMA beyond the underlying AWS resources your agents consume. Pricing is subject to change at general availability" (AWS What's New).
Two years of doing SEO taught me one thing about pricing queries: the word "free" sends more people to the wrong conclusion than any other. Somebody who searches "bedrock managed agents pricing" wants the number their finance team is going to see, and that number is never $0. The developer docs spell out what "underlying resources" covers: "You incur charges for model inference and the AWS resources that your application uses," and the AgentCore example "can continue to incur charges when no BMA turn is running" (AWS docs).
If you're new to the service itself, my Bedrock Managed Agents guide covers how sessions, the exec server and execution environments fit together. This one is only about the bill.
Here's every line that can show up, in one table:
| Cost line | Preview price | Billable unit | When it applies |
|---|---|---|---|
| BMA service fee | $0 | n/a | Always (subject to change at GA) |
| OpenAI model tokens | In-Region rate, e.g. $0.22 in / $1.32 out for GPT-5.6 Luna | Per 1M input, cached and output tokens | Every turn |
| AgentCore Runtime v2 CPU | $0.1276 per vCPU-hour | Per second, active CPU only | If AgentCore is your execution environment |
| AgentCore Runtime v2 memory | $0.0169 per GB-hour | Per second, idle memory reclaimed after 120 seconds | If AgentCore is your execution environment |
| NAT gateway | $0.045 per hour + $0.045 per GB | Per hour (partial hours bill as full) | AWS's AgentCore sample stack |
| S3, S3 Files, CloudWatch | Standard AWS rates | Storage and requests | Skills, outputs and logs |
| Self-hosted compute | Whatever your host costs | Your existing bill | If you run codex exec-server yourself |
That's the whole shape of it. The lines don't all matter equally, though, so the rest of this post is about the ones that do.
Line 1: model tokens, and the 10% in-Region premium
Every OpenAI model on Bedrock lists its prices on its own model card. The detail that matters for BMA is in the preview limitations: cross-Region inference isn't supported. Which means every BMA request is paying the in-Region rate, and AWS says it plainly: "Commercial In-Region prices include a 10% fee over OpenAI rates" (GPT-5.6 Luna card).
Here are the Standard-tier, short-context rates (272K input tokens or fewer) per million tokens for the OpenAI models I'd consider for an agent, all from their Bedrock cards on AWS's OpenAI models page:
| Model | Input (BMA rate) | Cache write | Cache read | Output (BMA rate) | OpenAI's own rate (in / out) |
|---|---|---|---|---|---|
| GPT-6 Luna | $0.11 | $0.1375 | $0.011 | $0.55 | $0.10 / $0.50 |
| GPT-5.6 Luna | $0.22 | $0.275 | $0.022 | $1.32 | $0.20 / $1.20 |
| GPT-6.1 Sol | $2.20 | $2.75 | $0.11 | $11.00 | $2.00 / $10.00 |
| GPT-5.6 Terra | $2.20 | $2.75 | $0.22 | $13.20 | $2.00 / $12.00 |
| GPT-5.6 Sol | $4.40 | $5.50 | $0.44 | $22.00 | $4.00 / $20.00 |
| GPT-6 Astra | $11.00 | $13.75 | $1.10 | $55.00 | $10.00 / $50.00 |
Four details on those cards will change your estimate more than the table suggests:
- Long context doubles input. Once a request goes over 272,000 input tokens, the long-context rate applies "to the full request," not just the overflow (GPT-6.1 Sol card). On GPT-5.6 Sol that's $8.80 input and $33.00 output. An agent dragging a big context window through every turn will hit this quickly.
- No Priority or Flex discounts. Every card above says Priority and Flex tiers aren't supported, so there's no cheaper batch-style tier to drop into for overnight jobs. Standard is all you get.
- Caching isn't uniform. The GPT-6.1 Sol card lists cache prices but says "Explicit prompt caching is not supported for this Bedrock model." Check the card for the exact model before you count on the 90% cache-read discount.
- The newest models live in one Region. On the
bedrock-mantleendpoint BMA uses, GPT-6 Luna and GPT-6.1 Sol are available only inus-east-1. If you deploy in Oregon or Ohio, you're on the GPT-5.6 family.
The newest generation is also the cheapest per token here: GPT-6 Luna is half the price of GPT-5.6 Luna on both input and output. If a Luna was going to be your pick anyway, test GPT-6 Luna first. My GPT-6 Luna pricing and GPT-6.1 Sol pricing posts compare them against the rest of OpenAI's lineup.
Line 2: AgentCore Runtime hours
BMA doesn't run your tools itself; they run on compute that you provide: either your own host, or Amazon Bedrock AgentCore Runtime, which is the default in AWS's example. If you use AgentCore, you pay its runtime rates (AgentCore pricing):
| AgentCore Runtime resource | Consumption rate | Committed baseline |
|---|---|---|
| v2 microVM CPU | $0.1276 per vCPU-hour | $0.0997 per vCPU-hour (launching by October 2026) |
| v2 microVM memory | $0.0169 per GB-hour | $0.0132 per GB-hour (launching by October 2026) |
| v1 microVM CPU | $0.0895 per vCPU-hour | n/a |
| v1 microVM memory | $0.00945 per GB-hour | n/a |
| Runtime instances (EC2) | EC2 On-Demand rate | plus a 12% management fee |
This billing model is kinder to agents than it first looks. AWS says CPU "scales to zero during I/O wait (waiting for LLM responses, tool / API calls, or database queries)," and on v2 "idle memory is reclaimed automatically" after 120 seconds. Since an agent spends most of its life waiting on the model, most of its wall-clock time never bills for CPU at all.
AWS's own pricing page has a worked example that's close to a support agent: 1 million sessions a month, 10 minutes each, 90% I/O wait, 1 vCPU and up to 2.5 GB. Its total is $0.006703 per session, or $6,703 a month for the million. And that's runtime only, before you've paid for a single token.
One thing to watch: the BMA AgentCore example sets the runtime's idle timeout and maximum lifetime to 28,800 seconds, which is eight hours. Idle memory gets reclaimed, sure, but a session you forget about is still a session.
Line 3: the NAT gateway that bills while your agent sleeps
AWS's AgentCore example stack creates a VPC with private subnets, an S3 gateway endpoint and one NAT gateway. At Amazon VPC pricing in US East, a NAT gateway costs $0.045 per hour plus $0.045 per GB processed, and partial hours bill as full hours.
Across a 730-hour month that's $32.85 before your agent does anything. For a production team that's small money. On a dev account, where someone tried the sample on a Friday and then forgot about it, it's a line that keeps showing up every month until somebody runs the cleanup steps. That's the reason AWS's own docs suggest running the example in a development account.
The other storage lines (S3 buckets for skills and outputs, S3 Files mounts, CloudWatch logs) are normally cents at test volume. Since they scale with how much your agent writes, it's worth checking them on your AWS bill once you're in production.
What one agent task actually costs
This is the worked example I use. One task, where the agent reads some code and docs and then writes a result. That's 200,000 input tokens and 20,000 output tokens, with no caching, 10 minutes on a 2 vCPU / 4 GB AgentCore session where the CPU is busy 20% of the time.
- Runtime: 120 busy seconds x 2 vCPU x $0.1276 per hour = $0.0085, plus 600 seconds x 4 GB x $0.0169 per hour = $0.0113. Call it $0.02, and less if v2 reclaims idle memory.
- Tokens on GPT-5.6 Luna: 0.2M x $0.22 + 0.02M x $1.32 = $0.070.
- Tokens on GPT-5.6 Sol: 0.2M x $4.40 + 0.02M x $22.00 = $1.32.

If you take one thing from this post, take that chart. The runtime is the smallest bar on the chart. Going from GPT-6 Luna to GPT-6 Astra on the same task is a 100x jump in token cost, while the AgentCore runtime stays at about two cents. If you're trying to cut a BMA bill, the model choice is the lever. Tuning vCPUs barely moves anything.
Here's how that scales across a month on GPT-5.6 Luna, with the NAT gateway left on:
| Tasks per month | Tokens | AgentCore runtime | NAT gateway | Monthly total | Per task |
|---|---|---|---|---|---|
| 1,000 | $70 | $20 | $33 | $123 | $0.123 |
| 10,000 | $704 | $198 | $33 | $935 | $0.093 |
| 100,000 | $7,040 | $1,977 | $33 | $9,050 | $0.091 |

At low volume, the fixed NAT gateway makes up over a quarter of the bill. By 100,000 tasks it's noise, and tokens are almost 80%. So the advice flips depending on scale: small teams should self-host or tear down the sample stack between tests, and big teams should spend their effort on model choice and prompt caching (making the repeated context cacheable).
You can plug your own numbers into the calculator below, which uses the same rates as the tables above:
Is the 10% AWS premium worth paying?
I'd expect this is the question most buyers are really asking. The answer is more favorable to AWS than the headline suggests. Here's the same 10-minute GPT-5.6 Luna task, priced four ways:
| Setup | Tokens | Runtime | Per task |
|---|---|---|---|
| OpenAI Agents API, self-hosted sandbox | $0.064 (OpenAI rate) | $0 | $0.064 |
| Bedrock Managed Agents, self-hosted | $0.070 (in-Region rate) | $0 | $0.070 |
| Bedrock Managed Agents on AgentCore | $0.070 | about $0.020 | $0.090 |
| OpenAI Agents API, hosted 4 GB container | $0.064 | $0.06 | $0.124 |

The hosted-container line uses OpenAI's published rate of $0.12 per 20-minute session for a 4 GB container, billed by the minute with a 5-minute minimum (OpenAI pricing). My Agents API pricing post goes through every container size.
So the AWS premium on this task is about $0.006. That's real money at a million tasks a month ($6,400 or so), but it's smaller than the gap between OpenAI's hosted container and AgentCore's I/O-aware billing. Say you'd otherwise run on OpenAI's hosted sandbox: then BMA on AgentCore can come out cheaper per task, not more expensive. If you'd self-host either way, BMA costs you exactly the 10%.
What the premium buys is the part that matters to a security review: the agent runtime, model inference and your tools all stay inside your AWS account, under AWS's contract. The procurement logic was put well by one Hacker News commenter:
"A lot of companies already have data processing agreements and compliance sign-off for using AWS. Many are hesitant to send their data to AI startups with an incentive to train their models and a history of being.... loose with how they intake training data. Even when they do give assurances otherwise. AWS is more trusted in this aspect. If this ends up similar to Claude on Bedrock, it's the same price."
It didn't turn out to be quite the same price, but it's close, 10% more for the in-Region guarantee. If your company already has AWS paperwork signed and would need months to approve a new AI vendor, that 10% is probably the cheapest compliance work you'll buy this year.
How BMA pricing compares to other managed agent runtimes
Every major lab sells a managed agent loop now, and each charges for the runtime in its own way. Tokens are always extra, at whatever that vendor's model rates are:
| Runtime | Runtime fee | Billable unit | Models |
|---|---|---|---|
| Bedrock Managed Agents | $0 in preview, plus your compute | n/a | OpenAI on Bedrock |
| Claude Managed Agents | $0.08 per session-hour | Time in running status | Claude only |
| OpenAI Agents API | $0.03 to $1.92 per 20-minute container | Per minute, 5-minute minimum | OpenAI only |
| Gemini Managed Agents | Compute not billed in preview | n/a | Gemini only |
| AgentCore harness | No harness fee, plus AgentCore Runtime | Per second CPU and memory | Any Bedrock, OpenAI, Gemini or LiteLLM-compatible model |
Right now two of these are free at the runtime layer, and both of them are previews. That's the pattern to notice: launch pricing in this category is generous, and none of the labs has said what its preview will cost once it's generally available.
The AgentCore harness is the one AWS-native option that's already generally available, model-agnostic and has no separate harness fee. Unless you're committed to OpenAI's harness specifically, it gets you the same runtime billing and more model choice. Claude Managed Agents adds a runtime charge but ships more features today: memory stores, multiagent and vendor sandboxes. My Anthropic API pricing post has the token side of that comparison, and the OpenAI Agents API alternatives roundup covers the wider field.
There's also a strategic angle that one LinkedIn post put better than I could:
"Near-zero switching costs between frontier models on the same bill sounds like a buyer's market. It is, for now. When you can swap Claude for GPT-5.5 with a one-line code change, models start looking interchangeable, and the platform hosting them all owns the customer relationship and the pricing power."
That's a reason to keep your agent code portable, not a reason to avoid BMA.
Hidden costs to budget for
The rate cards themselves are clear enough. What catches teams out is the stuff that isn't on any rate card:
- The GA fee nobody has announced. AWS hasn't published a post-preview price or date. If you want a placeholder, Claude's $0.08 per session-hour would add about $0.013 to the 10-minute task above. Treat that as a budgeting scenario, not a prediction.
- Context growth across turns. Each turn resends the conversation so far, so a 20-turn task can cost far more than 20 copies of the first turn. Caching helps on the models where it's supported, and the 272K long-context cliff hurts on the ones where it isn't.
- The things BMA doesn't include yet. The preview has no built-in long-term memory, so AWS says to "provision and authorize any application-specific datastore separately" (preview limitations). If you add AgentCore Memory, it's $0.25 per 1,000 new events plus $0.75 per 1,000 stored long-term records a month (AgentCore pricing).
- Human review. For actions with external effects, the security docs tell you to enforce authorization and any human-in-the-loop review yourself. That's engineering time, not an AWS line item.
- Engineering time, full stop. It's the biggest line of all, and it never shows up on the AWS bill.
On that last point, it's worth hearing from AgentCore's own critics:
"There's not really a good solution, as AgentCore runtime sucks and is expensive. You basically have to build this yourself because nobody is solving for self-hosted managed infra for agents, and we don't really have the time to build this sort of system on top of building our actual product."
On the numbers above, I'd disagree that the runtime is expensive. The second half of that comment, though, is the real cost of any managed runtime: building the product is still on you.
Is Bedrock Managed Agents worth it for a support agent?
Here's where I see the math go wrong most often. Someone prices a support agent on BMA, sees about nine cents per conversation and decides it'll be dirt cheap. The token math is right, too. But it leaves out the helpdesk integration, retrieval over your knowledge base, escalation rules, testing against real tickets and the on-call engineer who maintains it.
I've watched this play out in both directions at eesel for years. One churned mid-market customer, who left after a broken integration and slow support, told the eesel team "long term we will just build our own, which is so possible now with AI." And the other direction: an engineering lead at a hardware company with a 300+ article knowledge base explained why they chose to buy:
"We could try to write our own LLM application but we didn't want to invest our time into that. We wanted something that we would not have to maintain."
Both of those are reasonable calls. If you have engineers who want to own an AI agent and you need OpenAI models inside AWS, BMA's pricing won't be the thing that stops you. If you're comparing BMA against hiring a ready-made agent, compare the full cost of building vs buying instead of the token bill alone. My guide to building support agents lists every piece you'd be signing up for.
eesel: the support agent with a fixed price
If what you're actually pricing is an AI helpdesk teammate, eesel takes the token meter out of the budget. It plugs into Zendesk, Freshdesk, Gorgias and the rest of your helpdesk and learns from your past tickets and help center. Before it answers a live customer, it gets run against hundreds of your historical tickets in simulation.
The pricing is a credit plan rather than a token bill: a free plan with 100 credits, then paid plans from $299 a month for 500 credits, where one ticket or chat handled is one credit, with every feature and unlimited seats. Finance can put that number in a spreadsheet without guessing at context growth, and on the customer support side, Gridwise saw 73% of its tier-1 requests resolved in the first month.
And if the reason you were looking at BMA is that you want to drive agents from a terminal or a script, the eesel CLI does that for the same teammate you see in the dashboard. You can connect integrations, edit its instructions, approve or deny pending actions and read its activity, with JSON output and --dry-run on writes. It can be driven by coding agents like Claude Code, Codex and Cursor too, which I covered in my AI agent CLI post.
Try eesel free and see what a support agent costs when the plumbing is already done.
Frequently Asked Questions
How much does Bedrock Managed Agents cost?
Is Bedrock Managed Agents free during the preview?
Why does Bedrock Managed Agents pricing cost more than OpenAI's API?
What is the cheapest way to run Bedrock Managed Agents?
How does Bedrock Managed Agents pricing compare to Claude Managed Agents?
Will Bedrock Managed Agents pricing change at general availability?
Should I build a support agent on Bedrock Managed Agents to save money?

Article by
Kurnia Kharisma
Kurnia is a software engineer and writer at eesel AI with two years of SEO experience, writing about AI tools, helpdesk software, and customer support. He pairs a developer's understanding of how these products are built with search-driven research into what actually ranks and resonates with the people searching for them.








