
What is Amazon Bedrock Managed Agents?
Amazon Bedrock Managed Agents, powered by OpenAI, is a managed runtime for stateful AI agents on AWS. You create a session with a model, instructions, tools and an execution environment, send it messages, and the service runs the agent loop and the model calls. In AWS's words, it's "built on a customized version of OpenAI's Agents API engineered to be AWS-native" (AWS What's New), and the product page says it "combines OpenAI models with the Codex harness and Amazon Bedrock AgentCore" (AWS).
I ship eesel's integrations and APIs, so I've wired up this exact kind of loop by hand more than once. And the first thing to know is one the marketing page kind of buries: "managed" here covers the conversation and the reasoning loop, not the machine your agent works on. That machine, you run yourself.
It took a while to get here, and here's the timeline, all from primary sources:
| Date | What happened | Source |
|---|---|---|
| Feb 27, 2026 | Amazon and OpenAI announce a partnership, including a "Stateful Runtime Environment" on Bedrock | Amazon |
| Apr 28, 2026 | BMA announced in limited preview, alongside OpenAI models and Codex on Bedrock | OpenAI |
| Jun 1, 2026 | OpenAI models and Codex on Bedrock go GA; BMA still "coming soon" | AWS ML blog |
| Sep 29, 2026 | BMA opens as a public preview in 3 US Regions | AWS What's New |
The named launch customers are Box and Salesforce. In Amazon's launch post, Box CTO Ben Kus says it brings together OpenAI's models "with the scale, security, and infrastructure of AWS," and Salesforce pairs it with its Headless 360 product on the AWS product page.
If you've read my guide to the OpenAI Agents API, the shape will feel familiar: same harness lineage as OpenAI Codex (its pricing is separate), same agent, session and events objects. The difference is mostly about who hosts the loop and who signs the requests.
How Bedrock Managed Agents works
AWS's docs break the service into six parts (AWS docs):
- Session: a stateful conversation with an agent. It names a model, instructions, tools, an IAM role and an execution environment.
- Turn: the work done in response to one message: reasoning, tool calls and output.
- Execution environment: compute you provide, where commands and local tools actually run.
- Exec server: the
codex exec-serverprocess that connects your environment to BMA over an outbound connection. - Items and events: items are the durable record of what happened; events are a live stream of progress.
- Session role: the IAM role BMA assumes to call the model (and, if you use it, to start AgentCore Runtime).
AWS has its own diagram of how those pieces talk to each other:

What surprised me was the box on the bottom right. On Claude Managed Agents or OpenAI's hosted Agents API, you can let the vendor run a sandbox for you. On BMA, there's no vendor sandbox option at all. Tools "execute in the compute environment that you provide" (AWS docs), so the exec server, the workspace, the network rules and any MCP servers are on you.

So that's the trade, the same one I keep seeing in every best AI agents comparison. You give up convenience and get a clean security story: model inference, the agent runtime and your tools all stay inside AWS, and the AWS product page says exactly that: "The agent runtime and model inference remain inside AWS" (AWS).
The API is small
The preview REST surface is seven operations on the regional bedrock-mantle endpoint, all signed with AWS SigV4 rather than an OpenAI API key (API reference): create, list, retrieve and delete a session, submit events, stream events, and list items. You send a message by posting an agent.session.input.message event, then read results from the items endpoint, 1 to 100 per page (sessions guide).
There are two details in the docs that will save you an afternoon:
- A successful submit returns an empty body. It means "accepted," not "done." Completion shows up in session state, items and events.
- An
idlesession isn't proof the task worked. AWS says to check the output items and, for commands, the exit code. Cancelling a turn also "does not undo side effects from tools that already completed."
Skills and tools
A skill is a folder with a SKILL.md file in it, placed under one of the session's capability directories (up to 32 paths). Tools come from STDIO MCP servers that run inside your environment (my AI agent MCP server guide explains the pattern), and you can restrict each server with an allowed_tools list (skills and tools). If you've used skills in Claude Code, the format is close enough to feel at home.
AWS is blunt about the security side: "Treat instructions from documents, websites, and tool responses as untrusted input" (security docs). That's prompt injection advice, and it hits a bit differently once the agent has a real shell on a real host.
Where your agent actually runs: self-hosted or AgentCore
You get two execution environments, and the choice decides most of your setup work and part of your bill (AWS docs):
| Self-hosted compute | AgentCore Runtime | |
|---|---|---|
| What you provide | A host, a workspace, network access, and a running codex exec-server | An AgentCore Runtime with the exec server and its adapter baked into an ARM64 container |
| What AWS's example stack creates | IAM roles only (no host) | Runtime, VPC with private subnets, 1 NAT gateway, versioned S3 buckets for skills and outputs, S3 Files mounts |
| Who starts the exec server | You, in a second terminal | BMA activates the Runtime for you |
| Example time limits | You decide | 28,800 seconds (8 hours) idle and max lifetime |
| Good for | Dev machines, existing containers, quick tests | Managed per-session compute that stays in your account |
AgentCore is AWS's broader agent platform, and it's the default compute for BMA. AWS's pitch is that as your agents grow, you can lean on more of it: authorization, agent and tool discovery, observability and evaluation (AWS).
If you just want to see something work, the self-hosted path is fastest. The AgentCore path is the one you'd actually ship, and it's also where the costs that don't show up on the "no additional charge" line start to stack up.
What the preview includes, and what it leaves out
If you only read one section twice before committing, make it this one. AWS's preview limitations page is unusually clear about the boundaries:

- Input is text only. The documented session input surface is text.
- No subagents, no code mode. AWS says not to enable them in a session config at all, so subagent orchestration patterns are off the table for now.
- No long-term memory. "The preview does not provide a built-in long-term memory integration; provision and authorize any application-specific datastore separately."
- No cross-Region inference. Which matters for price, as you'll see below.
- No customer-managed KMS key for service-managed session data.
- No console. Everything goes through the API and AWS's sample bundle.
- No dedicated turn-list API. You correlate turns through item
turn_idvalues.
There's a gap here that's easy to miss. The launch messaging says the managed runtime "handles inference, memory, and skills" (AWS) and that each agent "supports human approval before consequential actions," per the same preview announcement. On both, the developer docs say something narrower. Memory means context inside one session, and for actions with external effects the security guide says to "enforce authorization and any required human review in the application or tool implementation" (security docs). So the human-in-the-loop gate is something you build, not something you switch on.
For a preview none of this is unusual. What it does mean is that the version you can use today is closer to "a managed loop plus a good sample repo" than to a finished agent platform. OpenAI's own BMA page adds a warning worth taking literally: "Shared concepts don't imply identical API contracts or feature availability." Don't copy an OpenAI Agents API request and expect it to work on Bedrock.
How much Bedrock Managed Agents costs
AWS's pricing line is short: "During preview, there is no additional charge for BMA beyond the underlying AWS resources your agents consume. Pricing is subject to change at general availability" (AWS What's New). The docs spell out what "underlying resources" means: model inference plus whatever AWS resources your app uses, and the AgentCore example "can continue to incur charges when no BMA turn is running" (AWS docs).

Line 1: model tokens, at the in-Region rate
Every OpenAI model on Bedrock has its price on its model card. The catch is that BMA's preview doesn't support cross-Region inference profiles, so you pay the in-Region rate, and AWS notes that "Commercial In-Region prices include a 10% fee over OpenAI rates" (GPT-5.6 Luna card). Here are the short-context rates (272K input tokens or fewer) per million tokens:
| Model | Input (in-Region) | Cached input | Output (in-Region) | Global rate, same as OpenAI's (in / out) |
|---|---|---|---|---|
| GPT-5.6 Luna (BMA example default) | $0.22 | $0.022 | $1.32 | $0.20 / $1.20 |
| GPT-6 Luna | $0.11 | $0.011 | $0.55 | $0.10 / $0.50 |
| GPT-6.1 Sol | $2.20 | $0.11 | $11.00 | $2.00 / $10.00 |
| GPT-5.6 Terra | $2.20 | $0.22 | $13.20 | $2.00 / $12.00 |
| GPT-5.6 Sol | $4.40 | $0.44 | $22.00 | $4.00 / $20.00 |
| GPT-6 Astra | $11.00 | $1.10 | $55.00 | $10.00 / $50.00 |
Every rate comes from that model's Bedrock card, all linked from AWS's OpenAI models page. AWS doesn't publish a fixed list of BMA-supported models, and the GPT-6.1 Sol card notes that explicit prompt caching isn't supported for that model on Bedrock, so check your model before you count on cache discounts. Long-context requests (over 272K input, a fraction of the 1M-plus context window) cost more again. GPT-5.6 Sol jumps to $8.80 input and $33.00 output.
Line 2: runtime hours
If you use AgentCore Runtime, you pay its compute rates. On v2 microVMs that's $0.1276 per vCPU-hour and $0.0169 per GB-hour, billed per second, and AWS says CPU isn't billed while the agent waits on I/O such as the model's reply (AgentCore pricing). Self-hosted, you just pay whatever your host already costs.
Line 3: the NAT gateway nobody mentions
AWS's AgentCore sample stack creates one NAT gateway. At the published US East rate that's $0.045 per hour plus $0.045 per GB processed, and partial hours bill as full hours (Amazon VPC pricing). Left running for a month (730 hours), that's about $33 a month before your agent does a single thing. For an enterprise that's small money. Still, it's exactly the kind of line that shows up on a dev account three months after someone forgot to run the cleanup steps.
A worked example
Say one agent task reads a codebase and some docs: 200,000 input tokens, 20,000 output tokens, about 10 minutes on a 2 vCPU / 4 GB AgentCore session.
- On GPT-5.6 Luna: $0.044 input + $0.026 output = about $0.07 in tokens.
- On GPT-5.6 Sol: $0.88 input + $0.44 output = about $1.32 in tokens.
- Runtime ceiling: 2 vCPU x $0.1276 + 4 GB x $0.0169 = $0.32 per hour if the CPU were busy the whole time, so roughly $0.05 for 10 minutes, and less in practice because waiting on the model isn't billed for CPU.
The takeaway: model choice moves the bill about 19x; the runtime barely moves it. At low volume it's the fixed costs (NAT gateway, storage) you'll notice, and at high volume it's the tokens. If you want the same math for OpenAI's hosted version, my Agents API pricing post covers its per-20-minute container rates, and my AWS pricing guide covers the rest of the AWS bill.
How it compares to other managed agent runtimes
Every major lab now sells some version of "we'll run the agent loop for you." Where they really differ is where two things live: the loop and the tools.

Claude Managed Agents and OpenAI's Agents API also offer self-hosted sandboxes, so the top-left cluster is their default rather than their only mode. Even then, though, the orchestration stays with the lab: on Anthropic's version, tool inputs and outputs still flow to Anthropic's control plane, and on OpenAI's, choosing a self-hosted sandbox only moves tool execution. BMA is the only option here that keeps OpenAI's harness, OpenAI's models and your tools all inside AWS.
| Bedrock Managed Agents | Claude Managed Agents | OpenAI Agents API | Gemini Managed Agents | AgentCore harness | |
|---|---|---|---|---|---|
| Status | Public preview, 3 US Regions | Beta | Public beta (Sep 10, 2026) | Public preview (May 19, 2026) | Generally available |
| Models | OpenAI on Bedrock | Claude only | OpenAI only | Gemini only | Bedrock, OpenAI, Gemini, any LiteLLM provider |
| Where tools run | Your host or AgentCore only | Anthropic sandbox or yours | OpenAI sandbox, yours, partner providers, or none | Google sandbox (4 CPU / 16 GB) | microVM per session in your account |
| Runtime fee | None during preview | $0.08 per session-hour while running | $0.03 to $1.92 per 20-min container; none self-hosted | Compute not billed in preview | $0.1276 per vCPU-hour + $0.0169 per GB-hour |
| Long-term memory | No (preview) | Memory stores | Not published | No (files kept 7 days) | AgentCore Memory |
| MCP | STDIO inside your environment | Remote servers + tunnels | Remote or local | Remote HTTP | Via AgentCore Gateway |
| Subagents | No | Yes (multiagent) | Yes | Not published | Not published |
A couple of things in that table are worth a second look.
The AgentCore harness is BMA's closest rival, and it's also from AWS. It's GA, it supports "any model provided by Amazon Bedrock, OpenAI, Google Gemini, or any LiteLLM-compatible provider," and there's no separate harness charge (AgentCore harness). AWS has also moved the original Bedrock Agents into maintenance mode: it was renamed Bedrock Agents Classic and closed to new customers on July 30, 2026, with AWS pointing new builds at the harness (AWS docs). So unless you're committed to OpenAI's harness specifically, I'd start there.
Claude Managed Agents is the more complete product today. It has memory stores, multiagent orchestration, vendor sandboxes and remote MCP, and it's also reachable through Anthropic's Claude Platform on AWS. On that route, though, it's Anthropic and not AWS that processes the data. If your requirement is "AWS is the only processor," BMA (or Claude in Amazon Bedrock, which I covered in Claude Code on Bedrock) is the cleaner answer. For a wider list, see my OpenAI Agents API alternatives roundup.
How to get started with Bedrock Managed Agents
The fastest route is AWS's self-hosted example. If your AWS CLI is already set up, plan on an hour. AWS's self-hosted tutorial has the full commands:
- Install the tools. Node.js 20+, AWS CLI v2,
curlwith SigV4 support,jq, and Codex CLI 0.154.0 or later (it includescodex exec-server). - Pick a Region and endpoint.
us-east-1,us-west-2orus-east-2, with the endpointhttps://bedrock-mantle.<region>.api.aws. - Download the example bundle and the Codex binary that matches your host (always Linux ARM64 for AgentCore).
- Deploy the IAM roles with
npm ci,npx cdk bootstrapandnpm run deployfrom theself-hosted/folder. You get a client role and a session role that BMA assumes. - Create a session, attach the exec server, submit a turn. The bundle's numbered scripts (
0.create-session.sh,1.attach-exec-server.sh,2.submit-turn.sh,3.read-result.sh) walk you through it.
Most first runs break on the IAM model. There are three identities: the caller (needs BMA permissions plus iam:PassRole on the session role), the session role (trusted by bedrock-mantle.amazonaws.com and allowed bedrock-mantle:CreateInference for your model), and the identity your execution environment uses (security docs). The troubleshooting page is worth bookmarking before you start, especially the bit about checking AWS_PROFILE in each terminal. Also, use a dedicated workspace, because "the agent can use the files, tools, and permissions available to that environment."
What builders are saying
Public preview only opened this week, so hands-on feedback is still thin. What exists is launch reactions from people who live in AWS, and they mostly agree on the appeal:
"Bedrock Managed Agents (limited preview): AWS runs OpenAI's agent harness, and all inferences run through Bedrock. Basically, you'd use the OpenAI SDK against AWS-owned infrastructure, and your data stays in AWS.
AgentCore Runtime is the only one that's GA right now, so it's the only viable option if you need something for production."
The compliance angle comes up again and again in the Hacker News thread on the original launch:
"This would be a nice compliance win. One less sub-processor and all our data is already on AWS so less worrying about sending it off somewhere else"
The skeptics are worth hearing too. Analyst Mitch Ashley put the lock-in question plainly:
"The question for enterprise architects is whether AgentCore stays open enough to govern non-AWS execution, or quietly becomes the lock-in seam."
And one Hacker News commenter, talking about the compute layer BMA uses by default, didn't hold back:
"There's not really a good solution, as AgentCore runtime sucks and is expensive. You basically have to build this yourself because nobody is solving for self-hosted managed infra for agents, and we don't really have the time to build this sort of system on top of building our actual product."
That last line, "on top of building our actual product," is the whole story of managed agent runtimes. They shrink the infrastructure job, but the product job is still yours.
Who should use Bedrock Managed Agents (and who shouldn't)
Use it if you're already deep in AWS, your security or procurement team has signed off on AWS but not on a new AI vendor, and you specifically want OpenAI's harness and models. The IAM-per-agent model, CloudTrail logging and "nothing leaves the account" story are the real value, and they're hard to get anywhere else with OpenAI models.
Wait if you need subagents, long-term memory, image input, or anything outside three US Regions. Those aren't in the preview, and there's no GA date.
Skip it if you aren't tied to OpenAI models (the AgentCore harness is GA and model-agnostic), or if you want the vendor to run the sandbox too (Claude Managed Agents or OpenAI's hosted Agents API).
And think hard if your actual goal is a business agent, like one that answers support tickets. A support agent is a product, not a runtime. I run into this pattern a lot. In eesel's own churn notes, several customers, including an AR/construction-tech firm and a DTC beauty brand, left to build their support agent directly on an LLM API. Others went the opposite way. An engineering lead at a Bitcoin-ATM hardware company with a 300+ article Confluence knowledge base told the eesel team why they chose to buy:
"We could try to write our own LLM application but we didn't want to invest our time into that. We wanted something that we would not have to maintain."
BMA makes the "build" path shorter. It doesn't remove the helpdesk integration, the knowledge retrieval, the escalation logic, or the testing. And those are the parts that take months.
eesel for teams who want the agent, not the plumbing
Bedrock Managed Agents is infrastructure. eesel is the employee. Specifically, eesel is an AI teammate platform where you hire ready-to-work teammates for defined jobs, and for support that's the AI helpdesk teammate: it plugs into Zendesk, Freshdesk, Gorgias and the rest of your helpdesk in minutes, learns from your past tickets and help center, and gets run against hundreds of your historical tickets in a simulation before it ever replies to a live customer. On eesel's own customer stats, Gridwise saw 73% of its tier-1 tickets resolved in the first month.
If the reason you were looking at BMA is that you like driving agents from a terminal, eesel has that too. The eesel CLI (@eesel/cli) operates the same teammate and workspace you see in the dashboard: connect integrations, edit the agent's standing instructions, approve or deny pending actions with eesel approvals, and read every run with eesel activity. Every command prints JSON, writes support --dry-run so you can see the exact call before it's sent, and headless auth works in CI. Coding agents like Claude Code, Codex and Cursor can drive it, and every workspace also exposes an MCP server. It's the same idea as BMA's API-first design, pointed at a finished support agent instead of an empty loop.
Try eesel free. The pricing starts with a free plan of 100 credits, then paid plans from $299 a month for 500, and one ticket or chat handled is one credit.
Frequently Asked Questions
What is Amazon Bedrock Managed Agents?
How much does Bedrock Managed Agents cost?
Is Bedrock Managed Agents generally available?
What's the difference between Bedrock Managed Agents and the OpenAI Agents API?
Does Bedrock Managed Agents have memory?
Which models does Bedrock Managed Agents support?
Should I use Bedrock Managed Agents for customer support?

Article by
Rama Adi
Rama is a software engineer at eesel AI with two years of experience writing about B2B SaaS, AI tools, and customer support technology. Based in Bali, Indonesia, he brings a developer's perspective to product comparisons — cutting through marketing copy to what the integrations and APIs actually do.








