Codex Security Cloud explained: how OpenAI's security agent works

Kira
Written by

Kira

Katelin Teen
Reviewed by

Katelin Teen

Last edited October 1, 2026

Expert Verified
Hand-drawn illustration of a developer at a laptop connected to a smiling cloud that links a repository, a shielded sandbox box, and a findings list, with a second person looking on

What is Codex Security Cloud?

Codex Security is OpenAI's application security agent inside OpenAI Codex. It finds and confirms vulnerabilities, and also proposes the fix. Codex Security Cloud is the version that runs in Codex cloud against your connected GitHub repositories. You install it as a plugin and point it at a repo, then it keeps working whether your machine is on or not.

OpenAI's DevDay recap puts it plainly: Codex "investigates findings, removes duplicates and prepares fixes in the cloud, even with your laptop closed." It's available in research preview on the web and in the Codex app.

It was one of several DevDay launches, alongside always-on OpenAI Dots and the GPT-6.1 Sol model.

Codex Security Cloud findings dashboard with tabs for Overview, Findings, Repositories, Scans and Configure, a pipeline from Discovered to Done, and a Needs attention table with Fix buttons, as taken from OpenAI's DevDay recap
Codex Security Cloud findings dashboard with tabs for Overview, Findings, Repositories, Scans and Configure, a pipeline from Discovered to Done, and a Needs attention table with Fix buttons, as taken from OpenAI's DevDay recap

This didn't appear out of nowhere. The product has a longer history, and it is worth knowing because the numbers OpenAI quotes come from the earlier stages:

DateWhat shippedWho could use it
Oct 30, 2025Aardvark, "an agentic security researcher powered by GPT-5"Private beta, select partners
Mar 6, 2026Renamed Codex Security, research preview in Codex webPro, Enterprise, Business, Edu, free for the first month
Jul 2026Public CLI and TypeScript SDK on GitHub, Apache-2.0Anyone can install; scans need Codex Security access
Sep 29, 2026Codex Security Cloud plugin, Daybreak Blue includedPro, Business, Enterprise, Edu

The track record OpenAI points to is real. Aardvark found 92% of known and synthetically introduced vulnerabilities in its "golden" test repos. In the 30 days before the March launch, the beta cohort's scans covered more than 1.2 million commits and turned up 792 critical and 10,561 high-severity findings. OpenAI also says false positive rates fell by more than 50% across all repositories, and 14 CVEs have been assigned from its open-source work, including reports to OpenSSH and GnuTLS, and also Chromium.

How Codex Security Cloud works

The pipeline has four stages, and the Cloud FAQ lays them out in order. I've built agent pipelines like this at eesel, so for me the interesting part is the order, where validation sits before anything reaches a human.

Hand-drawn pipeline of five cards: threat model, scan repo or commits, reproduce in sandbox, proposed patch, and you open draft PR, with the first four bracketed as running in the cloud
Hand-drawn pipeline of five cards: threat model, scan repo or commits, reproduce in sandbox, proposed patch, and you open draft PR, with the first four bracketed as running in the cloud
  1. Analysis. Codex reads the repo and writes a threat model: entry points, trust boundaries, auth assumptions, risky components.
  2. Scanning. A Repository scan reviews everything once. Commit changes monitors new commits and can look back over existing history.
  3. Validation. For each likely issue, it tries to reproduce the problem in a clean container, running commands or tests and attaching logs as evidence. Findings that reproduce get marked validated. Ones that don't stay unvalidated, with the attempt still logged.
  4. Remediation. You get guidance and, when one can be generated, a "minimal actionable diff" with file and line context.

A few behaviors are worth knowing before you trust the output. It's language-agnostic, though OpenAI notes that quality depends on how well the model reasons about your language and framework. It doesn't need a build step to find issues, but it may try to build inside the container to reproduce one. And every job runs in an "ephemeral Codex container with session-scoped tools" which is torn down when the job ends.

Most teams will underuse the threat model. Codex drafts it from your code, and then it guides every future commit scan and how findings get ranked. OpenAI's threat model guide says to edit it when your architecture changes or "when findings miss the areas you care about." In practice I'd edit it on day one, because the model knows your code but doesn't know that the billing service is the thing your auditors ask about.

Four ways to run Codex Security, and which one is Cloud

People get confused at this point, so here is the map. OpenAI ships the same scanner through four surfaces, and only one of them is called "Cloud."

Hand-drawn 2x2 grid of Codex Security surfaces: Security Cloud plugin for GitHub repos in Codex cloud, Security plugin for a local repo in the desktop app, CLI and SDK for terminal and CI jobs, and Security Review for PR comments, all joined by a center label reading one scanner
Hand-drawn 2x2 grid of Codex Security surfaces: Security Cloud plugin for GitHub repos in Codex cloud, Security plugin for a local repo in the desktop app, CLI and SDK for terminal and CI jobs, and Security Review for PR comments, all joined by a center label reading one scanner
SurfaceWhere it runsBest for
Codex Security Cloud pluginCodex cloud, on connected GitHub reposAlways-on repo scans and commit monitoring
Codex Security pluginA Codex task on your machine, via the desktop Security workbenchScanning a local repo or one folder, deep scans
CLI and TypeScript SDKYour terminal or CI, npx @openai/codex-securityBulk scans across many repos, CI gates, SARIF export
Codex Security ReviewGitHub pull requestsA security-focused pass on each PR, @codex security review

The Security Review piece pairs well with Cloud. It goes deeper than Codex's general code review on security risks, and you can trigger it when a PR opens or on every push, or alongside code review. By default, automatic reviews post only High and Critical findings to the PR. If you already use Codex's GitHub integration, that's a settings toggle, not a new tool.

The CLI is worth a look even if you plan to live in Cloud. It's where the plumbing shows. You get --max-cost for an estimated spend limit and --fail-on-severity high for CI, and also findings false-positive to record why a finding doesn't apply. That note gets fed back as context to future scans, though it doesn't suppress the rule. It's the same reason why I like a real CLI on any agent, since the AI agent CLI pattern makes the agent scriptable instead of a dashboard you have to babysit.

How to set up Codex Security Cloud

Setup takes five steps per OpenAI's setup guide, and you need Codex cloud to be configured already for your workspace.

Scrolling capture of OpenAI's Codex Security Cloud setup documentation, from installing the plugin through monitoring new commits
  1. Install the plugin. Open Plugins in ChatGPT on the web or desktop, search for Codex Security Cloud, install and enable it, then open Security Cloud.
  2. Connect GitHub. Select New scan, then Connect GitHub if prompted, and grant access to the repos you want scanned.
  3. Start a repository scan. Choose the repo, pick a compatible Cloud environment (or create one), leave What to scan on Repository, and select Start scan.
  4. Review findings. Open Findings to see affected code, validation evidence and remediation guidance. Where you see Fix with Codex, generate a patch, review it, then select Create draft pull request.
  5. Turn on commit monitoring. Start a New scan, choose Commit changes, and select Create. Under Monitoring settings you can change the environment, set how many days of history to review, pause monitoring, and edit the threat model under Project context.

If the plugin doesn't show up at all, the docs give one answer only: "check with your workspace administrator." That leads into the messiest part of the launch.

Who can use Codex Security Cloud?

On paper, it's four plans. OpenAI's DevDay recap says Cloud is "Available to all Pro, Business, Enterprise and Edu users on desktop and web." Plus isn't on the list, and the Security Review docs say outright it's "not available on Plus."

Here's the wrinkle: on the same day, the feature matrix on OpenAI's Codex pricing page still marks "Codex Security for connected GitHub repositories" as available on Enterprise / Education only, with Pro and Business shown as unavailable. One of those two pages is out of date, and my bet is on the matrix because the launch post is newer. Still, I wouldn't promise a Pro or Business user access until they see the plugin in their own marketplace.

PlanPriceCodex Security Cloud (per DevDay)Codex cloud
Plus$20/monthNot availableYes
Pro$100, $200 or $500/monthYesYes
Business$20/user/month annual, $25 monthlyYesYes
Enterprise and EduContact salesYesYes
API key onlyAPI ratesNo (no cloud features)No

Prices are from OpenAI's Codex pricing page. One more detail is that an API key can run the CLI, but the API Key option has "No cloud-based features," so Cloud needs a ChatGPT plan.

What does Codex Security Cloud cost?

There's no Codex Security line item anywhere. Scans run as Codex cloud work, so they draw from the same pool as everything else. Per the Security Review docs, security reviews "consume included Codex allowance or ChatGPT credits." Once your plan's included usage runs out, Plus and Pro users can buy credits, while Business, Edu and Enterprise plans on flexible pricing buy workspace credits, as covered in my Codex pricing guide.

The number that matters is the model rate. OpenAI's credit table states "Daybreak Blue uses GPT-5.6 Sol credit rates," which is 100 credits per million input tokens, 10 per million cached input and 500 per million output.

Hand-drawn bar chart of output credits per million tokens: GPT-6 Luna 12.5, GPT-6 Sol 250, Daybreak Blue 500 circled, and Daybreak Red 1,875, with a note that Blue bills at GPT-5.6 Sol rates
Hand-drawn bar chart of output credits per million tokens: GPT-6 Luna 12.5, GPT-6 Sol 250, Daybreak Blue 500 circled, and Daybreak Red 1,875, with a note that Blue bills at GPT-5.6 Sol rates

That's twice the per-token rate of GPT-6 Sol (50 / 5 / 250) and 40x GPT-6 Luna on output. So the reduced-refusal access isn't free, and you pay the older GPT-5.6 Sol price for it.

Daybreak Red, the specialist model that needs separate approval, sits at 312.5 / 31.25 / 1,875. In fairness to Blue, the local plugin's deep-scan docs already recommend gpt-5.6-sol for the best scan quality, so it is roughly the rate where a serious scan would run anyway.

To make it concrete, here is some illustrative arithmetic, not a measured scan. Say a full-repo scan reads 5 million input tokens, half of them cached, and writes 400,000 output tokens on Daybreak Blue:

LineTokensRate (credits per 1M)Credits
Fresh input2.5M100250
Cached input2.5M1025
Output0.4M500200
Total475

In terms of scale, OpenAI says a typical GPT-5.6 Sol task uses 5 to 30 credits. A repo scan with validation is many tasks' worth of work, and commit monitoring keeps going. If you run the CLI with an OpenAI API key instead, billing follows OpenAI API pricing, not credits. Real users of the CLI have reported it the same way: one said a scan "ate through 25% of my weekly credits," another that a failed run cost about $13. The CLI's --max-cost limit helps, but OpenAI's CLI FAQ is clear that it's "an estimate, not a hard spending cap."

My advice is to run one repository scan on a mid-sized repo and check your usage dashboard before and after it. Only then switch on commit monitoring across the org.

Why the bundled Daybreak Blue matters most

If you read only one section, make it this one. The most common complaint about Codex Security before Cloud wasn't false positives, it was refusals.

Daybreak is OpenAI's program for defenders. Daybreak Blue gives "access to flagship models with reduced refusals for authorized defensive workflows" such as vulnerability discovery, secure code review, threat modeling and patch validation. Normally you get it by applying through Trusted Access for Cyber, and the approval isn't guaranteed. The tiers are covered in detail in my GPT-5.6-Cyber post.

Without it, the standard model's cyber guardrails fire on exactly the work that a security agent exists to do. When OpenAI open-sourced the CLI in July, the Hacker News thread was filled with people who hit that wall:

Hacker News

"Thing is, you WILL encounter refusals with Sol doing anything remotely adjacent to security work. Which for Codex Security is kinda... problematic."

Another user ran it on a small open-source library, and after watching it for 41 minutes got "This content was flagged for possible cybersecurity risk" at the end:

Hacker News

"it ran for over 40 minutes and during that time I had no idea what was happening, thought it was frozen or in a bad state. Also, it ate through 25% of my weekly credits :("

Codex Security Cloud "includes access to models offered through Daybreak Blue without a separate Daybreak application," per the DevDay recap. That's the real upgrade. One honest caveat is that the launch is only two days old, and I haven't found anyone yet who confirms refusals are gone on Cloud specifically. The Hacker News launch thread had zero comments when I checked.

What early users have said so far

Hands-on reports are still mostly about the CLI and the earlier plugin, since those run the same scanner. The praise is real, even if it's early. Simon Willison, who previewed it in April, put it this way:

"I've been previewing this in Codex for a few weeks - it's very good! Had some great results from it having it run security reviews against code written using other models"

The worries fall into three buckets:

  • Cost and resilience. Beyond the refusal stories, one user's scan hit their account's rate limit and gave up after a minute, costing about $13. An OpenAI team member replied in the thread that retries and resume were still to come (Hacker News).
  • Code leaving the building. This isn't an offline scanner. As an OpenAI team member explained on Hacker News, code and context are sent to OpenAI's hosted model, and "If your company doesn't allow source code to leave its environment, you shouldn't run this against that codebase."
  • Coverage claims. A Hacker News commenter, relaying curl maintainer Daniel Stenberg's posts, said both Codex Security and Anthropic's Claude Mythos came back with zero issues on curl, before another tool's scan led to six CVEs. That's secondhand, so treat it as a warning about relying on only one scanner and not as a benchmark.

Where it fits next to the tools you already run

OpenAI's own FAQ answers the replacement question in one word: "Does it replace SAST? No. Codex Security complements SAST." Rule-based scanners give you broad and deterministic coverage, while Codex Security adds reasoning about how your code fits together and sandbox validation on top. And because it's an AI scan, results can vary between runs even with the same configuration.

Here's how the closest options compare, from each vendor's own pages:

ToolPublished priceTests findings?Proposes patches?
Codex Security CloudIncluded Codex usage, then creditsYes, reproduces in a sandboxYes, you open the draft PR
Claude Code SecurityNo price; limited preview for Enterprise and TeamRe-examines each finding to "prove or disprove" itYes, with human approval
GitHub Code Security$30 per active committer/monthNot stated (CodeQL static analysis)Yes, Copilot Autofix
SnykFree; Team from $25/monthRe-scans each fix candidateYes, Snyk Agent Fix
SemgrepFree to 10 contributors; Teams from $30/contributor/monthNot statedRemediation listed

The pairing I'd actually run is to keep your rule-based scanner (CodeQL through GitHub's paid plans, or Snyk or Semgrep) as the floor, and then add Codex Security Cloud for the reasoning-heavy bugs that a rule can't express. If your team lives in Claude Code, its /security-review command and GitHub Action are the nearest equivalent.

My Claude Code review covers how that agent behaves day to day. If you're still choosing a coding agent at all, my best AI coding assistant tools list is the place to start, and Claude Code's GitHub integration shows the PR-side setup.

For the wider field, see the OpenAI Codex alternatives roundup and the GPT-5.6-Cyber alternatives list.

Five things to check before you switch it on

These come straight from the docs. Each one has bitten someone, or will.

  1. PR comments are as public as your PR. Per the Security Review docs, findings posted to a pull request "inherit that pull request's GitHub visibility." On a public repo, anyone can read a vulnerability write-up before you've fixed it. Set the reporting threshold accordingly, or keep findings in Codex.
  2. A missing finding isn't a fixed finding. The CLI FAQ says "A missing finding or scan comparison alone doesn't prove that a fix worked." Rerun the original scan and recheck the specific issue.
  3. Watch coverage, not just findings. Scans report coverage as complete, partial or unknown. A clean report with partial coverage means "didn't look," not "nothing there."
  4. Patches are proposals. Codex never auto-applies a fix or edits your PR branch, which is good. Keep it that way in your process too, and run your tests on every draft PR it opens.
  5. It's a research preview. Behavior, limits and plan availability can change. The plugin has shipped six releases between August 21 and September 24, 2026, so pin your expectations to the version in front of you.

eesel, for the jobs a security agent doesn't own

Codex Security Cloud is a good example of what a ready-to-work AI teammate looks like, with one job and the right tools for it, plus evidence before it asks a human to act. That's the same shape I build at eesel, only for different jobs. eesel is an AI teammate platform, and today you can hire two teammates: an AI helpdesk teammate that joins Zendesk, Freshdesk, Gorgias or Front, and an AI blog writer for content and SEO. The best AI teammates roundup shows how others approach the same idea.

The validate-first idea maps over directly. Codex reproduces a vulnerability before showing it to you; eesel runs the helpdesk teammate against hundreds of your past tickets in a simulation before it touches a live customer. And the gap between "validated" and "shippable" is real in support too. In one real-traffic trial on an e-commerce Zendesk inbox, the teammate hit 93% triage accuracy and 100% spam detection, yet agents sent only 12% of drafts as-is and rewrote the rest. Accurate isn't the same as ready, which is why actions outside a teammate's rules wait for human approval and every run lands in a shared activity log.

eesel Reports view for a Zendesk teammate showing task volume, trigger events by type, and approval usage per tool
eesel Reports view for a Zendesk teammate showing task volume, trigger events by type, and approval usage per tool

If the Codex Security CLI is what appealed to you, eesel has the same kind of surface. The eesel CLI lets you operate your teammate and workspace from a terminal: eesel approvals list shows what's waiting on a human, eesel activity lists every run, and scripts can automate the rest. Each workspace also works as an MCP server, so Codex or Claude Code can drive the same teammate you see in the dashboard. My AI teammates explainer goes deeper on the model.

Pricing is public: a free plan with 100 credits, then teammate plans from $299/month for 500 credits, where one ticket or chat is one credit. See the pricing page, or try eesel and watch the helpdesk teammate answer your real past tickets the same afternoon.

Frequently Asked Questions

What is Codex Security Cloud?
Codex Security Cloud is a plugin in OpenAI's Codex that scans connected GitHub repositories in the cloud, monitors new commits, tries to reproduce each likely vulnerability in an isolated container, and proposes a patch you can turn into a draft pull request. It's the hosted version of OpenAI Codex's security agent, launched as a research preview at DevDay on September 29, 2026.
Who can use Codex Security Cloud?
OpenAI's DevDay recap says it's available to Pro, Business, Enterprise and Edu users on desktop and web. It isn't on Plus. OpenAI's own Codex pricing matrix still lists Codex Security for connected GitHub repos as Enterprise only, so check with your workspace admin. My ChatGPT pricing guide covers what each plan costs.
How much does Codex Security Cloud cost?
There's no separate price. Scans draw on your plan's included Codex usage, then ChatGPT credits. The bundled Daybreak Blue model bills at GPT-5.6 Sol rates: 100 credits per million input tokens and 500 per million output tokens. See my Codex pricing breakdown for how credits work.
Does Codex Security Cloud replace SAST tools like CodeQL or Snyk?
No, and OpenAI says so directly in its FAQ: Codex Security complements static analysis rather than replacing it. Rule-based scanners give you broad, repeatable coverage; Codex Security adds reasoning and sandbox validation on top. GitHub Copilot Autofix on CodeQL alerts is the closest rule-based pairing.
Does Codex Security Cloud fix vulnerabilities automatically?
No. It generates a proposed patch for findings where it can, but it never applies it or pushes to your branch. You review the diff and select Create draft pull request yourself. That human gate is the same pattern I'd want from any AI agent that touches production.
What is Daybreak Blue in Codex Security Cloud?
Daybreak Blue is OpenAI's reduced-refusal model access for authorized defensive security work. Codex Security Cloud includes it without a separate Daybreak application, which matters because users of the standalone CLI reported scans that ran 30 to 40 minutes and then got refused. My GPT-5.6-Cyber post explains the Blue and Red tiers.
Is my code safe with Codex Security Cloud?
Each scan runs in an ephemeral Codex container that's torn down after the job, but your code is still sent to OpenAI's hosted models. Business and Enterprise data isn't used for training by default. If your policy says source code can't leave your environment, this tool isn't for that repo. ChatGPT Enterprise adds retention and residency controls.
How is Codex Security Cloud different from the Codex Security CLI?
They share the same scanner. The CLI (@openai/codex-security) runs from your terminal or CI and can use an API key, while Codex Security Cloud runs in Codex cloud against connected GitHub repos and keeps working when your laptop is closed. The AI agent CLI guide explains why a CLI surface matters for scripting.

Share this article

Kira

Article by

Kira

Kira is a writer at eesel AI with a Computer Science background and over a year of hands-on experience evaluating AI-powered customer service tools. She focuses on breaking down how helpdesk platforms and AI agents actually work so that support teams can make better buying decisions.

Related Posts

All posts →
Hand-drawn illustration of a developer on a laptop and a colleague looking at a dashboard with a code repository, a security shield, a usage meter and a stack of credit coins
Trending

Codex Security Cloud pricing 2026: plans, credits, and what a scan costs

Codex Security Cloud pricing has no line item. Scans draw from your ChatGPT plan, then credits at Daybreak Blue rates. My real scan worked out to about 89 credits.

Kurnia KharismaKurnia KharismaOct 1, 2026
Hand-drawn illustration of a reviewer holding a scorecard with three ticks and a question mark, next to a cloud-connected code repository and a magnifying glass over a bug inside a shielded sandbox box
Trending

Codex Security Cloud review: I ran its scanner on a buggy app

My Codex Security Cloud review: I planted 5 bugs in a small app and ran OpenAI's scanner. It found all 5, wrote a 55KB report, and cost $3.75 to $6.77.

Rama AdiRama AdiOct 1, 2026
Hand-drawn illustration of a developer comparing five shield-shaped security scanners lined up on a code repository, each with a magnifying glass over a bug
Alternatives

8 best Codex Security Cloud alternatives in 2026 (compared)

The 8 best Codex Security Cloud alternatives in 2026, from Claude Security to Semgrep and Strix, compared on how they prove a bug is real and what they cost.

KiraKiraOct 1, 2026
Hand-drawn illustration of a person holding a green ChatGPT key card in front of a row of open doors, with a weekly usage meter in the corner
Trending

Sign in with ChatGPT: how it works, which apps support it, and the limits

Sign in with ChatGPT lets Plus and Pro users spend their plan inside 16 partner apps. How the two permissions work, the weekly caps, and what partners still charge.

Rama AdiRama AdiOct 1, 2026
Illustrated hero banner of AI agents working inside separate sandbox boxes while a shield-shaped watchdog monitors them from outside, for a post on the NVIDIA Open Agent Safety Platform
Trending

NVIDIA Open Agent Safety Platform: what it is, how it works, and who needs it

The NVIDIA Open Agent Safety Platform is two things: OpenShell, a free runtime you can install today, and Sentry, a hardware watchdog most teams won't touch.

KiraKiraSep 29, 2026
One plugin package feeding several different AI coding agents at once
Trending

Agent Plugins: the new open standard for AI agent extensions

Agent Plugins 1.0.0 shipped on 6 August 2026 with AWS, Cursor, Microsoft, OpenAI and Vercel behind it. Here is what it standardizes, and what it leaves out.

Rama AdiRama AdiAug 6, 2026
Hand-drawn illustration of two developers at a laptop below a cloud holding a friendly robot, connected to an identity shield, a locked database and a chip, with the AWS logo on an orange circle
Trending

Amazon Bedrock Managed Agents explained: OpenAI's agent harness inside your AWS account

Bedrock Managed Agents runs OpenAI's agent harness on AWS while your tools stay on your own compute. How it works, what the preview leaves out, and what it costs.

Rama AdiRama AdiOct 1, 2026
Hand-drawn illustration of a storefront of business tool icons with a buyer handing over a coin, representing OpenAI Marketplace
Trending

OpenAI Marketplace: how it works, who's on it, and the fine print

OpenAI Marketplace lets enterprises spend part of their OpenAI commitment on 32 partner tools. Here's how the money flows, who's listed, and what's unpublished.

Kurnia KharismaKurnia KharismaOct 1, 2026
Hand-drawn illustration of three round, smiling characters working at their own desks and laptops, all connected by dotted lines to one cloud server
Trending

OpenAI Dots explained: what always-on agents do, cost, and where they stop

OpenAI Dots are always-on ChatGPT agents with their own cloud computer. Here's how they work, who gets one, what they cost, and the limits worth knowing first.

KiraKiraOct 1, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free