
What does the OpenAI Decisions API cost right now?
Nothing you can actually pay for yet. OpenAI announced the Decisions API at DevDay on September 29, 2026, and the DevDay 2026 recap describes it as an API that focuses "Luna's intelligence on a specific set of user-defined questions with finite pre-defined answers". In practice you send text or images, you get back an answer from your own list, and then you use that answer to classify content, route requests or pick an agent's next step.
I build integrations and APIs at eesel, so naturally the first thing I did was to try calling it. Here is what is public and what isn't, as I checked it again on October 2:
| Question | Answer today | Where I checked |
|---|---|---|
| Is there a price? | No Decisions row on the pricing page | OpenAI API pricing |
| Billing unit (per call, per question, per token)? | Unpublished | Recap, pricing page, changelog |
| Is there documentation? | No guide or API reference page; /guides/decisions returns 404 | API guides index |
| Can a normal API key use it? | No. POST /v1/decisions returns HTTP 403, "Decision API is not enabled for this user." | My own API calls, Oct 1 and Oct 2 |
| Which model runs it? | GPT-6 Luna | OpenAI Developers on X |
| Latency claim? | "Less than a few hundreds of milliseconds end to end" (OpenAI staff post, not a docs figure) | Tibo on X |
| General availability date? | "Broad release planned in the coming days" | DevDay 2026 recap |
The 403 is worth a second look though. Nearby paths like /v1/decisions/create return 404, so /v1/decisions is a real and live route, only sitting behind a feature flag. The gate fires before the request body even gets checked, which means the errors don't leak the request shape, and they don't leak the price either.
The support example OpenAI gives itself is the one that matters for this post:
"Send text or images as context. For example, supply a support request and the teams it could go to. The API returns a selection your app can use. Preview access is limited to selected API customers for testing. Broad release planned in the coming days"
What's the closest public price to the Decisions API?
GPT-6 Luna, because that's the model sitting underneath. I want to be careful on this point: OpenAI has not said the Decisions API bills at Luna rates. It might bill per decision or discount the output, or it might do none of that. Still, Luna is what the API runs on, so its rate card is the cost floor OpenAI is working from.
Here is the full GPT-6 Luna rate card from the API pricing page, per 1M tokens, for prompts up to 272K tokens:
| Tier | Input | Cached input | Cache writes | Output |
|---|---|---|---|---|
| Standard | $0.10 | $0.01 | $0.125 | $0.50 |
| Batch | $0.05 | $0.005 | $0.0625 | $0.25 |
| Flex | $0.05 | $0.005 | $0.0625 | $0.25 |
| Fast | $0.20 | $0.02 | $0.25 | $1.00 |
A few rules on the Luna model page end up changing the bill more than the headline rate does:
- Prompts over 272K input tokens bill at 2x input and 1.5x output for the whole request.
- Batch and Flex cost 50% of Standard. Fast mode costs 2x. Luna has no Ultrafast tier; only GPT-6 Astra does.
- Regional data residency and FedRAMP endpoints add 10% for models released after March 5, 2026, per the pricing page.
- Structured outputs aren't priced separately. You pay Luna's token rates, nothing more.
If you want the deeper story on tiers and on the long-context trap, my GPT-6 Luna pricing post goes line by line, and OpenAI API pricing covers the rest of the OpenAI models.
What does one routing decision actually cost?
Since I couldn't call the Decisions API, I went for the next best thing and built the exact job it's made for on the endpoint I was able to call. That meant twenty support tickets, each one with a correct answer I wrote by hand, sent to the Responses API with a strict JSON schema of allowed answers. Every call had to answer these questions at once:
- Which queue? billing, shipping, technical, account, security or other
- What priority? urgent, normal or low
- Safe to auto-reply? yes or no
The tickets were the kind a real queue sees on a normal week. There was a double charge with a chargeback threat and a "where is my order", a team-wide SSO outage, a GDPR deletion request and a phishing report, plus a Spanish-language refund, some spam, and one prompt injection telling the model to file itself as low-priority and auto-reply. I ran every ticket twice through four setups, which comes to 40 calls per setup and 160 in total.
| Setup | Cost per 1,000 tickets | Right queue | Right priority | Right auto-reply call | All 3 right | Median time |
|---|---|---|---|---|---|---|
| Luna, reasoning none | $0.047 | 40/40 | 34/40 | 33/40 | 29/40 | 1.48s |
| Luna, reasoning low | $0.069 | 40/40 | 32/40 | 37/40 | 29/40 | 1.73s |
| Luna, reasoning medium | $0.089 | 40/40 | 32/40 | 38/40 | 30/40 | 2.34s |
| GPT-6.1 Sol, reasoning low | $1.03 | 40/40 | 30/40 | 39/40 | 29/40 | 2.17s |

A few things jumped out at me from this.
First, queue routing is a solved problem at this price. Every setup got the queue right on all 40 calls, and that includes the injection ticket, which every model filed under billing even though it was told to pick "other". If the only question you plan to ask the Decisions API is "which team owns this?", then Luna already answers it today for under five cents per thousand tickets.
Second, the frontier model doesn't buy you much of anything here. GPT-6.1 Sol cost about 22x more than Luna with no reasoning and got the same 29 of 40 perfect answers. It stood out on the auto-reply call, where it was the best, but it was the worst at priority.
Third, my measured times sat between 1.5 and 2.3 seconds from a laptop, with the network included. That is the bar any "few hundred milliseconds" claim needs to clear, and speed is the one thing the Decisions API could change that pricing on its own can't.
One caveat on the caching side. My instructions block was about 350 tokens, and cached tokens came back as 0 on every call, so none of these numbers include the 90% cached-input discount. A longer policy prompt that does get cached would cost you less per call than the input share here suggests.
Why does reasoning effort change the bill?
Because thinking is billed as output, and output is the expensive side of Luna's rate card. Every setup read the same 356 input tokens per ticket. With reasoning off, Luna wrote a 23-token answer, while at medium effort it wrote about 25 tokens of answer plus 82 tokens of thinking that you never get to see.

That is how a 107-token response manages to nearly double the cost of a call with 356 tokens of input. It's also the reason the billing unit OpenAI picks matters more than the rate itself. If the Decisions API bills per decision or makes output free, the thinking tax simply drops out of your forecast.
This exact point is a big part of why Jev's pricing landed so well with developers:
"I just love the simplicity of having only an input price. Input is pretty easy to estimate and calculate upfront, which makes the cost of running something at scale much more predictable. With LLMs, even with JSON schema constraints and structured output, the actual cost can still be hard to predict because of varying output lengths and, especially, unpredictable reasoning costs."
The useful takeaway here is that, whatever OpenAI announces, you should check whether reasoning is on by default. Luna's default reasoning.effort is medium, according to the Luna model page. So if you route tickets on Luna today and never set it, you are already paying the medium price.
What will the Decisions API have to beat on price?
Jev, mostly. TypeSafe launched it on September 15, two weeks ahead of DevDay, as a model that gives back typed answers and probabilities instead of text. Its Models page lists $0.042 per 1M input tokens and says "Output tokens are free." The Hacker News thread under the DevDay recap had already done the comparison within hours:
"They say it's built on Luna, which costs $0.10M/in, vs Jev which only costs $0.04M/in, which is interesting ..."
Here is what 1M routing decisions would cost across the options I'd consider, assuming 500 input tokens and 10 output tokens for each one, with no caching and no reasoning. The rates come from each vendor's own pricing page, and Google's Gemini rates are the ones most likely to move, since the 3.8 Flash promo ends December 31. Claude Haiku 4.5 is from Anthropic's pricing page.
| Option | Input per 1M | Output per 1M | Image input | 1M decisions |
|---|---|---|---|---|
| TypeSafe Jev | $0.042 | Free | No, text only | $21 |
| GPT-6 Luna, Batch | $0.05 | $0.25 | Yes | $27.50 |
| GPT-6 Luna, Standard | $0.10 | $0.50 | Yes | $55 |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | Yes | $140 |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | Yes | $175 |
| Gemini 3.8 Flash (promo to Dec 31) | $0.75 | $3.75 | Yes | $412.50 |
| Claude Haiku 4.5 | $1 | $5 | Yes | $550 |
| OpenAI Decisions API | Unpublished | Unpublished | Yes | Unknown |
The gap between Jev and Luna Standard is $34 per million decisions, which for most support teams is a rounding error. A team handling 20,000 tickets a month would spend about $1.10 on Luna Standard and $0.42 on Jev, and neither of those numbers belongs in a budget meeting.
Price isn't the only gap between them, either. Jev's docs say it takes "Text only", so no image input, and context is capped at 64k tokens. The Decisions API accepts text or images, per the DevDay recap. When your tickets arrive with screenshots of error messages or photos of a broken parcel, that difference matters a lot more than four cents does.
My Jev alternatives post covers the rest of the field, and in the Jev review you'll find my hands-on read of the model itself.
Work out your own routing bill
Plug in your ticket volume and prompt size to see what the routing step would cost you on each option. The rates are list prices from the vendor pages above. The Decisions API isn't in there, because there's nothing to plug in for it yet.
Run your real volume through it and the point more or less makes itself: at helpdesk scale, the routing step costs dollars a year on any of the cheap models. That's why I'd pick on accuracy and input types, and also on how much glue code you end up writing, not on the token price.
What are the hidden costs in a Luna-based decision?
The per-token rate is the small number in all of this. The ones below are what move a real bill, based on Luna's published rules. If the Decisions API ends up billing on Luna tokens, every one of them carries over, and if it bills per decision instead, some of them will go away.
| Cost driver | What it does to the bill | Source |
|---|---|---|
| Default reasoning effort | Luna defaults to medium; in my test that nearly doubled cost per call | Luna model page |
| Long prompts | Over 272K input tokens, the whole request bills at 2x input, 1.5x output | Luna model page |
| Data residency | +10% on regional and FedRAMP endpoints | API pricing |
| Fast mode | 2x Standard; not available with EU data residency for Luna | Using GPT-6 guide |
| Images | Count as input tokens at Luna's rates; screenshots add up fast | API pricing |
| Rate limits | Tier 1 is 500 requests a minute, so a busy queue needs a higher usage tier | Luna model page |
There's one more thing to say on the image line. On September 25, OpenAI fixed a bug that had "degraded image understanding" in GPT-6 Sol and Luna, per the API changelog, and it recommends rerunning image evals. If you tested image-based routing on Luna before that date, it's worth testing again before you trust the results.
Why is a wrong answer the real cost?
This is the part of the test I would show to a support lead. The same call answered two very different questions, and the stakes on them weren't close at all.

A wrong queue costs you a reassignment, not much more. A wrong "yes, safe to auto-reply" costs you a bot answering an angry customer or a phishing victim on its own, or even a GDPR request. Luna with no reasoning made that second mistake 7 times in 40 calls. On the injection ticket, which demanded an auto-reply, it said yes once out of two runs. Medium effort cut the misses down to 2, and Sol cut them to 1.
So the cheapest setup per call stops being the cheapest per month once a person has to clean up after it. Spending $0.042 more per 1,000 tickets for medium effort is, in my view, the best money in this whole post. One Hacker News commenter put the same idea in support terms:
"If they release AGI and it costs $1 and 5 seconds to decide "is the customer asking for a refund", then that's a terrible use case for AGI if another tool can do it with 95% accuracy for $0.002 and 50ms."
I agree with the first half of that. The only change I'd make is that on the auto-reply question, 95% isn't the bar. The fix most teams land on is letting the cheap model route everything, and then sending only the "is it safe to answer alone?" call through a slower and more careful check, with a person in the loop whenever it isn't sure. That's how I think about AI ticket triage and any AI triage tool in general, and it's covered step by step in how to automate ticket triage.
Should you wait for Decisions API pricing or build now?
Build now, for most teams. Here is how I'd split it up:
- You route a few thousand tickets a month. Use Luna with structured outputs today and set reasoning on purpose. The bill is cents. You can swap to the Decisions API later if its speed or price turns out better, since the change is one endpoint.
- You need sub-second answers, like a live chat or an agent's next step. Wait for the Decisions API, or test Jev in the meantime. My Luna calls took about 1.5 seconds, which is fine for email but slow for a conversation.
- Your inputs are screenshots or photos. Skip Jev, since it's text only. Luna or the Decisions API are the options, and rerun your image evals after the September 25 fix.
- You want routing done inside Zendesk or Freshdesk, not an API. You don't need a decisions endpoint at all. What you need is the thing that uses it. My AI ticket classification guide is the place to start. For a specific helpdesk, there's a roundup of Zendesk classification apps, and a separate walkthrough of Freshdesk auto-triage.
The skeptics in the DevDay thread also had a fair point about the timing:
"It's another Jev copy, like we've seen so many over the last few weeks. But with no benchmarks or price comparison, which likely means it doesn't compare that well."
I wouldn't go quite that far. OpenAI shipped Luna's own price card on day one, and a missing price for a gated preview is pretty normal. But "no price, no docs, no benchmarks" is still a good reason not to plan a roadmap around it this week.
eesel for ticket routing
The Decisions API is infrastructure. It picks an answer from your list, and then the rest is on you to write, from the helpdesk connection and the tags to the fallbacks and the "send to a human" rule, plus the dashboard that shows what it did. eesel is the employee that does that job for you. Its AI helpdesk teammate joins your Zendesk queue (see the Zendesk integration) or your Freshdesk one, learns from your help center and past tickets, and routes, tags and replies, with escalation rules you write in plain English.

The part that matters most after the test above is this one: I'd never point an auto-reply at a live queue without replaying it first. eesel runs a simulation over hundreds of your past tickets and scores its answers against what your team actually sent, so you see the "safe to auto-reply" mistakes before a customer does. And if you came here because you'd rather work from code, that side is covered too. The eesel CLI runs the same teammate and workspace from a terminal: eesel instructions edits the routing rules, eesel activity shows every ticket it touched, and eesel approvals lets a person sign off on actions before they happen. Every command prints JSON and supports --dry-run, so scripts and coding agents like Claude Code or Cursor can drive it, and each workspace also works as an MCP server.
Pricing is per ticket, not per token: a ticket or chat is one credit, plans start at $299 for 500 credits, and there's a free tier with 100 credits and no card. Try eesel on a slice of your own queue and see how it routes.
Frequently Asked Questions
How much does the OpenAI Decisions API cost?
Is the OpenAI Decisions API free during the preview?
Is the Decisions API priced per decision or per token?
How does Decisions API pricing compare to Jev?
What is the cheapest way to route support tickets with OpenAI today?
Does reasoning effort change Decisions API pricing?
Do I need the Decisions API to auto-route tickets in my helpdesk?

Article by
Rama Adi
Rama is a software engineer at eesel AI with two years of experience writing about B2B SaaS, AI tools, and customer support technology. Based in Bali, Indonesia, he brings a developer's perspective to product comparisons — cutting through marketing copy to what the integrations and APIs actually do.








