OpenAI Decisions API: what it is, how it works, and what's new

Kira
Written by

Kira

Katelin Teen
Reviewed by

Katelin Teen

Last edited October 1, 2026

Expert Verified
Hand-drawn illustration of a person sending a support ticket into a small router box that sends it to one of three teammates

What is the OpenAI Decisions API?

The Decisions API is a new OpenAI endpoint for one narrow job: choosing between options you define in advance. Here's how OpenAI's DevDay 2026 recap describes it:

"Decisions API enables real-time decision-making by focusing Luna's intelligence on a specific set of user-defined questions with finite pre-defined answers. Developers supply context using text or images, and get back answers they can use to classify content, route requests, or choose an agent's next action."

Scrolling capture of OpenAI's DevDay 2026 recap page, which introduces the Decisions API alongside the other launches, as taken from OpenAI

I build AI agents for a living, and most of the model calls inside one of them don't write anything at all. They are small forks in the road, like which team owns this ticket or does this need a human. A general model can answer those fine, only it goes the slow way, by generating text first. The Decisions API is OpenAI betting that this kind of call should get its own lane.

It sits beside the bigger DevDay launches, like the always-on OpenAI Dots agents, the GPT-6.1 Sol model and Codex Security Cloud.

It got less stage time than those. The keynote segment is under a minute, starting around 22:15 in the keynote, and the same week also brought ChatGPT Space. Still, if you run a support queue, this is the launch that maps the closest onto your daily work.

How does the Decisions API work?

You send in three things and you get one thing back. The flow, as far as OpenAI has described it, looks like this:

Hand-drawn flow of three cards: context such as ticket text or a screenshot, your questions with fixed answers like Billing, Shipping, Technical or Other, and one answer back, used to classify content, route requests or pick a next action
Hand-drawn flow of three cards: context such as ticket text or a screenshot, your questions with fixed answers like Billing, Shipping, Technical or Other, and one answer back, used to classify content, route requests or pick a next action
  1. Context. Text or images, so a ticket body or a chat transcript, or a product photo the customer attached.
  2. Questions with finite answers. You define each question, plus the full list of answers it is allowed to return.
  3. A selection back. One answer out of your list, and then your code acts on it.

The support example actually came from OpenAI's developer account itself, in a thread from @OpenAIDevs:

"Send text or images as context. For example, supply a support request and the teams it could go to. The API returns a selection your app can use."

Under the hood it runs on GPT-6 Luna, which OpenAI calls its "most efficient model for focused, high-volume tasks" on the Luna model page. How the keynote explains the speed is that a pre-defined set of options is what lets Luna answer in a fraction of a second. Nothing has to be written out word by word, so there's less to wait for.

If you plan to build on it, the parts OpenAI hasn't explained yet matter about as much:

  • The request and response schema, since there is no API reference page.
  • Whether one call can carry several questions, and how many answers each question can have.
  • Whether the response carries a confidence score. Press coverage says yes, but I found no OpenAI page or post that mentions one.
  • The billing unit and rate limits, also whether Batch or Flex apply.

One Hacker News reader took a guess at the mechanism in the DevDay thread: "I think they simply use the LLMs softmax scores (uncalibrated confidence)". It's a plausible reading, though nobody has confirmed it.

Can you use the OpenAI Decisions API yet?

Only if OpenAI picked you. The recap says it's "available in limited preview today with a broad release planned in the coming days", and the OpenAI Developers thread adds that "preview access is limited to selected API customers for testing."

I went and checked what this means in practice. POST https://api.openai.com/v1/decisions is a real route, and a standard key gets this answer from it:

JSON
{"error":{"message":"Decision API is not enabled for this user.","type":"invalid_request_error","param":null,"code":null}}

That is an HTTP 403, which means a feature gate rather than a missing route. Nearby paths like /v1/decisions/create and /v1/beta/decisions do return 404. The gate fires even on an empty body, so the errors also don't leak anything about the request shape. I got the same 403 on October 1, then again on October 2, three days after the "coming days" promise.

The paper trail is thin as well. As of October 2:

What I checkedResult
Docs guide (/api/docs/guides/decisions)404
API reference (/api/reference/decisions)404
API changelog, Sep 29 entriesNo Decisions API entry
API pricing pageNo Decisions API row
Standalone announcement postNone, only the DevDay recap
GET /v1/modelsNo model ID containing "decision"

So for now the Decisions API is a promise and a gated endpoint. That is pretty normal for a preview, but it does mean you can't load test it or price it yet, and you can't read its limits either.

What does the Decisions API add over structured outputs?

One HN commenter asked the obvious question within hours of the launch:

Hacker News

"It seems a bit silly since OpenAI LLMs already can output structured data."

It's a fair point. You can already make Luna answer from a fixed list today: use the Responses API, set reasoning to none, then pass a strict JSON schema where the only field is an enum of your answers. Below is the exact request I ran while the Decisions API stayed gated:

JSON
{"model":"gpt-6-luna","reasoning":{"effort":"none"},
 "input":[{"role":"developer","content":"Route the support ticket. Which queue?"},
          {"role":"user","content":"I was charged twice for my order #4471"}],
 "text":{"format":{"type":"json_schema","name":"route","strict":true,
   "schema":{"type":"object","properties":{"queue":{"type":"string","enum":["billing","shipping","technical","other"]}},
   "required":["queue"],"additionalProperties":false}}}}

It came back {"queue":"billing"} on all three runs, using 60 input tokens and 12 output tokens. That works out to $0.000012 a call at Luna's Standard rates. Anyone who used OpenAI function calling will know the idea, just with a tighter leash, and it's also the setup behind most AI ticket routing for SaaS today.

There is a catch though, which a commenter raised in an earlier HN thread about OpenAI and Jev:

Hacker News

"It's not guaranteed to be correct: it's guaranteed to be formatted in a particular way. You can get the same thing with grammars on any LLM."

So the format problem is solved already. What the Decisions API promises on top is speed, plus whatever accuracy tuning "focusing Luna's intelligence" turns out to mean. On what's public so far, the two compare like this:

Decisions APILuna + structured outputs
StatusLimited preview, selected customersAvailable now
InputText or imagesText and images
Output"A selection" from your answersJSON that matches your schema
Speed"Less than a few hundreds of milliseconds end to end" (OpenAI staff claim)1.46s median in my test
PriceNot published$0.10 input / $0.50 output per 1M tokens
Confidence scoreNot confirmedNone by default
DocsNone yetStructured outputs guide

How fast is the Decisions API?

The only written speed claim from OpenAI comes from a staff member, there's nothing in the docs. Thibault Sottiaux, who works at OpenAI, posted on launch day:

"Decisions API, for lightning fast constrained decision making powered by Luna. Supports visual inputs, and tuned to be able to make decisions in less than a few hundreds of milliseconds end to end."

To see what that would be beating, I routed 20 support tickets through Luna twice per setup and asked for a queue and a priority, also whether the ticket was safe to auto-reply. That's 160 calls in all, timed end to end from a laptop, so the network round trip is counted in.

Hand-drawn bar chart of time per decision: Jev 70 to 500 ms and the Decisions API a few hundred ms are vendor claims drawn as dashed bars, while my tests measured Luna with no reasoning at a 1.46 second median and Luna at medium effort at 2.33 seconds
Hand-drawn bar chart of time per decision: Jev 70 to 500 ms and the Decisions API a few hundred ms are vendor claims drawn as dashed bars, while my tests measured Luna with no reasoning at a 1.46 second median and Luna at medium effort at 2.33 seconds
Setup (my test)MedianFastestSlowest
GPT-6 Luna, reasoning none1.46s0.95s2.79s
GPT-6 Luna, reasoning low1.62s0.95s3.20s
GPT-6 Luna, reasoning medium2.33s1.44s5.75s
GPT-6.1 Sol, reasoning low2.17s1.52s4.98s

If OpenAI's claim holds up, the Decisions API would be roughly five times faster than my best Luna setup. Where that gap matters, and where it doesn't:

  • Email and ticket triage. Not really. Nobody notices if a ticket got tagged in 300ms or in 1.5 seconds.
  • Live chat. Yes. A 1.5 second pause before the bot has even decided who should answer, that adds up over a conversation.
  • Agents. Here is where it matters the most. An agent built with something like OpenAI AgentKit that makes 20 small choices per task waits 30 seconds on Luna, versus a few seconds at the claimed speed.

Some press coverage shows a "150 ms vs 1.6 s" chart. I couldn't find those numbers on any OpenAI page or post, so I wouldn't plan around them before the docs ship.

How does it compare with TypeSafe Jev?

It's hard to talk about this launch without bringing up TypeSafe Jev. TypeSafe launched Jev on September 15 as a model built only for typed decisions, then the Decisions API showed up two weeks after. One HN commenter said it bluntly in the DevDay thread: "Decisions API is a validation for Jev and the entire space it created."

Going by what each company has published, they line up like this:

OpenAI Decisions APITypeSafe Jev
AccessLimited previewOpen to everyone since Sep 27
InputText or imagesText only, per the Jev models page
Question typesQuestions with fixed answersChoice, Score and Noul (true or false probability)
OutputA selectionChoice plus probabilities and confidence
PriceNot published$0.042 per 1M input tokens, output free
Speed"Less than a few hundreds of milliseconds""70ms-500ms" end to end, per the launch post
Options per questionNot publishedUp to 255 per Choice
Rate limitsNot published100K tokens/s or 40 requests/s

For most teams it comes down to two rows. Jev doesn't accept images, so if your queue is full of screenshots and damage photos, that points you to OpenAI. On the other side Jev returns a probability with each answer, and that's exactly what you need for deciding when not to act.

My Jev review covers how it held up in testing. For the rate math, the Jev pricing breakdown has it, including the faster Jev Ultrafast tier.

On price, the early Luna comparison already leans toward Jev:

Hacker News

"They say it's built on Luna, which costs $0.10M/in, vs Jev which only costs $0.04M/in, which is interesting ..."

Keep in mind that's Luna's list price, not a Decisions API price. OpenAI could price the endpoint quite differently once it opens.

What can you use the Decisions API for in support?

OpenAI named three jobs, and each of them has a clear support version:

  • Classify content. Tag a ticket by intent, product or sentiment, spot the spam, or tell a refund request apart from a return.
  • Route requests. Send a ticket to billing, shipping or technical, or to a language or tier queue. This is the classic intelligent routing, the job OpenAI used as its own example. The same pattern works for ecommerce routing.
  • Choose an agent's next action. Decide if it should look up the order, ask a clarifying question, reply, or escalate to a person, normally based on intent detection.

The first two are what most teams mean when they say ticket triage, and typed decisions already work well there. Most helpdesks ship some native version too, from Freshdesk auto-triage to a long list of Zendesk classification apps. On one real-traffic trial with a jewelry e-commerce store doing about 1,000 tickets a month on Zendesk and Shopify, eesel's typed calls hit 93% triage accuracy and caught 100% of the spam with zero false positives, and spam was 22% of that inbox.

My own test showed where the trouble begins. Every setup picked the right queue 40 of 40 times. On "is this safe to auto-reply?" though, Luna with no reasoning got 33 of 40, and even GPT-6.1 Sol only got 39 of 40. Routing is the easy call; knowing when not to act is the hard one. A wrong queue costs you a few minutes, while a wrong auto-reply goes straight out to a customer.

Buyers draw the line in the same place. A CX lead at a supplements brand doing about 7,000 tickets a month on Gorgias told eesel on a sales call that they couldn't check every AI reply by hand, so the AI had to stay out of anything it wasn't sure about:

"I need an AI who is only handling the tickets that it's confident to handle and all the other ones, leave them alone."

A decisions endpoint gives you the answer. Unless it returns a calibrated confidence as well, it doesn't tell you when to leave a ticket alone, so that rule is still yours to build. Before you trust any classifier with an action, it's worth reading up on false positives in AI tagging.

What will the Decisions API cost?

Nobody outside OpenAI knows yet. There is no price row, and OpenAI hasn't said if it bills per token, per call or per question. For now the one public yardstick is Luna's own rate card:

GPT-6 Luna tierInput per 1MCached input per 1MOutput per 1M
Standard$0.10$0.01$0.50
Batch / Flex$0.05$0.005$0.25
Fast$0.20$0.02$1.00
Scrolling capture of the GPT-6 Luna model page showing its pricing, limits and supported endpoints, as taken from OpenAI

At those rates my 160-call test cost $0.047 per 1,000 tickets with no reasoning, and $0.089 at medium. Even a 10,000-ticket month comes out well under a dollar. The full math and hidden costs, plus the Jev price fight, are in my Decisions API pricing post, and the broader OpenAI API pricing guide has every model's rate. On pricing alone, the decision is never the expensive part of a support stack.

Should you build on the Decisions API now?

Not yet, unless you're in the preview. You can still build the same feature today and swap the endpoint in later on. This is how I'd decide it:

Hand-drawn decision tree: if you're routing tickets this month and not writing code, use a helpdesk AI teammate; if you are writing code and need images or want to stay on OpenAI, use Luna with a strict schema today and switch when the Decisions API opens; otherwise use Jev, which is text only and open now
Hand-drawn decision tree: if you're routing tickets this month and not writing code, use a helpdesk AI teammate; if you are writing code and need images or want to stay on OpenAI, use Luna with a strict schema today and switch when the Decisions API opens; otherwise use Jev, which is text only and open now
  • You're writing code and staying on OpenAI. Ship Luna with a strict enum schema and reasoning set to none or low. Keep the questions and the answer list in one place, then moving over to /v1/decisions stays a small change. The Luna alternatives post covers other small models if you want a fallback, like Gemini 3.5 Flash-Lite.
  • You need sub-second answers on text. Test Jev now. It is open and priced, and it returns probabilities too. The Jev alternatives list covers the rest of that field.
  • Your inputs are images. Stay with OpenAI. Luna takes images today, and per the API changelog OpenAI fixed an image-encoding bug on September 25 that had "degraded image understanding", so rerun any older image evals.
  • You want routing inside Zendesk or Freshdesk, not an API. Then you can skip the endpoint question completely. Start with my guide on how to automate ticket triage, or the roundup of the best AI for ticket triage. Zendesk teams can also compare Zendesk Intelligent Triage.

On timing, the skeptics do have a point:

Hacker News

"It's another Jev copy, like we've seen so many over the last few weeks. But with no benchmarks or price comparison, which likely means it doesn't compare that well."

I'd be less harsh about it. A gated preview with no price is normal, and the image input is a real difference. Still, "no docs, no price, no benchmarks" is a good enough reason not to build a roadmap on it this week.

eesel for confident ticket routing

The Decisions API is infrastructure. It picks an answer from your list, and everything around that answer you still have to build yourself, like the helpdesk connection and the tags, the "send to a human" rule, and a log of what it did. eesel is the teammate that already does that job. Its AI helpdesk teammate joins your queue in Zendesk or Freshdesk, learns from your help center and past tickets, then routes, tags and replies, with escalation rules you write in plain English.

eesel activity view filtered to a Zendesk instance, listing resolved and pending conversations the AI teammate handled
eesel activity view filtered to a Zendesk instance, listing resolved and pending conversations the AI teammate handled

According to my test the part that matters most is the "safe to auto-reply" call, and that is where eesel puts its effort. Before it touches a live queue, eesel replays hundreds of your past tickets and scores its answers against what your team actually sent, so the risky calls show up for you before a customer sees them. Any ticket it isn't sure about goes to a person.

If you came here because you would rather build from code, the eesel CLI runs the same teammate and workspace from a terminal. eesel instructions edits the routing rules, eesel activity lists every ticket it touched, and eesel approvals lets a person sign off on an action before it happens. Every command prints JSON and supports --dry-run, so scripts and coding agents like Claude Code or Cursor can drive it, and each workspace also works as an MCP server.

Pricing is per ticket, not per token: a ticket or chat is one credit, plans start at $299 for 500 credits, and the free tier gives you 100 credits with no card. Try eesel on a slice of your queue, and see which of the tickets it's confident enough to take.

Frequently Asked Questions

What is the OpenAI Decisions API?
It's an API that takes some context, like a support ticket or an image, plus questions you write with a fixed list of possible answers, and returns one answer per question. OpenAI announced it at DevDay on September 29, 2026 and says it runs on GPT-6 Luna. It's built for jobs like ticket classification, routing and picking an agent's next step.
Is the OpenAI Decisions API available yet?
Only in limited preview. OpenAI says access is limited to selected API customers, with a broad release planned. On October 2, 2026 a standard API key still got a 403 "Decision API is not enabled for this user" error, and there's no docs page yet. Until it opens, structured outputs on Luna do the same job.
How much does the OpenAI Decisions API cost?
OpenAI hasn't published a price. The closest yardstick is GPT-6 Luna at $0.10 input and $0.50 output per 1M tokens, which worked out to about $0.047 per 1,000 routed tickets in my test. My Decisions API pricing breakdown has the full math.
How is the Decisions API different from structured outputs?
Structured outputs already force a model to answer from an enum you define, and that works today. What the Decisions API adds is speed: OpenAI staff say it's tuned to decide in under a few hundred milliseconds end to end, while Luna with structured outputs took a 1.46 second median in my test. For email ticket triage that gap barely matters, but in live chat you'd feel it.
Is the OpenAI Decisions API a copy of TypeSafe Jev?
It does the same kind of job. TypeSafe Jev launched two weeks earlier and also answers fixed-choice questions in milliseconds. The differences so far: Jev is open to everyone and text only, with a published price of $0.042 per 1M input tokens and free output. The Decisions API accepts images but has no public price or docs yet.
Can I use the OpenAI Decisions API to route support tickets?
Yes, routing is OpenAI's own example: send a support request plus the teams it could go to, and get back a selection. You still have to connect it to your helpdesk, decide what happens on low confidence, and log every call. An AI helpdesk agent like eesel handles that part inside Zendesk or Freshdesk without code.
Does the Decisions API return a confidence score?
OpenAI hasn't said. Some press coverage mentions confidence, but no OpenAI page, doc or post does. If you need a confidence signal today, Jev returns probabilities with each choice, and a guide to reducing false positives covers how to gate actions on it.

Share this article

Kira

Article by

Kira

Kira is a writer at eesel AI with a Computer Science background and over a year of hands-on experience evaluating AI-powered customer service tools. She focuses on breaking down how helpdesk platforms and AI agents actually work so that support teams can make better buying decisions.

Related Posts

All posts →
Hand-drawn illustration of support tickets flowing into a router that sends them to one team, with a price tag and a cost meter
Trending

OpenAI Decisions API pricing in 2026: what it costs before OpenAI says

OpenAI Decisions API pricing isn't published yet. I priced the same routing job on GPT-6 Luna, Jev and others, and found the real cost is wrong answers.

Rama AdiRama AdiOct 2, 2026
OpenAI logo connected to six outlined squares
Guides

OpenAI Embeddings API: how semantic search actually works

Learn how the OpenAI Embeddings API supports semantic search and retrieval, what a support knowledge workflow still needs, and how to test it before relying on results.

Rama AdiRama AdiOct 12, 2025
Illustration comparing alternatives to the OpenAI Agents API for building AI agents
Trending

The 8 best OpenAI Agents API alternatives in 2026

The best OpenAI Agents API alternatives in 2026, split into the two real choices: swap the model provider, or swap the orchestration layer and keep model choice.

Kurnia KharismaKurnia KharismaSep 11, 2026
Illustration of the OpenAI Agents API and the agent loop of reasoning and tool calls
Trending

OpenAI Agents API: what it is and how to build an agent in 2026

A plain-English guide to the OpenAI Agents API: how the agent loop works, the four runtimes, the hosted tools, and how to actually build and ship one.

Rama AdiRama AdiSep 11, 2026
Illustration of the OpenAI Agents API pricing model with API tokens and tool costs
Trending

OpenAI Agents API pricing: what building an agent actually costs in 2026

OpenAI's Agents API has no single price. You stack model tokens, hosted-tool calls, and container costs, then pay for every turn of the loop. Here's the real math.

Rama AdiRama AdiSep 11, 2026
Text and image inputs moving through risk flags to a human reviewer
Guides

OpenAI Moderation API: build a safer support review flow

Learn what OpenAI Moderation API signals mean, how to route flagged support content safely, and how eesel CLI helps test the support workflow around them.

Rama AdiRama AdiOct 12, 2025
Blue gradient graphic reading Realtime API GA and OpenAI
Guides

OpenAI Realtime API: a current guide to live voice support

Learn when the OpenAI Realtime API fits a live voice-support experience, how to choose a session and transport, and what to test before callers rely on it.

Rama AdiRama AdiOct 12, 2025
A person viewing connected user and assistant message threads
Guides

OpenAI Threads API: conversation state after Assistants

Learn why OpenAI conversation state now belongs in the Responses and Conversations APIs, what your application still owns, and how to test an eesel teammate safely.

Rama AdiRama AdiOct 12, 2025
A base network, curated examples, checklist, and refined network
Guides

OpenAI Fine-Tuning API: what to do as it winds down

Learn OpenAI's current fine-tuning status, how to decide between training and support configuration, and how to test a safer path before customer replies change.

Rama AdiRama AdiOct 12, 2025

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free