
What is a decision model?
A decision model takes two things: a state (the raw input, like a support email, a JSON blob or a screenshot) and a set of typed questions about that state. It sends back one typed answer per question, each with a probability. No paragraph, no JSON you have to parse, no "Sure! Here's my analysis."
TypeSafe's docs put the problem it solves well: when you ask an LLM to make a judgment your code will consume, "you are coercing a text-generation system into outputting structured decisions, then parsing the results back." A decision model skips the text entirely.

Almost every model in the category has settled on the same three question types, which TypeSafe calls primitives:
| Question type | What it asks | What comes back | Support example |
|---|---|---|---|
| Noul | Is this statement true? | One probability, 0 to 1 | "Is this request urgent?" |
| Choice | Which option fits? | The pick, a probability per option, a confidence | "Billing, technical or sales?" |
| Score | Where on this scale? | A level, a probability per level, a confidence | "No impact, minor, major or critical?" |
You can mix all three in one request, and every question is evaluated against the same state in parallel. TypeSafe says "adding questions barely changes the response time," and the Strands Decider README makes the same point: the text is read once and "each question adds only its own tokens."
Here's what that looks like in practice, from Cloudflare's own launch example. One call asks three questions about a single support message:
curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef \
-X POST \
-H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" \
-d '{
"model": "clef",
"state": "Checkout has been failing for every customer for the last hour.",
"questions": {
"urgent": { "type": "noul", "instructions": "Is this support request urgent?" },
"team": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {
"billing": "Payments, invoices, and refunds",
"technical": "Outages, errors, and configuration",
"sales": "Plans and upgrades"
}
},
"severity": {
"type": "score",
"instructions": "How severe is the customer impact?",
"criteria": ["No impact", "Minor", "Major", "Critical"]
}
}
}'
If you've done ticket triage by hand, that request will look familiar. It's urgency, team and severity, the same three fields a human triager fills in, and the same split behind urgency and impact in most helpdesks.
How does a decision model work under the hood?
This is the part I find most interesting, because the trick is simpler than the marketing suggests.
Most open decision models start from an ordinary open LLM (usually a Qwen or Gemma base) and change how it answers. Cloudflare's launch post describes Clef running "a prefill-only pass, then scores the valid schema choices in parallel." Strands removed the language-model head entirely, so Strands Decider 2B physically can't generate text, and its launch blog says the scoring head it added is "just over a million total parameters."
Put plainly, three things happen:
- The model reads the state and questions once. That's the "prefill" step every LLM does before writing. Decision models stop there.
- A small head scores every allowed answer. Instead of generating "technical" letter by letter, it scores billing, technical and sales against what it just read, all at once.
- The scores are trained to be honest. TypeSafe calls its training method Reinforcement Learning for Calibrated Decisions (RLCD). Cloudflare built its own version of RLCD for Clef and pairs it with a Brier loss, a standard penalty for overconfident probabilities (Cloudflare).
That third step is the real product. A confidence of 0.9 is only useful if answers at 0.9 are right about 90% of the time. Strands reports that on unseen short classification tasks, "answers at a confidence of 0.9 or more are right about 95% of the time." That's why decision models are pitched as a fit for ticket prioritization and other calls where you want a safe cutoff. That's the property that lets you write if confidence > 0.9: auto-route, else: send to a human.
Because nothing is generated, output costs drop to zero or close to it. Compare that with function calling on an LLM, where every brace and quote mark is a billed output token. The same goes for JSON mode. Jev charges nothing for output, and Clef lists no output price at all. That's also why Jev's homepage leans so hard on the "more like code" framing: one request, one typed response, every time.
Decision models vs LLMs vs classic classifiers
The loudest pushback on Hacker News is that this is just a classifier with new branding. One commenter summed it up as "decision model = classifier." They're not wrong about the math. Where they're wrong is about what changed for the person shipping it.
"LLM as judges - generalized, but too slow. ... traditional classifiers, very specific 1 trick pony, super fast and cheap once built. ... decision models/jev - are generic, you can throw them at most generic classification problems, and they are good enough."
That's the honest version. Here's how the three options compare for a typical AI ticket routing job:
| LLM with a strict schema | Decision model | Classic trained classifier | |
|---|---|---|---|
| New labels | Change the prompt | Change the request | Collect data, retrain |
| Output | Text you parse | Typed answer + probability | Label + probability |
| Typical latency | 1 to 3 seconds (my Luna test: 1.3 to 2.6s) | 39 to 524 ms median | Under 10 ms |
| Output token cost | Billed (Luna: $0.50/M) | Free or not billed | None (self-hosted) |
| Calibrated confidence | Weak after post-training | The core feature | Possible, hard to get right |
| Can write a reply | Yes | No | No |
| Best for | Open-ended work, replies, reasoning | Many small zero-shot calls | One fixed, high-volume task |
The Luna latency figures come from a test I ran on GPT-6 Luna with reasoning off for my Decisions API pricing post. The decision-model range is the Cloudflare benchmark below, minus Laya. The classifier row reflects Laya, a 421M encoder that behaves much like a classic classifier and that Cloudflare clocked at 5.8 ms.
The column I'd focus on is "New labels." Another commenter put the practical argument better than any vendor: "The ability to knock out any arbitrary classification problem in minutes instead of in a week is a big deal." If your support team adds a new product line and a new routing category every quarter, retraining a classifier every quarter is the hidden cost a decision model removes.
Which decision models can you use today?
The category moved absurdly fast. The community-run Jev Decision Index launched on September 22 with 31 open entrants and listed 70 open models by its September 28 update. Most are hobby reproductions. Here are the ones with a real vendor or a notable score behind them:
| Model | Maker | Size | Open weights | Hosted | Price | Inputs | Decision Index |
|---|---|---|---|---|---|---|---|
| Jev | TypeSafe | Undisclosed | No | TypeSafe API, Workers AI | $0.042/M input, output free | Text | 57.9 |
| Clef | Cloudflare | 27B | Yes, Apache 2.0 | Workers AI | $0.24/M input | Text, JSON, images, video | 61.2 (self-reported) |
| Clef-flash | Cloudflare | 9B | Yes, Apache 2.0 | Workers AI | $0.09/M input | Text, JSON, images, video | 57.1 (self-reported) |
| pplx-decider-v1-27b | Perplexity | 26B | Yes | Not verified | Self-host | Text, images | 56.4 |
| Strands Decider 2B | AWS Strands Labs | 1.9B | Yes, Apache 2.0 | No | Free (self-host) | Text (images experimental) | Not on board |
| Bespoke-Nimble-9B-v2 | Bespoke Labs | 9B adapter | Yes | No | Self-host | Text | 39.6 |
| Kev 9B | Jared Palmer | 9B adapter | Yes, Apache 2.0 | No | Self-host | Text | 38.5 |
| Tev1-4B | Together AI | 4B | Yes, license not final | Together serverless | Not captured | Text | 29.0 |
| Laya | Convai Innovations | 421M | Yes, Apache 2.0 | No | Self-host | Text, 100+ languages | 6.0 |
| Decisions API | OpenAI | Built on GPT-6 Luna | No | Limited preview | Not published | Text, images | Not on board |
Sources: the Decision Index, Cloudflare's self-reported board, and each model's own page linked in the sections below. The Index is chance-corrected across 38 benchmarks, so 0 means random guessing.
The four I'd actually look at first are below.
TypeSafe Jev
Jev is the original and still the reference everyone benchmarks against. It's a closed model from TypeSafe AI, a San Francisco team led by Diogo Almeida, who co-invented the RLHF and InstructGPT methods behind ChatGPT. On September 27, TypeSafe dropped its waitlist, so anyone can now create an account.

The headline price in the launch post is $0.042 per million input tokens with free output, and TypeSafe claims end-to-end response times of 70 to 500 ms. It's text-only with a 32K context. Two things to know before you commit: Jev's architecture and weights aren't public, and TypeSafe itself says it "can't prove it isn't subsidized."
For cost math at your volume, see my Jev pricing breakdown. If you want the hands-on view, I covered it in my Jev review. There's also a separate lower-latency tier, which I wrote up as Jev Ultrafast.
Cloudflare Clef and Clef-flash
Clef is the open model that claims to beat Jev. Cloudflare built it on Qwen3.8-27B (Clef-flash on Qwen3.5-9B), released both under Apache 2.0, and made them fully Jev-API compatible, so swapping is mostly a URL change.

Clef's two real advantages over Jev are vision (it reads images and video) and a 64K context. On Cloudflare's numbers, it also wins 3 of 4 of TypeSafe's own workflow evals, including customer service, where all three models land within a point of each other (Clef 76.3, Clef-flash 77, Jev 76.0). The caveat is that both Clef scores are Cloudflare grading its own homework: they aren't reproduced on the upstream board, and Cloudflare skipped two benchmarks. It's also about 5.7x Jev's per-token price. I ran the per-decision math in my Cloudflare Clef pricing post.
Strands Decider 2B
Strands Decider 2B is AWS Strands Labs' entry, written up by Marc Brooker, Mike Chambers and Fabio Nonato de Paula. It's the small, runs-on-your-laptop option: 1.9B parameters on a Qwen3.5-2B base, with weights, training data and training scripts all released.

It runs on CPU, NVIDIA GPUs or Apple silicon, with a median of 115 ms on an RTX 3090 and 153 ms on an M3 Pro for short prompts, per the GitHub repo. On JevBench it scores 0.762 on the newest version. I like how candid the launch post is: it says the single-pass design makes it "significantly worse at solving complex problems than reasoning models" and unsuited for chatbots or summaries. Two practical notes: there's no hosted version (Hugging Face shows no inference provider), and its local server has no authentication, so it's for experiments, not production traffic. If you already build on the Strands harness, it plugs into a tool-call hook as a proceed/deny gate.
OpenAI Decisions API
OpenAI announced the Decisions API at DevDay on September 29, with the DevDay recap describing it as "focusing Luna's intelligence on a specific set of user-defined questions with finite pre-defined answers." Seven days later, I still find no docs page, no reference schema, no price row and no changelog entry. The /v1/decisions endpoint exists, but a standard key gets an HTTP 403 saying the API is not enabled for this user.
So for now, the only checkable OpenAI number is GPT-6 Luna's token rate: $0.10 per million input and $0.50 per million output. If you're on OpenAI today, a Luna call with a strict enum schema does the same job; my OpenAI Decisions API post walks through that workaround.
How fast are decision models?
Speed is where the category earns its name. Cloudflare measured median latency across 43 benchmarks for six decision models on its own stack:

Two honest caveats on that chart. First, Jev's 524 ms is a hosted round trip over the internet, while the open models were measured on Cloudflare's own GPUs, so it isn't a like-for-like race. Second, Laya is blazing fast but scores 6.0 on the Decision Index; speed without accuracy is just a fast wrong answer.
For comparison, an LLM doing the same job is measured in seconds. One developer who swapped a mock support agent's guardrail layer from an LLM to Jev reported "194ms vs 5,825ms. $125 vs $5,894 per million reviews" on a small 51-case test. That gap is the whole pitch for putting a decision model in front of an AI helpdesk agent as a guardrail.
Are decision models accurate?
This is where I'd push back on the hype. Decision models are good enough for a lot of jobs, but they're not the most accurate option on the table.
On JevBench, the Strands evaluation notes describe the top of the board as "reasoning models (GPT-6 Luna 0.996), then Jev rebuilds on larger torsos, mostly 12B and up (0.83-0.93)." A reasoning LLM like Luna or Claude Sonnet 5.5 still wins on accuracy. Decision models win on cost, speed and calibration, and you trade a few points of accuracy for that.
Accuracy also drops as the number of labels grows. Laya's own card shows it scoring 0.425 on BANKING77 (77 labels) against Jev's 0.870 (Laya). If your helpdesk has 60 routing categories, test on all 60, not a tidy demo of four.
And the confidence scores, the whole selling point, are not settled science yet:
"for production routing, you almost always want to threshold on confidence ('auto-handle above 0.9, route to a human below'), and that only works if the probabilities mean what they say. A 96% model with overconfident outputs is operationally worse than a 94% model with honest ones."
One tester asked Jev for the result of a fair die roll and got back roughly 80% confidence on "1" (wat10000). Another showed an open Jev copy flipping to "100% legitimate" when the email itself said it was legitimate (FooBarWidget). Decision models can't hallucinate a paragraph, but they can absolutely be confidently wrong or talked into an answer.
The Strands model card gives the right default: "measure on your own traffic before you trust a threshold." It's the same lesson behind stopping AI hallucinations in support: test on real tickets, not demos.
Which decision model fits your job?
Pick the situation closest to yours.
That last tab isn't a joke. One Reddit commenter put it perfectly: "half the routing i've written turned out to be an if statement wearing a costume." Plain Zendesk routing rules still cover a lot of ground.
Where do decision models fit in customer support?
Support is the use case every vendor reaches for first, and for good reason. Auto-triage is mostly a stack of small, bounded calls. Cloudflare's launch example is a support ticket. TypeSafe's own workflow evals include a customer service set of 5,330 rows.
Intent classification is a choice question. So is spam detection. Urgency is a score, and "should this go to a human?" is a yes/no, which is the heart of AI escalation. Even sentiment analysis is a score on a fixed scale. They're all bounded questions with a fixed set of answers. That's exactly the shape a decision model is built for.

Here's what I've learned from years of putting AI on live support queues at eesel: the sorting is the easy half. On one cross-validated trial on an e-commerce company's real Zendesk inbox, AI triage hit 93% accuracy and caught 100% of the spam, which made up 22% of the inbox. On the same trial, only 12% of AI-drafted replies were sent as-is, and 7% had a factual error. The ticket classification call was nearly solved. The reply was where the work was.
The confidence score is the piece buyers have been asking for all along. A CX lead handling 7,000 tickets a month told eesel's team on a sales call:
"I need an AI who is only handling the tickets that it's confident to handle and all the other ones, leave them alone."
That's a decision model's core idea in one sentence. It's also why eesel simulates every rollout against a team's historical tickets before anything goes live: a confidently wrong answer is worse than no answer, and the only way to know where your threshold sits is to test on real traffic.
One practical point from an MSP owner is worth hearing before you rebuild your stack around speed. In an r/SmallMSP thread, they said the model wasn't the bottleneck: "I suspect that Autotask is my weak link when it comes to speed." Shaving 2 seconds off a routing call doesn't matter much if the ticket then sits in a queue for 40 minutes.
So where does a decision model belong? Inside a workflow you're building yourself, as the cheap first pass. If you're writing code and processing millions of events, a decision model is a great building block for routing, moderation or guardrails, and my roundup of Jev alternatives covers the build-it-yourself options. If you're a support team that wants tickets tagged, routed, escalated and answered in your existing helpdesk, a raw model is the start of a project, not the end of one. You'd still be building the tagging, the escalation process, the knowledge lookup and the reply.
For the full DIY path, start with my guide to automating ticket triage. If you'd rather buy than build, my roundup of AI triage tools compares the ready-made options.
Try eesel for triage and replies
eesel's AI helpdesk teammate does both halves of the job decision models only do one of. It joins your existing queue (there's a Zendesk integration, plus Freshdesk and Gorgias), tags and routes tickets, escalates what it shouldn't touch, and writes the reply from your help center and past tickets. Each automation even has a "Model strength" setting, with a lighter model "for simple jobs like tagging or routing" (automations docs) and the stronger one for drafting, so the cheap decision and the hard reply already run on different models.

The confidence idea carries through too. Every action (a tag, a reply, a refund) is set to auto, needs approval or disabled, and events your instructions don't cover are skipped and logged with a reason. You can run a simulation on your own past tickets before anything goes live.
If you came here as a developer, the eesel CLI is the bridge. Run npx @eesel/cli, and you can connect a helpdesk, set up an automation with plain-language routing instructions via eesel automations, and approve or deny held actions with eesel approvals, all from a terminal. Every command returns JSON, so Claude Code, Cursor or Codex can drive the whole setup, and --dry-run checks a write before it happens. You get the routing logic you'd have built around a decision model, plus the part that answers the customer, without writing the glue.
Plans start at $299 a month for 500 credits on eesel pricing, where one credit is one ticket, with no per-seat fees, and there's a free tier with 100 credits. Try eesel on your own tickets and see where your confidence line actually sits.
Frequently Asked Questions
What is a decision model in AI?
A decision model is a small AI model that answers typed questions (yes/no, pick one, or score on a scale) about a piece of input and returns each answer with a probability, instead of writing text. TypeSafe Jev started the category in September 2026, and decision models are now used for jobs like ticket triage. Sorting a ticket by topic, also called intent classification, is a typical job.
How are decision models different from LLMs?
An LLM generates text token by token, so you have to parse its answer. A decision model reads the input once and scores the allowed answers in a single pass, so there is no text to generate and output is effectively free. The trade-off is that decision models can't write replies, code or summaries, which is why a support team still needs an AI helpdesk agent or a person for the reply.
Are decision models just classifiers?
Mathematically they are close: they pick from labels and return probabilities. What is new is that they work zero-shot on labels you define at request time, so you don't need to collect data and train a model for each new ticket classification problem.
How much do decision models cost?
Jev costs $0.042 per million input tokens with free output, Cloudflare Clef costs $0.24 and Clef-flash $0.09 per million input tokens, and Strands Decider 2B is free to run yourself. OpenAI has not published a price for its Decisions API. My Jev pricing post has the full math. For Clef, see the Clef pricing breakdown.
Which decision model is the best?
On the community Decision Index, Jev scores 57.9 and no independently tested open model beats it yet. Cloudflare reports Clef at 61.2 on its own copy of the board, but that score is self-reported. For images, Clef; for offline or CPU use, Strands Decider 2B; for the cheapest hosted text decisions, Jev.
Can I use a decision model for ticket routing?
Yes. Ticket routing, urgency and escalation calls are the textbook use case, and Cloudflare's own launch example routes a support message by team and severity. You still need code or a helpdesk tool to act on the label, which is what an AI helpdesk teammate handles end to end.
Is the OpenAI Decisions API available?
OpenAI announced the Decisions API at DevDay 2026 in limited preview. As of October 6, 2026 there is no public documentation, schema or price, and the endpoint returns a not-enabled error for standard API keys. My OpenAI Decisions API post tracks what is known.
Do decision models replace an AI support agent?
No. A decision model makes one bounded call, like which team or how urgent. It doesn't look up an order, apply a refund policy or write a reply. Decision models are a building block for automating ticket triage, not a replacement for the teammate that resolves the ticket.

Article by
Kira
Kira is a writer at eesel AI with a Computer Science background and over a year of hands-on experience evaluating AI-powered customer service tools. She focuses on breaking down how helpdesk platforms and AI agents actually work so that support teams can make better buying decisions.








