
Microsoft-Decision-1 alternatives at a glance
| Model | Vendor | Input per 1M tokens | Output | Context per call | Images | How you run it | Weights and license | Decision Index v0.3.1 (Full) | Data and compliance notes | Free allowance |
|---|---|---|---|---|---|---|---|---|---|---|
| Microsoft-Decision-1 (baseline) | Microsoft | $0.042 | Free | 32,768 | No | Azure Foundry, OpenRouter | Closed | Not listed | Azure SLA; regions not published | General Azure $200, 30 days |
| Jev 1.13 | TypeSafe | $0.042 | Free | 64k per request, 32k state + longest question | No | TypeSafe API, Cloudflare Workers AI | Closed | 60.11 (reference, tied 2nd) | Zero data retention on the Cloudflare listing | None documented |
| OpenAI Decisions API | OpenAI | $0.10 | Free | 1,050,000 (2x input price past 272K) | Yes | OpenAI API, public beta | Closed | Not listed | ZDR and HIPAA for eligible customers, US and EU residency | None |
| Clef-flash | Cloudflare | $0.038 | Free | 24,576 | Yes | Workers AI, or self-host | Open, Apache 2.0 | 47.61 (#33) | Cloudflare account terms | 10,000 Workers AI neurons a day |
| Clef | Cloudflare | $0.24 | Free | 65,536 | Yes | Workers AI, or self-host | Open, Apache 2.0 | 53.08 (#19) | Cloudflare account terms | 10,000 Workers AI neurons a day |
| Perplexity Decider v1.1 | Perplexity | Your GPU cost | n/a | 8,192 per decision | Yes | Self-host, about 49 GiB of weights | Apache 2.0, authenticated download | 62.75 (#1) | Stays on your infrastructure | Free download |
| Rune 26B-A4B v3 | Surogate (Invergent) | Your GPU cost | n/a | 262k | Yes | Self-host with the surogate engine | Apache 2.0, contact-info gate | 57.43 (#10) | Stays on your infrastructure | Free download |
| Strands Decider 2B | AWS Strands Labs | Your GPU cost | n/a | Tested at 4,096 | Optional, untrained | Self-host, fits a 3090 | Apache 2.0 | 21.75 (#84) | Stays on your infrastructure | Free download |
| eesel (not a model) | eesel | $299 a month for 500 tickets or chats | n/a | Whole ticket and history | Yes | Plugs into your helpdesk | Hosted product | n/a | HIPAA on Enterprise | 100 free credits |
Prices and context come from each vendor's own docs, checked October 11, 2026. Decision Index scores come from the Jev Decision Index board, version 0.3.1, updated October 10.
Why look past Microsoft-Decision-1?
I wrote up what the model does in my Microsoft-Decision-1 overview and what it costs in the pricing breakdown. It is a strong release. Microsoft's own table puts it at 83.5% average accuracy across 36 benchmarks with an 85 ms median on Foundry. But a few gaps keep coming up:
- No open weights. It is a post-trained Qwen3.5-9B, but the Hugging Face page Microsoft would use returns 404. You rent it or you don't use it. If you liked the Qwen family for self-hosting, this one is closed.
- Text only, 32K context. Screenshots, scanned forms and very long threads need another model.
- Thin Azure paperwork. It is GA in the catalog, but it has no row on Azure's Foundry pricing page yet, and no published regions, quotas or batch rate. Teams buying through Azure models usually want those before they commit.
- No independent score yet. It is not on the Jev Decision Index, so every accuracy number so far is Microsoft's own.
That last point drove most of the launch-day reaction. The top reply under Satya Nadella's launch post was blunt:
"How is Jev not ranked? That's literally the #1 model people are going to compare this to."
Microsoft later updated the post to add Jev to the accuracy and calibration table. Still, the question underneath it is the right one. Decision models are a crowded field now, and the decision model you pick should depend on where it runs and what it reads, not on launch-day charts.

How I picked these alternatives
I come at this from the search side. My job at eesel is mostly figuring out what buyers are actually asking when they type a query, and "Microsoft-Decision-1 alternatives" two days after launch splits cleanly into price, hosting and input type. So I kept to models that do the same job: you send a state plus fixed answer options, and you get calibrated probabilities back in one pass.
Every model below either appears in Microsoft's own comparison table or ranks on the Decision Index, or both. I left out general chat models with JSON mode, since function calling on a big model is a different cost curve entirely. Each entry uses the same five headings so you can scan across them.
How much text can each one read per call?
Context is the spec that bites first in support work. A single ticket with order history, a policy excerpt and three back-and-forth replies can pass 10,000 tokens.

Perplexity Decider is missing from that chart on purpose. Its card says the text and all images of one decision must fit in 8,192 tokens, and longer inputs are rejected rather than truncated. Top score, smallest window.
1. Jev 1.13 (TypeSafe)

What it is
Jev is the model that started the category. TypeSafe serves it through one endpoint, POST /v1/systemone, and it scores every question you send against a shared state in parallel. I covered it in depth in my Jev overview.
Where it beats Microsoft-Decision-1
- Bigger requests. Jev takes 64k tokens per request, with 32k for the state plus the longest single question, per TypeSafe's models page. Microsoft caps at 32,768.
- Better calibration. Even on Microsoft's own table, Jev scores 93.7 on calibration against Microsoft's 92.2.
- More places to buy it. Besides TypeSafe's API, Jev is on Cloudflare Workers AI at the same price, with free cached input and zero data retention.
Where it falls short
- Text only, like Microsoft. Images must be turned into text first.
- TypeSafe warns that "Rate limits are adjusting dynamically" and can change without notice. Today's limit is 100K tokens and 80 requests per second.
- Microsoft's table has it slightly behind on accuracy, 82.3% to 83.5%.
One Hacker News tester ran both through their own suite the day after launch:
"In my minimal suite it was cheaper (by 0.72x) but higher latency (283ms vs 369ms p50) than Jev. Results were very comparable across all scenarios I measure, first of these that I have tested that actually justifies its existence as a commercial release."
Same list price, different bills: tokenizers count differently, so test on your own tickets.
Pricing
$0.042 per million input tokens, output free, no documented free tier. My Jev pricing post works through cost per decision, and Jev Ultrafast covers the speed tier.
Best for
Teams that want a Microsoft-Decision-1 alternative with the same price and API shape, longer inputs, and no Azure dependency.
2. OpenAI Decisions API (gpt-6-luna)

What it is
OpenAI's take on the same idea, released in public beta on October 6. You call POST /v1/decisions with gpt-6-luna, an input of text and images, and a list of questions of type predicate, choice or score. My Decisions API guide has the walkthrough.
Where it beats Microsoft-Decision-1
- Images. It reads text and images in one call. Images must be inline base64 data URLs.
- Enormous context. It inherits GPT-6 Luna's 1,050,000-token window, so a whole ticket history fits.
- Compliance options. OpenAI's guide lists Zero Data Retention and HIPAA for eligible customers, plus US and EU data residency.
Where it falls short
- It costs about 2.4x more per token than Microsoft.
- Still beta. OpenAI says it expects GA "in the coming weeks".
- Prompts over 272K tokens bill at 2x input, and regional processing premiums apply.
- OpenAI publishes no millisecond latency, only "about 10x faster than the Responses API". Microsoft's table measured it at 300 ms median and 79.4% accuracy.
Pricing
$0.10 per million input tokens, with no output, cache-read or cache-write charges. That rate lives in the Decisions guide, not yet on the main OpenAI API pricing page. More in my Luna pricing post.
Best for
Teams already on OpenAI who need screenshots or long threads scored, or who need HIPAA paperwork in place.
3. Cloudflare Clef-flash

What it is
Cloudflare's small decision model, a 9B built on the same Qwen3.5-9B base as Microsoft's. It runs on Workers AI as @cf/cloudflare/clef-flash, and the weights are on Hugging Face under Apache 2.0. My Clef overview covers both Clef sizes.
Where it beats Microsoft-Decision-1
- Cheapest hosted rate. $0.038 per million input tokens, down from $0.09 a week ago.
- Images. Cloudflare's page lists vision support, with images billed as input tokens.
- Free daily allowance. Workers AI gives 10,000 neurons a day on Free and Paid plans. By my math that is roughly 2.9M Clef-flash input tokens a day.
- Open weights, so you can move it in-house later.
Where it falls short
- The context window shrank to 24,576 tokens with the price cut, smaller than Microsoft's.
- It ranks 33rd on the Decision Index (47.61), well below Clef and Jev.
- It is not in Microsoft's comparison table, so there is no shared benchmark between the two.
Pricing
$0.038 per million input tokens and no output charge. Past the free neurons, Workers AI bills $0.011 per 1,000 neurons.
Best for
High-volume text or image classification where cost per call matters most and inputs stay short.
4. Cloudflare Clef

What it is
The bigger Clef: a 27B multimodal decision model post-trained from Qwen3.8-27B, hosted on Workers AI and published under Apache 2.0.
Where it beats Microsoft-Decision-1
- Twice the context. 65,536 tokens.
- Images and video in the state, per Cloudflare's model card.
- An independent score. It is on the Decision Index at 53.08 Full, and its public-benchmark-only score of 61.71 is the highest public number among the hosted options here.
Where it falls short
- At $0.24 per million input tokens, it costs nearly 6x Microsoft's rate.
- Its private-test score drops it to 19th, so the public number flatters it.
- Cloudflare's self-reported median was 209 ms, slower than Microsoft's claimed 85 ms.
Pricing
$0.24 per million input tokens, no output charge, and the same 10,000 free neurons a day (about 458K Clef input tokens). My Clef pricing post has the math.
Best for
Teams on Cloudflare who need longer multimodal inputs and will pay more per call for them.
5. Perplexity Decider v1.1 (27B)

What it is
Perplexity's open decision model, a full fine-tune of Qwen3.8-27B with a separate decision head. Perplexity says v1.1 gained most of its improvement from "the lifting of the causal mask" and more training data. You download it from Hugging Face and run it with the included code.
Where it beats Microsoft-Decision-1
- Top independent score. It ranks first on Decision Index v0.3.1 at 62.75, ahead of Jev's 60.11.
- You own the weights. Apache 2.0, nothing leaves your servers.
- Images, passed as file paths, PIL images or base64 URLs.
Where it falls short
- Hard 8,192-token limit per decision, text and images combined.
- You need a CUDA GPU with room for about 49 GiB of weights plus working memory, and an authenticated Hugging Face account to download.
- No hosted API: the card says no inference provider serves it.
- Default causal inference does not reproduce the evaluated setup, so you must use Perplexity's own code or a careful export.
Pricing
Free to download. Your cost is GPU time. On cost, one Hacker News commenter put it well:
"Open weight model pricing isn't really up to the company that built them (unless they are also serving them, in which case their API price may differ). It's really up to the serving cost and/or pricing of whoever is actually serving the model."
Best for
ML teams with spare GPU capacity who want the most accurate open model and keep inputs short.
6. Surogate Rune 26B-A4B v3

What it is
Invergent's open decision model, a full fine-tune of Gemma 4 26B-A4B (8 of 128 experts active per token). It answers choice, noul and score questions and is served fastest by the open-source surogate engine.
Where it beats Microsoft-Decision-1
- Long inputs. The card claims a 262k context, eight times Microsoft's.
- Images, with the vision tower kept. Invergent's own image test scored 83.4% at 1,120 image tokens.
- Optional thinking. With
"thinking": true, questions it is unsure about reason before answering. Invergent estimates that lifts its index by about 1.8 points, at about 5 seconds per thinking question. - On Microsoft's own table it scored 79.7%, close to OpenAI's Luna.
Where it falls short
- Downloads sit behind a contact-information form.
- Its default probabilities run overconfident. The card recommends reading answers at temperature 2 to fix calibration.
- Invergent measured a 388 ms median per request on one RTX PRO 6000, slower than Microsoft's claim.
- 10th on the Decision Index (57.43), behind Perplexity and Jev.
Pricing
Free to download under Apache 2.0. You pay for the hardware.
Best for
Self-hosters who need long or image-heavy inputs and can tune serving themselves.
7. Strands Decider (AWS Strands Labs)

What it is
A small open decision model from AWS's Strands Labs, built for agent steps like tool selection, guardrails and routing. The reference model is a 2B on Qwen3.5, and four Gemma 4 sizes followed on October 10 and 11. I covered the launch in my Strands Decider post.
Where it beats Microsoft-Decision-1
- Runs on modest hardware. The 2B hit a 115 ms median on a single RTX 3090 in Strands' own test.
- Free and open, Apache 2.0, and it fits inside the wider Strands harness for agents.
- Optional image input with
--vision, though Strands notes nothing was trained on images.
Where it falls short
- Accuracy is far lower. Microsoft's table shows 54.8% on only 23 of 36 benchmarks, and the Decision Index ranks it 84th (21.75).
- It was tested at 3,072 and 4,096-token windows.
- There is no hosted API anywhere, Bedrock included.
Pricing
Free. Your own GPU or CPU.
Best for
Cheap, local agent guardrails and tool picks where speed on small hardware beats top accuracy.
8. eesel (if you want the decision and the work)

What it is
Not a model. The seven options above are infrastructure. eesel is the employee: an AI helpdesk teammate that joins your Zendesk, Freshdesk or Gorgias queue, makes the routing call, and then writes the reply or takes the action. It learns from your past tickets and help center.
Where it beats Microsoft-Decision-1
- No pipeline to build. No thresholds, webhooks or helpdesk API glue.
- It acts on the answer. Tags, priority, escalation and the reply itself.
- Risky actions can wait for approval, so a human signs off before anything irreversible.
Where it falls short
- Far more expensive per ticket than a single model call, because it includes the work.
- It lives in your helpdesk. If you need decisions inside your own app's code path, use a model.
Pricing
eesel's plans start at $299 a month for 500 credits, where one ticket or one chat is one credit however long it runs. 5,000 credits is $1,749 a month. There is a free start with 100 credits and no card.
Best for
Support teams whose real goal is tickets sorted and answered, not a model to wire up.
Which Microsoft-Decision-1 alternative should you pick?

My short version:
- Same job, no Azure: Jev at the same price. Compare the two on your own data first.
- Lowest bill: Clef-flash, if your inputs fit in 24,576 tokens.
- Images or very long threads: OpenAI's Decisions API hosted, Rune v3 self-hosted.
- Best open accuracy: Perplexity Decider v1.1, if 8,192 tokens is enough.
- Small and local: Strands Decider for agent guardrails.
For support teams, the use case is almost always ticket classification and routing. That is where the probability score earns its keep: you only automate what the model is confident about. On real queues the hard part comes after. In one example from our own data, a Web3 infrastructure company's AI matched a cold contact-list sales pitch against past Zendesk tickets, classified it as spam, and drafted a polite decline as an internal note instead of trying to answer it. A decision model gives you the "spam" label. Something still has to read the history and write the note.
Watch false positives in tagging too. A 1% error rate is cheap per call and expensive when it routes a refund to sales. My roundup of the best AI for triage compares tools that handle this out of the box.
Skip the model, keep the decision
If you're comparing Microsoft-Decision-1 alternatives because you want tickets routed, tagged and prioritized, you may not need a model at all. eesel works like a new hire for your helpdesk: it plugs into Zendesk or Freshdesk in minutes, reads your past tickets, makes the same triage decisions a decision model would, and then answers the customer. You can simulate it on historical tickets before it touches a live one.

If you'd rather drive it from code, the eesel CLI runs the same teammate the dashboard does. Every command prints JSON, so you, a script, or a coding agent like Claude Code, Codex or Cursor can set it up. eesel integrations connects your helpdesk, eesel automations enable <platform> <key> --instructions "..." writes a triage rule in plain words, and eesel approvals lists held actions to approve or deny. More on why that matters in my agent CLI piece.
Try eesel free with 100 credits and point it at a week of your tickets. If it routes them as well as your decision pipeline would, you've saved yourself the build.
Frequently Asked Questions
What are the best Microsoft-Decision-1 alternatives?
The closest swap is TypeSafe's Jev, which costs the same $0.042 per million input tokens. OpenAI's Decisions API adds images and a huge context window, Cloudflare's Clef-flash is the cheapest hosted option at $0.038, and Perplexity Decider, Rune v3 and Strands Decider are open-weight Microsoft-Decision-1 alternatives you run yourself.
Is there a cheaper alternative to Microsoft-Decision-1?
Yes. Cloudflare's Clef-flash lists $0.038 per million input tokens against Microsoft's $0.042, with free output on both. The open-weight models cost nothing to download, but you pay for the GPU. My Clef pricing breakdown covers the free daily allowance on Workers AI.
Is Jev better than Microsoft-Decision-1?
On Microsoft's own 36-benchmark table, Microsoft-Decision-1 scored 83.5% average accuracy to Jev's 82.3%, while Jev scored higher on calibration (93.7 vs 92.2). Prices are identical. Jev reads 64k tokens per request against Microsoft's 32,768. My Jev review has more on how it behaves in practice.
Is there an open-source alternative to Microsoft-Decision-1?
Microsoft has not released weights, so you cannot self-host it. Perplexity Decider v1.1, Surogate Rune 26B-A4B v3, Cloudflare's Clef models and AWS's Strands Decider all ship under Apache 2.0. Perplexity Decider v1.1 currently tops the community Decision Index.
Which Microsoft-Decision-1 alternative supports images?
OpenAI's Decisions API, Cloudflare Clef and Clef-flash, Perplexity Decider and Rune v3 all accept images. Microsoft-Decision-1 and Jev are text only. If screenshots show up in your ticket triage flow, that alone narrows the list.
How does Microsoft-Decision-1 compare to the OpenAI Decisions API?
OpenAI charges $0.10 per million input tokens on gpt-6-luna, about 2.4x Microsoft's rate, also with free output. In exchange you get images, a 1,050,000-token window and ZDR and HIPAA options for eligible customers. The API is still in public beta. See my Decisions API pricing post.
Do I need a decision model to route support tickets?
Only if you are building routing into your own product or pipeline. If the goal is tickets sorted and answered inside your helpdesk, an AI helpdesk agent does the decision and the work without code. My guide on how to automate ticket triage walks through both paths.

Article by
Kurnia Kharisma
Kurnia is a software engineer and writer at eesel AI with two years of SEO experience, writing about AI tools, helpdesk software, and customer support. He pairs a developer's understanding of how these products are built with search-driven research into what actually ranks and resonates with the people searching for them.








