Microsoft-Decision-1 pricing (2026): $0.042 per million tokens, explained

Rama Adi
Written by

Rama Adi

Katelin Teen
Reviewed by

Katelin Teen

Last edited October 10, 2026

Expert Verified
Microsoft-Decision-1 pricing hero banner in Microsoft blue, a developer sending inputs to a model that returns scored options

What Microsoft-Decision-1 actually is

Microsoft-Decision-1 is a small model that does not write text. You give it a situation and a fixed set of answer options, and it returns a calibrated probability for each option. Microsoft launched it on October 9, 2026, in Microsoft Foundry and through OpenRouter, with a launch post by Satya Nadella on X.

It belongs to a fast-growing category of decision models, alongside TypeSafe Jev, Cloudflare Clef and Strands Decider. Microsoft built it by post-training Alibaba's open-weight Qwen3.5-9B (see Qwen pricing for the family), and says it will "soon rebase it on other models, including Microsoft AI (MAI) and OpenAI."

Microsoft Foundry catalog page for Microsoft-Decision-1 showing the GA lifecycle, json output and a 32768 context window, as taken from Microsoft Foundry
Microsoft Foundry catalog page for Microsoft-Decision-1 showing the GA lifecycle, json output and a 32768 context window, as taken from Microsoft Foundry

The quick facts from the Foundry model card:

  • Lifecycle: "Generally available (GA)", version 1, sold "Direct from Azure"
  • Context window: 32,768 tokens
  • Input / output: text in, JSON out. No explanations or rationales
  • Question types: yes/no, multiple choice, rating, and rubric-based grading
  • Weights: not released. The Hugging Face path returns a 404

I build integrations and APIs for a living, and this shape is familiar. Most of what an AI support agent does all day is not writing. It is answering small questions: is this billing or technical, is it urgent, is the customer angry, is the draft safe to send. A model that only answers those questions, quickly and with a confidence score, is useful. The price question is whether that usefulness is cheap enough to call on every single ticket.

How much does Microsoft-Decision-1 cost?

Microsoft states it in one sentence in the launch post: "Input tokens cost $0.042 USD per million tokens. Output tokens are free."

ItemPriceNotes
Input tokens$0.042 per 1MEverything you send: the ticket, the questions, the options
Output tokensFreeThe probabilities that come back
Free tierNone model-specificAzure's general $200 / 30-day credit applies to new accounts
Batch discountNot published
Cached input discountNot published
Provisioned throughput (PTU)Not publishedCatalog mentions PTUs only in generic "Direct from Azure" text
OpenRouter$0.042 in / $0 outSingle provider: Azure

Free output matters more than it looks. With an LLM, a classification answer costs you output tokens, and output is usually priced 5x input or more. Here, the answer is free, so the bill is driven entirely by how much text you send.

Where you pay

There are two ways in, and they price the same:

  1. Microsoft Foundry. Billed through your Azure subscription. Microsoft's Learn docs say models sold by Azure are "billed through your Azure subscription, covered by Azure service-level agreements, and supported by Microsoft" (Microsoft Learn).
  2. OpenRouter. Same $0.042 input and $0 output, served by Azure behind the scenes. OpenRouter's page shows a weighted average paid of $0.04152 per million input tokens over the first few days (OpenRouter).
OpenRouter model page for Microsoft-Decision-1 listing $0.042 input and $0 output per 1M tokens, a 33K context and an Oct 9, 2026 release, as taken from OpenRouter
OpenRouter model page for Microsoft-Decision-1 listing $0.042 input and $0 output per 1M tokens, a 33K context and an Oct 9, 2026 release, as taken from OpenRouter

If your company already runs on Azure, or already buys Azure OpenAI models, Foundry is the obvious path: one invoice, existing procurement, and the Azure SLA. OpenRouter is the quicker path for a prototype, since you skip the Azure setup entirely.

Microsoft-Decision-1 vs Jev vs Clef vs OpenAI: price per token

Microsoft's launch post compares cost only against GPT-6 Sol. That is where the "200 times less expensive" line comes from. It is a fair number, but it compares a narrow decision model to a flagship general model, which is like comparing a calculator's price to a laptop's.

The comparison that matters is against the other decision models. Here are the live rate cards as of October 11, 2026:

ModelVendorInput per 1MOutput per 1MContext
Clef-flashCloudflare$0.038none listed24,576
Microsoft-Decision-1Microsoft$0.042free32,768
Jev 1.13TypeSafe$0.042free64k per request
Decisions API (gpt-6-luna)OpenAI$0.10free1.05M
ClefCloudflare$0.24none listed65,536
Strands Decider 2BAWS (open weights)self-host onlyself-host only3k to 4k tested
Bar chart of input price per 1M tokens: Clef-flash $0.038, Microsoft-Decision-1 $0.042, Jev $0.042, OpenAI Decisions API $0.10, Clef $0.24
Bar chart of input price per 1M tokens: Clef-flash $0.038, Microsoft-Decision-1 $0.042, Jev $0.042, OpenAI Decisions API $0.10, Clef $0.24

Three things jump out:

  • It ties Jev to the cent. Microsoft matched Jev pricing exactly, which is not an accident in a market where Jev is the default.
  • It undercuts OpenAI by 2.4x. OpenAI's Decisions API went into public beta on October 6 at $0.10 per million input tokens.
  • It is no longer the cheapest. Cloudflare cut Clef-flash to $0.038, with a smaller 24,576-token context. That is about 10% less than Microsoft-Decision-1. My earlier Clef pricing breakdown covers how Workers AI bills it.

The developer crowd noticed the missing comparison right away. The top reply under Nadella's launch post asked the obvious question:

"How is Jev not ranked? That's literally the #1 model people are going to compare this to."

Microsoft later updated the post to add Jev to its accuracy and calibration tables, but the cost chart still only shows GPT-6 Sol. On Hacker News, one commenter put the price critique bluntly:

Hacker News

"Looks like yet another non-price-competitive Jev competitor. Microsoft only compares the price of theirs to GPT Sol(!), not GPT Terra, or GPT Luna (which is what OpenAI's Jev wannabe is based on), and certainly not Jev (4/10 the cost of Luna). I can't remember when a new product created So many competitors so quickly. What is very clear is that everyone is saying "Doh!", slapping themselves on the forehead, and scrambling to get a slice of this obvious-in-retrospect massive pie. What no-one appears to have done yet is to come close to Jev on pricing!"

That comment was written before the price was widely spotted. In fact, Microsoft did come close to Jev on pricing: it matched it.

Per token is not per decision

Here is the twist that makes rate cards less useful than they look. Each model tokenizes and formats your request differently, so the same ticket can cost a different number of tokens on each one. The only hands-on comparison I found came from a Hacker News user who ran Microsoft-Decision-1 against Jev in their own test suite:

Hacker News

"In my minimal suite it was cheaper (by 0.72x) but higher latency (283ms vs 369ms p50) than Jev. Results were very comparable across all scenarios I measure, first of these that I have tested that actually justifies its existence as a commercial release. Qwen tunes are nice and all but either price or performance makes each example I have tested not viable unless you are able to cheaply self-host and fine-tune further."

Same rate card, different real bill. That is one person's suite, not a benchmark, but it is the right instinct: run your own tickets through both and compare the invoice, not the price page.

What a decision really costs

Since output is free, the only number you need is how many input tokens each call carries. That is the ticket text, plus your questions and their option descriptions.

Three rows showing cost per 1,000 decisions: a 300-token short message costs $0.013, a 1,000-token full ticket costs $0.042, and a 4,000-token long thread with order data costs $0.168
Three rows showing cost per 1,000 decisions: a 300-token short message costs $0.013, a 1,000-token full ticket costs $0.042, and a 4,000-token long thread with order data costs $0.168

Microsoft's own estimate is that classifying one million texts costs about $11 on Microsoft-Decision-1 versus about $2,434 on GPT-6 Sol. Working backwards at $0.042 per million tokens, that implies about 260 input tokens per text in their test. Support tickets are usually longer than that, and a long thread with order history attached can run to several thousand tokens.

Here is what a support team actually pays, assuming two decision calls per ticket (say, one for routing and one for "safe to auto-reply?"):

TeamTickets / monthTokens per callMonthly tokensMicrosoft-Decision-1OpenAI Decisions APIClef
Small team2,0001,0004M$0.17$0.40$0.96
Mid-size10,0001,50030M$1.26$3.00$7.20
Large queue100,0002,000400M$16.80$40.00$96.00

Clef figures are at list price, before Cloudflare's daily free allowance. Either way, a 100,000-ticket queue costs less than $20 a month to score. At that point the model is not your cost. The engineering time to build, test and maintain the pipeline around it is.

One more lever: a single call can carry several named questions about the same ticket (OpenRouter). If you ask "which team", "how urgent" and "is it spam" in one request, you pay for the ticket text once instead of three times. Bundling is the single easiest way to cut the bill, and the cheapest piece of LLM optimization you will ever do.

Estimate your own Microsoft-Decision-1 costs

Plug in your volume, average tokens per call, and how many calls you make per ticket. The calculator prices all five hosted decision models at their list rates.

Pricing gotchas to watch for

The rate card is simple. The fine print around it is thin, and that is the gotcha. A few things to check before you build on it:

  • It is not on Azure's pricing page yet. The catalog's "View pricing" link lands on the Foundry Models pricing page, which lists other Microsoft models but no Microsoft-Decision-1 row as of October 11. The $0.042 figure comes from the launch post and OpenRouter.
  • No published regional or provisioned rates. On the same Azure page, Microsoft's MAI-DS-R1 costs 1.1x more on Regional deployments than on Global. Nothing confirms whether Microsoft-Decision-1 follows the same pattern, so check before you pick a data-residency option.
  • No region list or quotas. The model is missing from Microsoft's region availability page, and rate limits are not documented.
  • Your chat SDK will not work. On OpenRouter it runs on a separate Decisions endpoint, and "chat completions SDKs will not work with it." At launch, one developer who deployed it in Foundry wrote that "it doesn't say what to POST or what to expect back" (wkcheng, Hacker News). Budget integration time.
  • The weights change under you. OpenRouter notes that "weights are updated continually while the API shape stays the same." If you set a confidence threshold like "auto-route above 0.85", re-check it after updates.
  • The 32,768-token cap. Long email threads with attachments pasted in can exceed it. You will need to trim, and trimming logic is engineering time too.

None of these are dealbreakers. They are the normal state of a model that is two days old. But "GA" on the model card and "not on the pricing page" on Azure is a combination worth a quick email to your Azure rep before you commit volume.

Is it accurate enough to justify the price?

Price only matters if the answers are right. Microsoft's numbers are strong, but they are Microsoft's own tests. In its 36-benchmark comparison (147,137 questions), Microsoft reports 83.5% average accuracy, ahead of Jev 1.13.0 at 82.3% and GPT-6 Luna Decisions at 79.4%. It ranks third on calibration, at 92.2 against Jev's 93.7.

ModelAvg accuracyMedian latencyCalibration
Microsoft-Decision-183.5%85 ms (p95 125 ms)92.2
Jev 1.13.082.3%240 ms93.7
GPT-6 Luna Decisions79.4%300 ms89.9
GPT-6 Sol (reference)not ranked3.01 snot scored

Source: Microsoft launch post, all figures Microsoft-run. Microsoft's latency was measured through Foundry in the same region; others are from JevBench.

Two independent checks are worth knowing. OpenRouter's own measurement shows a p50 latency of 0.19 seconds, slower than Microsoft's 85 ms figure, which makes sense since it includes network hops. And Microsoft-Decision-1 does not yet appear on the community-run Jev Decision Index leaderboard, where Perplexity's decider and Jev hold the top two spots.

If you threshold on probabilities, calibration matters more than accuracy. One Hacker News user summed up that trade-off: "if you need calibration better use JEV" (nowittyusername). The gap is 1.5 points, which is small, but it is in Jev's favor.

Microsoft also shared internal results: its Microsoft Copilot team found the model "competitive with GPT5.6 Luna and 100 times faster," and Xbox Research found it "200 times less expensive" than GPT-6 Sol on over 10,000 feedback items. Useful signals, but again self-reported. If you are weighing it against a small general model, my GPT-5.6 Luna write-up has that side.

Is Microsoft-Decision-1 worth it for support teams?

At $0.042 per million tokens, the model's price stops being a factor. The real question is whether you want to build the system around it.

Flow of three cards: a ticket arrives, Microsoft-Decision-1 scores it for about $0.00004 by answering billing, urgent and safe to auto-reply, then someone still does the work of writing the reply, issuing the refund and updating the order
Flow of three cards: a ticket arrives, Microsoft-Decision-1 scores it for about $0.00004 by answering billing, urgent and safe to auto-reply, then someone still does the work of writing the reply, issuing the refund and updating the order

What it gives you is the sorting step of ticket triage: which queue, how urgent, is it spam, is this AI draft safe to send.

What you still need to build is everything around it: pulling the ticket through your helpdesk API, writing the questions (these triage prompt templates are a decent start), setting the thresholds, routing based on the answer, and then actually resolving the ticket.

That last step is the expensive one, and it is where I have seen the real errors live. In one real-traffic trial on an e-commerce Zendesk inbox, eesel's triage hit 93% accuracy and caught 100% of spam, while only 12% of the AI drafts were sent with no edits. Sorting is the easy half. Answering well is the hard half.

The probability score is still the most useful thing about these models. A CX lead running 7,000 tickets a month put the requirement plainly on a sales call with eesel:

"I need an AI who is only handling the tickets that it's confident to handle and all the other ones, leave them alone."

A calibrated probability is exactly the tool for that. So my verdict:

  • Use Microsoft-Decision-1 if you have engineers, you already run on Azure, and you want triage, routing or guardrail checks inside your own product or pipeline.
  • Use Jev or Clef-flash if you are cost-optimizing at huge volume and are not tied to Azure. The price gap is small, so test accuracy on your own data first.
  • Skip building it if what you actually want is tickets sorted and answered inside your helpdesk. That is a product, not a model call, and you can automate ticket triage without writing code. See my roundup of the best AI for ticket triage.

eesel for the decision and what comes after it

If you are pricing Microsoft-Decision-1 because you want tickets routed, tagged and prioritized, consider skipping the build. eesel is an AI teammate for your helpdesk. It plugs into Zendesk, Freshdesk, Gorgias and others in a few minutes, learns from your past tickets and help center, and makes the same Zendesk triage and escalation calls a decision model would. Then it writes the reply, updates the order, or hands off to a human when it is not confident.

eesel activity feed showing Zendesk tickets worked by the AI agent with pending and resolved statuses
eesel activity feed showing Zendesk tickets worked by the AI agent with pending and resolved statuses

Pricing is per ticket, not per token. eesel's plans start at $299 a month for 500 credits, where one ticket or one chat is one credit, no matter how many steps the agent takes inside it. At 5,000 credits it is $1,749 a month, or about 35 cents per ticket. That is far more than a decision call, because it includes the work, not just the sorting. Compared with a human agent's cost per ticket, it is a different story; my AI vs human cost breakdown runs those numbers, and Zendesk AI pricing shows what the native add-on route costs. Before going live, you can simulate the agent against your historical tickets to see what it would have done, which is how I would test any of these models anyway.

If you are a developer, the eesel CLI gives you the same teammate from a terminal. npx @eesel/cli init chat-bubble --site <url> sets up a workspace, eesel integrations connects your helpdesk, eesel automations wires up routing rules, and eesel approvals lets you approve or deny actions the agent held for review. It is the same agent as the dashboard, and every command prints JSON, so Claude Code, Cursor or a CI script can drive it. eesel mcp token also exposes the workspace as an MCP server. That is the practical difference from a decision model: the model gives your code a probability, while the CLI gives your code an AI agent that acts on it.

You can try eesel free with 100 credits, no card. Connect your helpdesk, point it at a week of tickets, and compare its routing to what your own decision pipeline would do.

Frequently Asked Questions

How much does Microsoft-Decision-1 cost?

Microsoft-Decision-1 costs $0.042 per million input tokens, and output tokens are free. That is the whole Microsoft-Decision-1 pricing card today: no tiers, no seat fee. For a sense of scale, it matches Jev pricing to the cent.

Is Microsoft-Decision-1 free to use?

No. There is no free tier specific to the model. A new Azure account gets Azure's general $200 credit for 30 days, which would cover millions of decisions at this rate. If you only need tickets sorted without building anything, AI ticket classification tools are another route.

Is Microsoft-Decision-1 cheaper than Jev?

On the rate card, no. Both list $0.042 per million input tokens with free output. Cloudflare's Clef-flash is now slightly cheaper at $0.038. One Hacker News tester reported Microsoft-Decision-1 came out cheaper than Jev in their own suite, which suggests token counts differ per model, so test on your own inputs. My Jev alternatives roundup covers the wider field.

How does Microsoft-Decision-1 pricing compare to the OpenAI Decisions API?

OpenAI's Decisions API charges $0.10 per million input tokens on gpt-6-luna, also with free output. That makes Microsoft-Decision-1 about 2.4x cheaper per token. The Decisions API pricing breakdown has the details.

What does one decision cost on Microsoft-Decision-1?

It depends on how much text you send. A 1,000-token support ticket costs about $0.000042 to score, or 4.2 cents per 1,000 tickets. A 4,000-token thread with order data costs four times that. Microsoft's own estimate is about $11 to classify one million short texts. See ticket triage for what those decisions are usually for.

Can I buy Microsoft-Decision-1 through OpenRouter?

Yes. OpenRouter lists it at the same $0.042 input and $0 output, served by Azure. It uses OpenRouter's Decisions API endpoint, not the chat completions endpoint, so your existing chat SDK code will not work with it as-is.

Does Microsoft-Decision-1 have batch or cached-input discounts?

Not that Microsoft has published. The model is not yet listed on Azure's Foundry Models pricing page, and there is no published rate for provisioned throughput, batch or cached input. If you plan to commit to reserved capacity, ask your Azure account team for a quote first. For broader cost planning, see AI support agent cost.

Is Microsoft-Decision-1 worth it for customer support teams?

For routing and triage inside your own product, yes: the price is low enough that it stops being a line item. The cost that matters is the work after the decision, like writing the reply or issuing the refund. That is the part an AI helpdesk agent like eesel handles.

Share this article

Rama Adi

Article by

Rama Adi

Rama is a software engineer at eesel AI with two years of experience writing about B2B SaaS, AI tools, and customer support technology. Based in Bali, Indonesia, he brings a developer's perspective to product comparisons — cutting through marketing copy to what the integrations and APIs actually do.

Related Posts

All posts →
Hand-drawn hero banner with the Microsoft logo on a blue band, a decision flowchart splitting into two outcomes and joining at a check mark, and three people at a desk reviewing it
Trending

Microsoft-Decision-1: what it is, how it works, and where it fits

Microsoft-Decision-1 scores fixed answer options instead of writing text. How it works, how to call it, what its benchmarks claim and where it fits in support.

KiraKiraOct 11, 2026
Cloudflare Clef hero banner in Cloudflare orange, two people looking at a decision model connected to users, websites, devices and a list of options
Trending

What is Cloudflare Clef? Cloudflare's decision model, explained

Cloudflare Clef explained: what the decision model does, how Clef differs from Clef-flash, how to run it on Workers AI or Ollama, and where it fits in ticket routing.

KiraKiraOct 7, 2026
Hand-drawn hero banner of a person feeding questions into a switchboard that routes them into labeled lanes, each with a confidence dial
Trending

Decision models explained: Jev, Clef, Strands Decider and the new AI category

Decision models return typed answers with confidence scores instead of text. What they are, how they work, every model you can use today, and where they fit in support.

KiraKiraOct 6, 2026
Hand-drawn hero banner of a person at a laptop sending documents into a small decision box that sorts them into three labelled trays while a second person looks on
Trending

Strands Decider 2B: AWS's free decision model, tested against the hype

Strands Decider 2B is a free, open 1.9B decision model from AWS's Strands Labs. What it does, how accurate and fast it is, what it costs to run, and where it fits.

Rama AdiRama AdiOct 8, 2026
Cloudflare Clef pricing hero banner in Cloudflare orange, two people comparing the cost of two decision models
Trending

Cloudflare Clef pricing (2026): $0.24 per million tokens, explained

Cloudflare Clef pricing broken down: $0.24 per million input tokens for Clef, $0.09 for Clef-flash, no output charge, a daily free allowance, and what a decision really costs.

Kurnia KharismaKurnia KharismaOct 6, 2026
TypeSafe Jev pricing hero banner in rose and off-white, showing a low token cost per million
Trending

TypeSafe Jev pricing (2026): $0.042 per million tokens, output free

TypeSafe Jev pricing broken down: $0.042 per million input tokens, output free, no plan tiers yet, and what a System One model actually costs to run in production.

Kurnia KharismaKurnia KharismaSep 22, 2026
Illustration of two people reviewing a Cohere Embed 5 pricing dashboard with the Cohere logo
Trending

Cohere Embed 5 pricing: what Pro and Fast really cost in 2026

Cohere Embed 5 pricing is $0.12 per 1M tokens for Pro and $0.08 for Fast, with images at $0.40. Here is the full table, Model Vault math, and the storage bill nobody prices in.

Rama AdiRama AdiOct 1, 2026
Hand-drawn illustration of two people looking at a price tag and a speed gauge next to the Mistral logo
Trending

Mistral Large 4 pricing: API rates, service tiers, and the real cost per task

Mistral Large 4 pricing is $0.68 in and $2.09 out per 1M tokens on sale. Regional and Priority tiers are priced off the $1.36 / $4.18 list price, not the sale.

Rama AdiRama AdiOct 8, 2026
Hand-drawn illustration of two people working with a voice waveform that branches into audio, storage, and voice profile options, representing Eleven v4 pricing
Trending

Eleven v4 pricing in 2026: API, app credits, and agent minutes explained

Eleven v4 pricing explained: $0.08 per 1K characters at list, $0.04 for Turbo, a launch discount ending October 12, and how credits and agent minutes work.

Kurnia KharismaKurnia KharismaOct 7, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free