Microsoft-Decision-1: what it is, how it works, and where it fits

Kira
Written by

Kira

Katelin Teen
Reviewed by

Katelin Teen

Last edited October 10, 2026

Expert Verified
Hand-drawn hero banner with the Microsoft logo on a blue band, a decision flowchart splitting into two outcomes and joining at a check mark, and three people at a desk reviewing it

What is Microsoft-Decision-1?

Microsoft-Decision-1 is a decision-scoring model. You hand it a piece of text (Microsoft and OpenRouter call it the state) and one or more questions with a fixed set of answers. It hands back a probability for every answer. That's the whole interface.

Microsoft announced it on October 9, 2026 in a post on its Command Line blog by Achint Srivastava, VP of Software Engineering in the Office of the CTO. The post says it's built for "routing, classification, prioritization, verification, and workflow control" (Microsoft). Satya Nadella posted the launch the same day, saying Microsoft is already testing it internally on incident response, quality control and scientific discovery (Satya Nadella on X).

Microsoft Command Line blog post introducing Microsoft-Decision-1 by Achint Srivastava, dated 2026.10.09, as taken from Microsoft's announcement
Microsoft Command Line blog post introducing Microsoft-Decision-1 by Achint Srivastava, dated 2026.10.09, as taken from Microsoft's announcement

If the category is new to you, my explainer on decision models covers the background. The short version: TypeSafe Jev started the category in September 2026, Cloudflare, AWS and OpenAI followed within weeks, and Microsoft-Decision-1 is the first entry from a hyperscaler that sells the model on its own cloud.

What it isn't matters as much. Microsoft's model card says it's "not designed for text generation, open-ended question answering, conversation, translation, or summarization" (Foundry catalog). It can't write a reply, a summary or a line of code. It picks.

How does Microsoft-Decision-1 work?

I spend my days building the agent side of eesel, so this is the part I went looking for first. The mechanics are simpler than the launch post makes them sound.

Microsoft took Alibaba's open-weight Qwen3.5-9B and post-trained it "for fast, single-pass decision scoring," and says it "will soon rebase it on other models, including Microsoft AI (MAI) and OpenAI" (Microsoft). Training data is "publicly available datasets subject to Microsoft's Open Data process, plus synthetic data created by the team" (Foundry catalog). The Qwen base is the same family behind Cloudflare's Clef-flash and AWS's Strands Decider.

Hand-drawn flow: a state card with the message "Help! My payouts have been failing for 3 days." and three questions (is_urgent yes or no, department billing, technical or sales, frustration Calm, Frustrated or Very angry) goes into a Microsoft-Decision-1 box labelled one pass and 32K tokens in, which outputs a JSON card with a probability bar per question and a crossed-out speech bubble labelled no explanation
Hand-drawn flow: a state card with the message "Help! My payouts have been failing for 3 days." and three questions (is_urgent yes or no, department billing, technical or sales, frustration Calm, Frustrated or Very angry) goes into a Microsoft-Decision-1 box labelled one pass and 32K tokens in, which outputs a JSON card with a probability bar per question and a crossed-out speech bubble labelled no explanation

In plain terms, three things happen in one call:

  1. It reads the state and every question once. Up to 32,768 tokens of text go in. Microsoft calls this "single-pass scoring" (Foundry catalog).
  2. It scores every allowed answer. Instead of writing "billing" letter by letter, it puts a probability on billing, technical and sales at the same time.
  3. It returns JSON and stops. There's no rationale. The card says plainly that the model "does not generate explanations or rationales."

The question types match what Jev users already know. On OpenRouter, one request can carry several named questions about the same state (OpenRouter):

Question typeWhat it asksWhat comes backSupport example
noulIs this true?One probability from 0 (no) to 1 (yes)"Does this message convey urgency?"
choiceWhich option fits?The pick plus a probability per option"Billing, technical or sales?"
scoreWhere on this scale?A position on an ordered rubric"Calm, Frustrated or Very angry?"

Microsoft's catalog lists a few extras on top: rubric-based grading of AI responses and agent actions, groundedness checks against supplied evidence, safety checks, and "explicit abstention," meaning an option like "cannot tell" when the evidence is thin (Foundry catalog). That last one is underrated. A model that can say "not enough here to decide" is far easier to put in front of a real queue.

Here's the OpenRouter sample request, lightly trimmed. It's a support ticket, which tells you who Microsoft and OpenRouter think the first buyers are:

TypeScript
const decision = await openrouter.alpha.decisions.create({
  decisionsRequest: {
    model: "microsoft/microsoft-decision-1",
    state: "Help! My payouts have been failing for 3 days.",
    questions: {
      is_urgent: {
        type: "noul",
        instructions: "Does this message convey urgency?",
        criteria: { true: "Explicitly time-sensitive", false: "No urgency expressed" }
      },
      department: {
        type: "choice",
        instructions: "Which team should handle this?",
        criteria: {
          billing: "Payments, invoicing, refunds",
          technical: "Bugs, outages, integrations",
          sales: "Pricing, upgrades, new accounts"
        }
      },
      frustration: {
        type: "score",
        instructions: "How frustrated is the customer?",
        criteria: ["Calm", "Frustrated", "Very angry"]
      }
    }
  }
});

Your code then reads the answers and decides what to do, for example escalating when is_urgent is above 0.8 and the department is billing. That threshold logic is the point. Compare it with JSON mode on a general LLM, where you pay for every output token and still get a label with no honest confidence attached.

How do you access Microsoft-Decision-1?

There are two doors, and they behave differently.

Microsoft Foundry, "Direct from Azure." The catalog lists Microsoft-Decision-1 as Version 1 with a lifecycle of "Generally available (GA)," text in, JSON out and a 32,768-token context window (Foundry catalog). Microsoft's Learn docs put it under models sold directly by Azure, which "are billed through your Azure subscription, covered by Azure service-level agreements, and supported by Microsoft" (Microsoft Learn). If you're already buying Azure OpenAI models, this is the familiar procurement path.

Microsoft Foundry catalog page for Microsoft-Decision-1 showing the Direct from Azure badge, Version 1, and quick facts listing Generally available (GA), text input, json output and a 32768 context window, as taken from Microsoft Foundry
Microsoft Foundry catalog page for Microsoft-Decision-1 showing the Direct from Azure badge, Version 1, and quick facts listing Generally available (GA), text input, json output and a 32768 context window, as taken from Microsoft Foundry

OpenRouter's Decisions API. No Azure account needed. The model ID is microsoft/microsoft-decision-1, it runs on POST https://openrouter.ai/api/alpha/decisions, and OpenRouter warns that "chat completions SDKs will not work with it" (OpenRouter). Azure is the only provider behind it, so every request still lands on Microsoft's infrastructure.

OpenRouter model page for Microsoft: Microsoft-Decision-1 showing $0.042 / $0 per 1M input and output price, a 33K context and an October 9, 2026 release date, as taken from OpenRouter
OpenRouter model page for Microsoft: Microsoft-Decision-1 showing $0.042 / $0 per 1M input and output price, a 33K context and an October 9, 2026 release date, as taken from OpenRouter
Microsoft FoundryOpenRouter
StatusGA, Version 1Live since October 9, 2026
AccountAzure subscriptionOpenRouter key
BillingYour Azure billOpenRouter credits
SLAAzure SLAs applyNot an Azure SLA
Request shapeNot publicly documented yetDocumented Decisions API
Model updatesVersion 1 in the catalog"Weights are updated continually while the API shape stays the same"
Measured speed85 ms p50, Microsoft's own test0.19 s best-provider p50, OpenRouter's measurement

Two rows in that table would change how I'd ship this. First, Foundry is the safer production home, but the request format isn't public yet. One Hacker News user who deployed it on day one wrote:

Hacker News

"Maybe I'm just not looking in the right place, but I cannot find what the API shape looks like. I've even deployed this model via Foundry and it doesn't say what to POST or what to expect back."

When I checked on October 11, there was still no Learn how-to page for the model. Second, OpenRouter's note that weights update continually means a threshold you tuned last month can drift under you. If you route on confidence > 0.9, re-check it after updates.

What do Microsoft's benchmarks claim?

Microsoft's launch post leads with a 36-benchmark comparison covering 147,137 questions "kept blind from training" (Microsoft). Here's the core of the table it published:

ModelAverage accuracyMedian latencyCalibration (100 = perfect)
Microsoft-Decision-183.5%85 ms p50, 125 ms p9592.2
Jev 1.13.082.3%240 ms93.7
Quyet-1.0-Large81.9%380 ms93.1
GPT-6 Luna Decisions79.4%300 ms89.9
H2O-Lightning-4B v1.177.2%210 ms91.8
Strands-Decider 2B (23 of 36 benchmarks)54.8%Not measuredNot scored
GPT-6 Sol (reference)Not ranked3.01 sNot scored

Source: Microsoft's announcement. Microsoft also reports that the model changes its answer on 1.3% of perturbations on average, with zero flips when options are paraphrased, reversed or shuffled. Safety testing covered 5,250 requests across 11 benchmarks.

Read the footnotes before you quote those numbers. Every competitor's latency comes from the JevBench v1.6.1 adjusted median, checked October 7. Microsoft-Decision-1's 85 ms was "measured through Foundry in the same region." That's not the same race. And Microsoft comes third on calibration, 1.5 points behind Jev, which matters more than the accuracy lead if you plan to auto-act above a confidence threshold.

Jev wasn't in the table at first. The most-liked reply under Nadella's launch post, with 335 likes, was:

"How is Jev not ranked? That's literally the #1 model people are going to compare this to."

The announcement now carries an editor's note saying it "was updated from the original to add benchmarks for Jev on accuracy and calibration." Credit to Microsoft for fixing it fast. My Jev review has more on the model it's being measured against.

The internal results are thinner but interesting. Microsoft says its Copilot team found it "competitive with GPT5.6 Luna and 100 times faster," and its XBOX Research team ran more than 10,000 feedback items through it at "over 14 times faster and 200 times less expensive" than GPT-6 Sol (Microsoft). Those are Microsoft's teams grading Microsoft's model, so I'd treat them as use-case signals, not proof.

What has been measured independently?

Not much yet, which is normal two days after launch. Here's everything I could check from outside Microsoft.

OpenRouter publishes its own latency for every model it serves. Over three days it measured a best-provider p50 of 0.19 seconds, an average p50 of 0.27 seconds and a p99 of 0.67 seconds, with 99.90% availability (OpenRouter).

Hand-drawn bar chart of median latency: Microsoft's own test on Foundry in the same region 85 ms, OpenRouter best provider 190 ms, OpenRouter average 270 ms, with a note that OpenRouter's p99 is 670 ms
Hand-drawn bar chart of median latency: Microsoft's own test on Foundry in the same region 85 ms, OpenRouter best provider 190 ms, OpenRouter average 270 ms, with a note that OpenRouter's p99 is 670 ms

That gap doesn't mean Microsoft's number is wrong. It's measured next to the model, while OpenRouter's includes a hop through a router. But if you call it over the public internet, plan for roughly 200 to 300 ms, not 85. For a single routing call on a ticket, that's still instant. For an agent that chains 20 decisions, Microsoft's own math applies: "adding just 100 milliseconds to each of 20 sequential decisions adds two seconds to the overall workflow."

The only hands-on comparison I found came from a Hacker News commenter who ran it against Jev:

Hacker News

"In my minimal suite it was cheaper (by 0.72x) but higher latency (283ms vs 369ms p50) than Jev. Results were very comparable across all scenarios I measure, first of these that I have tested that actually justifies its existence as a commercial release."

That's one person's small suite, but it lines up with the shape of the story: roughly Jev-level answers, at a Jev-level price.

What's missing: the community-run Jev Decision Index updated to v0.3.1 on October 10 with 117 entries, and Microsoft-Decision-1 isn't one of them. Jev sits at 60.11 on that board's new Full score, so there's a clear bar waiting. Usage is early too. OpenRouter's decision rankings this week show Jev 1.13 at 999M requests, GPT-6 Luna Decisions at 11.7M, and Microsoft-Decision-1 outside the top 10.

What are the limits and unknowns?

This is the section I'd read twice before putting it in production. Some of these are deliberate scope choices and some are just gaps in launch-week docs.

  • No explanations. It returns scores, never reasons. If a support lead asks "why did this go to billing?", the model can't say.
  • Text only, 32K context. No images or screenshots, unlike Clef or OpenAI's Decisions API. A long email thread with quoted history can hit the cap.
  • No open weights. The Hugging Face URL returns a 404, so self-hosting and fine-tuning aren't options.
  • No published regions, rate limits or quotas. It's absent from Azure's region availability page, and the catalog's license tab didn't render for signed-out visitors, so I couldn't read the license terms.
  • Not on Azure's price sheet yet. The $0.042 rate is in the announcement and on OpenRouter, but the Foundry Models pricing page has no row for it. I broke down what that means in my Microsoft-Decision-1 pricing post.
  • Out-of-scope uses. Microsoft says it "should not be used as the sole basis for decisions involving credit, employment, housing, insurance, education, healthcare, legal rights" (Foundry catalog).

The first point drew the sharpest comment in the HN thread, which hit 195 points and 65 comments:

Hacker News

"Using the output of a "decision" model without insight into the reasoning for a given decision seems very trusting."

I agree, with one twist. You don't need the model to explain itself if you log what it saw, what it picked and with what confidence, and send the low-confidence cases to a person. That's the same discipline I'd apply to stopping AI hallucinations in a support bot: test on real history, gate the risky calls, keep a trail.

How does Microsoft-Decision-1 compare with Jev, Clef and the rest?

Microsoft-Decision-1 lands in a crowded field. The quick version of who's who, checked October 11, 2026:

ModelMakerInput price per 1MContextInputsOpen weightsWhere you run itJev Decision Index v0.3.1
Microsoft-Decision-1Microsoft$0.042, output free32,768TextNoFoundry, OpenRouterNot listed
Jev 1.13TypeSafe$0.042, output free64K per requestTextNoTypeSafe API, Workers AI, OpenRouter60.11 (joint 2nd)
ClefCloudflare$0.2465,536Text, JSON, images, videoYes, Apache 2.0Workers AI53.08 (19th)
Clef-flashCloudflare$0.03824,576Text, JSON, images, videoYes, Apache 2.0Workers AI47.61 (33rd)
OpenAI Decisions APIOpenAI$0.10, output freeGPT-6 Luna's 1,050,000Text, imagesNoOpenAI API, public betaNot listed
Strands Decider 2BAWS Strands LabsFree to self-hostRun at 3,072 to 4,096TextYes, Apache 2.0Your own hardware21.75 (84th)

Sources: each vendor's own docs (for example Cloudflare's Clef page and OpenAI's Decisions guide) and the Decision Index data from October 10.

A few things jump out. Microsoft matched Jev's price exactly, and only Clef-flash is cheaper per token. Clef and OpenAI's API take images, which Microsoft doesn't. And OpenAI calls its yes/no type predicate rather than noul, so moving between them isn't a pure URL swap. I covered that API's costs in my Decisions API pricing breakdown.

The price comparison in Microsoft's own post drew pushback. Microsoft's cost chart puts Microsoft-Decision-1 at about $11 per million classified texts against about $2,434 for GPT-6 Sol. One HN commenter pointed out what's missing:

Hacker News

"Microsoft only compares the price of theirs to GPT Sol(!), not GPT Terra, or GPT Luna (which is what OpenAI's Jev wannabe is based on), and certainly not Jev (4/10 the cost of Luna)."

Fair point. Against a general model like GPT-6 Luna, any decision model looks cheap. Against its real peers, Microsoft-Decision-1 is priced level, not lower. If you're weighing the whole field, my Jev alternatives roundup goes model by model.

Is Microsoft-Decision-1 the right pick for your job?

Pick the situation closest to yours.

Microsoft-Decision-1 fit check

What does your job look like?

Strong fit

This is the case Microsoft-Decision-1 was built to win. It's GA in Foundry, billed on your Azure subscription and covered by Azure SLAs, at the same $0.042 per million input tokens as Jev. No new vendor review.

Watch for: the Foundry request format and regional availability aren't published yet, so confirm both before you commit a launch date.

Look elsewhere

Microsoft-Decision-1 is text-only. Cloudflare's Clef and Clef-flash read images and video, and OpenAI's Decisions API takes images as inline base64.

Watch for: Clef costs $0.24 per million input tokens, about 5.7x Microsoft's rate.

Pair it with something

The model returns scores and never a rationale. Log the input, the options and the probabilities yourself, and hand low-confidence cases to a person or an LLM that can explain its answer.

Watch for: an audit trail is your job, not the model's.

Out of scope

Microsoft says the model should not be the sole basis for decisions about credit, employment, housing, insurance, education, healthcare or legal rights. Keep a human making those calls.

Watch for: "sole basis" still allows it as one input, but your compliance team will want that written down.

Look elsewhere

There are no public weights. Strands Decider 2B and Cloudflare's Clef models are Apache 2.0 and can run on your own hardware.

Watch for: Strands has no hosted version, so you own the serving.

Fits the sorting half

Urgency, team and sentiment are fixed-option questions, and it answers them in milliseconds. Someone still has to tag the ticket, move it and write the reply. An AI helpdesk teammate does both halves inside Zendesk, Freshdesk or Gorgias.

Watch for: test any threshold on your own past tickets before it touches live traffic.

Based on Microsoft's model card and announcement, OpenRouter's model page and each rival's own docs, checked October 11, 2026.

If raw speed is your main criterion, my Jev Ultrafast write-up covers TypeSafe's low-latency tier, the closest rival to Microsoft's speed pitch.

What does Microsoft-Decision-1 mean for support teams?

Support is the use case everyone reaches for first, and with good reason. Ticket triage is a pile of small, bounded questions with fixed answers. Intent classification is a choice question. Urgency and sentiment analysis are scores on a fixed scale. "Should a human see this?" is a yes/no, the heart of AI escalation. Microsoft-Decision-1 can answer all of those in one call for a fraction of a cent.

Here's what I'd want every support lead to know before they get excited, though: sorting is the part that already works well.

Hand-drawn diagram: a new ticket splits into two columns. The decision model column, Sorting (the easy half), lists tag, route, urgency and spam or not, with 93% accurate at the bottom. The AI teammate column, Answering (the hard half), lists look up the order, check the policy, write the reply and act in the helpdesk, with 12% sent as-is at the bottom
Hand-drawn diagram: a new ticket splits into two columns. The decision model column, Sorting (the easy half), lists tag, route, urgency and spam or not, with 93% accurate at the bottom. The AI teammate column, Answering (the hard half), lists look up the order, check the policy, write the reply and act in the helpdesk, with 12% sent as-is at the bottom

Those numbers come from a cross-validated trial I've seen on real Zendesk traffic at an e-commerce company. AI triage was 93% accurate and caught 100% of the spam, which made up 22% of the inbox, with zero false positives. Drafted replies were pointed in the right direction 88% of the time. But only 12% went out without edits, and 7% contained a factual error. The label was easy. The answer was the job.

That matches what eesel's buyers say they actually want. A CX lead running 7,000 tickets a month put it this way on a call:

"The AI will never be able to answer 100% of the questions, but if it tries and just answers 'sorry I don't know this,' I cannot go and check all my 7,000 tickets to see if the AI actually made a good answer."

What they're asking for is a confident decision first (is this one I can handle?) and an action second. A decision model gives you the first part as a number. Turning that number into a tag in Zendesk, a reassignment, a refund held for approval and a reply in your brand voice is everything around it. That's why I'd treat Microsoft-Decision-1 as a component for teams building their own AI ticket routing, not a support tool on its own.

If you are building it yourself, a sensible order is:

  1. Start with the questions your triagers already answer. Team, urgency, spam, ticket prioritization level. Write them as choice and score questions with the same labels your helpdesk uses.
  2. Replay a month of past tickets. Compare the model's picks against what your team actually did, per label, not as one average.
  3. Set a threshold per question. Auto-route above it, send to a person below it, and log both.
  4. Keep plain rules where they work. If "refund" in the subject always means billing, Zendesk routing rules are free and never drift.
  5. Decide who writes the reply. That's the part a decision model can't do, and the part your customers actually see.

For a broader look at the options, my roundup of AI triage tools compares what's built into each helpdesk today, and the guide to automate ticket triage walks the setup end to end.

Try eesel for ticket triage and replies

Microsoft-Decision-1 is infrastructure: a fast, cheap way for your code to make a call. eesel is the employee that makes the call and then does the work. The AI helpdesk teammate joins your existing queue through the Zendesk integration, Freshdesk or Gorgias. It tags, routes and escalates, then writes the reply from your help center and past tickets.

The split this post keeps coming back to is already built in. Each automation has a "Model strength" setting, with "a lighter one for simple jobs like tagging or routing" and the usual model for drafting (eesel docs). Every action can be set to Auto, Needs approval or Disabled, and events the agent decides to skip are logged with a reason. Before anything goes live, you can simulate the agent against your own past tickets and see where it would have matched your team's replies.

eesel activity view for a Zendesk agent listing conversations it worked with Pending and Resolved statuses and links to each Zendesk ticket, next to the eesel chat panel
eesel activity view for a Zendesk agent listing conversations it worked with Pending and Resolved statuses and links to each Zendesk ticket, next to the eesel chat panel

If you'd rather wire it up from a terminal, the eesel CLI runs the same teammate the dashboard does. Every command prints JSON, so Claude Code, Codex or Cursor can drive the whole setup. A command like eesel automations enable <platform> <key> --instructions "..." sets up a triage rule in plain words, eesel approvals lists held actions to approve or deny, and --dry-run checks a write before it happens. I wrote more about why that matters in my piece on the agent CLI.

Pricing is a fixed monthly credit plan: one ticket is one credit however much work happens inside it, starting at $299 a month for 500 credits, and you get 100 free credits to try it with no card. Connect your helpdesk, point it at last month's tickets, and see how it sorts them and what it would have written back. Try eesel.

Frequently Asked Questions

What is Microsoft-Decision-1?

Microsoft-Decision-1 is a decision model Microsoft released on October 9, 2026. You give it some text and a set of fixed answer options, and it returns a calibrated probability for each option as JSON instead of writing text. It is one of a fast-growing group of decision models built for routing, classification and verification calls.

What is Microsoft-Decision-1 built on?

Microsoft says it post-trained Alibaba's open-weight Qwen3.5-9B for single-pass decision scoring and plans to rebase the model on Microsoft AI (MAI) and OpenAI models later. The Qwen family also sits under Cloudflare's Clef-flash and AWS's Strands Decider.

How do I use Microsoft-Decision-1?

There are two routes: deploy it from the Microsoft Foundry catalog, where it is sold Direct from Azure and billed to your Azure subscription, or call it through OpenRouter's Decisions API with the model ID microsoft/microsoft-decision-1. Chat completions SDKs do not work with it, so plan for a typed request the way you would for TypeSafe Jev.

How much does Microsoft-Decision-1 cost?

Microsoft-Decision-1 costs $0.042 per million input tokens and output tokens are free, the same headline rate as Jev. Azure's Foundry pricing page has no row for it yet. For the rival's cost math, see my Jev pricing breakdown.

Is Microsoft-Decision-1 better than Jev?

On Microsoft's own 36-benchmark table, Microsoft-Decision-1 averages 83.5% accuracy against 82.3% for Jev 1.13.0, while Jev scores higher on calibration (93.7 vs 92.2). No independent board has scored Microsoft-Decision-1 yet, so treat the lead as a claim. My Jev review covers what Jev does well.

Can Microsoft-Decision-1 triage support tickets?

Yes. Urgency, team and sentiment calls are exactly the fixed-option questions it is built for, and OpenRouter's own sample request is a support message. You still need something to apply the tag, move the ticket and write the reply, which is the work of an AI helpdesk teammate. See automating ticket triage for the full flow.

Does Microsoft-Decision-1 explain its decisions?

No. Microsoft's model card says it is text-only and does not generate explanations or rationales. If you need a reason with each decision, for an audit trail or a human escalation, you have to log the inputs and options yourself or use an LLM for that step.

Are Microsoft-Decision-1 weights open?

No. There is no Hugging Face repository for Microsoft-Decision-1, and Microsoft has not published a license or self-hosting option. If you need open weights, Cloudflare's Clef and AWS's Strands Decider are Apache 2.0.

Share this article

Kira

Article by

Kira

Kira is a writer at eesel AI with a Computer Science background and over a year of hands-on experience evaluating AI-powered customer service tools. She focuses on breaking down how helpdesk platforms and AI agents actually work so that support teams can make better buying decisions.

Related Posts

All posts →
Hand-drawn hero banner of a person at a laptop sending documents into a small decision box that sorts them into three labelled trays while a second person looks on
Trending

Strands Decider 2B: AWS's free decision model, tested against the hype

Strands Decider 2B is a free, open 1.9B decision model from AWS's Strands Labs. What it does, how accurate and fast it is, what it costs to run, and where it fits.

Rama AdiRama AdiOct 8, 2026
Hand-drawn hero banner of a person feeding questions into a switchboard that routes them into labeled lanes, each with a confidence dial
Trending

Decision models explained: Jev, Clef, Strands Decider and the new AI category

Decision models return typed answers with confidence scores instead of text. What they are, how they work, every model you can use today, and where they fit in support.

KiraKiraOct 6, 2026
Cloudflare Clef hero banner in Cloudflare orange, two people looking at a decision model connected to users, websites, devices and a list of options
Trending

What is Cloudflare Clef? Cloudflare's decision model, explained

Cloudflare Clef explained: what the decision model does, how Clef differs from Clef-flash, how to run it on Workers AI or Ollama, and where it fits in ticket routing.

KiraKiraOct 7, 2026
Microsoft-Decision-1 pricing hero banner in Microsoft blue, a developer sending inputs to a model that returns scored options
Trending

Microsoft-Decision-1 pricing (2026): $0.042 per million tokens, explained

Microsoft-Decision-1 pricing explained: $0.042 per million input tokens, free output, what a decision really costs, and how it stacks up against Jev, Clef-flash and OpenAI.

Rama AdiRama AdiOct 11, 2026
Hand-drawn hero banner in PostHog amber showing a hedgehog butler weighing inputs before picking an outcome, with two developers looking on
Trending

PostHog Jeeves: the open decision model that thinks before it picks

PostHog Jeeves is an open 9B decision model that writes a reasoning chain before it answers. Here is what it beats Jev at, what it costs you in latency, and where it fits.

KiraKiraOct 1, 2026
A person holding a shield beside floating content cards and a checkmark panel, with the Mistral mark on an orange background
Trending

Shieldstral: accuracy is settled, packaging decides

Shieldstral ties a 20B model on text safety at 3B. The top four guard models sit inside 1.6 F1 points, so what actually picks your guard is hosting, licence, reasons, and how many calls one message costs.

Rama AdiRama AdiAug 18, 2026
A person talking to a lifelike lip-synced AI video avatar on a screen, illustrating a Gemini 3.8 Live Avatar review
Trending

Gemini 3.8 Live Avatar review: is the talking AI face worth it?

A hands-on review of Google's Gemini 3.8 Live Avatar: what the lip-synced AI face does well, the real per-minute cost, and where it actually fits support.

Rama AdiRama AdiSep 27, 2026
A person talking to a lifelike AI video avatar on a screen, illustrating Gemini 3.8 Live Avatar
Trending

Gemini 3.8 Live Avatar: what it actually does for customer support

Google's Gemini 3.8 Live Avatar puts a lip-synced AI face on enterprise agents. I break down how it works, the real per-minute cost, and where it fits support.

KiraKiraSep 27, 2026
Illustrated hero banner for TypeSafe Jev, the ultrafast System One AI model, with a speed gauge
Trending

Is Jev really ultrafast? TypeSafe's System One model, tested

TypeSafe calls Jev an ultrafast System One model at 70-500ms a decision. Here is what the speed claim really means, where it holds up, and where it does not.

Rama AdiRama AdiSep 22, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free