7 Microsoft-Decision-1 alternatives in 2026, compared on price and fit

Kurnia Kharisma
Written by

Kurnia Kharisma

Katelin Teen
Reviewed by

Katelin Teen

Last edited October 11, 2026

Expert Verified
Hand-drawn hero banner in Microsoft blue: a developer points at a row of five option cards, the first one picked, next to the Microsoft logo

Microsoft-Decision-1 alternatives at a glance

ModelVendorInput per 1M tokensOutputContext per callImagesHow you run itWeights and licenseDecision Index v0.3.1 (Full)Data and compliance notesFree allowance
Microsoft-Decision-1 (baseline)Microsoft$0.042Free32,768NoAzure Foundry, OpenRouterClosedNot listedAzure SLA; regions not publishedGeneral Azure $200, 30 days
Jev 1.13TypeSafe$0.042Free64k per request, 32k state + longest questionNoTypeSafe API, Cloudflare Workers AIClosed60.11 (reference, tied 2nd)Zero data retention on the Cloudflare listingNone documented
OpenAI Decisions APIOpenAI$0.10Free1,050,000 (2x input price past 272K)YesOpenAI API, public betaClosedNot listedZDR and HIPAA for eligible customers, US and EU residencyNone
Clef-flashCloudflare$0.038Free24,576YesWorkers AI, or self-hostOpen, Apache 2.047.61 (#33)Cloudflare account terms10,000 Workers AI neurons a day
ClefCloudflare$0.24Free65,536YesWorkers AI, or self-hostOpen, Apache 2.053.08 (#19)Cloudflare account terms10,000 Workers AI neurons a day
Perplexity Decider v1.1PerplexityYour GPU costn/a8,192 per decisionYesSelf-host, about 49 GiB of weightsApache 2.0, authenticated download62.75 (#1)Stays on your infrastructureFree download
Rune 26B-A4B v3Surogate (Invergent)Your GPU costn/a262kYesSelf-host with the surogate engineApache 2.0, contact-info gate57.43 (#10)Stays on your infrastructureFree download
Strands Decider 2BAWS Strands LabsYour GPU costn/aTested at 4,096Optional, untrainedSelf-host, fits a 3090Apache 2.021.75 (#84)Stays on your infrastructureFree download
eesel (not a model)eesel$299 a month for 500 tickets or chatsn/aWhole ticket and historyYesPlugs into your helpdeskHosted productn/aHIPAA on Enterprise100 free credits

Prices and context come from each vendor's own docs, checked October 11, 2026. Decision Index scores come from the Jev Decision Index board, version 0.3.1, updated October 10.

Why look past Microsoft-Decision-1?

I wrote up what the model does in my Microsoft-Decision-1 overview and what it costs in the pricing breakdown. It is a strong release. Microsoft's own table puts it at 83.5% average accuracy across 36 benchmarks with an 85 ms median on Foundry. But a few gaps keep coming up:

  • No open weights. It is a post-trained Qwen3.5-9B, but the Hugging Face page Microsoft would use returns 404. You rent it or you don't use it. If you liked the Qwen family for self-hosting, this one is closed.
  • Text only, 32K context. Screenshots, scanned forms and very long threads need another model.
  • Thin Azure paperwork. It is GA in the catalog, but it has no row on Azure's Foundry pricing page yet, and no published regions, quotas or batch rate. Teams buying through Azure models usually want those before they commit.
  • No independent score yet. It is not on the Jev Decision Index, so every accuracy number so far is Microsoft's own.

That last point drove most of the launch-day reaction. The top reply under Satya Nadella's launch post was blunt:

"How is Jev not ranked? That's literally the #1 model people are going to compare this to."

Microsoft later updated the post to add Jev to the accuracy and calibration table. Still, the question underneath it is the right one. Decision models are a crowded field now, and the decision model you pick should depend on where it runs and what it reads, not on launch-day charts.

Hand-drawn 2x2 quadrant with Hosted API on the left, You host it on the right, Text only at the bottom and Text + images at the top: Microsoft-Decision-1 and Jev sit in hosted text only; OpenAI Decisions API, Clef-flash and Clef in hosted with images; Perplexity Decider, Rune v3 and Strands Decider in self-hosted with images
Hand-drawn 2x2 quadrant with Hosted API on the left, You host it on the right, Text only at the bottom and Text + images at the top: Microsoft-Decision-1 and Jev sit in hosted text only; OpenAI Decisions API, Clef-flash and Clef in hosted with images; Perplexity Decider, Rune v3 and Strands Decider in self-hosted with images

How I picked these alternatives

I come at this from the search side. My job at eesel is mostly figuring out what buyers are actually asking when they type a query, and "Microsoft-Decision-1 alternatives" two days after launch splits cleanly into price, hosting and input type. So I kept to models that do the same job: you send a state plus fixed answer options, and you get calibrated probabilities back in one pass.

Every model below either appears in Microsoft's own comparison table or ranks on the Decision Index, or both. I left out general chat models with JSON mode, since function calling on a big model is a different cost curve entirely. Each entry uses the same five headings so you can scan across them.

How much text can each one read per call?

Context is the spec that bites first in support work. A single ticket with order history, a policy excerpt and three back-and-forth replies can pass 10,000 tokens.

Hand-drawn bar chart of context window per call in tokens: Clef-flash 24,576, Microsoft-Decision-1 32,768, Jev 64k, Clef 65,536, Rune v3 262k, OpenAI Decisions API 1,050,000 with a break mark showing the bar is cut short
Hand-drawn bar chart of context window per call in tokens: Clef-flash 24,576, Microsoft-Decision-1 32,768, Jev 64k, Clef 65,536, Rune v3 262k, OpenAI Decisions API 1,050,000 with a break mark showing the bar is cut short

Perplexity Decider is missing from that chart on purpose. Its card says the text and all images of one decision must fit in 8,192 tokens, and longer inputs are rejected rather than truncated. Top score, smallest window.

1. Jev 1.13 (TypeSafe)

TypeSafe docs Models page listing Jev 1.13 at $42 per billion or $0.042 per million tokens, 100K tokens per second and 80 requests per second, a 64k context per request and text-only input, as taken from TypeSafe's docs
TypeSafe docs Models page listing Jev 1.13 at $42 per billion or $0.042 per million tokens, 100K tokens per second and 80 requests per second, a 64k context per request and text-only input, as taken from TypeSafe's docs

What it is

Jev is the model that started the category. TypeSafe serves it through one endpoint, POST /v1/systemone, and it scores every question you send against a shared state in parallel. I covered it in depth in my Jev overview.

Where it beats Microsoft-Decision-1

  • Bigger requests. Jev takes 64k tokens per request, with 32k for the state plus the longest single question, per TypeSafe's models page. Microsoft caps at 32,768.
  • Better calibration. Even on Microsoft's own table, Jev scores 93.7 on calibration against Microsoft's 92.2.
  • More places to buy it. Besides TypeSafe's API, Jev is on Cloudflare Workers AI at the same price, with free cached input and zero data retention.

Where it falls short

  • Text only, like Microsoft. Images must be turned into text first.
  • TypeSafe warns that "Rate limits are adjusting dynamically" and can change without notice. Today's limit is 100K tokens and 80 requests per second.
  • Microsoft's table has it slightly behind on accuracy, 82.3% to 83.5%.

One Hacker News tester ran both through their own suite the day after launch:

Hacker News

"In my minimal suite it was cheaper (by 0.72x) but higher latency (283ms vs 369ms p50) than Jev. Results were very comparable across all scenarios I measure, first of these that I have tested that actually justifies its existence as a commercial release."

Same list price, different bills: tokenizers count differently, so test on your own tickets.

Pricing

$0.042 per million input tokens, output free, no documented free tier. My Jev pricing post works through cost per decision, and Jev Ultrafast covers the speed tier.

Best for

Teams that want a Microsoft-Decision-1 alternative with the same price and API shape, longer inputs, and no Azure dependency.

2. OpenAI Decisions API (gpt-6-luna)

OpenAI developer docs page for the Decisions API, describing it as turning text and images into decisions 10x faster than the Responses API, with a public beta notice naming gpt-6-luna as the only model, as taken from OpenAI's docs
OpenAI developer docs page for the Decisions API, describing it as turning text and images into decisions 10x faster than the Responses API, with a public beta notice naming gpt-6-luna as the only model, as taken from OpenAI's docs

What it is

OpenAI's take on the same idea, released in public beta on October 6. You call POST /v1/decisions with gpt-6-luna, an input of text and images, and a list of questions of type predicate, choice or score. My Decisions API guide has the walkthrough.

Where it beats Microsoft-Decision-1

  • Images. It reads text and images in one call. Images must be inline base64 data URLs.
  • Enormous context. It inherits GPT-6 Luna's 1,050,000-token window, so a whole ticket history fits.
  • Compliance options. OpenAI's guide lists Zero Data Retention and HIPAA for eligible customers, plus US and EU data residency.

Where it falls short

  • It costs about 2.4x more per token than Microsoft.
  • Still beta. OpenAI says it expects GA "in the coming weeks".
  • Prompts over 272K tokens bill at 2x input, and regional processing premiums apply.
  • OpenAI publishes no millisecond latency, only "about 10x faster than the Responses API". Microsoft's table measured it at 300 ms median and 79.4% accuracy.

Pricing

$0.10 per million input tokens, with no output, cache-read or cache-write charges. That rate lives in the Decisions guide, not yet on the main OpenAI API pricing page. More in my Luna pricing post.

Best for

Teams already on OpenAI who need screenshots or long threads scored, or who need HIPAA paperwork in place.

3. Cloudflare Clef-flash

Cloudflare Workers AI model page for clef-flash, a 9B multimodal decision model, showing a 24,576-token context window, vision support and $0.038 per million input tokens, as taken from Cloudflare's docs
Cloudflare Workers AI model page for clef-flash, a 9B multimodal decision model, showing a 24,576-token context window, vision support and $0.038 per million input tokens, as taken from Cloudflare's docs

What it is

Cloudflare's small decision model, a 9B built on the same Qwen3.5-9B base as Microsoft's. It runs on Workers AI as @cf/cloudflare/clef-flash, and the weights are on Hugging Face under Apache 2.0. My Clef overview covers both Clef sizes.

Where it beats Microsoft-Decision-1

  • Cheapest hosted rate. $0.038 per million input tokens, down from $0.09 a week ago.
  • Images. Cloudflare's page lists vision support, with images billed as input tokens.
  • Free daily allowance. Workers AI gives 10,000 neurons a day on Free and Paid plans. By my math that is roughly 2.9M Clef-flash input tokens a day.
  • Open weights, so you can move it in-house later.

Where it falls short

  • The context window shrank to 24,576 tokens with the price cut, smaller than Microsoft's.
  • It ranks 33rd on the Decision Index (47.61), well below Clef and Jev.
  • It is not in Microsoft's comparison table, so there is no shared benchmark between the two.

Pricing

$0.038 per million input tokens and no output charge. Past the free neurons, Workers AI bills $0.011 per 1,000 neurons.

Best for

High-volume text or image classification where cost per call matters most and inputs stay short.

4. Cloudflare Clef

Cloudflare Workers AI model page for clef, a 27B multimodal decision model, showing a 65,536-token context window, vision support and $0.24 per million input tokens, as taken from Cloudflare's docs
Cloudflare Workers AI model page for clef, a 27B multimodal decision model, showing a 65,536-token context window, vision support and $0.24 per million input tokens, as taken from Cloudflare's docs

What it is

The bigger Clef: a 27B multimodal decision model post-trained from Qwen3.8-27B, hosted on Workers AI and published under Apache 2.0.

Where it beats Microsoft-Decision-1

  • Twice the context. 65,536 tokens.
  • Images and video in the state, per Cloudflare's model card.
  • An independent score. It is on the Decision Index at 53.08 Full, and its public-benchmark-only score of 61.71 is the highest public number among the hosted options here.

Where it falls short

  • At $0.24 per million input tokens, it costs nearly 6x Microsoft's rate.
  • Its private-test score drops it to 19th, so the public number flatters it.
  • Cloudflare's self-reported median was 209 ms, slower than Microsoft's claimed 85 ms.

Pricing

$0.24 per million input tokens, no output charge, and the same 10,000 free neurons a day (about 458K Clef input tokens). My Clef pricing post has the math.

Best for

Teams on Cloudflare who need longer multimodal inputs and will pay more per call for them.

5. Perplexity Decider v1.1 (27B)

Hugging Face model card for perplexity-ai/pplx-decider-v1.1-27b under an Apache 2.0 license, with a table showing its overall Decision Index of 61.56 against Jev's 57.9, as taken from Hugging Face
Hugging Face model card for perplexity-ai/pplx-decider-v1.1-27b under an Apache 2.0 license, with a table showing its overall Decision Index of 61.56 against Jev's 57.9, as taken from Hugging Face

What it is

Perplexity's open decision model, a full fine-tune of Qwen3.8-27B with a separate decision head. Perplexity says v1.1 gained most of its improvement from "the lifting of the causal mask" and more training data. You download it from Hugging Face and run it with the included code.

Where it beats Microsoft-Decision-1

  • Top independent score. It ranks first on Decision Index v0.3.1 at 62.75, ahead of Jev's 60.11.
  • You own the weights. Apache 2.0, nothing leaves your servers.
  • Images, passed as file paths, PIL images or base64 URLs.

Where it falls short

  • Hard 8,192-token limit per decision, text and images combined.
  • You need a CUDA GPU with room for about 49 GiB of weights plus working memory, and an authenticated Hugging Face account to download.
  • No hosted API: the card says no inference provider serves it.
  • Default causal inference does not reproduce the evaluated setup, so you must use Perplexity's own code or a careful export.

Pricing

Free to download. Your cost is GPU time. On cost, one Hacker News commenter put it well:

Hacker News

"Open weight model pricing isn't really up to the company that built them (unless they are also serving them, in which case their API price may differ). It's really up to the serving cost and/or pricing of whoever is actually serving the model."

Best for

ML teams with spare GPU capacity who want the most accurate open model and keep inputs short.

6. Surogate Rune 26B-A4B v3

Hugging Face model card for surogate/rune-26b-a4b-GGUF, a Gemma 4 based decision model under Apache 2.0, showing a contact-information gate and the Surogate Rune banner with choice, noul and score question types, as taken from Hugging Face
Hugging Face model card for surogate/rune-26b-a4b-GGUF, a Gemma 4 based decision model under Apache 2.0, showing a contact-information gate and the Surogate Rune banner with choice, noul and score question types, as taken from Hugging Face

What it is

Invergent's open decision model, a full fine-tune of Gemma 4 26B-A4B (8 of 128 experts active per token). It answers choice, noul and score questions and is served fastest by the open-source surogate engine.

Where it beats Microsoft-Decision-1

  • Long inputs. The card claims a 262k context, eight times Microsoft's.
  • Images, with the vision tower kept. Invergent's own image test scored 83.4% at 1,120 image tokens.
  • Optional thinking. With "thinking": true, questions it is unsure about reason before answering. Invergent estimates that lifts its index by about 1.8 points, at about 5 seconds per thinking question.
  • On Microsoft's own table it scored 79.7%, close to OpenAI's Luna.

Where it falls short

  • Downloads sit behind a contact-information form.
  • Its default probabilities run overconfident. The card recommends reading answers at temperature 2 to fix calibration.
  • Invergent measured a 388 ms median per request on one RTX PRO 6000, slower than Microsoft's claim.
  • 10th on the Decision Index (57.43), behind Perplexity and Jev.

Pricing

Free to download under Apache 2.0. You pay for the hardware.

Best for

Self-hosters who need long or image-heavy inputs and can tune serving themselves.

7. Strands Decider (AWS Strands Labs)

Hugging Face model card for StrandsAgents/strands-decider-2B-qwen3.5-v1-2610 under Apache 2.0, describing it as a small, fast decision model for agentic AI, as taken from Hugging Face
Hugging Face model card for StrandsAgents/strands-decider-2B-qwen3.5-v1-2610 under Apache 2.0, describing it as a small, fast decision model for agentic AI, as taken from Hugging Face

What it is

A small open decision model from AWS's Strands Labs, built for agent steps like tool selection, guardrails and routing. The reference model is a 2B on Qwen3.5, and four Gemma 4 sizes followed on October 10 and 11. I covered the launch in my Strands Decider post.

Where it beats Microsoft-Decision-1

  • Runs on modest hardware. The 2B hit a 115 ms median on a single RTX 3090 in Strands' own test.
  • Free and open, Apache 2.0, and it fits inside the wider Strands harness for agents.
  • Optional image input with --vision, though Strands notes nothing was trained on images.

Where it falls short

  • Accuracy is far lower. Microsoft's table shows 54.8% on only 23 of 36 benchmarks, and the Decision Index ranks it 84th (21.75).
  • It was tested at 3,072 and 4,096-token windows.
  • There is no hosted API anywhere, Bedrock included.

Pricing

Free. Your own GPU or CPU.

Best for

Cheap, local agent guardrails and tool picks where speed on small hardware beats top accuracy.

8. eesel (if you want the decision and the work)

eesel activity view for a Zendesk agent listing the conversations it worked, with Pending and Resolved statuses and links to each Zendesk ticket, next to the eesel chat panel
eesel activity view for a Zendesk agent listing the conversations it worked, with Pending and Resolved statuses and links to each Zendesk ticket, next to the eesel chat panel

What it is

Not a model. The seven options above are infrastructure. eesel is the employee: an AI helpdesk teammate that joins your Zendesk, Freshdesk or Gorgias queue, makes the routing call, and then writes the reply or takes the action. It learns from your past tickets and help center.

Where it beats Microsoft-Decision-1

  • No pipeline to build. No thresholds, webhooks or helpdesk API glue.
  • It acts on the answer. Tags, priority, escalation and the reply itself.
  • Risky actions can wait for approval, so a human signs off before anything irreversible.

Where it falls short

  • Far more expensive per ticket than a single model call, because it includes the work.
  • It lives in your helpdesk. If you need decisions inside your own app's code path, use a model.

Pricing

eesel's plans start at $299 a month for 500 credits, where one ticket or one chat is one credit however long it runs. 5,000 credits is $1,749 a month. There is a free start with 100 credits and no card.

Best for

Support teams whose real goal is tickets sorted and answered, not a model to wire up.

Which Microsoft-Decision-1 alternative should you pick?

Hand-drawn decision tree starting from Leaving Microsoft-Decision-1?, asking Need images?: yes leads to Hosted: OpenAI or Clef-flash and Own GPUs: Rune v3; no leads to Same price, API: Jev, Cheapest: Clef-flash and Small GPU: Strands Decider; a separate card below reads No code: AI helpdesk teammate
Hand-drawn decision tree starting from Leaving Microsoft-Decision-1?, asking Need images?: yes leads to Hosted: OpenAI or Clef-flash and Own GPUs: Rune v3; no leads to Same price, API: Jev, Cheapest: Clef-flash and Small GPU: Strands Decider; a separate card below reads No code: AI helpdesk teammate

My short version:

  • Same job, no Azure: Jev at the same price. Compare the two on your own data first.
  • Lowest bill: Clef-flash, if your inputs fit in 24,576 tokens.
  • Images or very long threads: OpenAI's Decisions API hosted, Rune v3 self-hosted.
  • Best open accuracy: Perplexity Decider v1.1, if 8,192 tokens is enough.
  • Small and local: Strands Decider for agent guardrails.

For support teams, the use case is almost always ticket classification and routing. That is where the probability score earns its keep: you only automate what the model is confident about. On real queues the hard part comes after. In one example from our own data, a Web3 infrastructure company's AI matched a cold contact-list sales pitch against past Zendesk tickets, classified it as spam, and drafted a polite decline as an internal note instead of trying to answer it. A decision model gives you the "spam" label. Something still has to read the history and write the note.

Watch false positives in tagging too. A 1% error rate is cheap per call and expensive when it routes a refund to sales. My roundup of the best AI for triage compares tools that handle this out of the box.

Skip the model, keep the decision

If you're comparing Microsoft-Decision-1 alternatives because you want tickets routed, tagged and prioritized, you may not need a model at all. eesel works like a new hire for your helpdesk: it plugs into Zendesk or Freshdesk in minutes, reads your past tickets, makes the same triage decisions a decision model would, and then answers the customer. You can simulate it on historical tickets before it touches a live one.

eesel Zendesk integration page showing the help center, macros and past tickets it learns from, plus triggers that run on every customer ticket message
eesel Zendesk integration page showing the help center, macros and past tickets it learns from, plus triggers that run on every customer ticket message

If you'd rather drive it from code, the eesel CLI runs the same teammate the dashboard does. Every command prints JSON, so you, a script, or a coding agent like Claude Code, Codex or Cursor can set it up. eesel integrations connects your helpdesk, eesel automations enable <platform> <key> --instructions "..." writes a triage rule in plain words, and eesel approvals lists held actions to approve or deny. More on why that matters in my agent CLI piece.

Try eesel free with 100 credits and point it at a week of your tickets. If it routes them as well as your decision pipeline would, you've saved yourself the build.

Frequently Asked Questions

What are the best Microsoft-Decision-1 alternatives?

The closest swap is TypeSafe's Jev, which costs the same $0.042 per million input tokens. OpenAI's Decisions API adds images and a huge context window, Cloudflare's Clef-flash is the cheapest hosted option at $0.038, and Perplexity Decider, Rune v3 and Strands Decider are open-weight Microsoft-Decision-1 alternatives you run yourself.

Is there a cheaper alternative to Microsoft-Decision-1?

Yes. Cloudflare's Clef-flash lists $0.038 per million input tokens against Microsoft's $0.042, with free output on both. The open-weight models cost nothing to download, but you pay for the GPU. My Clef pricing breakdown covers the free daily allowance on Workers AI.

Is Jev better than Microsoft-Decision-1?

On Microsoft's own 36-benchmark table, Microsoft-Decision-1 scored 83.5% average accuracy to Jev's 82.3%, while Jev scored higher on calibration (93.7 vs 92.2). Prices are identical. Jev reads 64k tokens per request against Microsoft's 32,768. My Jev review has more on how it behaves in practice.

Is there an open-source alternative to Microsoft-Decision-1?

Microsoft has not released weights, so you cannot self-host it. Perplexity Decider v1.1, Surogate Rune 26B-A4B v3, Cloudflare's Clef models and AWS's Strands Decider all ship under Apache 2.0. Perplexity Decider v1.1 currently tops the community Decision Index.

Which Microsoft-Decision-1 alternative supports images?

OpenAI's Decisions API, Cloudflare Clef and Clef-flash, Perplexity Decider and Rune v3 all accept images. Microsoft-Decision-1 and Jev are text only. If screenshots show up in your ticket triage flow, that alone narrows the list.

How does Microsoft-Decision-1 compare to the OpenAI Decisions API?

OpenAI charges $0.10 per million input tokens on gpt-6-luna, about 2.4x Microsoft's rate, also with free output. In exchange you get images, a 1,050,000-token window and ZDR and HIPAA options for eligible customers. The API is still in public beta. See my Decisions API pricing post.

Do I need a decision model to route support tickets?

Only if you are building routing into your own product or pipeline. If the goal is tickets sorted and answered inside your helpdesk, an AI helpdesk agent does the decision and the work without code. My guide on how to automate ticket triage walks through both paths.

Share this article

Kurnia Kharisma

Article by

Kurnia Kharisma

Kurnia is a software engineer and writer at eesel AI with two years of SEO experience, writing about AI tools, helpdesk software, and customer support. He pairs a developer's understanding of how these products are built with search-driven research into what actually ranks and resonates with the people searching for them.

Related Posts

All posts →
Hand-drawn hero banner with the Microsoft logo on a blue band, a decision flowchart splitting into two outcomes and joining at a check mark, and three people at a desk reviewing it
Trending

Microsoft-Decision-1: what it is, how it works, and where it fits

Microsoft-Decision-1 scores fixed answer options instead of writing text. How it works, how to call it, what its benchmarks claim and where it fits in support.

KiraKiraOct 11, 2026
Microsoft-Decision-1 pricing hero banner in Microsoft blue, a developer sending inputs to a model that returns scored options
Trending

Microsoft-Decision-1 pricing (2026): $0.042 per million tokens, explained

Microsoft-Decision-1 pricing explained: $0.042 per million input tokens, free output, what a decision really costs, and how it stacks up against Jev, Clef-flash and OpenAI.

Rama AdiRama AdiOct 11, 2026
Hand-drawn hero banner of a person at a laptop sending documents into a small decision box that sorts them into three labelled trays while a second person looks on
Trending

Strands Decider 2B: AWS's free decision model, tested against the hype

Strands Decider 2B is a free, open 1.9B decision model from AWS's Strands Labs. What it does, how accurate and fast it is, what it costs to run, and where it fits.

Rama AdiRama AdiOct 8, 2026
Cloudflare Clef hero banner in Cloudflare orange, two people looking at a decision model connected to users, websites, devices and a list of options
Trending

What is Cloudflare Clef? Cloudflare's decision model, explained

Cloudflare Clef explained: what the decision model does, how Clef differs from Clef-flash, how to run it on Workers AI or Ollama, and where it fits in ticket routing.

KiraKiraOct 7, 2026
Hand-drawn hero banner of a person feeding questions into a switchboard that routes them into labeled lanes, each with a confidence dial
Trending

Decision models explained: Jev, Clef, Strands Decider and the new AI category

Decision models return typed answers with confidence scores instead of text. What they are, how they work, every model you can use today, and where they fit in support.

KiraKiraOct 6, 2026
Hand-drawn hero banner in PostHog amber showing a hedgehog butler weighing inputs before picking an outcome, with two developers looking on
Trending

PostHog Jeeves: the open decision model that thinks before it picks

PostHog Jeeves is an open 9B decision model that writes a reasoning chain before it answers. Here is what it beats Jev at, what it costs you in latency, and where it fits.

KiraKiraOct 1, 2026
Hand-drawn illustration of a person speaking to a phone, glasses, watch, smart speaker and laptop, with the Cactus logo on a red background
Trending

Cactus Whistle: what the 16.9 MB speech-to-text model can really do

Cactus Whistle is a free 16.9 MB speech-to-text model that runs on a phone CPU. Here's where it beats Whisper base, where it loses, and what to build with it.

Rama AdiRama AdiOct 9, 2026
Cloudflare Clef pricing hero banner in Cloudflare orange, two people comparing the cost of two decision models
Trending

Cloudflare Clef pricing (2026): $0.24 per million tokens, explained

Cloudflare Clef pricing broken down: $0.24 per million input tokens for Clef, $0.09 for Clef-flash, no output charge, a daily free allowance, and what a decision really costs.

Kurnia KharismaKurnia KharismaOct 6, 2026
Hand-drawn illustration of a person sending a support ticket into a small router box that sends it to one of three teammates
Trending

OpenAI Decisions API: what it is, how it works, and what's new

The OpenAI Decisions API picks one answer from a fixed list, fast. Here's how it works, what's still gated, and how close you can get with GPT-6 Luna today.

KiraKiraOct 2, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free