Strands Decider 2B: AWS's free decision model, tested against the hype

Rama Adi
Written by

Rama Adi

Katelin Teen
Reviewed by

Katelin Teen

Last edited October 8, 2026

Expert Verified
Hand-drawn hero banner of a person at a laptop sending documents into a small decision box that sorts them into three labelled trays while a second person looks on

What is Strands Decider?

Strands Decider 2B is a small decision model from Strands Labs, the experimental arm of AWS's open-source Strands Agents project. The launch post lists Marc Brooker, Mike Chambers and Fabio Nonato de Paula as authors, and everything ships under Apache-2.0.

Strands Decider 2B launch post on the Strands blog dated October 1, 2026, describing it as a small, open source decision model, as taken from Strands
Strands Decider 2B launch post on the Strands blog dated October 1, 2026, describing it as a small, open source decision model, as taken from Strands

A decision model doesn't write text. Think of it as a small language model with the talking removed. It answers questions you define, picked from options you define, and gives you a probability for each. The category took off after TypeSafe Jev launched in September, and the Strands team says so plainly: it is "a type of model that has been gaining a lot of attention since TypeSafe AI's launch of Jev."

I build eesel's integrations and helpdesk API surface, and the triage call (which team, how urgent, is this spam) is the step I've wired into more helpdesks than any other. So I read this release the way I'd read a part I might drop into a pipeline: what does it return, how fast, where does it break, and what's left for me to build.

Here is the short spec sheet:

SpecStrands Decider 2B
ReleasedOctober 1, 2026
Parameters1.9 billion
Base modelQwen3.5-2B-Base, LoRA rank 16 + pointer head
Current reference modelstrands-decider-2B-hobson-v21 (v19 stays published)
Question typesnoul (yes/no), choice (one of N), score (ordered scale)
OutputOne answer per question, with a calibrated probability
Context4,096 tokens by default
LicenseApache-2.0 (weights, code, data, recipe)
Hosted optionNone
HardwareCUDA, Apple MPS or MLX, CPU

Every row comes from the project's GitHub README and its Hugging Face model card.

How does Strands Decider work?

The trick is subtraction. Strands takes a normal language model, keeps its "torso" (the layers that understand text), and throws away the head that would normally generate words. In its place sits a tiny pointer head, just over a million parameters, that scores each answer option. The README sums it up as "One forward pass, no generation, no decoding loop."

Strands Decider 2B architecture diagram: a decoder LLM's language-modelling head is discarded, and the same Qwen3.5-2B torso with a rank-16 LoRA feeds a pointer head that scores each option, as taken from Strands
Strands Decider 2B architecture diagram: a decoder LLM's language-modelling head is discarded, and the same Qwen3.5-2B torso with a rank-16 LoRA feeds a pointer head that scores each option, as taken from Strands

Because it never generates, it can't wander off format. If you ask for one of billing, sales or retail, you get one of those three, every time. That's the same promise as function calling on a big LLM (or JSON mode, its looser cousin), but cheaper, because there's no token-by-token loop and no JSON to parse.

It also means more questions cost very little. The model reads your text once, and "each question adds only its own tokens." So asking "which team?", "is it urgent?" and "how frustrated is the writer?" in one call is close to the price of asking one.

The Hugging Face card's own demo is a support ticket, which tells you exactly who AWS thinks the buyers are:

Hugging Face model card for strands-decider-2B-hobson-v19 showing a CLI example that asks which team should handle a failing-payouts message, whether it is urgent, and how frustrated the writer is, as taken from Hugging Face
Hugging Face model card for strands-decider-2B-hobson-v19 showing a CLI example that asks which team should handle a failing-payouts message, whether it is urgent, and how frustrated the writer is, as taken from Hugging Face

To run it as a service, you start the bundled server and post to it:

Bash
pip install strands-decider
strands-decider serve StrandsAgents/strands-decider-2B-hobson-v21 --port 8000

The weights download from Hugging Face on first use. Then POST /v1/systemone with a state and a questions object, and you get back per-question answers, token usage (output_tokens: 1) and latency_ms. One thing to know before you get excited about swapping it in for Jev: the docs say JevBench's adapter runs against this endpoint unmodified, then add, "Compatibility with the Jev API itself is not verified" (docs/inference.md).

The vendor is also upfront about limits. The launch post says the one-pass design "makes it significantly worse at solving complex problems than reasoning models", and that it is "unsuited for coding, chatbots, document summarization." The model card adds that long, multi-step documents are its weak spot.

How accurate is Strands Decider?

On the JevBench public set (231 tasks), the numbers from the repo are:

Measurementv21 (reference)v19 (launch)
Accuracy0.762 (176/231)0.723 (167/231)
Brier score (lower is better)0.3230.342
Calibration error (ECE)0.0640.052
Easy / standard / hard tiers1.000 / 0.931 / 0.5501.000 / 0.875 / 0.505

The sentence I respect most in the whole release is the one where the team talks down its own upgrade. Across six training runs each, the v19 and v21 recipes "average 172.8 and 172.3 JevBench tasks: the same accuracy." v21's higher headline number is mostly seed luck, and they say so. That honesty is worth more than the extra four points.

Hand-drawn bar chart of JevBench public-set accuracy: GPT-6 Luna 0.996, Jev rebuilds on 12B+ torsos 0.83 to 0.93, Strands Decider v21 0.762 highlighted, then the 2B class: decision-2b 0.753, system-one-open 0.732, Strands Decider v19 0.723, Mapika decider-2b 0.710, Open-Jev 2B 0.645
Hand-drawn bar chart of JevBench public-set accuracy: GPT-6 Luna 0.996, Jev rebuilds on 12B+ torsos 0.83 to 0.93, Strands Decider v21 0.762 highlighted, then the 2B class: decision-2b 0.753, system-one-open 0.732, Strands Decider v19 0.723, Mapika decider-2b 0.710, Open-Jev 2B 0.645

Where does that leave it on the board? On the JevBench v1.4.2 leaderboard (89 systems, September 25), v19 sat 50th overall and 3rd of 33 in the 2B class, behind FlyMy.AI's decision-2b (0.753) and Gemma-based system-one-open (0.732), per the evaluation notes. The top of the board belongs to reasoning models, with GPT-6 Luna at 0.996, and then Jev rebuilds on 12B-plus torsos at 0.83 to 0.93. If you want the whole field side by side, the Jev alternatives roundup ranks them.

So the honest framing: Strands Decider is a strong small model, not a strong model. It trades about 23 points of accuracy against a frontier LLM for speed, price and the ability to run on your own box.

A note if you look at the training-history chart in the launch post: it labels two comparison points "strands-decider-2b v10/v11". Marc Brooker corrected that on Hacker News: those are a different model called "decider-2B", "which isn't ours."

Do the confidence scores mean anything?

Calibration is the whole selling point. A confidence score is only useful if it's honest: a label is only useful for automation if you can say "act above 0.9, send to a human below." The repo claims that on unseen short classification tasks, "answers at a confidence of 0.9 or more are right about 95% of the time."

That's encouraging, with two caveats the model card itself raises: the confidence bands were set on short classification only, so "measure on your own traffic before you trust a threshold," and the model reads questions less closely than documents, so a reworded question often gets the same answer. Early users have found a third:

Hacker News

"From what i can tell from my (limited) experiments, there is a lot of variation caused by simply reordering options, plus a heavy bias toward the first option listed"

The fix is boring and it works: shuffle option order in your tests, and score the model on a few hundred of your own labelled tickets before you trust any number. Brooker himself suggested on HN that at 2B you can fine-tune on your questions, or simply recalibrate without fine-tuning.

How fast is Strands Decider on real hardware?

The launch post's headline figure is a 115 ms median on an RTX 3090. That's real, and it's a GPU number. Here is the wider picture, with every figure taken from the repo, its pull requests or Hacker News:

HardwareLatency per questionSource
RTX 3090 (GPU)median 115 ms, p95 299 msREADME
M3 Pro (Apple silicon)warm median 153 ms under 300 tokens, 234 ms across all tasksREADME
M4 Pro (CPU only)0.3 to 0.5 s once loadedPR #46
Xeon 8375C, 8 threadsabout 1.6 s short input, 16 to 18 s at ~3,000 tokensdocs/inference.md
Graviton 4, bf16 (PR not merged)1.15 sPR #48
Scatter chart of Strands Decider 2B v18 latency against prompt length on an RTX 3090: p50 106 ms and p95 296 ms, with latency rising from about 100 ms under 1,000 tokens to around 300 ms near 3,000 tokens, as taken from Strands
Scatter chart of Strands Decider 2B v18 latency against prompt length on an RTX 3090: p50 106 ms and p95 296 ms, with latency rising from about 100 ms under 1,000 tokens to around 300 ms near 3,000 tokens, as taken from Strands

The pattern matters more than any single number. Latency grows roughly in line with input length, so a long email thread costs far more than a one-line chat message. And on the kind of CPU box most teams spin up by default, a decision takes seconds. One HN user saw about half a second on six cores; another measured 1,732 ms in a CPU-only Docker container. If the decision sits in front of a customer waiting on a chat reply, or inside an intelligent routing step, you want a GPU or Apple silicon.

How much does Strands Decider cost?

The model is free. Running it is not, and there's no one to pay for it but you.

  • No hosted endpoint. The Hugging Face page says "This model isn't deployed by any Inference Provider," and there is no Bedrock listing. Bedrock only appears in the launch post as the LLM in the agent example.
  • No authentication. The local server binds to 127.0.0.1 and, in the card's words, "has no authentication: use it for local experiments." Putting it behind a load balancer, auth and monitoring is your job, the same chores covered in the Strands harness alternatives for agent hosting.
  • No official llama.cpp build yet. Brooker explained on HN that you can't convert just the LoRA, because "there's a whole pointer head and some custom layers." A community GGUF port exists, as does an ONNX one, but both are unofficial.

Here's how it stacks up against the two hosted options most teams compare it to:

Strands Decider 2BTypeSafe JevCloudflare Clef
WeightsOpen, Apache-2.0ClosedOpen, Apache-2.0
Size1.9BNot disclosed27B (Clef-flash 9B)
Hosted APINoneYes, incl. Workers AIWorkers AI
PriceFree, plus your hardware$0.042 per 1M input tokens, output free$0.24 per 1M input (flash $0.09)
Context4,096 default32,00064K hosted
Fine-tune it yourselfYes, full recipeNoRL tuning via Cloudflare team
Deep diveThis postJev reviewClef explained

Jev's rate comes from its Workers AI listing, and Clef's from Cloudflare's model page. The full math lives in the Clef pricing breakdown.

At Jev's rate, a million 500-token decisions cost about $21 in input tokens. That's the bar a self-hosted GPU has to beat once you add the engineer-hours to keep it running. For most teams it won't (the same math as any self-hosted model, see Qwen pricing), until volume is very high or the data can't leave your network. That last case, regulated data, is where Strands Decider is easiest to justify.

What did AWS actually release?

This is where Strands Decider stands apart from Jev, and it's the part I'd point a builder at. The launch post says the release includes "all the training data and scripts we used to build the model."

The strands-labs/strands-decider repository on GitHub, Apache-2.0, with folders for configs, data, docs, evaluation, research, training and src, and a recent commit releasing v21, as taken from GitHub
The strands-labs/strands-decider repository on GitHub, Apache-2.0, with folders for configs, data, docs, evaluation, research, training and src, and a recent commit releasing v21, as taken from GitHub
  • Training data. Hugging Face datasets, ContractNLI and MuSiQue, synthetic rows checked by open-weight models, and the output distributions of a frozen Qwen3.5-4B teacher. Each source and its license is listed in data/sources.md.
  • The recipe. One script, training/recipe.sh all. A retrain takes about 11 hours on one RTX 3090 or about 1 h 10 min on eight H100s.
  • A preregistered research log. Every training run states its predictions and failure conditions before it runs. I've rarely seen a model release that does this, and it's the reason I trust the "v21 is the same as v19" admission.

That makes Strands Decider less a Jev competitor and more a kit for building your own triage classifier. If your categories are specific (your product lines, your refund rules, your escalation tiers), retraining on a few thousand of your own labelled ticket classification rows is the path the authors themselves point to. If fine-tuning is new to you, the recipe is one of the cleaner places to learn it.

The repo is moving fast, too. Pull requests merged on October 7 and 8 add support for Gemma 4 torsos (E2B, E4B, 12B and 26B-A4B) and a soup command that averages several training runs into one checkpoint. The PRs name five upcoming checkpoints, and Brooker wrote on HN that "we've got some new ones coming this week." None have published benchmark numbers yet, so I'd wait for them before planning around them.

What are developers saying about Strands Decider?

The Hacker News launch thread passed 270 points within a day. The tone is warmer than the Jev threads were, mostly because the weights are open. The clearest positive report came from someone running it on-device:

Hacker News

"This is a great model. I've been running it on device in chrome extension to filter things like email. It is just the right mix of size, capability and speed to make it generally useful for adhoc bulk classification tasks."

Co-author Marc Brooker framed where he sees it going, which is a useful mental model for anyone building agents:

Hacker News

"There are a ton a ways to use models like this. Model routing is a popular emerging one. The one that really interests me is hybrid agentic workflows - filling the gap between deterministic workflows (e.g. AWS StepFunctions) and fully agentic workflows that use frontier intelligence."

The pushback is fair and specific. One thread argued that strong benchmark calibration "does not establish reliability on unfamiliar production inputs" (soltanov). Another commenter noted that Cloudflare's Clef runs on llama.cpp, while Strands needs its own CLI and server for now:

Hacker News

"clef from cloudflare runs on llama.cpp - being locked-in to strands cli would be a bummer and will slow down adoption."

That matches my read. The model is good for its size, the tooling is young, and the people getting value right now are developers comfortable hosting their own inference.

Where does Strands Decider fit in a support queue?

The obvious use is ticket triage: which team, how urgent, is it spam. It's classic intent classification, and decision models are a good fit for that call. This is where I'd set expectations carefully.

Hand-drawn two-column infographic: In the box lists open weights under Apache-2.0, a local server, the training recipe and data, and a label with confidence; You still build lists hosting and auth, your own threshold, tests on your tickets, the action itself and the written reply; a bracket joins both columns as one triage system
Hand-drawn two-column infographic: In the box lists open weights under Apache-2.0, a local server, the training recipe and data, and a label with confidence; You still build lists hosting and auth, your own threshold, tests on your tickets, the action itself and the written reply; a bracket joins both columns as one triage system

A label is a fraction of a triage system. Someone still has to host the model, pick a threshold, test it on real tickets, then actually tag, assign or escalate the ticket in Zendesk or Freshdesk. And after all that, the customer still needs an answer.

Here's what that looks like when the confidence threshold is doing its job:

Hand-drawn support triage flow: a new ticket goes into Strands Decider, which asks which team and whether it is urgent; confidence 0.9 and above leads to auto-route and tag, below 0.9 leads to human review, and both paths end at a highlighted card reading someone still writes the reply, with a note to tune on your own traffic
Hand-drawn support triage flow: a new ticket goes into Strands Decider, which asks which team and whether it is urgent; confidence 0.9 and above leads to auto-route and tag, below 0.9 leads to human review, and both paths end at a highlighted card reading someone still writes the reply, with a note to tune on your own traffic

That last box is the one teams underestimate. On a cross-checked eesel trial on a real e-commerce Zendesk inbox, AI triage hit 93% accuracy and caught 100% of spam (about 22% of the inbox), but only 12% of drafted replies were sent as-is. Sorting is the easy half; the reply is where the work is.

What buyers ask for is exactly what a calibrated decision model promises. A CX lead handling about 7,000 tickets a month put it this way on an eesel sales call:

"I cannot go and check all my 7,000 tickets to see if the AI actually made a good answer... I need an AI who is only handling the tickets that it's confident to handle and all the other ones, leave them alone."

So here's my recommendation by situation:

  • You're a developer building an agent or a custom pipeline, your data can't leave your network, and you have a GPU: Strands Decider is a great base. Retrain it on your own labels.
  • You want a decision call with zero hosting: start with Jev instead. Clef and OpenAI's Decisions API are the other hosted options.
  • You run a support team and want tickets sorted and answered: you don't need a decision model at all. You need something that makes the call and acts on it inside your helpdesk. The AI triage tools roundup covers that route, and the auto-triage glossary entry explains the terms.

eesel for triage, routing and the reply

If the job you're picturing for Strands Decider is "sort my support queue," eesel's AI helpdesk teammate does that job end to end. It plugs into Zendesk, Freshdesk or Gorgias in a few minutes (the Zendesk integration is the most common setup), learns from your past tickets and help center, then tags, routes, escalates and writes the reply.

eesel activity feed for a Zendesk instance listing tickets the AI teammate worked, marked Pending and Resolved
eesel activity feed for a Zendesk instance listing tickets the AI teammate worked, marked Pending and Resolved

The same split Strands Decider is built on, a light model for the quick call and a heavier one for the hard part, is already a setting. Each eesel automation has a "Model strength" option: the usual model for drafting and multi-step work, or a lighter one for jobs like tagging or ticket routing. Events that don't match your instructions are skipped and logged with a reason, which is the "leave them alone" behaviour that CX lead asked for. Every action can be set to Auto, Needs approval or Disabled in actions and approvals, so refunds can stay behind a human click while tagging runs on its own.

If you'd have reached for Strands Decider because you like working from a terminal, the eesel CLI gives you the same teammate from the command line. npx @eesel/cli works with no account (a temporary workspace with $5 of usage), every command prints JSON, and errors come back as one JSON line with a retryable flag, so Claude Code, Cursor or Codex can drive the whole setup. A developer or a coding agent can wire a triage job with eesel automations enable <platform> <key> --instructions "...", check a change first with --dry-run, and clear held actions with eesel approvals approve <id>. It's the same agent as the dashboard, and changes sync both ways. More on that in my look at the agent CLI space.

Before anything goes live, the simulation skill runs the agent over your real past tickets and scores each answer, so you see accuracy on your own traffic instead of a benchmark. The free plan includes 100 credits with no card, and paid plans start at $299 a month for 500 tickets or chats on eesel pricing, with no per-seat fee.

Try eesel on a slice of your queue and see how many tickets it would sort, and answer, this week.

Frequently Asked Questions

What is Strands Decider?

Strands Decider 2B is an open-source decision model released by AWS's Strands Labs team on October 1, 2026. It reads a piece of text plus typed questions (yes/no, pick one, or score on a scale) and returns each answer with a confidence score in one forward pass, instead of generating text. It is built for jobs like ticket triage, routing and guardrails.

Is Strands Decider free?

Yes. The weights, code, training data and recipe are all released under Apache-2.0, so Strands Decider costs nothing to download. You pay for the hardware you run it on, because there is no hosted endpoint on Bedrock or any Hugging Face inference provider. Compare that with Jev pricing or Clef pricing, which are hosted and billed per token.

How accurate is Strands Decider 2B?

The v21 reference model scores 0.762 on the JevBench public set (176 of 231 tasks), and v19 scores 0.723. That puts it near the top of the 2B size class, but reasoning LLMs like GPT-6 Luna still score 0.996 on the same tasks. The repo itself says v21 and v19 have the same accuracy once you average six training runs.

Can Strands Decider run on a CPU?

Yes, but expect seconds, not milliseconds. The repo measures about 1.6 seconds per question on an 8-thread Xeon for a short input, and 16 to 18 seconds at around 3,000 tokens. On an RTX 3090 GPU the median is 115 ms. For a live routing step, plan for a GPU or Apple silicon.

Strands Decider vs Jev: which should I use?

Pick Strands Decider if you need open weights, self-hosting or the option to fine-tune on your own questions. Pick TypeSafe Jev if you want a hosted API with no servers to run and a 32,000-token context. The Jev alternatives roundup covers the rest of the field.

Can I use Strands Decider for support ticket triage?

Yes. The Hugging Face card's own example routes a payout complaint to billing, sales or retail and flags urgency. You still need to host it, tune a confidence threshold on your own tickets, and build the step that tags, routes or replies. An AI helpdesk teammate like eesel handles that whole loop inside your helpdesk.

Is there a Strands Decider API on AWS Bedrock?

Not as of October 8, 2026. The only server is a local one (strands-decider serve) that binds to 127.0.0.1 with no authentication. Bedrock appears in the launch post only as the LLM in the agent example. If you want a hosted decision API, look at OpenAI's Decisions API, Jev or Clef.

Share this article

Rama Adi

Article by

Rama Adi

Rama is a software engineer at eesel AI with two years of experience writing about B2B SaaS, AI tools, and customer support technology. Based in Bali, Indonesia, he brings a developer's perspective to product comparisons — cutting through marketing copy to what the integrations and APIs actually do.

Related Posts

All posts →
Hand-drawn hero banner of a person feeding questions into a switchboard that routes them into labeled lanes, each with a confidence dial
Trending

Decision models explained: Jev, Clef, Strands Decider and the new AI category

Decision models return typed answers with confidence scores instead of text. What they are, how they work, every model you can use today, and where they fit in support.

KiraKiraOct 6, 2026
Cloudflare Clef hero banner in Cloudflare orange, two people looking at a decision model connected to users, websites, devices and a list of options
Trending

What is Cloudflare Clef? Cloudflare's decision model, explained

Cloudflare Clef explained: what the decision model does, how Clef differs from Clef-flash, how to run it on Workers AI or Ollama, and where it fits in ticket routing.

KiraKiraOct 7, 2026
Hand-drawn hero banner in PostHog amber showing a hedgehog butler weighing inputs before picking an outcome, with two developers looking on
Trending

PostHog Jeeves: the open decision model that thinks before it picks

PostHog Jeeves is an open 9B decision model that writes a reasoning chain before it answers. Here is what it beats Jev at, what it costs you in latency, and where it fits.

KiraKiraOct 1, 2026
Illustrated hero banner for TypeSafe Jev, the ultrafast System One AI model, with a speed gauge
Trending

Is Jev really ultrafast? TypeSafe's System One model, tested

TypeSafe calls Jev an ultrafast System One model at 70-500ms a decision. Here is what the speed claim really means, where it holds up, and where it does not.

Rama AdiRama AdiSep 22, 2026
TypeSafe Jev pricing hero banner in rose and off-white, showing a low token cost per million
Trending

TypeSafe Jev pricing (2026): $0.042 per million tokens, output free

TypeSafe Jev pricing broken down: $0.042 per million input tokens, output free, no plan tiers yet, and what a System One model actually costs to run in production.

Kurnia KharismaKurnia KharismaSep 22, 2026
TypeSafe Jev hero banner in rose and off-white, illustrating a fast typed-decision model
Trending

TypeSafe Jev review: the 'System One' model that gives AI the properties of code

A hands-on TypeSafe Jev review: what the System One model actually does, whether the speed, price and 'can't hallucinate' claims hold, and where a typed-decision model fits real work.

Rama AdiRama AdiSep 21, 2026
Hand-drawn illustration of two people in a natural back-and-forth voice conversation, in OpenAI teal
Trending

GPT-Live-1 review: is OpenAI's full-duplex voice model worth it?

A hands-on GPT-Live-1 review: OpenAI's full-duplex voice model is the most natural I have used, but it delegates reasoning to a backend model and bills per minute. Here is the honest verdict.

KiraKiraSep 11, 2026
Two people having a natural conversation with an AI voice assistant, sound waves flowing between them
Trending

GPT-Live-1: OpenAI's full-duplex voice model, explained

What GPT-Live-1 actually is: OpenAI's full-duplex voice model that listens and speaks at once, now in ChatGPT and the API at $0.05 per minute.

KiraKiraSep 11, 2026
Two people talking across a table while an audio-visual AI model watches, listens and speaks in the same loop
Trending

SeedRealtime: what ByteDance's audio-visual model actually does

SeedRealtime is ByteDance's audio-visual full-duplex model. Here is what it does, what ByteDance published, and what you can actually call today.

KiraKiraAug 18, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free