
What Cloudflare Clef actually is
Clef and Clef-flash launched on October 1, 2026 as the first models trained by Cloudflare's Workers AI team. They are decision models, not chatbots. You send a piece of state (a ticket, a log line, a JSON record, an image) plus up to 64 typed questions, and you get back a probability for every allowed answer. No free-form text, nothing to parse.
That puts Clef in the same family as TypeSafe Jev, which kicked off the category in September at a famously low price. It's also Cloudflare's second notable AI launch in two months, after the Kitesurf browser engine in August. Cloudflare made Clef fully compatible with Jev's System One API, so switching is a change of endpoint and model name. The weights are open under Apache 2.0 on Hugging Face, which is a big part of the pricing story further down.

The three question types are the whole interface:
noul: a yes/no question. Returns the probability the answer is yes.choice: pick one option from a set you define. Returns the pick, a probability per option, and a confidence value.score: rate against an ordered rubric. Returns a probability-weighted score.
Cloudflare's own example is pure ticket triage work: "Is this support request urgent?", "Which team should handle this request?" and "How severe is the customer impact?" in one call. That is why I, writing for a support-tools company, care about it.
I come at this from the SEO side and the engineering side, and the search intent behind "Clef pricing" is a buying question: is this cheap enough to put on every ticket, every request, every event? At eesel I've watched that exact decision play out on live support queues for years. In one real-traffic trial on an e-commerce Zendesk inbox, AI triage hit 93% accuracy and 100% spam detection, while drafted replies carried a 7% factual error rate. Classification was the cheap, reliable part. Acting on it was where the cost and the risk lived. Keep that in mind as you read the numbers.
How much does Cloudflare Clef cost?
The full price list fits in one table. Both models bill only on input tokens through Workers AI pricing, which charges $0.011 per 1,000 neurons behind the scenes.
| Clef | Clef-flash | Jev (on Workers AI) | |
|---|---|---|---|
| Model ID | @cf/cloudflare/clef | @cf/cloudflare/clef-flash | typesafe/jev |
| Size | 27B | 9B | Not disclosed |
| Input price | $0.24 per M tokens | $0.09 per M tokens | $0.042 per M tokens |
| Output price | None listed | None listed | $0.00 |
| Neurons per M input tokens | 21,818 | 8,182 | Not listed |
| Context window (hosted) | 65,536 tokens | 65,536 tokens | 32,000 tokens |
| Vision input | Yes, up to 4 images | Yes, up to 4 images | No |
| Questions per request | 1 to 64 | 1 to 64 | Typed questions |
| Median latency (Cloudflare's test) | 209.3 ms | 38.8 ms | 524.1 ms |
| Open weights | Apache 2.0 | Apache 2.0 | No |
Jev's row comes from its own Workers AI model page, where Cloudflare lists it as a third-party model at $0.042 input, $0.00 output. So you can run all three on one Cloudflare bill and swap between them with a config change.

The free allowance, in decisions
Cloudflare gives every account 10,000 neurons per day at no charge, reset at 00:00 UTC. That allowance is worth $0.11 a day at list price. Converted to work:
- Clef: about 458,000 input tokens a day, or roughly 900 decisions at 500 tokens each.
- Clef-flash: about 1.22 million input tokens a day, or roughly 2,400 decisions at 500 tokens.
On the Workers Free plan, that is a hard ceiling. Cloudflare's pricing page says that once you exceed a limit, "further operations will fail with an error." To keep going you need Workers Paid, which carries a $5 monthly minimum per account. After that you pay $0.011 per 1,000 neurons above the free 10,000 each day.
There is no batch discount and no cached-input price on either Clef row, unlike some of the LLMs on the same page. Prepaid AI Gateway credits are a way to pay, not a discount.
Clef vs Jev vs an LLM: cost per decision
Here is where the "6x more than Jev" headline that went around launch week gets more interesting. Hacker News spotted it within hours:
"Pricing is $0.24/million input tokens which is ~6x compared to Jev. Clef-flash is at $0.09 which is way more competitive."
The per-token math is right: $0.24 is 5.7x $0.042, and Clef-flash is 2.1x. But per-token rates are a strange way to compare things that cost a fraction of a cent per call. So I priced 1,000 routing decisions on a 500-token ticket (state plus a three-question schema) across Workers AI's list prices. For the two LLMs I added 60 output tokens for a small JSON answer, since they bill output and decision models don't.

Two things jump out. First, Clef costs about the same per decision as gpt-oss-20b on the same platform ($0.12 vs $0.118 per thousand). You are not paying less than a small LLM for Clef. You are paying for a typed answer that comes back in about 200 ms with no parsing, instead of generated text you have to validate. Second, Clef-flash costs less than half of either gpt-oss model, roughly ties the small 3B to 8B Llama and Qwen models, and sits within $0.024 per thousand decisions of Jev.
Frontier models are a different price class entirely. My Claude pricing breakdown shows input rates measured in dollars per million, not cents, and the same goes for the GPT-5.6 price card. Even DeepSeek V4 Flash, a budget favorite, lists at $0.44 per million input tokens on Workers AI, almost double Clef.
Meanwhile the OpenAI Decisions API, the other big decision endpoint, still has no public price.
What actually drives your Clef bill
The rate card is fixed. What you control is how many tokens you send per call and how many calls you make. That first lever is the one teams get wrong.

Clef's 64K context window is a selling point against Jev's 32K, and it is also a temptation. It invites you to paste the whole ticket thread, the customer's order history and the macro library into state "for context." Every one of those tokens is billed. In my worked numbers below, moving from a 500-token summary to a 4,000-token thread multiplies the bill by exactly 8, and it does so on every model.
| Monthly decisions | Tokens per call | Clef | Clef-flash | Jev |
|---|---|---|---|---|
| 30,000 (about 1,000/day) | 500 | $0.30 | $0.00 | $0.63 |
| 300,000 (about 10,000/day) | 500 | $32.70 | $10.20 | $6.30 |
| 3,000,000 (about 100,000/day) | 500 | $356.70 | $131.70 | $63.00 |
| 30,000 | 4,000 | $25.50 | $7.50 | $5.04 |
| 300,000 | 4,000 | $284.70 | $104.70 | $50.40 |
| 3,000,000 | 4,000 | $2,876.68 | $1,076.72 | $504.00 |
Clef and Clef-flash columns subtract the 10,000 free neurons a day (assuming steady daily traffic) and exclude the $5 Workers Paid minimum. Jev is shown at list price because Cloudflare doesn't say whether third-party models draw on the free allowance. At low volume, Clef-flash is effectively free. At 3 million decisions a month the absolute numbers are still small next to a support team's payroll, but the 8x state multiplier is the difference between a rounding error and a line item.
The practical fix is boring: summarize or truncate before you call. Cloudflare's own schema notes that "Long text state is truncated to fit the model's token limit," so overstuffed state doesn't even get fully read past the window; you just pay for the attempt. If you are routing tickets, the subject line plus the first customer message is usually enough. The ticket triage automation guide covers what signal actually matters, and these triage prompt templates show how to boil a thread down first.
Estimate your own Clef costs
Plug in your daily volume and average tokens per call. The calculator applies the free daily allowance to Clef and Clef-flash and prices Jev and gpt-oss-20b at list for comparison.
Clef vs Clef-flash: which one to pay for
On paper, the bigger model should be the safe default. In Cloudflare's own numbers, it often isn't. Clef-flash is 2.7x cheaper per token and more than 5x faster at the median, and it wins several of the published benchmarks outright.

The launch blog by Michelle Chen publishes the full table. The rows that matter for a pricing decision:
| Benchmark | Clef | Clef-flash | Jev |
|---|---|---|---|
| BFCL (case exact) | 98.47 | 98.76 | 95.75 |
| BANKING77 (macro-F1) | 94.20 | 90.93 | 79.74 |
| CLINC150+OOS (macro-F1) | 97.43 | 66.77 | 89.27 |
| Home appliances (case exact) | 82.95 | 97.73 | 52.27 |
| Customer service workflow | 76.3 | 77.0 | 76.0 |
| Median latency | 209.3 ms | 38.8 ms | 524.1 ms |
Read the CLINC150 row twice. That benchmark includes 150 intents plus out-of-scope examples, and Clef-flash drops to 66.77 while Clef holds 97.43. So the rule I'd follow: use Clef-flash when your option list is short and stable (urgent or not, five teams, a severity scale), and pay for Clef when you're picking from dozens of intents or need to catch "none of the above" reliably. On the customer service workflow eval, the three models land within 1 point of each other, which tells you the price difference is buying speed, not support accuracy. For TypeSafe's side of the speed story, the Jev ultrafast test digs into its 70 to 500 ms claim.
These are vendor-run benchmarks. A commenter on the launch thread made the point that matters:
"Yeah I also thought it was strange their pareto frontier didn't include cost."
Run your own labelled tickets through both before committing. At 500 tokens per ticket, a 5,000-ticket test costs about $0.60 on Clef.
Pricing gotchas to watch for
Clef's pricing is simpler than most AI rate cards, but six details catch people out:
- The free plan stops, it doesn't bill. On Workers Free, going past 10,000 neurons in a day means errors until 00:00 UTC. Put production traffic on Workers Paid.
- Images ride the input meter. Clef accepts up to 4 images per request (4 MiB each, 13 MiB body), and the docs list no separate image price. Test what a screenshot-heavy ticket actually costs before you enable vision.
- Rate limits aren't spelled out. The Workers AI limits page doesn't name Clef specifically. Load test before you put it in a hot path.
- Hosted context is smaller than the weights. Cloudflare's hosted models run at 64K, while the Ollama build of Clef lists 256K. If you truly need long state, self-hosting is the way to get it.
- Fine-tuning has no price yet. The new RL fine-tuning service starts as a hands-on engagement with Cloudflare's forward-deployed engineers, via an interest form. Self-serve is promised later. Budget for a sales conversation, not a rate card.
- There's no output line to save on. Optimizations that help with LLMs, like shorter answers or caching, don't apply. Input size is the only dial.
Self-hosting Clef: the $0 option
Because both models are Apache 2.0, the cheapest Cloudflare Clef pricing tier is the one where you don't pay Cloudflare at all. The Hugging Face card shows Clef is post-trained from Qwen3.8-27B (Clef-flash from Qwen3.5-9B) and was tested on a single NVIDIA H200. The Ollama build of Clef is an 18GB download. Its Hugging Face card shows 5,416 downloads in the last month and 18 community quantizations already.

Self-hosting only wins at real volume. At 300,000 short decisions a month, hosted Clef-flash is about $10. No GPU you can rent beats that. It starts to make sense when you're in the millions of long calls a month, you need the 256K context, or your data can't leave your network. HN had a sharper version of the build-your-own argument:
"If you have a very narrow use case you can train a BERT based decision model on a laptop an hour if you have good data to train it on. It'll answer faster than the roundtrip to clef/jev and use <1gb memory"
That's fair if your categories never change. The reason decision models exist is that you can add a new team or a new question without retraining anything. For a support queue where routing rules shift every quarter, that flexibility is worth a few dollars a month. Open weights also come with a caveat one commenter flagged: the weights are permissive, but the training data and pipeline aren't published, so this is open weights rather than open source in the strict sense. If you're comparing open models more broadly, the Jev alternatives roundup covers the other routes.
Is Clef worth it for support teams?
If you're an engineering team building your own intelligent routing layer, yes. Clef-flash at $0.09 per million tokens is cheap enough to call on every inbound ticket, and the typed output removes a whole class of parsing bugs you'd get from an LLM. For routing alone, it's hard to argue with a sub-cent price.
But notice what you're buying: a probability. "Urgent: 0.94, team: technical, severity: 2.8." Everything after that is still your job. Someone has to write the code that moves the ticket, decide what to do at 0.6 confidence, draft the reply (the job of a support AI copilot), check it against your knowledge base, and escalate the weird ones. In that real-traffic trial I mentioned, the 93%-accurate triage was never the problem; the replies were.
A CX lead doing about 7,000 tickets a month put the actual requirement plainly on a sales call with eesel:
"I need an AI who is only handling the tickets that it's confident to handle and all the other ones, leave them alone."
A decision model gives you the confidence score that makes that possible. It doesn't give you the AI that handles the ticket. That's the gap between AI ticket classification and an AI helpdesk agent, and it's where the real cost lives.
A Clef decision is about $0.0001. A support engineer's week spent wiring routing rules, retries and fallbacks into a helpdesk is not. If you've ever set up Zendesk AI triage by hand, you know the routing rule is the easy 10%. The same is true of Freshdesk auto triage and every other queue.
If you've already compared best AI for ticket triage tools, Clef sits one layer below them: it's an ingredient for building one.
eesel for the decision and what comes after it
Clef is infrastructure. eesel is the employee. eesel is an AI teammate platform, and its AI helpdesk teammate joins your Zendesk, Freshdesk, Gorgias or Front queue, learns from your past tickets and help center, and makes the same calls Clef does as part of AI ticket routing (is this urgent, which team, can I answer it) and then acts: tags, routes, drafts, resolves, or hands off. Actions you mark as needing approval wait in an activity feed until a human says yes.

If you like Clef because it's programmable, you'll want the eesel CLI. It puts the same teammate from the dashboard into your terminal, and every command prints JSON, so Claude Code, Cursor or Codex can drive a full setup. npx @eesel/cli init chat-bubble --site <url> spins up a free workspace with no account. eesel integrations connects your helpdesk, eesel instructions sets standing rules like "only auto-reply above high confidence," eesel automations wires tagging and routing, and eesel approvals approve <id> clears held actions from a script. Changes sync both ways with the dashboard. The AI agent CLI post goes deeper.
Pricing is per ticket, not per token, which makes the AI vs human cost math easy to run: eesel pricing starts with 100 free credits and paid plans from $299 for 500 credits a month, where one credit covers a whole ticket or chat however long it runs. That's a different unit from Clef's because it's buying a different thing: the decision plus the work. If you'd rather build the triage layer yourself, Clef-flash is a great price. If you'd rather have tickets handled this week, try eesel on your own queue and see the drafts before you pay.
Frequently Asked Questions
How much does Cloudflare Clef cost?
Clef costs $0.24 per million input tokens on Workers AI, and Clef-flash costs $0.09. Neither model lists an output price, so input is the whole Cloudflare Clef pricing story. For comparison, see my Jev pricing breakdown.
Is Cloudflare Clef free to use?
Partly. Every Workers AI account gets 10,000 free neurons per day, which works out to roughly 900 short Clef decisions or 2,400 Clef-flash decisions a day. The weights are also Apache 2.0, so self-hosting costs nothing but your own hardware. If you just need tickets sorted without building anything, AI ticket classification tools are another route.
Why is Clef more expensive than Jev?
Per token, Clef is about 5.7x Jev's $0.042 rate, and Clef-flash is about 2.1x. Cloudflare's pitch is speed and accuracy: it reports a 209 ms median for Clef and 38.8 ms for Clef-flash against 524 ms for Jev. The TypeSafe Jev explainer covers the model Clef is compatible with.
What is the difference between Clef and Clef-flash pricing?
Clef (27B) is $0.24 per million input tokens and Clef-flash (9B) is $0.09, so Clef-flash is about 2.7x cheaper per token. In Cloudflare's own benchmarks Clef-flash is also faster, while Clef leads on harder intent sets like CLINC150. For most ticket triage traffic, test Clef-flash first.
Does Cloudflare charge for Clef output tokens?
No output price appears on the Clef model page or the Workers AI pricing table, which lists Clef as an input-only row. The answer is a set of probabilities, not generated text, so there is nothing long to bill on the way out. That is a real difference from an LLM, where output tokens usually cost several times more than input.
How much does Cloudflare Clef cost for ticket routing at scale?
At 3 million short decisions a month (about 500 tokens each), Clef runs about $357 and Clef-flash about $132 after the free allowance, plus the $5 Workers Paid minimum. If you send whole 4,000-token threads instead, multiply by eight. Compare that with what an AI support agent costs when the work after the decision is included.
Can I self-host Clef to avoid Cloudflare pricing?
Yes. Both models are open weights under Apache 2.0 on Hugging Face, and the Ollama build of Clef is an 18GB download. Cloudflare says it tested on a single NVIDIA H200, so budget for real GPU memory. Teams weighing hosted versus self-run models can also read the Qwen pricing guide, since Clef is built on Qwen.
Is Clef a replacement for an AI support agent?
No. Clef returns a decision such as "urgent, technical team, 0.92" and stops there; your code still has to route, reply and escalate. An AI helpdesk agent like eesel makes that decision and then acts on it inside Zendesk, Freshdesk or Gorgias.

Article by
Kurnia Kharisma
Kurnia is a software engineer and writer at eesel AI with two years of SEO experience, writing about AI tools, helpdesk software, and customer support. He pairs a developer's understanding of how these products are built with search-driven research into what actually ranks and resonates with the people searching for them.








