Cloudflare Clef pricing (2026): $0.24 per million tokens, explained

Kurnia Kharisma
Written by

Kurnia Kharisma

Katelin Teen
Reviewed by

Katelin Teen

Last edited October 6, 2026

Expert Verified
Cloudflare Clef pricing hero banner in Cloudflare orange, two people comparing the cost of two decision models

What Cloudflare Clef actually is

Clef and Clef-flash launched on October 1, 2026 as the first models trained by Cloudflare's Workers AI team. They are decision models, not chatbots. You send a piece of state (a ticket, a log line, a JSON record, an image) plus up to 64 typed questions, and you get back a probability for every allowed answer. No free-form text, nothing to parse.

That puts Clef in the same family as TypeSafe Jev, which kicked off the category in September at a famously low price. It's also Cloudflare's second notable AI launch in two months, after the Kitesurf browser engine in August. Cloudflare made Clef fully compatible with Jev's System One API, so switching is a change of endpoint and model name. The weights are open under Apache 2.0 on Hugging Face, which is a big part of the pricing story further down.

Cloudflare Workers AI model page for Clef showing the 65,536-token context window, vision support and $0.24 per M input tokens unit pricing, as taken from Cloudflare docs
Cloudflare Workers AI model page for Clef showing the 65,536-token context window, vision support and $0.24 per M input tokens unit pricing, as taken from Cloudflare docs

The three question types are the whole interface:

  • noul: a yes/no question. Returns the probability the answer is yes.
  • choice: pick one option from a set you define. Returns the pick, a probability per option, and a confidence value.
  • score: rate against an ordered rubric. Returns a probability-weighted score.

Cloudflare's own example is pure ticket triage work: "Is this support request urgent?", "Which team should handle this request?" and "How severe is the customer impact?" in one call. That is why I, writing for a support-tools company, care about it.

I come at this from the SEO side and the engineering side, and the search intent behind "Clef pricing" is a buying question: is this cheap enough to put on every ticket, every request, every event? At eesel I've watched that exact decision play out on live support queues for years. In one real-traffic trial on an e-commerce Zendesk inbox, AI triage hit 93% accuracy and 100% spam detection, while drafted replies carried a 7% factual error rate. Classification was the cheap, reliable part. Acting on it was where the cost and the risk lived. Keep that in mind as you read the numbers.

How much does Cloudflare Clef cost?

The full price list fits in one table. Both models bill only on input tokens through Workers AI pricing, which charges $0.011 per 1,000 neurons behind the scenes.

ClefClef-flashJev (on Workers AI)
Model ID@cf/cloudflare/clef@cf/cloudflare/clef-flashtypesafe/jev
Size27B9BNot disclosed
Input price$0.24 per M tokens$0.09 per M tokens$0.042 per M tokens
Output priceNone listedNone listed$0.00
Neurons per M input tokens21,8188,182Not listed
Context window (hosted)65,536 tokens65,536 tokens32,000 tokens
Vision inputYes, up to 4 imagesYes, up to 4 imagesNo
Questions per request1 to 641 to 64Typed questions
Median latency (Cloudflare's test)209.3 ms38.8 ms524.1 ms
Open weightsApache 2.0Apache 2.0No

Jev's row comes from its own Workers AI model page, where Cloudflare lists it as a third-party model at $0.042 input, $0.00 output. So you can run all three on one Cloudflare bill and swap between them with a config change.

Cloudflare Workers AI pricing page showing $0.011 per 1,000 neurons and a free allocation of 10,000 neurons per day on both Workers Free and Workers Paid, as taken from Cloudflare docs
Cloudflare Workers AI pricing page showing $0.011 per 1,000 neurons and a free allocation of 10,000 neurons per day on both Workers Free and Workers Paid, as taken from Cloudflare docs

The free allowance, in decisions

Cloudflare gives every account 10,000 neurons per day at no charge, reset at 00:00 UTC. That allowance is worth $0.11 a day at list price. Converted to work:

  • Clef: about 458,000 input tokens a day, or roughly 900 decisions at 500 tokens each.
  • Clef-flash: about 1.22 million input tokens a day, or roughly 2,400 decisions at 500 tokens.

On the Workers Free plan, that is a hard ceiling. Cloudflare's pricing page says that once you exceed a limit, "further operations will fail with an error." To keep going you need Workers Paid, which carries a $5 monthly minimum per account. After that you pay $0.011 per 1,000 neurons above the free 10,000 each day.

There is no batch discount and no cached-input price on either Clef row, unlike some of the LLMs on the same page. Prepaid AI Gateway credits are a way to pay, not a discount.

Clef vs Jev vs an LLM: cost per decision

Here is where the "6x more than Jev" headline that went around launch week gets more interesting. Hacker News spotted it within hours:

Hacker News

"Pricing is $0.24/million input tokens which is ~6x compared to Jev. Clef-flash is at $0.09 which is way more competitive."

The per-token math is right: $0.24 is 5.7x $0.042, and Clef-flash is 2.1x. But per-token rates are a strange way to compare things that cost a fraction of a cent per call. So I priced 1,000 routing decisions on a 500-token ticket (state plus a three-question schema) across Workers AI's list prices. For the two LLMs I added 60 output tokens for a small JSON answer, since they bill output and decision models don't.

Bar chart of cost per 1,000 decisions on a 500-token ticket at Workers AI list prices: Jev $0.021, Clef-flash $0.045, gpt-oss-20b $0.118, Clef $0.12, gpt-oss-120b $0.22
Bar chart of cost per 1,000 decisions on a 500-token ticket at Workers AI list prices: Jev $0.021, Clef-flash $0.045, gpt-oss-20b $0.118, Clef $0.12, gpt-oss-120b $0.22

Two things jump out. First, Clef costs about the same per decision as gpt-oss-20b on the same platform ($0.12 vs $0.118 per thousand). You are not paying less than a small LLM for Clef. You are paying for a typed answer that comes back in about 200 ms with no parsing, instead of generated text you have to validate. Second, Clef-flash costs less than half of either gpt-oss model, roughly ties the small 3B to 8B Llama and Qwen models, and sits within $0.024 per thousand decisions of Jev.

Frontier models are a different price class entirely. My Claude pricing breakdown shows input rates measured in dollars per million, not cents, and the same goes for the GPT-5.6 price card. Even DeepSeek V4 Flash, a budget favorite, lists at $0.44 per million input tokens on Workers AI, almost double Clef.

Meanwhile the OpenAI Decisions API, the other big decision endpoint, still has no public price.

What actually drives your Clef bill

The rate card is fixed. What you control is how many tokens you send per call and how many calls you make. That first lever is the one teams get wrong.

Diagram showing tokens per call times calls per month times $0.24 per million equals your bill, with tokens per call highlighted and the note that 500 versus 4,000 tokens is an 8x difference
Diagram showing tokens per call times calls per month times $0.24 per million equals your bill, with tokens per call highlighted and the note that 500 versus 4,000 tokens is an 8x difference

Clef's 64K context window is a selling point against Jev's 32K, and it is also a temptation. It invites you to paste the whole ticket thread, the customer's order history and the macro library into state "for context." Every one of those tokens is billed. In my worked numbers below, moving from a 500-token summary to a 4,000-token thread multiplies the bill by exactly 8, and it does so on every model.

Monthly decisionsTokens per callClefClef-flashJev
30,000 (about 1,000/day)500$0.30$0.00$0.63
300,000 (about 10,000/day)500$32.70$10.20$6.30
3,000,000 (about 100,000/day)500$356.70$131.70$63.00
30,0004,000$25.50$7.50$5.04
300,0004,000$284.70$104.70$50.40
3,000,0004,000$2,876.68$1,076.72$504.00

Clef and Clef-flash columns subtract the 10,000 free neurons a day (assuming steady daily traffic) and exclude the $5 Workers Paid minimum. Jev is shown at list price because Cloudflare doesn't say whether third-party models draw on the free allowance. At low volume, Clef-flash is effectively free. At 3 million decisions a month the absolute numbers are still small next to a support team's payroll, but the 8x state multiplier is the difference between a rounding error and a line item.

The practical fix is boring: summarize or truncate before you call. Cloudflare's own schema notes that "Long text state is truncated to fit the model's token limit," so overstuffed state doesn't even get fully read past the window; you just pay for the attempt. If you are routing tickets, the subject line plus the first customer message is usually enough. The ticket triage automation guide covers what signal actually matters, and these triage prompt templates show how to boil a thread down first.

Estimate your own Clef costs

Plug in your daily volume and average tokens per call. The calculator applies the free daily allowance to Clef and Clef-flash and prices Jev and gpt-oss-20b at list for comparison.

Clef vs Clef-flash: which one to pay for

On paper, the bigger model should be the safe default. In Cloudflare's own numbers, it often isn't. Clef-flash is 2.7x cheaper per token and more than 5x faster at the median, and it wins several of the published benchmarks outright.

Scatter chart of median latency against price per million input tokens: Clef-flash at 38.8 ms and $0.09, Clef at 209 ms and $0.24, Jev at 524 ms and $0.042
Scatter chart of median latency against price per million input tokens: Clef-flash at 38.8 ms and $0.09, Clef at 209 ms and $0.24, Jev at 524 ms and $0.042

The launch blog by Michelle Chen publishes the full table. The rows that matter for a pricing decision:

BenchmarkClefClef-flashJev
BFCL (case exact)98.4798.7695.75
BANKING77 (macro-F1)94.2090.9379.74
CLINC150+OOS (macro-F1)97.4366.7789.27
Home appliances (case exact)82.9597.7352.27
Customer service workflow76.377.076.0
Median latency209.3 ms38.8 ms524.1 ms

Read the CLINC150 row twice. That benchmark includes 150 intents plus out-of-scope examples, and Clef-flash drops to 66.77 while Clef holds 97.43. So the rule I'd follow: use Clef-flash when your option list is short and stable (urgent or not, five teams, a severity scale), and pay for Clef when you're picking from dozens of intents or need to catch "none of the above" reliably. On the customer service workflow eval, the three models land within 1 point of each other, which tells you the price difference is buying speed, not support accuracy. For TypeSafe's side of the speed story, the Jev ultrafast test digs into its 70 to 500 ms claim.

These are vendor-run benchmarks. A commenter on the launch thread made the point that matters:

Hacker News

"Yeah I also thought it was strange their pareto frontier didn't include cost."

Run your own labelled tickets through both before committing. At 500 tokens per ticket, a 5,000-ticket test costs about $0.60 on Clef.

Pricing gotchas to watch for

Clef's pricing is simpler than most AI rate cards, but six details catch people out:

  1. The free plan stops, it doesn't bill. On Workers Free, going past 10,000 neurons in a day means errors until 00:00 UTC. Put production traffic on Workers Paid.
  2. Images ride the input meter. Clef accepts up to 4 images per request (4 MiB each, 13 MiB body), and the docs list no separate image price. Test what a screenshot-heavy ticket actually costs before you enable vision.
  3. Rate limits aren't spelled out. The Workers AI limits page doesn't name Clef specifically. Load test before you put it in a hot path.
  4. Hosted context is smaller than the weights. Cloudflare's hosted models run at 64K, while the Ollama build of Clef lists 256K. If you truly need long state, self-hosting is the way to get it.
  5. Fine-tuning has no price yet. The new RL fine-tuning service starts as a hands-on engagement with Cloudflare's forward-deployed engineers, via an interest form. Self-serve is promised later. Budget for a sales conversation, not a rate card.
  6. There's no output line to save on. Optimizations that help with LLMs, like shorter answers or caching, don't apply. Input size is the only dial.

Self-hosting Clef: the $0 option

Because both models are Apache 2.0, the cheapest Cloudflare Clef pricing tier is the one where you don't pay Cloudflare at all. The Hugging Face card shows Clef is post-trained from Qwen3.8-27B (Clef-flash from Qwen3.5-9B) and was tested on a single NVIDIA H200. The Ollama build of Clef is an 18GB download. Its Hugging Face card shows 5,416 downloads in the last month and 18 community quantizations already.

Hugging Face model card for Cloudflare/clef showing the Apache 2.0 license, 27B parameters, BF16 tensors and the Qwen3.8-27B base model, as taken from Hugging Face
Hugging Face model card for Cloudflare/clef showing the Apache 2.0 license, 27B parameters, BF16 tensors and the Qwen3.8-27B base model, as taken from Hugging Face

Self-hosting only wins at real volume. At 300,000 short decisions a month, hosted Clef-flash is about $10. No GPU you can rent beats that. It starts to make sense when you're in the millions of long calls a month, you need the 256K context, or your data can't leave your network. HN had a sharper version of the build-your-own argument:

Hacker News

"If you have a very narrow use case you can train a BERT based decision model on a laptop an hour if you have good data to train it on. It'll answer faster than the roundtrip to clef/jev and use <1gb memory"

That's fair if your categories never change. The reason decision models exist is that you can add a new team or a new question without retraining anything. For a support queue where routing rules shift every quarter, that flexibility is worth a few dollars a month. Open weights also come with a caveat one commenter flagged: the weights are permissive, but the training data and pipeline aren't published, so this is open weights rather than open source in the strict sense. If you're comparing open models more broadly, the Jev alternatives roundup covers the other routes.

Is Clef worth it for support teams?

If you're an engineering team building your own intelligent routing layer, yes. Clef-flash at $0.09 per million tokens is cheap enough to call on every inbound ticket, and the typed output removes a whole class of parsing bugs you'd get from an LLM. For routing alone, it's hard to argue with a sub-cent price.

But notice what you're buying: a probability. "Urgent: 0.94, team: technical, severity: 2.8." Everything after that is still your job. Someone has to write the code that moves the ticket, decide what to do at 0.6 confidence, draft the reply (the job of a support AI copilot), check it against your knowledge base, and escalate the weird ones. In that real-traffic trial I mentioned, the 93%-accurate triage was never the problem; the replies were.

A CX lead doing about 7,000 tickets a month put the actual requirement plainly on a sales call with eesel:

"I need an AI who is only handling the tickets that it's confident to handle and all the other ones, leave them alone."

A decision model gives you the confidence score that makes that possible. It doesn't give you the AI that handles the ticket. That's the gap between AI ticket classification and an AI helpdesk agent, and it's where the real cost lives.

A Clef decision is about $0.0001. A support engineer's week spent wiring routing rules, retries and fallbacks into a helpdesk is not. If you've ever set up Zendesk AI triage by hand, you know the routing rule is the easy 10%. The same is true of Freshdesk auto triage and every other queue.

If you've already compared best AI for ticket triage tools, Clef sits one layer below them: it's an ingredient for building one.

eesel for the decision and what comes after it

Clef is infrastructure. eesel is the employee. eesel is an AI teammate platform, and its AI helpdesk teammate joins your Zendesk, Freshdesk, Gorgias or Front queue, learns from your past tickets and help center, and makes the same calls Clef does as part of AI ticket routing (is this urgent, which team, can I answer it) and then acts: tags, routes, drafts, resolves, or hands off. Actions you mark as needing approval wait in an activity feed until a human says yes.

eesel dashboard showing the Zendesk activity feed with tickets marked Pending and Resolved, plus Approved, Rejected and Pending filters for held actions
eesel dashboard showing the Zendesk activity feed with tickets marked Pending and Resolved, plus Approved, Rejected and Pending filters for held actions

If you like Clef because it's programmable, you'll want the eesel CLI. It puts the same teammate from the dashboard into your terminal, and every command prints JSON, so Claude Code, Cursor or Codex can drive a full setup. npx @eesel/cli init chat-bubble --site <url> spins up a free workspace with no account. eesel integrations connects your helpdesk, eesel instructions sets standing rules like "only auto-reply above high confidence," eesel automations wires tagging and routing, and eesel approvals approve <id> clears held actions from a script. Changes sync both ways with the dashboard. The AI agent CLI post goes deeper.

Pricing is per ticket, not per token, which makes the AI vs human cost math easy to run: eesel pricing starts with 100 free credits and paid plans from $299 for 500 credits a month, where one credit covers a whole ticket or chat however long it runs. That's a different unit from Clef's because it's buying a different thing: the decision plus the work. If you'd rather build the triage layer yourself, Clef-flash is a great price. If you'd rather have tickets handled this week, try eesel on your own queue and see the drafts before you pay.

Frequently Asked Questions

How much does Cloudflare Clef cost?

Clef costs $0.24 per million input tokens on Workers AI, and Clef-flash costs $0.09. Neither model lists an output price, so input is the whole Cloudflare Clef pricing story. For comparison, see my Jev pricing breakdown.

Is Cloudflare Clef free to use?

Partly. Every Workers AI account gets 10,000 free neurons per day, which works out to roughly 900 short Clef decisions or 2,400 Clef-flash decisions a day. The weights are also Apache 2.0, so self-hosting costs nothing but your own hardware. If you just need tickets sorted without building anything, AI ticket classification tools are another route.

Why is Clef more expensive than Jev?

Per token, Clef is about 5.7x Jev's $0.042 rate, and Clef-flash is about 2.1x. Cloudflare's pitch is speed and accuracy: it reports a 209 ms median for Clef and 38.8 ms for Clef-flash against 524 ms for Jev. The TypeSafe Jev explainer covers the model Clef is compatible with.

What is the difference between Clef and Clef-flash pricing?

Clef (27B) is $0.24 per million input tokens and Clef-flash (9B) is $0.09, so Clef-flash is about 2.7x cheaper per token. In Cloudflare's own benchmarks Clef-flash is also faster, while Clef leads on harder intent sets like CLINC150. For most ticket triage traffic, test Clef-flash first.

Does Cloudflare charge for Clef output tokens?

No output price appears on the Clef model page or the Workers AI pricing table, which lists Clef as an input-only row. The answer is a set of probabilities, not generated text, so there is nothing long to bill on the way out. That is a real difference from an LLM, where output tokens usually cost several times more than input.

How much does Cloudflare Clef cost for ticket routing at scale?

At 3 million short decisions a month (about 500 tokens each), Clef runs about $357 and Clef-flash about $132 after the free allowance, plus the $5 Workers Paid minimum. If you send whole 4,000-token threads instead, multiply by eight. Compare that with what an AI support agent costs when the work after the decision is included.

Can I self-host Clef to avoid Cloudflare pricing?

Yes. Both models are open weights under Apache 2.0 on Hugging Face, and the Ollama build of Clef is an 18GB download. Cloudflare says it tested on a single NVIDIA H200, so budget for real GPU memory. Teams weighing hosted versus self-run models can also read the Qwen pricing guide, since Clef is built on Qwen.

Is Clef a replacement for an AI support agent?

No. Clef returns a decision such as "urgent, technical team, 0.92" and stops there; your code still has to route, reply and escalate. An AI helpdesk agent like eesel makes that decision and then acts on it inside Zendesk, Freshdesk or Gorgias.

Share this article

Kurnia Kharisma

Article by

Kurnia Kharisma

Kurnia is a software engineer and writer at eesel AI with two years of SEO experience, writing about AI tools, helpdesk software, and customer support. He pairs a developer's understanding of how these products are built with search-driven research into what actually ranks and resonates with the people searching for them.

Related Posts

All posts →
TypeSafe Jev pricing hero banner in rose and off-white, showing a low token cost per million
Trending

TypeSafe Jev pricing (2026): $0.042 per million tokens, output free

TypeSafe Jev pricing broken down: $0.042 per million input tokens, output free, no plan tiers yet, and what a System One model actually costs to run in production.

Kurnia KharismaKurnia KharismaSep 22, 2026
Illustrated hero banner for TypeSafe Jev, the ultrafast System One AI model, with a speed gauge
Trending

Is Jev really ultrafast? TypeSafe's System One model, tested

TypeSafe calls Jev an ultrafast System One model at 70-500ms a decision. Here is what the speed claim really means, where it holds up, and where it does not.

Rama AdiRama AdiSep 22, 2026
TypeSafe Jev hero banner in rose and off-white, illustrating a fast typed-decision model
Trending

TypeSafe Jev review: the 'System One' model that gives AI the properties of code

A hands-on TypeSafe Jev review: what the System One model actually does, whether the speed, price and 'can't hallucinate' claims hold, and where a typed-decision model fits real work.

Rama AdiRama AdiSep 21, 2026
Hand-drawn hero banner of a person feeding questions into a switchboard that routes them into labeled lanes, each with a confidence dial
Trending

Decision models explained: Jev, Clef, Strands Decider and the new AI category

Decision models return typed answers with confidence scores instead of text. What they are, how they work, every model you can use today, and where they fit in support.

KiraKiraOct 6, 2026
Hand-drawn illustration of two developers at a laptop below a cloud holding a friendly robot, connected to an identity shield, a locked database and a chip, with the AWS logo on an orange circle
Trending

Amazon Bedrock Managed Agents explained: OpenAI's agent harness inside your AWS account

Bedrock Managed Agents runs OpenAI's agent harness on AWS while your tools stay on your own compute. How it works, what the preview leaves out, and what it costs.

Rama AdiRama AdiOct 1, 2026
Qwen 3.8 Flash Next launch banner
Trending

Qwen 3.8 Flash Next: Alibaba's open-weight Qwen4 preview, explained

Qwen 3.8 Flash Next is Alibaba's open-weight preview of the Qwen4 architecture. Here is what it is, what it costs, and whether it belongs in your stack.

KiraKiraAug 30, 2026
Two people searching a document index powered by Cohere Embed 5 embeddings
Trending

Cohere Embed 5: Pro vs Fast, benchmarks, pricing, and how to use it

Cohere Embed 5 explained: what Pro and Fast are, the index-with-Pro trick, how the benchmarks hold up, what it costs, and when the token price stops mattering.

KiraKiraOct 1, 2026
Illustration of two people reviewing a Cohere Embed 5 pricing dashboard with the Cohere logo
Trending

Cohere Embed 5 pricing: what Pro and Fast really cost in 2026

Cohere Embed 5 pricing is $0.12 per 1M tokens for Pro and $0.08 for Fast, with images at $0.40. Here is the full table, Model Vault math, and the storage bill nobody prices in.

Rama AdiRama AdiOct 1, 2026
Illustration of one large coordinator fish routing work across a school of smaller fish, as analysts look on
Trending

Sakana Fugu Max pricing: every rate, and what it really costs

A full breakdown of Sakana Fugu Max pricing: the $2/$6 flat token rates, the subscription tiers, the orchestration-token catch, and how the real bill compares to Sonnet 5, GPT-5.6, and Kimi K3.

Kurnia KharismaKurnia KharismaSep 14, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free