Cloudflare Clef Decision Models: Jev-Compatible, Benchmarked and Priced

TL;DR
Cloudflare's Clef is a 27B Apache-2.0 decision model on Workers AI at $0.24 per million input tokens, and Clef-flash is a 9B model at $0.09 with a 38.8 ms median decision, both Jev-compatible, with an RL fine-tuning service attached.
Cloudflare released Clef and Clef-flash on October 1: two decision models with Apache 2.0 weights, hosted on Workers AI at $0.24 and $0.09 per million input tokens. They take a state plus a schema of typed questions and return a probability for every allowed answer in a single forward pass. Clef is a 27B multimodal model post-trained from Qwen3.8-27B; Clef-flash is its 9B sibling, and Cloudflare's own run of the Jev Decision Index puts its median decision at 38.8 ms against Jev's 524.1 ms. Both speak Jev's SystemOne API, and the release pairs them with an RL fine-tuning service.
This post will be updated as details are confirmed.
What shipped#
@cf/cloudflare/clef(27B) and@cf/cloudflare/clef-flash(9B) on Workers AI, with weights on Hugging Face under Apache 2.0.- No text generation and nothing to parse: you send a
state(text or JSON) and up to 64 typed questions (noultrue/false,choicenamed options,scoreordered options); the response carries a probability for every option, so code can branch directly. - Non-autoregressive scoring: Clef does a prefill-only pass over the state, then a joint schema head scores every option of every question in parallel. Post-training adds a rank-256 LoRA and a Brier loss term to calibrate the probabilities, with an RL stage Cloudflare calls RLCD.
- 65,536-token context and a vision encoder on Clef. Cloudflare's comparison sheet puts Jev at 32k context and text-only input.
- The RL fine-tuning service: forward-deployed engineers now, self-serve platform later, built on AI Gateway (capture your traffic as a dataset), Workers AI (rollouts), Containers (RL sandbox), a new Trainer, and bring-your-own-model redeploy on Workers AI.
Official Sources#
| Resource | What it covers |
|---|---|
| Cloudflare announcement: Clef decision models | Launch, architecture, benchmarks, RL service |
| Workers AI docs: clef | API, pricing, limits |
| Workers AI docs: clef-flash | API, pricing, limits |
| Clef model card (Hugging Face) | Weights, usage, full Decision Index table |
| TypeSafe Jev: the First Decision-Only Model | Jev's launch pricing and workflow evals, verified September 16 |
Where Clef beats Jev, and where it does not#
Cloudflare's published run of the Jev Decision Index 0.2.1 (vendor-run, so score it like any launch deck):
- Clef leads BFCL case-exact accuracy (98.5 vs 95.8), BANKING77 macro-F1 (94.2 vs 79.7) and CLINC150+OOS macro-F1 (97.4 vs 89.3).
- Jev still leads GPQA Diamond (78.3 vs 48.0), MMLU-Pro (82.7 vs 65.9), When2Call accuracy (81.0 vs 72.4) and BRIGHT retrieval (47.5 vs 45.9 nDCG@10).
- On the four Typesafe workflow evals, Clef wins invoice processing and security incidents, Clef-flash wins customer service, and Jev holds agent trace observability.
- Median decision latency: Clef 209.3 ms, Clef-flash 38.8 ms, Jev 524.1 ms. p95: 238.6 ms, 122.4 ms, 536.0 ms.
Chart: Cloudflare, from the Clef announcement. Scores are Cloudflare's own run of the Jev Decision Index.
Pricing#
| Model | Parameters | Input price (per 1M tokens) | Median decision | p95 |
|---|---|---|---|---|
| Clef | 27B | $0.24 | 209.3 ms | 238.6 ms |
| Clef-flash | 9B | $0.09 | 38.8 ms | 122.4 ms |
| Jev (TypeSafe, verified September 16) | n/a | $0.042 | 524.1 ms | 536.0 ms |
Worked example: 1 million decisions with a 2,000-token state is 2 billion input tokens, which costs $480 on Clef, $180 on Clef-flash and $84 on Jev. The trade is latency against price: Clef-flash costs about 2.1x Jev's input rate for roughly 13x lower median latency, while full Clef costs about 5.7x and is still 2.5x faster. Workers AI lists no output charge for either model, because there is no generated text to bill.
Run it: the fastest path is Workers AI#
These are decision endpoints, not chat models, so there is no OpenCode model id to point at. The command below is copied from the Workers AI docs, and I did not run it here (no Cloudflare account in this environment):
curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef-flash \
-X POST \
-H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" \
-d '{
"model": "clef-flash",
"state": "Checkout has been failing for every customer for the last hour.",
"questions": {
"urgent": { "type": "noul", "instructions": "Is this support request urgent?" },
"team": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {
"billing": "Payments, invoices, and refunds",
"technical": "Outages, errors, and configuration",
"sales": "Plans and upgrades"
}
},
"severity": {
"type": "score",
"instructions": "How severe is the customer impact?",
"criteria": ["No impact", "Minor", "Major", "Critical"]
}
}
}'
The response gives answers.urgent as a probability, answers.team as a chosen option with per-option probabilities, and answers.severity as a weighted score. To self-host, the model card ships a systemone() helper that accepts the same request body as the hosted API; its tested stack is torch 2.11 with transformers 5.10.2 on a single H200 for the 27B. Limits to design around: 4 images per request at 4 MiB each, a 13 MiB request body cap, long states truncated to the 65,536-token window, and fine-tuning that is still a conversation with Cloudflare's forward-deployed engineers rather than a self-serve job.
What people are actually saying#
On the Hacker News thread (145 points, 50 comments at the time of writing), the excitement is real and so is the pushback:
- The strongest counter-case: one commenter argues this is open weights, not open source, because the weights are Apache 2.0 but the data and training pipeline are not published, so nothing here is reproducible from the announcement.
- Several commenters treat the category as easy to copy now that the interface exists; one summary is that any pretrained LLM can be adapted to work this way, and another notes a small BERT-style classifier can be trained on a laptop for a narrow case and avoid the network round trip entirely.
- The calibration argument is the most useful practitioner thread: training a general classifier is not the hard part, calibrated probabilities are, and then the counterpoint that calibration barely matters when the alternative was an LLM's raw softmax scores. Another commenter says they have run a classifier-first, LLM-fallback router in production for a year.
- On price, one commenter notes Clef's $0.24 per million input tokens is roughly 6x Jev, while Clef-flash at $0.09 is the competitive SKU. A second thread on r/LocalLLaMA linked the Hugging Face card within half an hour of the announcement.
The angle: 16 days to commodity#
TypeSafe opened Jev on September 15. Sixteen days later, the interface it defined has a larger-context, vision-capable, faster, API-compatible family from a company that already operates the edge runtime, gateway and containers much of this stack runs on. Cloudflare wins twice: it sells the inference per input token, and it sells the fine-tuning path, where a customer's own logged traffic through AI Gateway becomes the training set. TypeSafe loses the model-differentiation argument on everything except raw input price and a few reasoning-heavy evals, and independent decision-model vendors, including open replicas like Jeff, now have to compete on calibration, workflow evals and fine-tuning data rather than request shape. The second-order effect is the one to watch: decision endpoints are becoming a bundled feature of inference platforms, so the durable business moves up-stack into RL fine-tuning and evaluation, and the per-token price of a decision falls toward zero. For a builder, that is the good news: at $0.09 per million and 38.8 ms, routing decisions can sit in the hot path of an agent loop instead of behind a heuristic.
FAQ#
What is Cloudflare Clef?#
A decision model: you send a state and a schema of typed questions, and it returns a probability for every allowed option in one forward pass, with no generated text. It is available on Workers AI as @cf/cloudflare/clef (27B) and @cf/cloudflare/clef-flash (9B), with Apache 2.0 weights on Hugging Face.
How much does Clef cost?#
$0.24 per million input tokens for Clef and $0.09 for Clef-flash, per the Workers AI model pages. A million decisions with 2,000-token states costs $480 and $180 respectively. There is no output charge because there is no generated text.
Is Clef open source?#
The weights are Apache 2.0 on Hugging Face. The training data and pipeline are not published, so it is open weights rather than a fully reproducible open-source release, a distinction Hacker News commenters raised immediately.
Continue Reading#
- TypeSafe Jev: the First Decision-Only Model, Benchmarked and Priced - the model class Cloudflare is measuring itself against
- Jeff: Jev-Style Decision Models Trained at Home on One GPU - the open replication route with verified latency numbers
- How to Use Jev - every hosted surface that speaks the same request shape
- AI Model Routing and Orchestration Layer - where a decision endpoint fits in an agent stack
- Cloudflare Folds Workers AI Into AI Gateway - the control plane the fine-tuning service captures data through
Sources#
| Source | URL |
|---|---|
| Cloudflare announcement: Clef decision models (fetched October 1, 2026) | https://blog.cloudflare.com/clef-decision-models/ |
| Workers AI docs: clef | https://developers.cloudflare.com/workers-ai/models/clef/ |
| Workers AI docs: clef-flash | https://developers.cloudflare.com/workers-ai/models/clef-flash/ |
| Clef model card (Hugging Face) | https://huggingface.co/Cloudflare/clef |
| Hacker News thread (145 points, 50 comments) | https://news.ycombinator.com/item?id=49923692 |
| r/LocalLLaMA thread | https://www.reddit.com/r/LocalLLaMA/comments/1wv4zzi/clef_open_weights_decision_model_by_cloudflare/ |
| TypeSafe Jev pricing and launch (our coverage, September 16) | https://developersdigest.tech/blog/typesafe-jev-system-one-models-release-guide-2026 |
Get the next deep dive like this in your inbox
One email a week on News and the rest of the AI dev stack. Free.
Read next on AI coding tools
TypeSafe Jev: the First Decision-Only Model, Benchmarked
TypeSafe's Jev is the first System One model: typed decisions with calibrated probabilities, 70-500ms responses, and $0.042 per million input tokens.
7 min readJeff: Jev-Style Decision Models Trained at Home on One GPU
An independent project fine-tunes Qwen3.5 and Gemma 4 into 0.8B and 2B decision models that answer in 22-28 ms using Jev's request shape. The verified numbers, the run commands, and where the benchmark stops matching real work.
6 min readHow to Use Jev: Every Way to Call It, the Opus 5.5 Pairing, and Laya vs Kev vs Ollaya
TypeSafe's Jev dropped its waitlist on September 27. Every verified way to call it (API, Python SDK, llm CLI, Pydantic AI, Cloudflare, Vercel, OpenRouter), the Jev + Claude Opus 5.5 coding-agent pattern driving searches, and how the open alternatives Laya, Kev, and Ollaya compare.
10 min readNew here? Start with
Technical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.




