OpenAI Decisions API vs Jev vs Clef: Price, Speed, Which to Pick

TL;DR
OpenAI Decisions API costs $0.10 per million input tokens, 2.4x Jev. Pick Jev for cheap text decisions, Clef-flash for speed, Decisions API for images.
Pick TypeSafe's Jev for high-volume text decisions where price matters most, Cloudflare's Clef-flash when latency matters most, and OpenAI's Decisions API when the input is an image or you need OpenAI's data-residency and HIPAA terms. All three take a piece of state plus typed questions and return probabilities instead of text. The deciding numbers: Jev charges $0.042 per million input tokens, Clef-flash $0.09, and the Decisions API $0.10, so a million decisions over a 2,000-token state costs $84, $180 and $200 respectively. Output is free on all three.
OpenAI opened the Decisions API to every developer in public beta on October 6, three weeks after Jev launched and five days after Cloudflare shipped Clef. That makes "decision model" a three-vendor category, and the question in the OpenAI developer forum this week is the obvious one: is OpenAI's version worth switching to? This page answers it with list prices, the limits each vendor documents, and the first community tests.
Last updated: October 9, 2026. Prices, limits and request shapes were checked against each vendor's docs on that date.
Official Sources#
| Resource | What it confirms |
|---|---|
| OpenAI Decisions API guide | POST /v1/decisions, gpt-6-luna only, question types, $0.10 input pricing, image rules, ZDR and data residency |
| OpenAI public beta announcement | Public beta on October 6, "up to 10x faster" than Luna through the Responses API |
| OpenAI API pricing and the gpt-6-luna model page | gpt-6-luna rates on the Responses API, cached input and cache writes, the 272K long-context rate and the regional processing premium |
| TypeSafe docs: Models | Jev 1.13 price, rate limits, context, text-only input |
| Workers AI: clef and clef-flash | Clef pricing, 65,536-token context, vision, 64-question limit |
| Cloudflare Clef announcement | Apache 2.0 weights, Cloudflare's latency run against Jev |
| AI SDK Decisions docs | experimental_decide across TypeSafe, OpenAI, Anthropic and Google |
At a Glance#
| OpenAI Decisions API | TypeSafe Jev | Cloudflare Clef-flash / Clef | |
|---|---|---|---|
| Model | gpt-6-luna | jev-1.13.0 (jev-latest) | @cf/cloudflare/clef-flash (9B), @cf/cloudflare/clef (27B) |
| Input / MTok | $0.10 | $0.042 | $0.09 / $0.24 |
| Output | Free | Free | No output charge listed |
| Caching | None (no cache-read or cache-write charges) | None listed | None listed |
| Context | Not stated in the guide; long-context pricing applies | 64k per request, 32k for state plus longest question | 65,536 tokens |
| Images | Yes, inline base64 data URLs only | No, text only | Yes, up to 4 embedded images |
| Question types | predicate, choice, score | noul, choice, score | noul, choice, score (1 to 64 per request) |
| Weights | Closed | Closed | Apache 2.0 on Hugging Face |
| Status | Public beta, GA expected "in the coming weeks" | Open signup, limits "adjusting dynamically" | Live on Workers AI |
Two rows deserve a note. OpenAI's guide does not state a context window for /v1/decisions, but it does say "long-context input pricing multipliers apply," which implies requests can go past the standard tier. A forum member replied in the launch thread that "the standard context tier is up to 272K input tokens" and that "the 1M-token context window is also available"; that is a community answer, not documentation. On the gpt-6-luna model page, prompts over 272K input tokens are billed at 2x input for the full request. And Jev's published limits are 100,000 tokens per second and 80 requests per second, with TypeSafe warning that they can change without notice while it adds capacity.
Cost per Task#
The cleanest comparison is a routing or triage call: a 2,000-token state (a support ticket plus some account context) with a few questions, run a million times a month. That is 2 billion input tokens.
| Option | Input rate | 1M decisions at 2,000 tokens |
|---|---|---|
| TypeSafe Jev | $0.042 | $84 |
| Cloudflare Clef-flash | $0.09 | $180 |
| OpenAI Decisions API | $0.10 | $200 |
| Cloudflare Clef (27B) | $0.24 | $480 |
Two adders apply to the Decisions API on top of that $0.10: OpenAI's guide says "regional processing premiums and long-context input pricing multipliers apply." The gpt-6-luna model page puts regional processing at a 10% premium and long context (over 272K input tokens) at 2x input, so a request that uses regional processing would cost $0.11 per million, or $220 for the run above. A 2,000-token state never hits the long-context rate.
The Decisions API is about 2.4x Jev on list price, which matches what a developer in OpenAI's forum worked out: "It works out to ~2.4x that of Jev?"
The less obvious comparison is against gpt-6-luna on the ordinary Responses API. There, OpenAI's pricing page lists $0.10 input, $0.01 cached input and $0.50 output per million tokens. If 1,500 of your 2,000 tokens are a fixed instruction prefix that stays cached, and the model returns about 10 tokens of JSON, a call costs roughly $0.00007, or about $70 per million (plus a one-off cache write at $0.125 per million tokens each time the prefix is cached). That is cheaper than the Decisions API, because the Decisions API has no cache discount at all. One developer in the launch thread made exactly this point: "that is unfortunate as it makes classification tasks economically unattractive to switch from luna responses api to decisions api." What the Responses route does not give you is native per-option probabilities or the speed gain, so price is not the whole decision.
Speed#
Nobody has published an apples-to-apples latency test of all three yet, so treat every number here as indicative:
- OpenAI says the Decisions API decides "up to 10x faster than GPT-6 Luna through the Responses API." It publishes no millisecond figure.
- Cloudflare's own run of the Jev Decision Index puts the median decision at 38.8 ms for Clef-flash, 209.3 ms for Clef and 524.1 ms for Jev (p95: 122.4, 238.6 and 536.0 ms). Vendor-run, so score it like a launch deck.
- A forum test of the Decisions API against Jev reported, for multi-step conversational use, a median turn of "about 0.3 s against Luna's 1.6 s" in Jev's favour, while noting Luna "was more consistent per call." Another tester saw about 0.8 seconds for image inputs over a 3 Mbps connection.
If the decision sits in a hot path, such as choosing an agent's next action inside a loop, Clef-flash is the only option with a published sub-50 ms median. For background triage, all three are fast enough.
Accuracy and Calibration#
This is where the first week of testing is most useful, because all three vendors sell "calibrated probabilities" and the early evidence says the question type matters more than the vendor.
- Predicates look calibrated, choices do not. A developer ran 3,000 Decisions API requests on loaded-coin and marble-jar problems and posted the results. Predicate questions returned the exact probabilities (70% heads, 50% red). The same question as a
choicereturned 98% heads and 86% red, and moving red from first to last in the choice list dropped its probability to 73%. Their summary: "Changing the choice order changes the probabilities!" They foundscoreworse still on that toy test. - Jev has a documented version of the same weakness. TypeSafe's own Jev 1.13 jaggedness page lists a bias toward the first option in a Choice among its nine known failure modes, along with unreliable counting, math and date comparison.
- Jev led on judgment calls in one early comparison. In the launch thread, a developer who ran a few hundred word-game and conversational questions through both reported that Luna "made about three times as many confident wrong answers" on nuanced judgment calls, while the two were about equal on narrow yes/no factual checks. Their bottom line: "Jev is the better default." They also called it "a strong early signal rather than a general benchmark."
- Clef's benchmark wins are vendor-reported. Cloudflare's run has Clef ahead on tool-call and intent-classification sets and Jev ahead on reasoning-heavy sets like GPQA Diamond and MMLU-Pro. The full breakdown is in our Clef benchmark post.
The practical rule that falls out: phrase high-stakes decisions as yes/no questions (predicate on OpenAI, noul on Jev and Clef), set thresholds from your own labeled examples, and treat choice probabilities as a ranking, not a calibrated estimate, until you have measured them on your data.
Switching Cost#
Jev and Clef speak the same request shape: a state plus a map of named questions with criteria, sent to a SystemOne-style endpoint. Moving between them is mostly a URL and model change. The Decisions API is different: input instead of state, a questions array with a name on each item, predicate instead of noul, and choices or levels arrays instead of criteria.
curl https://api.openai.com/v1/decisions \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6-luna",
"input": "I was charged twice for my order.",
"questions": [{
"type": "choice",
"name": "department",
"instructions": "Which department should handle this complaint?",
"choices": [
{"value": "billing", "description": "Payments, invoices, and refunds."},
{"value": "technical", "description": "Problems using the product."},
{"value": "shipping", "description": "Delivery and tracking."},
{"value": "other", "description": "Requests outside these categories."}
]
}]
}'
That request is copied from OpenAI's guide; I did not run it here. If you want to avoid rewriting calls per vendor, the AI SDK's experimental_decide takes one question format and maps it to typeSafeAi.decisionModel('jev-latest'), openai.decisionModel('gpt-6-luna'), or language-model adapters for Anthropic and Google. The SDK labels the API experimental and says it "may change in patch releases," and it notes that the Anthropic and Google adapters return prompted estimates rather than native probability distributions. Our Jev how-to covers every other surface that speaks Jev's format, including Vercel AI Gateway and Pydantic AI.
Pick by Task#
| Task | Pick | Why |
|---|---|---|
| High-volume text triage and routing | Jev | Lowest input price ($0.042), and one early test found fewer confident wrong answers on judgment calls |
| Next-action choice inside an agent loop | Clef-flash | 38.8 ms median in Cloudflare's run, 9B weights you can also self-host |
| Image checks (damage, content policy, "does this screenshot contain code") | Decisions API or Clef | Jev is text only; OpenAI accepts inline base64 images, Clef up to 4 per request |
| Regulated data that already lives on OpenAI | Decisions API | ZDR and HIPAA for eligible customers, data residency in the US and Europe (EEA plus Switzerland) |
| You need to fine-tune on your own traffic | Clef | Apache 2.0 weights plus Cloudflare's RL fine-tuning service; Jev serves the same weights to every account |
| Cached, repeated prompts at the lowest cost | Luna on the Responses API | $0.01 cached input beats the uncached Decisions API, if you can live without native probabilities |
If you are already running Jev in production, there is no strong reason to move this week: the Decisions API is in beta, costs more per token, and every threshold you tuned would need recalibrating. If you are starting fresh on OpenAI and your inputs include images, it is the shortest path. The bigger story is the one we flagged when Clef launched: three vendors shipping the same primitive inside a month means the per-decision price is heading toward zero, and the lasting differences will be calibration, fine-tuning and where your data is allowed to go. For where a decision endpoint sits in a larger agent stack, see our model routing guide.
FAQ#
Is the OpenAI Decisions API cheaper than Jev?#
No. The Decisions API charges $0.10 per million input tokens with gpt-6-luna, and Jev charges $0.042, so OpenAI is about 2.4x the price. Neither charges for output, and neither lists a cache discount.
What model does the OpenAI Decisions API use?#
Only gpt-6-luna, called through the dedicated POST /v1/decisions endpoint. OpenAI's guide says it is the only model currently available and that the API is in public beta with general availability expected in the coming weeks.
Can the Decisions API read images?#
Yes. You combine input_text and input_image parts in a user message, and images must be inline base64 data URLs; hosted image URLs and file_id inputs are not supported. Jev is text only, while Clef accepts up to four embedded PNG, JPEG or WebP images per request.
Is Clef compatible with Jev?#
Yes, Clef uses Jev's SystemOne request shape (state, named questions, noul, choice and score), so switching is mostly an endpoint and model change. The OpenAI Decisions API uses a different shape and needs its own request code or an adapter such as the AI SDK's experimental_decide.
Which decision model is fastest?#
On published numbers, Clef-flash: Cloudflare's run puts its median decision at 38.8 ms against 524.1 ms for Jev. OpenAI publishes only a relative claim, "up to 10x faster" than Luna through the Responses API, so measure on your own workload before choosing on speed.
Sources#
| Source | URL |
|---|---|
| OpenAI Decisions API guide | developers.openai.com/api/docs/guides/decisions |
| OpenAI forum: Decisions API public beta announcement and replies | community.openai.com/t/decisions-api-is-now-available-in-public-beta/1403877 |
| OpenAI forum: toy experiments with the Decisions API | community.openai.com/t/results-from-toy-experiments-with-decisions-api/1404010 |
| OpenAI API pricing | developers.openai.com/api/docs/pricing |
| OpenAI gpt-6-luna model page | developers.openai.com/api/docs/models/gpt-6-luna |
| TypeSafe models and limits | docs.typesafe.ai/models |
| TypeSafe Jev 1.13 jaggedness | docs.typesafe.ai/model-jaggedness/jev-1.13 |
| Workers AI clef and clef-flash model pages | developers.cloudflare.com/workers-ai/models/clef/ |
| Cloudflare Clef announcement | blog.cloudflare.com/clef-decision-models/ |
| AI SDK Decisions | ai-sdk.dev/docs/ai-sdk-core/decisions |
Continue Reading#
- TypeSafe Jev Pricing, Rate Limits and Benchmarks (2026) - the full Jev picture: workflow evals, jaggedness list and every hosted surface
- Cloudflare Clef Decision Models: Jev-Compatible, Benchmarked and Priced - Cloudflare's benchmark breakdown and the RL fine-tuning service
- How to Use Jev: Every Way to Call It - the Jev plus Claude Opus 5.5 coding-agent pattern and the open alternatives
- OpenAI DevDay 2026 Recap: What Shipped and What to Test - where the Decisions API was first announced, next to GPT-6.1 Sol
- Jeff: Jev-Style Decision Models Trained at Home on One GPU - the local route if none of the hosted prices work for you
Get the next comparison like this in your inbox
One email a week on AI Models and the rest of the AI dev stack. Free.
Read next on AI coding tools
TypeSafe Jev Pricing, Rate Limits and Benchmarks (2026)
TypeSafe Jev costs $0.042 per million input tokens, output free, at 80 requests/second. Limits, benchmarks, failure modes and every way to call it, re-verified October 9.
7 min readCloudflare Clef Decision Models: Jev-Compatible, Benchmarked and Priced
Cloudflare's Clef is a 27B Apache-2.0 decision model on Workers AI at $0.24 per million input tokens, and Clef-flash is a 9B model at $0.09 with a 38.8 ms median decision, both Jev-compatible, with an RL fine-tuning service attached.
5 min readHow to Use Jev: Every Way to Call It, the Opus 5.5 Pairing, and Laya vs Kev vs Ollaya
TypeSafe's Jev dropped its waitlist on September 27. Every verified way to call it (API, Python SDK, llm CLI, Pydantic AI, Cloudflare, Vercel, OpenRouter), the Jev + Claude Opus 5.5 coding-agent pattern driving searches, and how the open alternatives Laya, Kev, and Ollaya compare.
10 min readNew here? Start with
Technical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.







