How to Use Jev: Every Way to Call It, the Opus 5.5 Pairing, and Laya vs Kev vs Ollaya

TL;DR
TypeSafe's Jev dropped its waitlist on September 27. Every verified way to call it (API, Python SDK, llm CLI, Pydantic AI, Cloudflare, Vercel, OpenRouter), the Jev + Claude Opus 5.5 coding-agent pattern driving searches, and how the open alternatives Laya, Kev, and Ollaya compare.
Jev is TypeSafe AI's decision-only model: you send it a piece of state plus typed questions, and it returns booleans, choices, or scores with calibrated probabilities, at $0.042 per million input tokens with free output. As of September 27, 2026 there is no waitlist: anyone can create an account at console.typesafe.ai, and Jev is also callable through Cloudflare Workers AI, Vercel AI Gateway, Pydantic AI, and Simon Willison's llm CLI. The most useful way to think about it is as the fast "System One" half of an agent, with a frontier model like Claude Opus 5.5 as the slow "System Two" half - which is exactly the pairing developers have been searching for this week.
We covered what TypeSafe shipped and the launch benchmarks in the Jev release guide. This post is the practical follow-up: how to call it today, how to wire it into a coding agent next to Opus 5.5, and whether an open alternative is good enough for your workload.
Why Everyone Is Searching for Jev#
We pulled Google Trends Explore for "jev" (worldwide, web search) on September 28, 2026. Trends numbers are relative interest on a 0-100 scale, not search counts, but the shape is clear:
- Past 30 days: interest sat at 0-1 until launch day, then climbed from 2 on September 15 to a peak of 100 on September 21 (the same day Simon Willison published his writeup), cooled to 26 by September 27, and ticked back up to 39 on a partial September 28, the day after the waitlist came down.
- Top rising queries, past 30 days (all "Breakout"): "jev typesafe", "what is jev ai", "how to use jev", "jev open source", "jev openrouter", "jev ai api", "jev architecture", "laya vs jev", and "codex".
- Top rising queries, past 7 days: "opus 5.5" and "jev engineering for coding agents" are both Breakout, followed by "laya vs jev" (+2,250%), "laya jev alternative" (+1,850%), "jev as a judge" (+500%), "kev jev" (+350%), and "jevbench" (+300%).
That is the outline of this post: how to use it, where it plugs into coding agents alongside Opus 5.5, and how it stacks up against the open alternatives.
Every Way to Call Jev#
All of these were checked against the linked docs on September 28, 2026.
| Surface | Model id | Best for |
|---|---|---|
| TypeSafe API | jev-latest (currently resolves to jev-1.13.0) | Production, direct billing |
Python SDK typesafe-sdk | jev-latest | Typed Python code |
| llm-typesafe plugin | typesafe/jev-latest, alias jev | Shell pipelines and quick tests |
| Pydantic AI | typesafe:jev-latest | Agents that already use Pydantic AI |
| Cloudflare Workers AI | typesafe/jev | Edge Workers, billed through Cloudflare |
| Vercel AI Gateway | typesafe-ai/jev | Apps already on the gateway |
| OpenRouter Jev Router | typesafe/jev-router | Using Jev to route chat requests to other models |
The raw API#
One endpoint, one state, any number of questions, all answered in a single call:
curl -X POST https://api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer $TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"state": "Hi, I have been trying to connect my Stripe account for 3 days and it keeps failing. I am losing sales. Please help ASAP.",
"model": "jev-latest",
"questions": {
"urgency": { "type": "noul", "instructions": "Does this message express urgency?" }
}
}'
Python SDK#
Install with pip install typesafe-sdk, set TYPESAFE_API_KEY, and use the three question types directly. This is the sync example from the SDK docs:
from typesafe_sdk import TypeSafeClient, Choice, Noul, Score
with TypeSafeClient() as client:
response = client.system_one(
state={"document": "I was charged twice. Please fix this ASAP."},
questions={
"billing": Noul(instructions="Is this ticket about billing?"),
"tone": Choice(
instructions="What is the customer's tone?",
criteria={"calm": None, "frustrated": None, "angry": None},
),
"urgency": Score(
instructions="How urgent is this ticket?",
criteria=["can wait", "this week", "today"],
),
},
)
print(response.nouls["billing"].noul)
print(response.choices["tone"].choice)
print(response.scores["urgency"].score)
From the terminal with llm#
Simon Willison's plugin is the fastest way to get a feel for the model:
llm install llm-typesafe
llm keys set typesafe
cat message.txt | llm -m jev \
-s 'Which team should handle this message?' \
-o answer_type choice \
-o criteria '{"billing":"Charges, invoices, payments","technical":"Product issues"}'
His writeup is also the best short critique of the model: a bare float can hide bias you cannot audit, so do not use it to rank people.
Cloudflare, Vercel, and OpenRouter#
On Workers AI, the call is env.AI.run('typesafe/jev', { state, questions }), at the same $0.042 input and $0 output. Vercel AI Gateway lists it as typesafe-ai/jev. OpenRouter is the odd one out: it does not expose Jev's decision API as a chat model, it lists Jev Router (typesafe/jev-router, added September 25), a router that uses Jev to pick a model and reasoning effort for each chat request. OpenRouter lists the router itself at $0; the model it routes to is what you pay for. If you searched "jev openrouter" hoping for typed decisions, use one of the other surfaces.
In your coding agent#
TypeSafe ships an agent skill that teaches Claude Code or Codex the API, the question types, and the design patterns:
claude plugin marketplace add typesafe-ai/skills
claude plugin install typesafe@typesafe-ai
For other agents: npx skills add typesafe-ai/skills --skill typesafe-ai. Jev is not in OpenCode's model registry as of today.
Jev + Opus 5.5: The Pattern Behind the Searches#
There is no official Jev and Claude Opus 5.5 integration. The two launched a week apart - Jev on September 15, Opus 5.5 on September 22 - and The Batch covered them as separate stories in the same issue. The connection is a pattern developers arrived at on their own, and Amit Patriwala's post Jev + Opus 5.5: Why Your Coding Agent Needs Two Kinds of Thinking is the clearest statement of it:
- Opus 5.5 is System Two. It reads the repository, plans, reasons through the migration, and writes the code.
- Jev is System One. It sits at the decision points: pick the next action from a fixed vocabulary (plan, search, edit, run, test, verify, iterate), judge whether a proposed shell command is safe, decide whether a failing test is a real bug or environment flakiness, and score whether the task meets its acceptance criteria.
The economics are why it works. A decision call over 2,000 tokens of state costs about $0.00008 on Jev. The same call on Opus 5.5 at $4/$20 costs about $0.009 even with a short answer - roughly 100x more, and seconds slower instead of milliseconds. An agent loop makes dozens of these branch decisions per task, and none of them need a paragraph of reasoning.
Here is a minimal sketch of a command gate, using only the SDK calls shown above. The threshold and the escalation path are yours to own in code:
from typesafe_sdk import TypeSafeClient, Noul
def command_needs_review(command: str, task: str) -> bool:
with TypeSafeClient() as client:
r = client.system_one(
state={"task": task, "proposed_command": command},
questions={
"destructive": Noul(
instructions="Could this command delete data, leak secrets, or change anything outside the repo?"
),
},
)
# noul is the probability of "yes"; escalate anything not clearly safe
return r.nouls["destructive"].noul > 0.2
Two cautions from the pattern's own advocates and from TypeSafe's docs. First, Jev makes decisions but does not grant permissions: your application code still enforces what the agent may run. Second, Pydantic AI's docs list the model's weak spots - unreliable arithmetic and counting, literal reading of negations, 255 options per choice, and a 32K-token context for state plus questions - so keep the state you pass it small and specific.
This is the same split we recommended in the model routing orchestration layer playbook, with a new price floor for the routing side.
Jev as a Judge#
The "jev as a judge" searches trace to a September 22 arXiv paper from Carnegie Mellon, JEV-as-a-Judge: Accept When Confident, Escalate When Unsure. Against sixteen generative and reward-model judges with blinded human adjudication, Jev came within three percentage points of the strongest LLM judge on ordinary preference and factuality judgments at 0.36% of the fee. The gap widened on tasks that require checking a derivation or resisting a persuasively written wrong answer. A cascade that accepts Jev's confident verdicts and escalates the rest kept 99% of the comparator's accuracy at lower cost. Our paper summary has the details. If you run evals on every agent trace, that cascade is the cheapest upgrade on this page.
Laya vs Kev vs Ollaya: The Open Alternatives#
The other half of the search interest is "laya vs jev" and "jev open source." Jev itself is closed and hosted. The open options:
| Option | What it is | License | Runs |
|---|---|---|---|
| Laya (Convai Innovations) | Non-autoregressive System 1 decision model on a ModernBERT-large encoder, 421M-parameter typed-decisions checkpoint, multilingual variant for 100+ languages | Apache 2.0 | Locally, via the laya Python package |
| Kev (Jared Palmer) | Decision models on Qwen base models at 0.8B, 4B, 9B, and 27B, serving the same /v1/systemone shape | Apache 2.0 | Apple Silicon for 0.8B up to 80GB GPUs for 27B |
| Ollaya | An Ollama-style runtime for decision models (Laya, Kev, winnow, decider, and more), TypeSafe API compatible so the official SDK works against it | Apache 2.0 | macOS, Windows, Linux, Docker |
| GLM-5.3-Flash logits method (Edgeless Systems) | Reads option probabilities from a general LLM's logits in one forward pass | Open source on GitHub | Anywhere you can get logits |
How they compare, from two independent benchmarks:
- JevBench (Benchmark Heaven, MIT-licensed harness, run September 27 on 842 decisions): capability scores of 64.7 for Jev 1.13.0, 49.9 for Laya, and 40.9 for Kev 4B, with an Imajev-4B model on top at 66.3. Cost per 1,000 decisions: $0.040 for Jev, $0.019 for Kev, and $0.0029 for Laya.
- Edgeless Systems' 29-dataset comparison (September 24): GLM-5.3-Flash with their method and Jev were statistically tied (each more accurate on 10 datasets, median gap 0.7 points in Jev's favor, p = 0.64), while Laya trailed by 13-15 points. Latency depended on geography: 180ms vs 264ms from Germany, 299ms vs 164ms from the US. Kev's own repo reports its 27B model close to Jev on unseen data and the smaller sizes about four points behind.
The honest read: Laya wins on cost and latency and runs anywhere, but gives up real accuracy. Kev at 27B closes most of the gap if you have the GPU. Jev is still the accuracy-per-dollar leader among hosted options, and the open stack is good enough that "we cannot send this data to a third party" is no longer a reason to skip decision models.
Which One Should You Use?#
| Situation | Pick |
|---|---|
| You want the most accurate hosted decision model today | Jev via the TypeSafe API or SDK |
| You are already on Workers or Vercel | Jev through Cloudflare (typesafe/jev) or the AI Gateway (typesafe-ai/jev) |
| You are building a coding agent on Opus 5.5 or GPT-6 | Jev at the branch points, the frontier model for planning and code |
| Data cannot leave your hardware | Kev (27B if you can) or Laya, served through Ollaya |
| You want one API for chat that picks the model for you | OpenRouter's Jev Router |
| You need images in the decision | A multimodal LLM with the logits method; Jev and Laya are text-only |
FAQ#
How do I use Jev?#
Create an account at console.typesafe.ai (there is no waitlist since September 27, 2026), set TYPESAFE_API_KEY, and call POST https://api.typesafe.ai/v1/systemone with a state and typed questions, or use the typesafe-sdk Python package. You can also call it through Cloudflare Workers AI, Vercel AI Gateway, Pydantic AI, or the llm-typesafe plugin.
How much does Jev cost?#
$0.042 per million input tokens, and output is free. Cloudflare Workers AI lists the same rates. TypeSafe has said the price may be subsidized and that it expects prices to go down.
Can Jev replace Claude Opus 5.5?#
No. Jev cannot write code, text, or plans; it only answers typed questions. The useful pattern is to pair them: Opus 5.5 plans and writes, Jev handles the high-volume branch decisions like next-action selection, command safety, and completion checks.
Is Jev on OpenRouter?#
OpenRouter lists Jev Router (typesafe/jev-router), which uses Jev to route chat requests to other models. It does not expose Jev's typed decision API. For typed decisions, use the TypeSafe API, Cloudflare, Vercel, or Pydantic AI.
What is the difference between Laya and Jev?#
Laya is an open-weight Apache 2.0 decision model from Convai Innovations that runs locally on a small encoder; Jev is TypeSafe's hosted, closed model. Independent benchmarks put Jev well ahead on accuracy (roughly 13-15 points on Edgeless Systems' datasets) while Laya is cheaper and faster.
Is there an open-source Jev?#
Not Jev itself. The closest open options are Kev (Apache 2.0, 0.8B to 27B) and Laya (Apache 2.0), which you can run with the Ollaya runtime using the same request format.
Sources#
Last updated: September 28, 2026
Continue Reading#
- TypeSafe Jev: the First Decision-Only Model Class, Benchmarked and Priced - the launch details, workflow evals, and pricing
- Claude Opus 5.5 Release Guide - the System Two half of the pairing, with the OpenCode setup
- GPT-6 Astra Release Guide - OpenAI's top model, and when frontier pricing is worth it
- LLM Router Comparison 2026 - where Jev Router fits among existing routers
- AI Model Routing and Orchestration Layer - the decomposition pattern that makes decision models pay off
Get the next comparison like this in your inbox
One email a week on AI Models and the rest of the AI dev stack. Free.
Read next on Claude Code
TypeSafe Jev: the First Decision-Only Model Class, Benchmarked and Priced
TypeSafe's Jev is the first 'System One' model: no text generation, typed decisions with calibrated probabilities, 70-500ms responses, and $0.042 per million input tokens. Benchmarks, pricing, and the API.
7 min readClaude Opus 5.5 Developer Guide: API Examples, Claude Code Setup, Pricing, and When to Use It
Claude Opus 5.5 (claude-opus-5-5) is Anthropic's new default Opus: $4/$20 per million tokens, $0.20 cache reads, 1M context, thinking always on with medium default effort. Runnable TypeScript and Python SDK examples, Claude Code setup, before/after prompts, and a decision guide vs Sonnet 5, Haiku 4.5, and Fable 5.1.
11 min readLLM Routers Compared: LiteLLM vs Portkey vs OpenRouter in 2026
A practical comparison of LLM routing tools - LiteLLM, Portkey, and OpenRouter - covering cost management, fallbacks, caching, and when to use each for production AI applications.
10 min readTechnical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.








