Jeff: Jev-Style Decision Models Trained at Home on One GPU

TL;DR
An independent project fine-tunes Qwen3.5 and Gemma 4 into 0.8B and 2B decision models that answer in 22-28 ms using Jev's request shape. The verified numbers, the run commands, and where the benchmark stops matching real work.
An independent project called Jeff shipped v1.1 on September 29: two fine-tunes of Qwen3.5 (0.8B and 2B) and a Gemma 4 variant that accept the same request shape as TypeSafe's Jev and return a calibrated probability for every option in a single forward pass. The 2B reaches 82.0 on the project's five-benchmark panel against Jev's published 83.0, the 0.8B answers in 22 ms on an RTX PRO 6000 and 28 ms on an M4 Max, and all training ran on one workstation GPU with synthetic data from an open model.
The point is not the panel score. The decision-model interface Jev defined now has an open, fine-tunable replica small enough to sit inside your own process. The useful question is where it is already good enough, and where the gap to real work is still the whole story.
What shipped#
Jeff is three models on Hugging Face with a serving stack: Jeff-Qwen3.5-0.8B, Jeff-Qwen3.5-2B, and Jeff-Gemma4-E2B, plus a chess fine-tune as a worked example. Code is MIT, weights Apache 2.0, and the project is explicit that it is a fork of the open AutoJev recipe and not affiliated with TypeSafe.
The v1.1 changes matter more than the version number: the Qwen models now choose among up to 254 options instead of 26, after a user report that v1.0 never picked an option past the 26th (the changelog credits the reporter). On the long-list test the 0.8B goes from 40.3 percent to 94.7 percent, calibration error drops from 0.049 to 0.021, and the release publishes the final-epoch checkpoint instead of the lowest-dev-loss one. The 2B still dipped from 83.1 to 82.0 overall, mostly on JudgeBench.
Run it#
From the project's README, on a CUDA machine (I did not run this here; the commands are copied from the repo):
uv sync --no-default-groups --extra cuda
uv run --no-default-groups hf download mstrasser/Jeff-Qwen3.5-0.8B --local-dir checkpoints/jeff-0.8b
JEFF_CHECKPOINT=checkpoints/jeff-0.8b PORT=8765 uv run --no-default-groups jeff-serve
The server exposes Jev's /v1/systemone shape: one state string plus any number of typed questions (choice, noul, score), each answered with per-option probabilities. The 0.8B is 1.7 GB of weights, the 2B is 4.2 GB, the Gemma variant is 9.3 GB. The README reports 22 ms per decision on an RTX PRO 6000 and 28 ms on an M4 Max via MLX, versus 463 ms on 32 CPU threads, so on Apple silicon this is local-runtime territory.
The benchmark headline and the field gap#
The panel is 4,599 questions across five public benchmarks, with Jeff's numbers measured locally and Jev's figures taken from published results on a different sample:
| Benchmark | Jeff 0.8B | Jeff 2B | Jev (published) |
|---|---|---|---|
| Overall (5 benchmarks) | 79.1 | 82.0 | 83.0 |
| Financial PhraseBank | 95.7 | 94.7 | 77.0 |
| RAGTruth | 85.6 | 87.7 | 77.3 |
| BBH | 64.9 | 68.7 | 94.3 |
| JevBench hard tier | 46.7 | 57.1 | 73.3 |
Jeff wins the classification and grounding rows and loses badly on the reasoning-heavy ones, which is what a 0.8B to 2B model should do: fast, calibrated choices between options you describe, not multi-step reasoning.
The Hacker News thread, at 569 points and over 200 comments, is where the panel meets production. The top practitioner report compared Jeff to Jev in their own classification cases and landed at 70 percent against 94, calling that "unacceptable." Another commenter ran job-ad classification beside Jev and Gemini 2.5 Flash Lite and found the 0.8B "completely useless" while the 2B still missed a core job-type label. One point on the panel; 24 in the field. That mismatch is the honest cost of zero-shot generality: the panel rewards tasks near the training distribution, and production punishes the tail.
The fine-tune escape hatch#
The counter-argument is in the same repo. A voice-navigation fine-tune on roughly 11,000 app-specific examples took about half an hour on one GPU and moved held-out accuracy from 31.7 percent to 95.8 percent. A chess fine-tune on 600,000 Lichess positions labelled by Stockfish, about 3.5 hours on one GPU, took held-out puzzle accuracy from 15.5 percent zero-shot to 55.8 percent. Base training takes about two hours for the 0.8B and 3.5 for the 2B on the same box.
That is a different product from Jev's. Jev sells a hosted model that generalizes; Jeff sells a local starting point you adapt, a small-model deployment you own outright. If your classification task is stable and high-volume, the fine-tune path is the one to price.
Why this matters beyond one repo#
The Latent Space move: if you have a classification workload, download the 0.8B, point jeff-serve at it, and measure it on 200 of your own examples before you argue about benchmarks. The repo's own cautions apply: describe options consistently and in words, ask independent questions together, and fine-tune when zero-shot does not hold.
The second-order effect: the small end of the decision-model market is being commoditized. Classification parity with a priced model is now a one-GPU result, which means the moat is not the interface but the data pipeline, the calibration, and the breadth of zero-shot generalization. Jev still holds the generality claim; the open ecosystem holds the price floor. Watch the routing layer absorb cheap local decision models as just another route target.
What people are actually saying#
- The field reports are the sharpest disagreement. On Hacker News, the top comparison put Jeff at 70 percent against Jev's 94 on their own classification cases, and a job-ad test found the 0.8B unusable for the task. One commenter called the panel-versus-field contrast the whole story with zero-shot classification.
- The counter-case is not a bigger LLM, it is embeddings. Several commenters argued a sentence embedder plus logistic regression, an SVM, or a small MLP matches these models on classic classification and runs in under a millisecond. One reported that pairing matching Jev and Laya on AG News, Emotion, MASSIVE Intent, and Banking77 while failing on reasoning datasets. If your task is not reasoning, that is the cheapest baseline to beat.
- Practitioners shipped things. One commenter built a working file sorter with Jeff, and another reported production classification with a Gemma 4 12B model predicting a single output token at 70-80 ms p90. The issue tracker shows the same pattern: the long-list bug was reported and fixed within a day.
- The architecture skeptics got airtime. Multiple commenters argued there is no architectural magic here (one summary: one-token prefill plus logit differences). The other camp's reply: at high volume, calibrated single-pass decisions hit a price and latency point hosted chat models do not.
- Even the author hedges. The README says the 2B is a worse game player than the 0.8B, that benchmark scores do not predict play, and that prompts matter enormously. The chess example ends with the model losing to Stockfish's weakest setting.
FAQ#
Is Jeff a Jev replacement?#
No. It is an unaffiliated fine-tune family that speaks Jev's request format. Zero-shot it matches on classification-style benchmarks and trails on reasoning; fine-tuned on your own examples it can beat the field reports above, at the cost of owning the training loop.
What hardware does Jeff need?#
The 0.8B runs on CPU (463 ms per decision on 32 threads), one GPU (22 ms on an RTX PRO 6000), and Apple silicon via MLX (28 ms on an M4 Max). Training it took about two hours on one RTX PRO 6000.
How is this different from schema-constrained chat?#
It never generates text: answers are read out of the final layer as probabilities, so there is nothing to parse and latency is one forward pass. The trade is that it cannot reason about the options, only choose among descriptions you write well.
Continue Reading#
- TypeSafe Jev: the First Decision-Only Model - what the hosted System One model does, priced, with the architecture explained
- How to Use Jev - every way to call it, and how the open alternatives compare
- Mistral Shieldstral 3B - a production example of a narrow small model doing classification work
- Local LLM Runtime for Coding Agents - picking the runtime once you decide to serve locally
- AI Model Routing and Orchestration - where a local decision model fits in your stack
Sources#
| Source | URL |
|---|---|
| firelex/jeff repository and README, fetched September 30, 2026 | https://github.com/firelex/jeff |
| Hacker News thread: Jeff - Jev-compatible 0.8B decision models, 569 points | https://news.ycombinator.com/item?id=49883844 |
| GitHub issue #1: option picking past the 26th (fixed in v1.1) | https://github.com/firelex/jeff/issues/1 |
| GitHub issue #2: serving-only install | https://github.com/firelex/jeff/issues/2 |
| Model card: Jeff-Qwen3.5-0.8B | https://huggingface.co/mstrasser/Jeff-Qwen3.5-0.8B |
| Model card: Jeff-Qwen3.5-2B | https://huggingface.co/mstrasser/Jeff-Qwen3.5-2B |
| Model card: Jeff-Gemma4-E2B | https://huggingface.co/mstrasser/Jeff-Gemma4-E2B |
Get the next deep dive like this in your inbox
One email a week on News and the rest of the AI dev stack. Free.
Read next on local and open-weight models
TypeSafe Jev: the First Decision-Only Model, Benchmarked
TypeSafe's Jev is the first System One model: typed decisions with calibrated probabilities, 70-500ms responses, and $0.042 per million input tokens.
7 min readHow to Use Jev: Every Way to Call It, the Opus 5.5 Pairing, and Laya vs Kev vs Ollaya
TypeSafe's Jev dropped its waitlist on September 27. Every verified way to call it (API, Python SDK, llm CLI, Pydantic AI, Cloudflare, Vercel, OpenRouter), the Jev + Claude Opus 5.5 coding-agent pattern driving searches, and how the open alternatives Laya, Kev, and Ollaya compare.
10 min readMistral Shieldstral: A 3B Open-Weight Policy-Adaptive Moderation Model That Beats Models 7x Its Size
Shieldstral is a 3B-parameter Apache 2.0 multimodal safety classifier that takes your moderation policy as a plain-language question at inference time, scores content 0-1 in a single forward pass, and runs on one 16GB GPU. It beats 12B-20B guard models on text safety and sets state of the art on multimodal benchmarks.
8 min readNew here? Start with
Technical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.








