AI MODELS
112 items
111 posts, 1 guide
Kolibri is an Apache 2.0 English-German MoE with 78B total and 3B active parameters and a 1M context. Benchmarks, the vLLM command, and who should run it.
Cloudflare's Clef is a 27B Apache-2.0 decision model on Workers AI at $0.24 per million input tokens, and Clef-flash is a 9B model at $0.09 with a 38.8 ms median decision, both Jev-compatible, with an RL fine-tuning service attached.
Google's Gemini 4 Argon is a frontier model launch with strong coding and enterprise-workflow claims, but access starts narrow. Here is what shipped, what is verified, and what to watch before planning around it.
An independent project fine-tunes Qwen3.5 and Gemma 4 into 0.8B and 2B decision models that answer in 22-28 ms using Jev's request shape. The verified numbers, the run commands, and where the benchmark stops matching real work.
OpenAI shipped GPT-6.1 Sol on September 29: near-Astra benchmark scores at a fifth of Astra's price, cache reads cut to $0.10 per million tokens, 1.05M context, and same-day Codex availability. The verified pricing, the Astra 6.1 safety hold, the community read, and how to run it.
Claude Sonnet 5.5 (claude-sonnet-5-5) is Anthropic's new mid-tier model: $2/$10 per million tokens, 70.6% on Terminal-Bench 4.0, 1M context, now GA in GitHub Copilot and on Vercel AI Gateway. The pricing math, the five breaking API changes, and where it fits next to Opus 5.5 and Sonnet 5.
GPT-6 Astra is OpenAI's max-capability model: 1.05M context, $10/$50 per million tokens, the first OpenAI model rated Critical for cyber, and new highs on Terminal-Bench 4.0 and OSWorld 2.0. What it is, what the benchmarks do and do not say, and the verified commands to try it.
A developer-first overview of the GPT-6 family: what Astra, Sol, and Luna are, what each costs per million tokens, the three new API features (async tool calling, mid-turn steering, cache-safe reasoning changes), and how to try each tier in Codex and the API today.
TypeSafe's Jev dropped its waitlist on September 27. Every verified way to call it (API, Python SDK, llm CLI, Pydantic AI, Cloudflare, Vercel, OpenRouter), the Jev + Claude Opus 5.5 coding-agent pattern driving searches, and how the open alternatives Laya, Kev, and Ollaya compare.
You cannot install Claude Code or Codex inside the Meta Muse app, but you can run both agents on Muse Spark 1.3 through the Meta Model API, use Muse Spark from OpenCode Go, and bring your Claude Code and Codex setup into Muse Code. Here are the verified configs.
Three frontier launches in 48 hours repriced the agentic workhorse tier: Grok 4.7 at $2/$6 (Sep 21), Claude Opus 5.5 at $4/$20 with $0.20 cache reads (Sep 22), and GPT-6 Sol at $2/$10 with Luna at $0.10/$0.50 (Sep 22). Same-day-verified rates, honest benchmark attribution, and a decision guide.
Claude Opus 5.5 (claude-opus-5-5) is Anthropic's new default Opus: $4/$20 per million tokens, $0.20 cache reads, 1M context, thinking always on with medium default effort. Runnable TypeScript and Python SDK examples, Claude Code setup, before/after prompts, and a decision guide vs Sonnet 5, Haiku 4.5, and Fable 5.1.
Google shipped two new audio-to-audio models on September 15: Gemini 3.8 Live (76.0 Speech-to-Speech Index, 1.18s first audio) and 3.8 Live Extended Thinking (82.6, #1), both with background tool execution. Benchmarks, per-minute pricing, and how to build with the Live API.
TypeSafe's Jev returns typed decisions at $0.042 per million input tokens and 70-500 ms latency. Benchmarks, pricing and alternatives, verified October 2, 2026.
Google shipped agentic video understanding on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite: the model decides which frames, audio, and transcripts to inspect instead of swallowing video at a fixed frame rate. Verified numbers: up to 88% fewer tokens, up to 66% lower cost, and up to 7% better accuracy on video benchmarks. Here is what changed and where the agentic loop still leaks.
Tencent's Hy4 preview ships 770B total parameters with 49B active under Apache 2.0 - a 1M-context text MoE with DeepSeek-style sparse attention, posted Terminal-Bench 85.4 and DeepSWE 64.3, and an OpenRouter price of $0.834/$2.501. Verified against the model card and the live OpenRouter page on August 31, 2026.
Google made Gemini Omni 1.1 Flash generally available today: 10-second scene-extension context, first and last frame interpolation, 360p drafts at a third of the cost, and 4K upscaling. Verified pricing: about $0.10 per second of 720p video.
DeepSeek shipped experimental vision for V4 Flash as deepseek-v4-flash-vision-exp. JPEG, PNG, GIF, and WebP; three input methods; 384 tokens per image. Here is the API contract and how to run it in OpenCode today.
OpenCode dropped Ox Alpha as a free stealth model on August 20, 2026: 1M context, multimodal, near-unlimited for about a week. Here is what is confirmed, where OpenCode and OpenRouter disagree on retention, and how to run it today.
Qwen3.8-27B is a 27B dense Apache-2.0 model that scores 61.7 on SWE-bench Pro and 42.2 on DeepSWE 1.1 - ahead of Opus 4.6 Max on both - while running on consumer hardware. Benchmarks, hardware math, and an honest when-to-use-it guide.

Get Smarter About AI Dev
New tutorials, open-source projects, and deep dives on coding agents - delivered weekly.