Skip to main content
Watch: I Asked Claude to Build Me a Business

AI MODELS

112 items

111 posts, 1 guide

Blog
Kolibri Release Guide: Aleph Alpha's 78B Open Model

Kolibri is an Apache 2.0 English-German MoE with 78B total and 3B active parameters and a 1M context. Benchmarks, the vLLM command, and who should run it.

Blog
Cloudflare Clef Decision Models: Jev-Compatible, Benchmarked and Priced

Cloudflare's Clef is a 27B Apache-2.0 decision model on Workers AI at $0.24 per million input tokens, and Clef-flash is a 9B model at $0.09 with a 38.8 ms median decision, both Jev-compatible, with an RL fine-tuning service attached.

Blog
Gemini 4 Argon: What Developers Can Actually Use

Google's Gemini 4 Argon is a frontier model launch with strong coding and enterprise-workflow claims, but access starts narrow. Here is what shipped, what is verified, and what to watch before planning around it.

Blog
Jeff: Jev-Style Decision Models Trained at Home on One GPU

An independent project fine-tunes Qwen3.5 and Gemma 4 into 0.8B and 2B decision models that answer in 22-28 ms using Jev's request shape. The verified numbers, the run commands, and where the benchmark stops matching real work.

Blog
GPT-6.1 Sol Release Guide: Near-Astra Agentic Work at $2/$10 and $0.10 Cache Reads

OpenAI shipped GPT-6.1 Sol on September 29: near-Astra benchmark scores at a fifth of Astra's price, cache reads cut to $0.10 per million tokens, 1.05M context, and same-day Codex availability. The verified pricing, the Astra 6.1 safety hold, the community read, and how to run it.

Blog
Claude Sonnet 5.5 Developer Guide: Pricing, Benchmarks, and the Five API Changes

Claude Sonnet 5.5 (claude-sonnet-5-5) is Anthropic's new mid-tier model: $2/$10 per million tokens, 70.6% on Terminal-Bench 4.0, 1M context, now GA in GitHub Copilot and on Vercel AI Gateway. The pricing math, the five breaking API changes, and where it fits next to Opus 5.5 and Sonnet 5.

Blog
GPT-6 Astra Release Guide: Benchmarks, $10/$50 Pricing, and How to Run It in Codex and OpenCode

GPT-6 Astra is OpenAI's max-capability model: 1.05M context, $10/$50 per million tokens, the first OpenAI model rated Critical for cyber, and new highs on Terminal-Bench 4.0 and OSWorld 2.0. What it is, what the benchmarks do and do not say, and the verified commands to try it.

Blog
GPT-6 in 7 Minutes: Astra, Sol, and Luna Explained for Developers

A developer-first overview of the GPT-6 family: what Astra, Sol, and Luna are, what each costs per million tokens, the three new API features (async tool calling, mid-turn steering, cache-safe reasoning changes), and how to try each tier in Codex and the API today.

Blog
How to Use Jev: Every Way to Call It, the Opus 5.5 Pairing, and Laya vs Kev vs Ollaya

TypeSafe's Jev dropped its waitlist on September 27. Every verified way to call it (API, Python SDK, llm CLI, Pydantic AI, Cloudflare, Vercel, OpenRouter), the Jev + Claude Opus 5.5 coding-agent pattern driving searches, and how the open alternatives Laya, Kev, and Ollaya compare.

Blog
How to Use Claude Code and Codex With Meta Muse: Every Setup That Actually Works

You cannot install Claude Code or Codex inside the Meta Muse app, but you can run both agents on Muse Spark 1.3 through the Meta Model API, use Muse Spark from OpenCode Go, and bring your Claude Code and Codex setup into Muse Code. Here are the verified configs.

Blog
GPT-6 Sol vs Claude Opus 5.5 vs Grok 4.7: The September 2026 Price War

Three frontier launches in 48 hours repriced the agentic workhorse tier: Grok 4.7 at $2/$6 (Sep 21), Claude Opus 5.5 at $4/$20 with $0.20 cache reads (Sep 22), and GPT-6 Sol at $2/$10 with Luna at $0.10/$0.50 (Sep 22). Same-day-verified rates, honest benchmark attribution, and a decision guide.

Blog
Claude Opus 5.5 Developer Guide: API Examples, Claude Code Setup, Pricing, and When to Use It

Claude Opus 5.5 (claude-opus-5-5) is Anthropic's new default Opus: $4/$20 per million tokens, $0.20 cache reads, 1M context, thinking always on with medium default effort. Runnable TypeScript and Python SDK examples, Claude Code setup, before/after prompts, and a decision guide vs Sonnet 5, Haiku 4.5, and Fable 5.1.

Blog
Gemini 3.8 Live and 3.8 Live Extended Thinking: Google's Audio-to-Audio Voice Models, Benchmarked and Priced

Google shipped two new audio-to-audio models on September 15: Gemini 3.8 Live (76.0 Speech-to-Speech Index, 1.18s first audio) and 3.8 Live Extended Thinking (82.6, #1), both with background tool execution. Benchmarks, per-minute pricing, and how to build with the Live API.

Blog
TypeSafe Jev: Decision-Only Model Pricing and Benchmarks (2026)

TypeSafe's Jev returns typed decisions at $0.042 per million input tokens and 70-500 ms latency. Benchmarks, pricing and alternatives, verified October 2, 2026.

Blog
Gemini's Agentic Video Understanding Cuts Video Tokens by 88%: How It Works and What Breaks

Google shipped agentic video understanding on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite: the model decides which frames, audio, and transcripts to inspect instead of swallowing video at a fixed frame rate. Verified numbers: up to 88% fewer tokens, up to 66% lower cost, and up to 7% better accuracy on video benchmarks. Here is what changed and where the agentic loop still leaks.

Blog
Tencent Hy4 Preview: 770B Open MoE, Agentic Benchmarks, and Where It Fits in 2026

Tencent's Hy4 preview ships 770B total parameters with 49B active under Apache 2.0 - a 1M-context text MoE with DeepSeek-style sparse attention, posted Terminal-Bench 85.4 and DeepSWE 64.3, and an OpenRouter price of $0.834/$2.501. Verified against the model card and the live OpenRouter page on August 31, 2026.

Blog
Gemini Omni 1.1 Flash Goes GA: Scene Extension, Keyframe Control, and 4K in the Gemini API

Google made Gemini Omni 1.1 Flash generally available today: 10-second scene-extension context, first and last frame interpolation, 360p drafts at a third of the cost, and 4K upscaling. Verified pricing: about $0.10 per second of 720p video.

Blog
DeepSeek V4 Flash Vision Exp: Experimental Vision, Limits, and How to Run It in OpenCode

DeepSeek shipped experimental vision for V4 Flash as deepseek-v4-flash-vision-exp. JPEG, PNG, GIF, and WebP; three input methods; 384 tokens per image. Here is the API contract and how to run it in OpenCode today.

Blog
Ox Alpha on OpenCode: The Free Stealth Model, Specs, Privacy Split, and How to Run It

OpenCode dropped Ox Alpha as a free stealth model on August 20, 2026: 1M context, multimodal, near-unlimited for about a week. Here is what is confirmed, where OpenCode and OpenRouter disagree on retention, and how to run it today.

Blog
Qwen3.8-27B vs Opus 4.6 Max: The Laptop-Sized Model That Beat a Frontier Flagship on Agentic Benchmarks

Qwen3.8-27B is a 27B dense Apache-2.0 model that scores 61.7 on SWE-bench Pro and 42.2 on DeepSWE 1.1 - ahead of Opus 4.6 Max on both - while running on consumer hardware. Benchmarks, hardware math, and an honest when-to-use-it guide.

Page 1 of 6Next
AI Development Stack

Get Smarter About AI Dev

New tutorials, open-source projects, and deep dives on coding agents - delivered weekly.

One email per weekReal code, not theoryFree forever