Skip to main content
Watch: I Asked Claude to Build Me a Business

AI MODELS

112 items

111 posts, 1 guide

Blog
Model Routing Recipes: Practical Config Patterns to Cut AI Spend

A code-heavy field guide to model routing. Real, runnable-style configs for tiering tasks by complexity, routing simple work to open-weights, reserving frontier models for hard reasoning, building failover chains, and keeping prompt caches warm with OpenRouter, LiteLLM, and Factory Router.

Blog
OpenRouter Fusion Makes Model Panels Real. Use Them Like Escalation, Not Autopilot

OpenRouter Fusion turns multi-model panels into an API feature. The useful lesson is not to run every prompt through more models. It is to define when a task deserves an expensive second opinion.

Blog
The Claude Tokenizer Change: What ~30% More Tokens Means for Your Bill

Anthropic's docs say the tokenizer introduced with Opus 4.7 can use up to 35% more tokens for the same text. Here is what that does to per-request cost, max_tokens, and cross-model comparisons.

Blog
Fable 5 with 1M Context: What Actually Works in Practice

Fable 5 1M context workflows that actually work: whole-repo reviews, log archaeology, multi-doc synthesis - plus the honest math on when RAG still wins.

Blog
Fable 5 Effort Levels Explained: low, medium, high, xhigh, max

Fable 5 effort levels explained: what low, medium, high, xhigh and max change, which models support each, what each costs, and where the defaults differ.

Blog
Handling Long-Running Fable 5 Requests: Timeouts, Streaming, and Background Patterns

Fable 5 long-running requests can run for many minutes per turn and hours per autonomous run. Here is how to configure client timeouts, streaming keepalive, batch polling, and background patterns so they actually finish.

Blog
The Fable 5 Orchestrator Playbook: One Smart Model Managing Cheap Workers

A practical playbook for running Claude Fable 5 as the orchestrator over Sonnet and Haiku workers, with verified cost math on when the premium pays off.

Blog
The Frontier Model Landscape, July 2026 Edition

A verified directory of the frontier AI models in July 2026 - Claude Fable 5, Opus 5, GPT-5.6 Sol/Terra/Luna, Sonnet 5, Gemini 3.1 Pro, Kimi K3, and DeepSeek V4 - with pricing checked against official docs.

Blog
How to Use Claude Fable 5: Every Access Path Explained

How to use Claude Fable 5 across every access path: claude.ai plans through June 22, the Claude API, Amazon Bedrock, Vertex AI, and Microsoft Foundry, with setup effort and first-prompt tips.

Blog
Is Claude Fable 5 Slow? Latency in Practice, and When It Matters

Claude Fable 5 latency measured: 109 seconds to first token at max effort vs 1.4s for Sonnet 4.6. When slow is fine, when it hurts, and how to route around it.

Blog
Migrating Off Retired GPT Models in 2026: A Working Checklist

Migrating off retired GPT models in 2026: the live retirement table, what maps to what, an eval-before-switch day plan, and when to jump providers.

Blog
Qwen 3.7 Max Developer Guide: 1M Context, $1.25/MTok, and Agent-First Architecture

Alibaba shipped Qwen 3.7 Max on May 19, 2026 with a 1M token context window, Anthropic-compatible API, and agent-first architecture. Here is what developers need to know about pricing, performance, and when to use it.

Blog
12 Ways Developers Are Actually Leveraging Claude Fable 5

Twelve documented Claude Fable 5 use patterns - agent orchestration, overnight runs, 1M-context refactors, effort tuning - each with a how-to seed and doc link.

Blog
Anthropic Model Names Explained: Fable, Mythos, Opus, Sonnet

Anthropic's model names: Haiku, Sonnet, Opus, then Fable and Mythos. What each tier means, the current model IDs and prices, and which one to pick.

Blog
Apple's LanguageModel Protocol: Xcode 27 Just Made Model Lock-In Optional

Apple shipped a LanguageModel protocol at WWDC 2026 that lets iOS and macOS developers swap between Claude, Gemini, and local models with a single dependency change. Here is what OS-level provider abstraction actually means for switching costs, moats, and your architecture decisions.

Blog
Fable 5 vs Opus 4.8: A Data-Driven Decision Guide for Engineering Teams

Fable 5 posts an 80.3% SWE-Bench Pro score and costs 2x Opus 4.8 - here is the task-profile scoring guide that tells you when the premium pays off.

Blog
Claude Mythos 5 Explained: What It Is, Who Can Access It, and Why It's Gated

Anthropic shipped two names for one architecture on June 9, 2026. Here is what separates Fable 5 from Mythos 5, who can actually get unrestricted access, and what developers should do right now.

Blog
The Model, IDE, CLI, and Agent Framework Changes That Actually Matter

The AI coding market is noisy. The changes that matter are easier to spot when you separate model capability, editor loops, terminal agents, background agents, agent frameworks, UI layers, context, security, and cost.

Blog
Models.dev Makes Model Routing Feel Like Infrastructure

The models.dev project is trending because AI teams need one boring source of truth for model specs, pricing, context windows, modalities, and tool support.

Blog
DeepSeek V4 Changes the Coding Agent Cost Equation

DeepSeek V4 is trending because it is close enough to frontier coding models at a much lower token price. The real question for developers is where cheap reasoning belongs in an agent stack.

PreviousPage 5 of 6Next
AI Development Stack

Get Smarter About AI Dev

New tutorials, open-source projects, and deep dives on coding agents - delivered weekly.

One email per weekReal code, not theoryFree forever