Skip to main content
Watch: I Asked Claude to Build Me a Business

AI MODELS

112 items

111 posts, 1 guide

Blog
Grok 4.5: xAI Releases Cursor-Trained Coding Model at $2/M Input Tokens

xAI launched Grok 4.5, trained on trillions of Cursor interaction tokens. At $2/M input pricing, it undercuts Claude and GPT while benchmarking near Opus 4.7 level.

Blog
Meta Muse Image: What Developers Can Actually Use Today

Meta's Muse Image is now in Meta AI, but it is not a public model API. Here is what the launch confirms, what remains preview-only, and how developers should evaluate it.

Blog
Meta Muse Spark 1.1 Developer Guide: First Paid Meta API for Agentic Tasks

Meta launches Muse Spark 1.1 through the new Meta Model API - a 1M-token-context model for personal agentic tasks with OpenAI-compatible endpoints, $20 free credits, and pricing that undercuts the competition.

Blog
Meta Launches Muse Spark 1.1: A Closed-Weights Agentic Model with Aggressive Pricing

Meta's first paid API model arrives with $1.25/M input tokens, 1M context window, and strong tool-use benchmarks. HN debates what it means for the open-weights company.

Blog
GLM 5.2 and the AI Margin Collapse Thesis

Martin Alderson's argument for why open-weights models like GLM 5.2 will compress frontier lab margins is sparking debate on HN. Here is what the thesis actually says, where HN agrees and disagrees, and why it matters for developers choosing models.

Blog
Claude Sonnet 5 Developer Guide: Migration, API, and Effort Levels

Everything developers need to migrate from Sonnet 4.6 to Sonnet 5 - three breaking API changes, the new effort parameter, tokenizer impact, and when to use each effort level. Verified against Anthropic's official docs on July 4, 2026.

Blog
Claude Sonnet 5 vs Sonnet 4.6: Should You Upgrade?

Claude Sonnet 5 lands near Opus 4.8 on some tasks for a fraction of the price - but a new tokenizer runs about 30 percent more tokens. Here is the upgrade decision for builders, with the numbers.

Blog
Running Fable 5 Agent Fleets in Production: The Operations Guide

Standing up a fleet of Fable 5 agents is the easy part. This is the operations layer - data retention rules, refusal-rate alerting, effort tuning, observability, and availability planning - that keeps the fleet running.

Blog
Fable 5 Is Back: The Anthropic Model the Government Switched Off

Anthropic's most capable model launched, got suspended by a US export-control order, and returned today. Here is what Fable 5 is, what changed on the way back, and whether builders should reach for it.

Blog
Fable 5 vs Opus 4.8: Which Should Orchestrate Your Agents?

The orchestrator is the most important model choice in an agent fleet. A fair head-to-head between Fable 5 and Opus 4.8 for that role, with a decision matrix by run length, budget, compliance, and refusal-handling tolerance.

Blog
GLM 5.2 in 9 Minutes: The Open-Weight Rival to GPT-5.5

A companion guide to the GLM 5.2 video: an open-weight model positioned against GPT-5.5, walked through with benchmarks, pricing, and a live OpenCode demo. Here is what the video covers and where to go deeper.

Blog
GPT-5.5 in 7 Minutes: Benchmarks, Codex Agents, Context Window, and Pricing

A companion guide to the GPT-5.5 video: OpenAI's newly released model rolling out to ChatGPT and Codex, reviewed through benchmarks, agent capabilities, context window, and pricing. Here is what the video covers and where to go deeper.

Blog
Claude Sonnet 5 Launch Analysis: The Most Agentic Sonnet Yet

Anthropic releases Claude Sonnet 5 with improved agentic capabilities, better tool use, and an introductory pricing deal. Here's what developers need to know.

Blog
Apertus: Europe's Answer to AI Sovereignty - and Why HN Is Skeptical

Switzerland's fully open foundation model promises transparent training data and EU compliance. The HN crowd has questions about actual performance.

Blog
Fugu Ultra's Frontier Performance Claim, Explained Without the Hype

Sakana says Fugu Ultra stands with Fable, Mythos, GPT-5.5, Gemini, and Opus by orchestrating models instead of being one giant model. Here is what the benchmarks show, what is novel, and what still needs proof.

Blog
Sakana Fugu and the Case for Not Betting Everything on One Proprietary Model

Sakana Fugu makes a timely argument for model routing: frontier performance should come from swappable systems, not a hard dependency on one proprietary API.

Blog
Sakana Fugu Ultra: The Model Router Making the Frontier Look Less Proprietary

Sakana Fugu Ultra is not just another giant model. It is a learned orchestration layer that routes work across expert models, matches frontier benchmark claims, and makes a serious case for multi-model AI systems.

Blog
How to Use GLM 5.2 and Other Custom Model Providers in Codex

Codex can use GLM 5.2, Mistral, local Ollama, or an internal proxy. Add a model_providers entry to config.toml with a base_url, env_key, and model.

Blog
The Router Era: Why Not Owning a Frontier Model Became an Advantage

No single model wins every task anymore, and the companies that never trained one - Factory, Devin, Perplexity, Cursor, OpenCode - are turning that into a moat. This is how model routing works, why open weights and neoclouds make it cheap, and the honest counter-argument.

Blog
GLM-5.2 Cost Math: When Open-Weights Coding Models Actually Save You Money

Z.ai's GLM-5.2 lands as a 753B open-weights coding model that beats GPT-5.5 on SWE-bench Pro for roughly one-sixth the per-token cost. Here is the real cost math, a worked cost-per-task example, and a when-to-use-which decision guide.

PreviousPage 4 of 6Next
AI Development Stack

Get Smarter About AI Dev

New tutorials, open-source projects, and deep dives on coding agents - delivered weekly.

One email per weekReal code, not theoryFree forever