Skip to main content
Watch: I Asked Claude to Build Me a Business

BENCHMARKS

31 items

31 posts

Blog
Terminal-Bench Shows Harness Scaling Is the Coding-Agent Benchmark Now

StateM pushes Terminal-Bench 2.1 to 95.3% raw accuracy by scaling the harness around the model. The lesson for coding-agent teams is that runbooks, state, and recovery loops now matter as much as model choice.

Blog
Apple SpeechAnalyzer vs Whisper: Independent Benchmark Shows Apple Winning on Accuracy

New benchmarks on 5,559 test utterances show Apple's iOS 26 SpeechAnalyzer API achieving 2.12% word error rate - beating all Whisper model sizes while running 3x faster.

Blog
Cursor Composer 2.5 Developer Guide 2026

Cursor shipped Composer 2.5 in May 2026 - a 1T parameter agentic coding model that matches Opus 4.7 and GPT-5.5 on benchmarks at roughly one tenth the cost. Here is everything you need to know to use it effectively.

Blog
GPT-5.5 Has a 3x Higher Hallucination Rate Than MIT-Licensed GLM-5.2

New benchmark data shows GPT-5.5 hallucinates 86% of the time when it does not know the answer - versus 28% for the open-weights GLM-5.2. The numbers challenge the assumption that bigger models equal more reliable output.

Blog
Claude Fable 5 vs GPT-5.6 Sol: Benchmarks, Pricing, and When Each Wins

Fable 5 is API-only at 2x GPT-5.6 Sol's price with a 15-point SWE-Bench Pro gap. Here is the decision framework for choosing between them in August 2026.

Blog
FrontierCode Benchmark Explained: Why AI Coding Quality Scores Are Wrong (And the Fix)

SWE-Bench has an 81% false-positive problem. FrontierCode replaces it with mergeability as the metric - and the scores are sobering for every AI coding tool on the market.

Blog
GPT-5.5 for Developers: A Production Field Guide

GPT-5.5 and 5.5 Pro hit the API on April 24. Here is what changes for builders: pricing, agentic tasks, tool-use, and the real benchmarks I ran the day it dropped.

Blog
Web Dev Arena: How to Test AI Coding Models on Real Frontend Work

Benchmarks are useful, but frontend work fails in places leaderboards barely measure. Here is how Web Dev Arena turns AI model comparison into a practical UI evaluation workflow.

Blog
Claude Sonnet 4.6: Approaching Opus at Half the Cost

Anthropic's Sonnet 4.6 narrows the gap to Opus on agentic tasks, leads computer use benchmarks, and ships with a beta million-token context window. Here's what actually changed.

Blog
Grok 4: xAI's Most Powerful AI Model

xAI has launched Grok 4, claiming the title of the world's most powerful AI model. With a $300/month Super Grok tier, saturated AMI benchmarks, and a coding model on the horizon, this is xAI's bigge...

Blog
xAI Grok 3 Launch: The Smartest AI on Earth?

xAI launched Grok 3 with 200,000 GPUs, outperforming GPT-4o, Sonnet 3.5, and DeepSeek R1 on reasoning benchmarks. Here is what the hardware, the benchmarks, and the new features actually mean for developers.

PreviousPage 2 of 2
AI Development Stack

Get Smarter About AI Dev

New tutorials, open-source projects, and deep dives on coding agents - delivered weekly.

One email per weekReal code, not theoryFree forever