Skip to main content
Watch: Claude Opus 5.5 Built an Entire 3D World

PRICING

44 items

43 posts, 1 tool

Blog
Cloudflare Clef Decision Models: Jev-Compatible, Benchmarked and Priced

Cloudflare's Clef is a 27B Apache-2.0 decision model on Workers AI at $0.24 per million input tokens, and Clef-flash is a 9B model at $0.09 with a 38.8 ms median decision, both Jev-compatible, with an RL fine-tuning service attached.

Blog
GPT-6.1 Sol Release Guide: Near-Astra Agentic Work at $2/$10 and $0.10 Cache Reads

OpenAI shipped GPT-6.1 Sol on September 29: near-Astra benchmark scores at a fifth of Astra's price, cache reads cut to $0.10 per million tokens, 1.05M context, and same-day Codex availability. The verified pricing, the Astra 6.1 safety hold, the community read, and how to run it.

Blog
Claude Cowork vs Claude Code (2026): Which One Should You Use, and What Does It Cost?

Claude Code is Anthropic's agent for software engineering; Claude Cowork is the same agent loop pointed at documents, spreadsheets, and research. Both come with every paid Claude plan from $17/month. Here is how to pick, with a verified September 2026 pricing table.

Blog
Codex Usage Limits and Pricing in 2026: Pro 100, Pro 200, Pro 500, Resets, and 'Model at Capacity' Errors

Codex is included in ChatGPT plans from Free to Enterprise. On September 29, 2026, OpenAI reopened Pro $200 with roughly half the old usage value, added Pro $500 with Astra Ultrafast, and left Plus at $20/month. Existing Pro $200 subscribers keep their old allowance through October 29, 2026. Here is how resets work, what each plan buys, and what to do about 'Selected model is at capacity'.

Blog
GPT-6 Sol vs Claude Opus 5.5 vs Grok 4.7: The September 2026 Price War

Three frontier launches in 48 hours repriced the agentic workhorse tier: Grok 4.7 at $2/$6 (Sep 21), Claude Opus 5.5 at $4/$20 with $0.20 cache reads (Sep 22), and GPT-6 Sol at $2/$10 with Luna at $0.10/$0.50 (Sep 22). Same-day-verified rates, honest benchmark attribution, and a decision guide.

Blog
Where to Run DeepSeek V4.1 Flash Free and Cheap

DeepSeek V4.1 Flash replaced V4 Flash. Official API: $0.15 input, $0.60 output off-peak. OpenRouter hosts start near $0.05. MIT weights are on Hugging Face.

Blog
Where to Run GLM-5.3 Free and Cheap: Flash, Batch and Provider Prices (2026)

GLM-5.3 routes compared, verified October 1, 2026: OpenRouter hosts list it from $0.12 per million input tokens, GLM-5.3-Flash costs $0.15 input and $0.50 output from Z.ai, batch drops to $0.45/$2.00, and the Coding Plan starts at $18 per month.

Blog
DeepSeek V4 Flash Is 90% Off Through Novita on Vercel AI Gateway: The Cost Math

DeepSeek V4 Flash routed to Novita on Vercel AI Gateway is 90% off for Pro customers through August 11, dropping the effective rate to $0.014 input / $0.028 output per million tokens. Here is the verified before/after math, the provider-pinning setup, and what a 10x cheap agent loop means for routing decisions.

Blog
Cursor Removes Dollar Costs From Its Usage Page: Token-Only Reporting Now

Cursor shipped a deliberate change on July 31 making the Usage page tokens-only for self-serve plans, removed the dollar Cost column, and zeroed per-request cost fields in the dashboard API - including for historical records. Staff confirmed the change is intentional and that the numbers are still tracked internally.

Blog
Budget AI Coding Models Compared September 2026: GPT-6 Luna vs V4.1-Flash vs Gemini 3.5 Flash vs Haiku 4.5

The cheap coding tier repriced again: GPT-6 Luna opened at $0.10/$0.50 and took the floor from DeepSeek, whose V4.1-Flash cut rates to $0.15/$0.60 off-peak. Gemini 3.5 Flash and Claude Haiku 4.5 hold the hosted middle. Prices verified September 26, 2026.

Blog
OpenAI Cuts GPT-5.6 Luna by 80%: The Price-Performance Frontier Just Shifted

OpenAI slashes GPT-5.6 Luna by 80% to $0.20/M input tokens, cuts Terra by 20%, adds Sol Fast mode at 2.5x speed, and reveals Sol autonomously optimized its own production kernels.

Blog
AI Model Routing Strategies for Cost-Effective Coding in 2026

A practical guide to routing between Claude Opus 5, Sonnet 5, Haiku 4.5, GPT-5.6 Sol/Terra/Luna, and Kimi K3 based on task complexity, cost budget, and latency requirements - with decision frameworks and code examples.

Blog
OpenAI Cuts GPT-5.6 Luna 80% and Terra 20%: The Cost-Per-Task Math for Agent Builders

Luna drops from $1/$6 to $0.20/$1.20 per million tokens, Terra from $2.50/$15 to $2/$12, and Sol gets a paid Fast mode. What the new floor means for agent economics, Codex quotas, and the competition.

Blog
Where to Access Kimi K3: Every Provider and Price Compared (2026)

Where to access Kimi K3: Moonshot's API at $3 input and $15 output per million tokens, OpenRouter hosts, inference clouds, and open weights.

Blog
How Much Should I Charge for a Website? A Practical Pricing Guide

A practical way to price website projects using scope, time, risk, and value, with real examples for landing pages, business sites, and custom builds.

Blog
GLM 5.2 and the AI Margin Collapse Thesis

Martin Alderson's argument for why open-weights models like GLM 5.2 will compress frontier lab margins is sparking debate on HN. Here is what the thesis actually says, where HN agrees and disagrees, and why it matters for developers choosing models.

Blog
Why Price Per 1M Tokens Is a Misleading Metric for LLM Costs

Comparing LLMs by token pricing alone can lead you to choose worse, more expensive models. Cost per task tells the real story.

Blog
The Economics of Agent Fleets: Fable 5 Orchestrators, Sonnet 5 Workers

One expensive orchestrator plus many cheap workers beats an all-frontier fleet for most workloads. Here is the decision-intent cost math with verified Fable 5, Sonnet 5, and Opus 4.8 prices, plus the Sonnet 5 tokenizer caveat that changes worker cost.

Blog
Claude in Microsoft Foundry on Azure: Developer Guide 2026

Claude is now GA in Microsoft Foundry on Azure with native billing, Entra ID auth, and GB300 Blackwell infrastructure. Here is the full developer setup - CCU pricing, SDK examples, deployment options, and what enterprise teams need to know.

Blog
AI's Affordability Crisis Is Really an Agent Cost Accounting Problem

A viral Hacker News thread about AI affordability points at the right problem, but developer teams need a more useful cost model: retries, cache misses, review time, routing, and failed loops.

Page 1 of 3Next
AI Development Stack

Get Smarter About AI Dev

New tutorials, open-source projects, and deep dives on coding agents - delivered weekly.

One email per weekReal code, not theoryFree forever