GPT-6 In 7 Minutes

TL;DR
Same-day-verified llm api pricing august 2026: Claude Fable 5, GPT-5.6 Sol/Terra/Luna, Claude Sonnet 5, Gemini 3.5 Flash, and DeepSeek V4 compared per million tokens, plus the caveats that change the math.
Direct answer
Same-day-verified llm api pricing august 2026: Claude Fable 5, GPT-5.6 Sol/Terra/Luna, Claude Sonnet 5, Gemini 3.5 Flash, and DeepSeek V4 compared per million tokens, plus the caveats that change the math.
Best for
Developers comparing real tool tradeoffs before choosing a stack.
Covers
Verdict, tradeoffs, pricing signals, workflow fit, and related alternatives.
Last updated: August 27, 2026
All prices verified August 27, 2026 against each provider's live pricing page (Google's pricing page was unreachable on August 27, so Gemini rows still carry the July 26 verification):
| Provider | Pricing | Models | Changelog / Docs |
|---|---|---|---|
| Anthropic | claude.com/pricing | Models overview | Release notes |
| OpenAI | openai.com/api/pricing | Models docs | Changelog |
| Google Gemini | ai.google.dev/pricing | Models index | Release notes |
| DeepSeek | api-docs.deepseek.com/pricing | V4 docs | API changelog |
Claude Sonnet 5 pricing became permanent (August 15). Anthropic removed the promotional framing from Sonnet 5: the $2/$10 per MTok rates are now the standard price, and the previously scheduled September 1 increase to $3/$15 will not occur (Anthropic pricing, verified August 15, 2026). This makes Sonnet 5 structurally the cheapest Claude model above Haiku, not just temporarily. Every "promo ends September 1" note in this post from July is now outdated - treat $2/$10 as the going rate. Sonnet 5 remains the strongest price-performance Claude in the lineup, and the Fable 5 orchestrator playbook economics improve accordingly.
DeepSeek peak/off-peak pricing gets a date and numbers (August 16, 16:00 UTC). DeepSeek's pricing page now lists the exact policy: off-peak rates at half of peak, peak hours 01:00-04:00 and 06:00-10:00 UTC (09:00-12:00 and 14:00-18:00 Beijing Time), effective August 16, 2026 at 16:00 UTC. New rates: V4 Flash at $0.22 input (cache miss) / $0.007 (cache hit) / $0.66 output off-peak, doubling to $0.44 / $0.014 / $1.32 during peak; V4 Pro at $0.66 / $0.022 / $1.98 off-peak, doubling to $1.32 / $0.044 / $3.96 during peak (DeepSeek pricing, verified August 15, 2026). The old flat $0.14/$0.28 Flash pricing dies August 16. The V4 Pro model also updated to version DeepSeek-V4-Pro-0813. The headline: even off-peak, DeepSeek output rates roughly double; the cheap-agent math changes materially for output-heavy loops, and every off-peak-scheduling assumption needs the new numbers.
GPT-5.6 Sol price cut (August 27). OpenAI's pricing page now lists GPT-5.6-Sol at $4/$20 per MTok (cached input $0.40, cache writes $5.00), down from the $5/$30 listed here in previous sections. Long-context rates move to $8/$30, and Fast mode to $8/$40. Promotional pricing is in effect at least through November 21, 2026 (OpenAI pricing, verified August 27, 2026). Sol at $20 output now sits below Claude Opus 5's $25 output and ties GPT-5.5 on input - the first time a 5.6-tier model undercuts Opus on both axes.
GPT-5.6 price cuts landed (July 30). OpenAI cut GPT-5.6 Luna 80% to $0.20/$1.20 per MTok and GPT-5.6 Terra 20% to $2/$12, and renamed Priority Processing to Fast mode (2x pricing, up to 2.5x speed) across the family (announcement, pricing page, both verified July 31, 2026). Luna now undercuts Claude Haiku 4.5 on input by 5x, and Terra ties Claude Sonnet 5's promotional input rate. The lower Luna and Terra rates also flow through to Codex and ChatGPT Work usage counting, so subscription quotas stretch further. Full breakdown in our price-cut analysis and budget tier comparison.
DeepSeek announced peak/off-peak pricing (July 31). DeepSeek's pricing page now warns that the API will adopt a peak/off-peak policy: 2x regular prices during peak hours (09:00-12:00 and 14:00-18:00 Beijing Time, UTC+8), effective date to be announced. The current $0.14/$0.28 V4 Flash rates below are pre-policy; budget models should assume off-peak prices are the floor and peak prices double. The V4 Flash model also updated to version DeepSeek-V4-Flash-0731 with 1M context, 384K max output, and Anthropic-format API support (DeepSeek pricing, verified July 31, 2026).
Claude Opus 5 launched (July 24) at $5/$25 per MTok. Anthropic's new flagship matches Opus 4.8 pricing while delivering near-Fable 5 intelligence on most benchmarks. It tops the Artificial Analysis Intelligence Leaderboard at 61, ahead of Fable 5 (60) and GPT-5.6 Sol (59). Available on the Claude API, Bedrock, Google Cloud, and Foundry. Fast mode (2x base pricing) supported on the Claude API only. Anthropic also lists Claude Mythos 5 (limited availability) at $10/$50, same rates as Fable 5. See the full Opus 5 vs Fable 5 comparison. This effectively resets the frontier tier: you now get Fable 5-competitive intelligence at Opus 4.8 prices.
DeepSeek deprecation deadline passed (July 24). The legacy deepseek-chat and deepseek-reasoner model names stopped serving requests on July 24. Any application still using these model identifiers receives errors. Migrate to deepseek-v4-flash ($0.14/$0.28 per MTok) or deepseek-v4-pro ($0.435/$0.87) if not already done. Our DeepSeek V4 migration guide covers the full transition.
GPT-5.6 family goes GA (July 9). OpenAI's new flagship line ships three tiers: Sol ($5/$30, matching GPT-5.5's rates), Terra ($2.50/$15 at launch, cut to $2/$12 on July 30), and Luna ($1/$6 at launch, cut to $0.20/$1.20 on July 30). All three inherit long-context pricing tiers that double rates above the standard window. Sol and GPT-5.5 now sit at parity on sticker price; routing decisions come down to benchmarks and cache behavior.
Claude Sonnet 5 pricing is now standard ($2/$10). Anthropic's August 15 pricing page lists Sonnet 5 at $2 input / $10 output with no end date and explicitly states the previously announced September 1 increase to $3/$15 will not occur. The promotional window is over in the sense that the price is now permanent. This undercuts every Opus tier and most competitors while staying well below GPT-5.6-Terra on output ($10 vs $12).
The tables below reflect August 27, 2026 verified rates.
Sticker prices for model APIs have never been this spread out. At the top of the August 2026 table, GPT-5.5-pro charges $180 per million output tokens. At the bottom, DeepSeek V4 Flash charges $0.66 per million output tokens off-peak (the peak/off-peak policy landed August 16). That is a 273x gap between two models you can call with nearly identical OpenAI-style request bodies.
This post is the raw API rate card: every current frontier and workhorse model from Anthropic, OpenAI, Google, and DeepSeek, with every number pulled from the live first-party pricing page and stamped with a verification date. If you are pricing coding tool subscriptions like Claude Code or Cursor plans instead, that is a different market with different math - see the companion post on AI coding tool pricing.
Four structural observations matter more than any single row: the DeepSeek output floor, the new GPT-5.6 tiering strategy, Anthropic's permanent Sonnet 5 rates, and the Claude tokenizer change that quietly skews per-MTok comparisons.
All prices in USD per million tokens (MTok), standard synchronous API. OpenAI and Anthropic rows verified August 27, 2026 against each provider's live pricing page (linked in the Official Sources table above); Gemini rows verified July 26, 2026 (page unreachable on July 31, August 15 and August 27); DeepSeek rows show the current off-peak rates now in effect, with the peak rates in the sub-table below (verified August 27, 2026).
| Model | Input | Cached input read | Output |
|---|---|---|---|
| Claude Fable 5 | $10.00 | $1.00 | $50.00 |
| Claude Mythos 5 (limited availability) | $10.00 | $1.00 | $50.00 |
| Claude Opus 5 | $5.00 | $0.50 | $25.00 |
| Claude Opus 4.8 | $5.00 | $0.50 | $25.00 |
| Claude Sonnet 5 | $2.00 | $0.20 | $10.00 |
| Claude Sonnet 4.6 | $3.00 | $0.30 | $15.00 |
| Claude Haiku 4.5 | $1.00 | $0.10 | $5.00 |
| GPT-5.6-Sol | $4.00 | $0.40 | $20.00 |
| GPT-5.6-Terra | $2.00 | $0.20 | $12.00 |
| GPT-5.6-Luna | $0.20 | $0.02 | $1.20 |
| GPT-5.5-pro | $30.00 | n/a | $180.00 |
| GPT-5.5 | $5.00 | $0.50 | $30.00 |
| GPT-5.4 | $2.50 | $0.25 | $15.00 |
| GPT-5.4-mini | $0.75 | $0.075 | $4.50 |
| GPT-5.4-nano | $0.20 | $0.02 | $1.25 |
| Gemini 3.1 Pro Preview | $2.00 (over 200K: $4.00) | $0.20 (over 200K: $0.40) plus storage | $12.00 (over 200K: $18.00) |
| Gemini 3.5 Flash | $1.50 | $0.15 plus storage | $9.00 |
| Gemini 3 Flash Preview | $0.50 | $0.05 plus storage | $3.00 |
| DeepSeek V4 Pro (off-peak) | $0.66 | $0.022 | $1.98 |
| DeepSeek V4 Flash (off-peak) | $0.22 | $0.007 | $0.66 |
DeepSeek peak/off-peak rates, in effect since August 16, 2026 at 16:00 UTC (off-peak hours are all times outside 01:00-04:00 and 06:00-10:00 UTC; peak = 2x off-peak; verified August 27, 2026):
| Model | Input (cache miss) | Input (cache hit) | Output |
|---|---|---|---|
| DeepSeek V4 Flash, off-peak | $0.22 | $0.007 | $0.66 |
| DeepSeek V4 Flash, peak | $0.44 | $0.014 | $1.32 |
| DeepSeek V4 Pro, off-peak | $0.66 | $0.022 | $1.98 |
| DeepSeek V4 Pro, peak | $1.32 | $0.044 | $3.96 |
Source pages, all verified August 27, 2026 (Gemini verified July 26):
A few quick reads off the table. DeepSeek V4 Flash output off-peak at $0.66 is roughly 76x cheaper than Claude Fable 5 output at $50, and 30x cheaper than GPT-5.6-Sol output at $20 (verified August 27, 2026 on the pages above); during peak hours that drops to 38x and 15x. The July 30 cuts reordered the cost-optimized tier: GPT-5.6-Luna at $0.20/$1.20 now undercuts Claude Haiku 4.5 ($1/$5) by 5x on input and 4x on output, and ties GPT-5.4-nano on input ($0.20) while beating it on output ($1.20 vs $1.25). Claude Sonnet 5 at $2/$10 is now the cheapest Claude model above Haiku at standard (non-promo) pricing, and its input rate ties GPT-5.6-Terra ($2/$12) after Terra's July 30 cut. The August 27 Sol cut to $4/$20 makes Sol's output ($20) cheaper than Opus 5's output ($25) while trailing on input ($4 vs $5). Gemini caching is still the odd one out: it bills a per-token cache rate plus an hourly storage fee ($1.00 per MTok per hour on the Flash models, $4.50 on 3.1 Pro Preview), so a cache you hold but rarely hit can cost more than no cache.
From the archive
Jun 11, 2026 • 11 min read
Jun 11, 2026 • 8 min read
Jun 11, 2026 • 8 min read
Jun 11, 2026 • 10 min read
Anthropic's pricing page carries a note that is easy to skim past: Opus 4.7 and later models, including Fable 5, use a new tokenizer that "may use up to 35% more tokens for the same fixed text" (Anthropic pricing docs, verified August 27, 2026). The models overview puts the typical figure at roughly 30% (verified August 27, 2026).
This matters for cross-provider comparisons. Fable 5 versus Opus 4.8 is apples to apples, since both use the new tokenizer. But if you compare either against Sonnet 4.5-era baselines, or against another provider using your old token counts, the same prompt now consumes up to a third more billable tokens. The honest comparison requires re-counting your actual prompts with Anthropic's token counting endpoint, not multiplying old counts by new rates. For workload-level math, see Fable 5 production cost modeling.
Claude Fable 5, Opus 4.8, Opus 4.7, Opus 4.6, and Sonnet 4.6 now include the full 1M token context window at standard pricing. Anthropic's pricing page states it plainly: a 900K-token request bills at the same per-token rate as a 9K-token request, and caching and batch discounts apply at standard rates across the full window (verified August 27, 2026).
Gemini still tiers. Gemini 3.1 Pro Preview doubles input from $2 to $4 and lifts output from $12 to $18 once a prompt crosses 200K tokens, and the cache rate doubles too (Gemini API pricing, verified July 26, 2026). Gemini 2.5 Pro tiers the same way at the same boundary. So the headline "Gemini Pro is cheaper than Sonnet" claim flips for genuinely long-context work: above 200K tokens, Gemini 3.1 Pro input ($4) exceeds Sonnet 4.6 input ($3), which does not tier at all. OpenAI tiers as well, framed as separate long-context rates: GPT-5.6-Sol rises from $4 input / $20 output to $8 / $30 in the long-context band, Terra from $2 / $12 to $4 / $18, Luna from $0.20 / $1.20 to $0.40 / $1.80, and GPT-5.5-pro goes from $30 / $180 to $60 / $270 (OpenAI pricing, verified August 27, 2026). Of the four providers, only Anthropic and DeepSeek charge one flat rate across the full window.
All four providers discount aggressively, but the shapes are different:
If your workload is a chat or agent loop that resends a large stable prefix every turn, effective cost is dominated by the cache read rate, not the headline input rate. That reshuffles the table substantially in DeepSeek's and Anthropic's favor.
Matching models by intended tier rather than by name, OpenAI/Anthropic/DeepSeek rows verified August 27, 2026, Gemini verified July 26:
| Tier | Cheapest | Mid | Premium |
|---|---|---|---|
| Max capability | - | Claude Fable 5 ($10 / $50) | GPT-5.5-pro ($30 / $180) |
| Frontier workhorse | Gemini 3.1 Pro Preview ($2 / $12 under 200K) | Claude Opus 5 ($5 / $25) | GPT-5.6-Sol ($4 / $20) / GPT-5.5 ($5 / $30) |
| Production mid-tier | Claude Sonnet 5 ($2 / $10) | GPT-5.6-Terra ($2 / $12) | Claude Sonnet 4.6 ($3 / $15) |
| Fast and cheap | DeepSeek V4 Flash ($0.22 / $0.66 off-peak) | GPT-5.6-Luna ($0.20 / $1.20) | Claude Haiku 4.5 ($1 / $5) |
Three rows shifted since July. Frontier workhorse: Opus 5 replaces Opus 4.8 at the same price ($5/$25) but with near-Fable 5 intelligence, making it a strong price-performance play - until the August 27 cut took GPT-5.6-Sol to $4/$20, undercutting Opus 5 on both axes. Production mid-tier: Claude Sonnet 5's $2/$10 rate is now permanent rather than promotional, so the tier's cheapest slot no longer expires on September 1. Fast and cheap: the July 30 cuts made GPT-5.6-Luna ($0.20/$1.20) the cheapest closed-provider option in the tier, undercutting Claude Haiku 4.5 ($1/$5) on both input and output, and only DeepSeek V4 Flash ($0.22/$0.66 off-peak; $0.44/$1.32 peak) sits below it. The max capability row remains unchanged: Fable 5 at $10/$50 is still the most capable model Anthropic offers, though Opus 5 now competes within a few percentage points on most benchmarks at half the price. There is a deeper dive in the Fable 5 cost-per-task analysis and the budget tier comparison.
Side projects and prototypes. DeepSeek V4 Flash ($0.22 / $0.66 off-peak) or Gemini 3 Flash Preview ($0.50 / $3.00, with a free tier) are hard to argue with. Among the US closed providers, GPT-5.6-Luna at $0.20 / $1.20 is now the cheapest input-and-output combo, edging out GPT-5.4-nano ($0.20 / $1.25). At these prices, model choice is a quality decision, not a budget one. The open-weights angle adds another dimension - see notes on DeepSeek's open-weights economics.
Production apps with steady traffic. The mid-tier triangle is now Claude Sonnet 5 ($2 / $10), GPT-5.6-Terra ($2 / $12), and Gemini 3.1 Pro Preview ($2 / $12 under 200K). Terra's July 30 cut pulled it level with the other two on input, and Sonnet 5's $2/$10 is no longer time-limited, so it wins the mid-tier on output. If your prompts regularly exceed 200K tokens, Gemini's tier flip moves it from cheapest to most expensive.
Agentic and long-horizon workloads. Output tokens dominate agent spend, and output is where the providers diverge most. After the August 27 cut, GPT-5.6-Sol ($4 / $20) now undercuts Claude Opus 5 ($5 / $25) on both input and output at the frontier workhorse tier. At the cheap end, Luna's $1.20 output versus Haiku's $5 makes OpenAI the cost winner for input-heavy agent loops - see the budget tier comparison for the per-task math. Fable 5 at $50 output only pays off when its quality reduces retries and total tokens. Batch APIs (50% off at Anthropic, OpenAI, and Google) are the easiest structural saving for any agent pipeline that is not latency-sensitive.
Teams routing across providers. With four providers publishing OpenAI-compatible or near-compatible APIs and a 273x output price spread, routing cheap-by-default with escalation on failure is increasingly the rational architecture. The LLM router comparison covers the tooling, and the routing strategies guide has the decision frameworks.
Compliance-constrained teams. Residency costs extra everywhere: Anthropic applies a 1.1x multiplier for US-only inference on Opus 4.6, Sonnet 4.6, and later models, and OpenAI charges a 10% uplift for regional data-residency processing on models released on or after March 5, 2026 (both verified August 27, 2026 on the pricing pages above).
Cheaper sticker prices do not automatically mean cheaper bills, and there are real reasons to stay put:
If none of those apply and your evals show parity, the spread in this table is too large to ignore.
DeepSeek V4 Flash, at $0.22 per million input tokens (cache miss, off-peak) and $0.66 per million output tokens, verified August 27, 2026 on DeepSeek's official pricing page. Cache hits drop input to $0.007 off-peak. During peak hours (01:00-04:00 and 06:00-10:00 UTC) rates double to $0.44 / $1.32 (cache hit $0.014) - still the cheapest table, just a different headline number. Among US closed providers, GPT-5.6-Luna at $0.20/$1.20 is the cheapest input-and-output combination after its July 30 price cut, edging GPT-5.4-nano ($0.20/$1.25), and Gemini 3 Flash Preview has a free tier.
No. Claude Fable 5, Opus 4.8, Opus 4.7, Opus 4.6, Sonnet 5, and Sonnet 4.6 include the full 1M token context window at standard per-token rates, with no premium above 200K tokens (verified August 27, 2026). Gemini 3.1 Pro Preview and Gemini 2.5 Pro still charge higher rates for prompts above 200K tokens. OpenAI's GPT-5.6 family also tiers: in the long-context band Sol rises from $4/$20 to $8/$30, Terra from $2/$12 to $4/$18, and Luna from $0.20/$1.20 to $0.40/$1.80.
Fable 5 costs $10 input / $50 output per MTok versus GPT-5.6-Sol at $4 / $20 (after the August 27 cut), so 2.5x on input and 2.5x on output. Against GPT-5.5-pro ($30 / $180), Fable 5 is actually the cheaper max-capability option by 3x on input and 3.6x more on output. All figures verified August 27, 2026.
GPT-5.6 is OpenAI's flagship model family, released July 9, 2026. It ships in three tiers: Sol ($5/$30 at launch, cut to $4/$20 on August 27), Terra ($2.50/$15 at launch, cut to $2/$12 on July 30), and Luna ($1/$6 at launch, cut to $0.20/$1.20 on July 30). All three support the same cache and batch discounts as GPT-5.5. After the August 27 cut, Sol at $4/$20 now undercuts GPT-5.5's $5/$30 on both axes; model selection comes down to benchmarks and your specific use case.
Claude models from Opus 4.7 onward, including Fable 5 and Sonnet 5, use a new tokenizer that can produce up to 35% more tokens for the same text compared with older Claude models, per Anthropic's pricing documentation (verified August 27, 2026). Token counts measured on older models do not transfer, so per-MTok rate comparisons against pre-4.7 baselines undercount real spend.
Anthropic, OpenAI, and Google all offer a 50% batch discount on input and output tokens (Anthropic and OpenAI verified August 27, 2026, Google verified July 26). DeepSeek's pricing page lists no batch tier, but its standard rates sit below the other providers' batch rates in most tiers anyway.
It does not - the promotion became the standard price. Anthropic's pricing page (verified August 27, 2026) states the $2/$10 per MTok rates for Sonnet 5 are now standard, and the previously scheduled increase to $3/$15 on September 1, 2026 will not occur. Budgets built on the July assumption that Sonnet 5 would get more expensive can stand.
Read next
The sub-$1.50 coding tier just got serious: DeepSeek V4 Flash 0731 posts frontier-adjacent agent scores at $0.14/$0.28 (peak/off-peak pricing from Aug 16), GPT-5.6 Luna dropped 80% to $0.20/$1.20, and Gemini 3.5 Flash and Claude Haiku 4.5 hold the hosted middle. Prices verified July 31 and August 15, 2026.
10 min readClaude Fable 5 vs Gemini: how Anthropic's $10/$50 API-only model compares to Gemini 3.1 Pro's $2/$12 preview on pricing, context, and benchmarks - and why Opus 5 changed the decision.
10 min readA practical guide to routing between Claude Opus 5, Sonnet 5, Haiku 4.5, GPT-5.6 Sol/Terra/Luna, and Kimi K3 based on task complexity, cost budget, and latency requirements - with decision frameworks and code examples.
10 min readTechnical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.
Open-source AI pair programming in your terminal. Works with any LLM - Claude, GPT, Gemini, local models. Git-aware ed...
View ToolOpen-source AI coding agent for terminal, desktop, and IDE. Works with 75+ LLM providers including Claude, GPT, Gemini,...
View ToolGoogle's frontier model family. Gemini 2.5 Pro has 1M token context and top-tier coding benchmarks. Gemini 3 Pro pushes...
View ToolAnthropic's smallest Claude 4.5 model. Near-frontier coding performance at one-third the cost of Sonnet 4 and up to 4-5x...
View ToolEvery coding agent in one window. Stop alt-tabbing between Claude, Codex, and Cursor.
View AppBeat the August 2026 Assistants API sunset. Paste old code, get Responses API.
View AppTurn a one-liner into a working Claude Code skill. From idea to installed in a minute.
View AppUse opus, sonnet, haiku, and best to switch models easily.
Claude CodeInteractive UI to switch models and effort sliders mid-session.
Claude CodeAdd gateway or custom models to the picker via environment variables.
Claude Code
Sign up for Runbear with the below link to receive a 25% discount! https://runbear.io/?utm_source=developerdigest Coupon Code: DEVELOPERDIGEST Valid until February 2025 In this video, I demonstrat...

Check out Zed here! https://zed.dev In this video, we dive into Zed, a robust open source code editor that has recently introduced the Agent Client Protocol. This new open standard allows...

In this video, we explore Kombai, a powerful platform transforming front-end development. Kombai enables seamless import of Figma designs and integration with widely used UI libraries and framework...

The sub-$1.50 coding tier just got serious: DeepSeek V4 Flash 0731 posts frontier-adjacent agent scores at $0.14/$0.28 (...

A practical guide to routing between Claude Opus 5, Sonnet 5, Haiku 4.5, GPT-5.6 Sol/Terra/Luna, and Kimi K3 based on ta...

How much of an AI session can you actually take with you? Store defaults, encrypted reasoning, opaque compaction, hidden...

OpenAI slashes GPT-5.6 Luna by 80% to $0.20/M input tokens, cuts Terra by 20%, adds Sol Fast mode at 2.5x speed, and rev...

Luna drops from $1/$6 to $0.20/$1.20 per million tokens, Terra from $2.50/$15 to $2/$12, and Sol gets a paid Fast mode....

Claude is now GA in Microsoft Foundry on Azure with native billing, Entra ID auth, and GB300 Blackwell infrastructure. H...

New tutorials, open-source projects, and deep dives on coding agents - delivered weekly.