
TL;DR
GPT-5.6 Sol gets a chat-focused retune with 68% fewer factual errors in OpenAI's internal eval, a new effort slider, and GPT-5.6 Luna becomes the default model for Free and Go users with unlimited text chats. What the API did not change and why the split matters.
OpenAI announced today that GPT-5.6 Sol, its frontier model, is being retuned for everyday ChatGPT conversations, while GPT-5.6 Luna becomes the default model for Free and Go users with unlimited text chats. Two numbers frame the update: OpenAI's internal evaluation found responses with at least one factual error were 68% less common with the new Sol and 62% less common with Luna compared to GPT-5.5 Instant on financial, medical, and legal prompts. The second half of the story is a tier shift: a model that cost $6 per million output tokens four weeks ago is now the free default.
| Change | Where | Users |
|---|---|---|
| GPT-5.6 Sol retuned for chat (focused answers, tighter formatting, fewer factual errors) | ChatGPT Chat | Plus, Pro |
| New effort slider (quick to deep reasoning) | ChatGPT web, mobile, desktop | Plus, Pro |
| GPT-5.6 Luna becomes the default model | ChatGPT | Free, Go |
| Unlimited text chats with Luna | ChatGPT | Free, Go (starts next week) |
| Think button for higher reasoning | ChatGPT | Free, Go (starts next week) |
| GPT-5.6 Sol powering Work and Codex | unchanged | - |
The retune is chat-specific: OpenAI explicitly says the version of Sol behind Work and Codex "is not changing as part of this release." So the API model id, pricing, and agent behavior you build against today are untouched. This is a product-surface update, not a model release - which is exactly why it is worth reading carefully.
The new Sol adapts answer length to the question, avoids unnecessary formatting, and gives a corrective answer instead of agreeing when agreement would not be useful. Plus and Pro users get a slider to set how much "thought" each response gets, and Instant and Thinking now feel like one model with different effort levels rather than two different personalities.
OpenAI's factuality numbers come from an internal evaluation, so treat them as a directional signal, not an independent benchmark. What matters for developers: the model family's behavior in chat is now tuned and productized separately from its behavior in agentic surfaces. That is a real departure from the GPT-5.5 era, where one behavior profile shipped everywhere.
From the archive
Aug 6, 2026 • 8 min read
Aug 6, 2026 • 10 min read
Aug 6, 2026 • 6 min read
Aug 5, 2026 • 7 min read
Free and Go users get GPT-5.6 Luna as the default this week, unlimited text chats starting next week, and a Think button that escalates harder questions to deeper reasoning. Limits remain on file uploads, images, and other tools.
This is the price story from July 30 - Luna dropped 80% to $0.20/$1.20 per million tokens - landing in the free product within a week. Luna keeps its 1,050,000-token context window, so the free tier now carries a million-token model as its default. OpenAI framed it as "more abundant intelligence": unlimited text chats at the efficiency tier is the consumer equivalent of what the API cut did for agent fleets.
Four takeaways:
Your API code did not change. Sol in the API, in Codex, and in Work is the same model at the same $5/$30 pricing. If you saw this headline and worried about a model swap under your agent, you do not need to.
Behavior is now a product surface decision. OpenAI can tune Sol one way for chat and keep it unchanged for agentic work. Expect more of this: the API may increasingly be the only place where "model behavior" is stable, because the consumer surfaces will keep absorbing these retunes.
The free tier is a testing ground. Unlimited Luna text chats with a Think button means millions of users are generating preference data on the efficiency tier. The 62% factuality improvement cited for Luna matters precisely because that tier is about to see the largest traffic OpenAI has ever routed through it.
Competitive pressure on consumer AI. Google's Gemini free tier and Anthropic's Claude free tier now answer to a free product running a frontier-adjacent model with a million-token context. Our GPT-5.6 vs Claude 5 model tier comparison and the budget model pricing landscape both need a footnote: the efficiency tier is now what consumers see first.
Chart: OpenAI, from the announcement post. The Think button is the free-tier escalation path to deeper reasoning.
The August update system card documents new under-18 guardrails: training against romantic roleplay, age-restricted challenges, and self-presentation as a substitute for real relationships, plus age-appropriate boundaries on sexual content, eating disorders, body-image risks, and graphic violence. The card also covers the eval changes behind the factuality numbers. If you evaluate OpenAI models on safety-relevant prompts, the card is the source of truth for what changed in this snapshot.
No. OpenAI states the version of Sol powering Work and Codex is not changing; the retune applies to the Chat experience in ChatGPT only.
This week Luna becomes the default for Free and Go users; unlimited text chats and the Think button arrive starting next week, subject to abuse guardrails.
No. Luna stays at $0.20 per million input and $1.20 per million output tokens from the July 30 cut, now also serving as the free-tier default.
It sets how much reasoning ChatGPT applies per response, from quick everyday answers to deeper planning, research, and coding work, on web, mobile, and desktop.
Read next
OpenAI slashes GPT-5.6 Luna by 80% to $0.20/M input tokens, cuts Terra by 20%, adds Sol Fast mode at 2.5x speed, and reveals Sol autonomously optimized its own production kernels.
7 min readLuna drops from $1/$6 to $0.20/$1.20 per million tokens, Terra from $2.50/$15 to $2/$12, and Sol gets a paid Fast mode. What the new floor means for agent economics, Codex quotas, and the competition.
8 min readA practical guide to choosing GPT-5.6 Sol, Terra, and Luna, using programmatic tool calling, caching, and the multi-agent beta in production.
8 min readTechnical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.
OpenAI's flagship. GPT-4o for general use, o3 for reasoning, Codex for coding. 300M+ weekly users. Tasks, agents, web br...
View ToolOpenAI's latest flagship model. Major leap in reasoning, coding, and instruction following over GPT-4o. Powers ChatGPT P...
View ToolUnified API for 200+ models. One API key, one billing dashboard. OpenAI, Anthropic, Google, Meta, Mistral, and more. Aut...
View ToolGoogle's frontier model family. Gemini 2.5 Pro has 1M token context and top-tier coding benchmarks. Gemini 3 Pro pushes...
View ToolInstall Ollama and LM Studio, pull your first model, and run AI locally for coding, chat, and automation - with zero cloud dependency.
Getting StartedApprove each action manually - the safest mode for new tasks.
Claude Code
Learn The Fundamentals Of Becoming An AI Engineer On Scrimba; https://v2.scrimba.com/the-ai-engineer-path-c02v?via=developersdigest OpenAI's New O1 Model and $200/Month ChatGPT Pro Tier: What's...

OpenAI AI has launched their first browser called ChatGPT Atlas, which incorporates ChatGPT for enhanced functionality. This browser allows users to interact with their documents using natural...

In this video, we delve into OpenAI's latest release, Codex, a cloud-based software engineering agent designed for various coding tasks. Unlike tools like Cursor or Windsurf, Codex integrates...

OpenAI slashes GPT-5.6 Luna by 80% to $0.20/M input tokens, cuts Terra by 20%, adds Sol Fast mode at 2.5x speed, and rev...

The "Building abundant intelligence" essay carries real engineering numbers: GPT-5.6 Sol cut serving costs 20%, speculat...

OpenAI is consolidating its desktop apps, not merging ChatGPT and Codex into one indistinguishable product. Here is how...

Meta released Muse Code, a terminal coding agent, and Muse Spark 1.2 on August 5, 2026. The model co-trains with the har...

Liquid AI shipped LFM2.5-2.6B on August 4, 2026: a 2.6B open-weight model trained for agentic work inside real harnesses...

OpenAI published the engineering story behind GPT-Live, its third-generation voice system: a full-duplex model with no t...

New tutorials, open-source projects, and deep dives on coding agents - delivered weekly.