GPT-6 in 7 Minutes: Astra, Sol, and Luna Explained for Developers
TL;DR
A developer-first overview of the GPT-6 family: what Astra, Sol, and Luna are, what each costs per million tokens, the three new API features (async tool calling, mid-turn steering, cache-safe reasoning changes), and how to try each tier in Codex and the API today.
GPT-6 is a three-model family from OpenAI: Astra is the top-capability tier ($10 input / $50 output per million tokens), Sol is the everyday reasoning tier ($2 / $10), and Luna is the high-volume tier ($0.10 / $0.50). All three share a 1,050,000-token context window, take text and images in, and return text. You can try them today in the OpenAI API as gpt-6-astra, gpt-6-sol, and gpt-6-luna, and in Codex through a ChatGPT plan.
That is the answer the GPT-6 In 7 Minutes video walks through in 7 minutes and 12 seconds. This post is the written version for developers: the tier table, the new API features, the numbers worth remembering, and the exact ways to switch a project over.
Last updated: September 28, 2026. Model specs and prices were checked against OpenAI's model pages that day. The video had no captions available when we wrote this, so video-specific details come from its published description and chapter list, and every number below was re-checked against OpenAI's own pages.
Official Sources#
| Source | Link |
|---|---|
| GPT-6 Astra announcement | openai.com/index/gpt-6-astra |
| GPT-6 Sol and Luna announcement | openai.com/index/introducing-gpt-6-sol-and-luna |
| Using GPT-6 (API guide) | developers.openai.com/api/docs/guides/latest-model |
| Astra, Sol, and Luna model pages | gpt-6-astra, gpt-6-sol, gpt-6-luna |
| System card | deploymentsafety.openai.com/gpt-6-astra |
The Three Tiers at a Glance#
| GPT-6 Astra | GPT-6 Sol | GPT-6 Luna | |
|---|---|---|---|
| API model id | gpt-6-astra | gpt-6-sol | gpt-6-luna |
| Input, per 1M tokens | $10.00 | $2.00 | $0.10 |
| Cached input, per 1M | $1.00 | $0.20 | $0.01 |
| Output, per 1M tokens | $50.00 | $10.00 | $0.50 |
| Context window | 1,050,000 | 1,050,000 | 1,050,000 |
| Max output | 128,000 | 128,000 | 128,000 |
| Reasoning effort | low to max | none to max | none to max |
| Best for | Long, tool-heavy, expensive-to-get-wrong work | Everyday coding and agents | Repeatable work at scale |
Astra also charges $12.50 per million tokens for cache writes. Requests that run past 272K tokens of context cost more; our Astra release guide has the long-context multipliers and the full benchmark table. OpenAI cut Sol and Luna prices by 50% against their GPT-5.6 promotional rates when it launched them, which is why the September price comparison with Opus 5.5 and Grok 4.7 looks so different from the summer one.
One detail that trips up migrations: Astra does not support the none reasoning effort, while Sol and Luna do. If a request uses none today, move it to low for Astra.
What the Video Covers#
The chapter list in the video description gives the shape of the seven minutes:
| Time | Chapter |
|---|---|
| 00:00 | GPT-6 launch overview |
| 00:20 | Computer use highlights |
| 00:58 | Benchmarks and alignment |
| 02:04 | Pricing and value |
| 03:07 | New API features |
| 03:25 | Context and multimodal |
| 03:50 | Showcase 3D examples |
| 04:30 | Safety and blog notes |
| 04:57 | Why computer agents matter |
| 05:50 | Speed gains and use cases |
| 06:54 | Wrap up and next video |
The through-line is computer use: a model that operates real software instead of only writing code about it. That is also why the 3D demos in the video matter. Building a detailed world in Unreal Engine is a test of an agent driving a large, stateful application. We built on that idea in two follow-up demos: a single-file Three.js world and a Blender scene that becomes a playable game.
The Numbers Worth Remembering#
All figures below are from OpenAI's launch announcement, where scores are the maximum at any effort level. They are vendor-reported, so treat them as a starting point for your own evals.
| Benchmark | GPT-6 Astra | GPT-5.6 Sol | Notes |
|---|---|---|---|
| Agents' Last Exam | 59.3% | 53.6% | Claude Opus 5 scored 55.5% |
| OSWorld 2.0 (offline set) | 72.6% at about 40 min per task | 65.7% at about 75 min per task | OpenAI's latency simulation |
| ScreenSpot-Pro (no tools) | 92.7% | 76.9% | Screen grounding |
| AutomationBench | 41.4% | 18.1% | Claude Fable 5.1 scored 31.4% |
| ARC-AGI-3 | 99.9% | 7.8% | Run with OpenAI's Responses API harness |
On the Codex side, OpenAI says it updated the harness alongside Astra and measured 1.9x faster task completion than the GPT-5.6 Sol experience on the Mind2Web benchmark.
One line in the video description deserves a precise reading. It mentions a 0% score on Exploit Gym. In OpenAI's table, the 0.0% figure is the "ExploitGym honeypot" alignment row, where lower is better: it measures whether a model facing an impossible task goes beyond its authorized scope, and GPT-5.6 Sol did so 48.2% of the time. It is not Astra's exploit success rate on the benchmark itself, which OpenAI reports as 42.4%. Astra is also the first OpenAI model rated Critical for cybersecurity, so the launch version refuses advanced offensive tasks like writing proof-of-concept exploits. The safety write-up covers what that means in practice.
Three New API Features#
The API guide lists what is new for GPT-6. These are the ones that change how you build agents:
- Async tool calling. Set
async: trueon a function or custom tool and the model keeps reasoning, calls other tools, or answers independent parts of the request while your application runs the tool. You return the result later using the originalcall_id. Your application still executes the tool and tracks pending work. - Mid-turn steering. You can send extra user instructions while the model is working, such as a correction or a changed requirement. Over a WebSocket connection, the Responses API preserves completed work and folds the update into a continuation.
- Reasoning changes that keep the cache. Add a
configuration_updateinput item to raise or lower reasoning effort without rewriting the prompt prefix. Keep request-levelreasoning.effortunchanged so the cached prefix survives.
For coding agents, the practical effect is fewer restarts. A long session can take a correction at minute 20 instead of being killed and replayed, and a cheap follow-up can drop to low effort without paying to re-read the whole context.
How to Try Each Tier#
In Codex. Astra, Sol, and Luna are available through ChatGPT plans that include them (Sol and Luna for Plus, Pro, Business, Enterprise, and Edu; Astra rolled out to Plus, Pro, Business, and Enterprise, with Enterprise admins switching it on per workspace). From the terminal, the -m flag picks the model:
codex -m gpt-6-astra
codex exec -m gpt-6-sol "Find why the integration tests in ./api are flaky and fix the root cause"
We confirmed the -m, --model flag against codex --help on codex-cli 0.155.0. New to the CLI? Start with OpenAI Codex in 7 Minutes.
In the API. Set model on a Responses API request and choose an effort level:
curl https://api.openai.com/v1/responses \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6-sol",
"input": "Summarize the failing tests in this log and propose a fix.",
"reasoning": { "effort": "medium" }
}'
Use the Responses API for anything with tools. The API guide says Astra supports Chat Completions, but its tool calling requires Responses, and Sol and Luna support function calling in Chat Completions only with reasoning_effort: "none".
Which Tier to Pick#
| If your task is | Start with |
|---|---|
| A long migration, a computer-use workflow, or a multi-hour agent run | Astra at high or xhigh |
| Daily coding, review, and refactors | Sol at medium |
| Classification, extraction, or many parallel sub-agents | Luna |
A cheap test that pays for itself: run one real task from your backlog on Sol and on Astra, log cost and wall-clock time, and only move to Astra where it removes retries. OpenAI's own numbers say Astra often uses fewer output tokens per task, but at five times Sol's per-token price, that has to show up in your logs before it counts.
Watch the Video#
GPT-6 In 7 Minutes is the fastest way to see the computer-use and 3D examples in motion, which a table cannot show. The chapter at 03:50 is the one to scrub to if you only have a minute.
FAQ#
What is GPT-6?#
GPT-6 is OpenAI's model family with three tiers: GPT-6 Astra (most capable), GPT-6 Sol (balanced), and GPT-6 Luna (efficient, high volume). Astra launched first, in early September 2026, and Sol and Luna followed on September 22.
How much does GPT-6 cost?#
Per million tokens: Astra is $10 input and $50 output, Sol is $2 and $10, and Luna is $0.10 and $0.50. Cached input is $1, $0.20, and $0.01 respectively.
How do I use GPT-6 in Codex?#
Pick the model in Codex on a ChatGPT plan that includes it, or pass -m gpt-6-astra, -m gpt-6-sol, or -m gpt-6-luna to the Codex CLI. Enterprise workspaces need an admin to enable Astra.
Does GPT-6 accept audio or video input?#
The API model pages list text and image input with text output for all three models. Audio and video are not listed as inputs.
What is async tool calling?#
It is a GPT-6 API feature where a tool marked async: true runs in your application while the model keeps working on independent parts of the request. You send the result back with the original call_id when it is ready.
Sources#
| Source | URL |
|---|---|
| GPT-6 Astra announcement (OpenAI), fetched September 28, 2026 | https://openai.com/index/gpt-6-astra/ |
| Introducing GPT-6 Sol and Luna (OpenAI), fetched September 28, 2026 | https://openai.com/index/introducing-gpt-6-sol-and-luna/ |
| Using GPT-6 API guide (OpenAI), fetched September 28, 2026 | https://developers.openai.com/api/docs/guides/latest-model |
| GPT-6 Astra model page, fetched September 28, 2026 | https://developers.openai.com/api/docs/models/gpt-6-astra |
| GPT-6 Sol model page, fetched September 28, 2026 | https://developers.openai.com/api/docs/models/gpt-6-sol |
| GPT-6 Luna model page, fetched September 28, 2026 | https://developers.openai.com/api/docs/models/gpt-6-luna |
| GPT-6 In 7 Minutes (YouTube, description and chapters) | https://www.youtube.com/watch?v=lLDFW9cjSsI |
Continue Reading#
- GPT-6 Astra Release Guide - the full benchmark table, Astra safeguards, and OpenCode commands
- GPT-6 Sol vs Claude Opus 5.5 vs Grok 4.7 - the September price war in one comparison
- GPT-6 World Building: One HTML File to an Interactive Three.js World - the 3D follow-up demo
- GPT-6 and Blender: Build a Playable 3D Game With Codex and AI Video - Astra driving Codex through Blender
- OpenAI Codex in 7 Minutes - the CLI basics before you switch models
Get the next deep dive like this in your inbox
One email a week on GPT-6 and the rest of the AI dev stack. Free.
Read next on AI coding tools
GPT-6 Astra Release Guide: Benchmarks, $10/$50 Pricing, and How to Run It in Codex and OpenCode
GPT-6 Astra is OpenAI's max-capability model: 1.05M context, $10/$50 per million tokens, the first OpenAI model rated Critical for cyber, and new highs on Terminal-Bench 4.0 and OSWorld 2.0. What it is, what the benchmarks do and do not say, and the verified commands to try it.
8 min readGPT-6 Sol vs Claude Opus 5.5 vs Grok 4.7: The September 2026 Price War
Three frontier launches in 48 hours repriced the agentic workhorse tier: Grok 4.7 at $2/$6 (Sep 21), Claude Opus 5.5 at $4/$20 with $0.20 cache reads (Sep 22), and GPT-6 Sol at $2/$10 with Luna at $0.10/$0.50 (Sep 22). Same-day-verified rates, honest benchmark attribution, and a decision guide.
9 min readGPT-6 World Building: One HTML File to an Interactive Three.js World
The latest Developers Digest demo has GPT-6 driving Codex to build a single-file Three.js world - five walkable rooms with AI-generated wall assets and even a museum room that teaches LLM concepts. Here is the verified toolchain and when this web-native 3D path wins.
7 min readNew here? Start with
Technical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.








