GPT-6 Astra Release Guide: Benchmarks, $10/$50 Pricing, and How to Run It in Codex and OpenCode

TL;DR
GPT-6 Astra is OpenAI's max-capability model: 1.05M context, $10/$50 per million tokens, the first OpenAI model rated Critical for cyber, and new highs on Terminal-Bench 4.0 and OSWorld 2.0. What it is, what the benchmarks do and do not say, and the verified commands to try it.
GPT-6 Astra is OpenAI's top-tier model: the max-capability member of the GPT-6 family, available in the API as gpt-6-astra at $10 per million input tokens and $50 per million output tokens, with a 1,050,000-token context window. OpenAI announced it on September 3, 2026 as a staged rollout, first to a limited set of organizations and then to ChatGPT Plus, Pro, Business, and Enterprise, the OpenAI API, Microsoft Azure, and AWS Bedrock. On September 22 OpenAI filled out the family with the cheaper GPT-6 Sol and GPT-6 Luna. If you have seen "Astra" all over your feeds this week, that is why: it is the model behind the DrivingBench real-car run, the 1941 Enigma break, and a lot of the long-horizon agent demos, and OpenAI's DevDay is tomorrow.
For developers the short version is: Astra is the model you reach for when a task is long, tool-heavy, and expensive to get wrong. For everyday coding, its sibling Sol is a fifth of the price.
Official Sources#
| Source | Link |
|---|---|
| OpenAI announcement (September 3, 2026, updated September 22) | openai.com/index/gpt-6-astra |
| API model page (fetched September 28, 2026) | developers.openai.com/api/docs/models/gpt-6-astra |
| System card | deploymentsafety.openai.com/gpt-6-astra |
| Safety overview | openai.com/index/safety-overview-gpt-6-astra |
| OpenCode docs | opencode.ai/docs |
What Shipped#
Astra is the reasoning-heavy flagship. Per the API model page:
- Model id:
gpt-6-astra - Context: 1,050,000 tokens, up to 128,000 output tokens
- Knowledge cutoff: April 30, 2026
- Modalities: text and image in, text out
- Endpoints: Responses, Chat Completions, and Batch
- Reasoning effort:
low,medium,high,xhigh, andmax - Tools: web search, file search, code interpreter, hosted shell, apply patch, skills, computer use, MCP, and tool search
Two product changes ship with it. In Codex, OpenAI says Astra can keep notes across context windows instead of repeatedly compacting a session into one summary, and earlier windows stay searchable; it is an experimental config option for now and OpenAI says it will become the default for Astra. And OpenAI updated the Codex harness for faster computer use, claiming 1.9x faster task completion than the GPT-5.6 Sol experience on Mind2Web.
Pro, Business, and Enterprise ChatGPT plans also get GPT-6 Astra Pro. Enterprise admins have to switch Astra on for their workspace; it is off by default.
Benchmarks#
Every number below is from OpenAI's own announcement table, where scores are the maximum at any effort level. Note the comparison column is GPT-5.6 Sol, not the new GPT-6 Sol, and Claude Opus 5.5 is not in the table because it shipped after Astra.
| Benchmark | GPT-6 Astra | GPT-5.6 Sol | Claude Fable 5.1 | Claude Opus 5 |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 57.9% | 37.3% | 55.8% | 52.6% |
| DeepSWE v1.1 | 74.1% | 72.7% | 67.4% | 73.7% |
| OSWorld 2.0 (offline set) | 72.6% | 65.7% | - | 70.2% |
| Agents' Last Exam | 59.3% | 53.6% | - | 55.5% |
| ARC-AGI-3 | 99.9% | 7.8% | - | 30.2% |
| Artificial Analysis Intelligence Index v4.1.1 | 61.2 | 60.9 | 65.7 | 63.1 |
| Artificial Analysis Coding Agent Index v1.4 | 67.0 | 65.1 | - | 68.1 |
Read this table the way you would read any vendor chart. The agentic rows (Terminal-Bench, OSWorld, Agents' Last Exam) are where Astra's lead is real and large. The two Artificial Analysis composite rows, which OpenAI included itself, tell a more modest story: on the independent intelligence index Astra trails Fable 5.1 and Opus 5, and on the coding-agent index it trails Opus 5. OpenAI also footnotes that ARC-AGI-3 was run with its Responses API harness, which changes two settings.
The efficiency claims matter as much as the scores. OpenAI says Astra hit its Terminal-Bench 4.0 result at roughly 63% lower estimated API cost per task than Fable 5.1, and on OSWorld 2.0 it finished tasks in about 40 minutes versus roughly 75 for GPT-5.6 Sol. Those are vendor latency simulations, not your workload, but they point at the same thing independent runs have shown: Astra tends to finish long tasks in fewer tokens. That shows up in DrivingBench too, where Astra was the first model to complete the real-car course.
Independent agent-level numbers point the same way on speed but not on price. On Artificial Analysis's Coding Agent Index (v1.5, checked September 28), Codex running Astra at max effort averaged $7.47 and 29.4 minutes per task, against $2.99 and 22.3 minutes for Codex on GPT-6 Sol and $13.00 and 1.1 hours for Claude Code on Opus 5.5. Astra is faster and cheaper per task than Opus 5.5 there, but Sol finishes the same benchmark tasks for less than half of Astra's cost.
Safety and the Cyber Rating#
Astra is the first OpenAI model to meet the Critical threshold for cybersecurity under the Preparedness Framework, which we flagged when OpenAI first disclosed it in August. In practice that means the launched version refuses advanced offensive tasks such as writing proof-of-concept exploits, while still doing secure code review and patching. OpenAI says broader defensive access goes through its Daybreak program, the same channel GPT-5.6-Cyber shipped through.
Two developer-facing consequences:
- The API can stop mid-task. OpenAI says extra safety checks can pause work in ChatGPT and Codex for your review, but in the API "the task will stop." Build retries and human hand-off into long Astra jobs.
- Reasoning is harder to monitor. OpenAI's own evaluations found Astra's written reasoning harder to monitor than GPT-5.6 Sol's, and says it is deploying production misalignment monitoring for Astra-class models. If your agent observability depends on reading chain-of-thought, test that assumption.
Pricing#
Verified on the API model page on September 28, 2026:
| GPT-6 Astra | GPT-6 Sol | GPT-6 Luna | |
|---|---|---|---|
| Input (per 1M tokens) | $10.00 | $2.00 | $0.10 |
| Cached input | $1.00 | $0.20 | $0.01 |
| Output (per 1M tokens) | $50.00 | $10.00 | $0.50 |
| Context window | 1,050,000 | 1,050,000 | 1,050,000 |
Astra also charges $12.50 per million for cache writes. Requests above 272K tokens pay 2x on input and cache rates and 1.5x on output, which works out to $20/$75. Fast mode runs up to 2x the speed of Standard at 2x the price. For the full cross-vendor comparison with Opus 5.5 and Grok 4.7, see our September price war breakdown.
A quick cost check: an agent turn that reads 60K tokens and writes 4K costs about $0.80 on Astra uncached, versus about $0.16 on Sol. With 80% of the input served from cache, the Astra turn drops to roughly $0.37. Astra only pays for itself when it removes retries, review cycles, or human time.
Run It in Codex#
Astra is available in Codex for ChatGPT plans that include it. The Codex CLI takes a model flag, which we confirmed against codex exec --help on codex-cli 0.155.0:
codex -m gpt-6-astra
codex exec -m gpt-6-astra "Find why the integration tests in ./api are flaky and fix the root cause"
Our own Astra e-commerce build walkthrough shows the model driving Codex end to end, including generated product assets.
Run It in OpenCode#
Astra is not on the OpenCode Go plan (only opencode-go/gpt-6-luna is listed there as of today), but it is listed for both the openai provider and OpenCode Zen in the models.dev registry that OpenCode reads. Install OpenCode with the official one-liner:
curl -fsSL https://opencode.ai/install | bash
Then connect a provider with /connect inside the TUI (an OpenAI API key, or OpenCode Zen), and run:
opencode run --model openai/gpt-6-astra --variant high "Refactor the payment module to use the new ledger API"
opencode run --model opencode/gpt-6-astra --variant max "Plan and execute the Postgres 17 upgrade in ./infra"
The --variant values map to the reasoning efforts above: low, medium, high, xhigh, and max. None of this site's automations run on Astra; at $10/$50 it is a flagpole model for us, not a background one.
Decision Guide#
| If you are... | Use |
|---|---|
| Running long, tool-heavy agent jobs (migrations, computer use, multi-hour Codex sessions) | GPT-6 Astra at high or xhigh |
| Doing everyday coding, review, and refactors | GPT-6 Sol at $2/$10 |
| Running high-volume classification, extraction, or sub-agent work | GPT-6 Luna, or a decision model like TypeSafe Jev if the output is a label or score |
| Wanting a second frontier opinion on the same task | Claude Opus 5.5 at $4/$20 |
| Doing offensive security research | Not launch-day Astra; apply through Daybreak |
A pattern worth testing: let Astra orchestrate and plan, and route bulk sub-tasks to Luna. The price gap between the two is 100x on input.
FAQ#
What is GPT-6 Astra?#
GPT-6 Astra is OpenAI's highest-capability GPT-6 model, announced September 3, 2026. It targets long-horizon agentic work, computer use, coding, and science, and is available in ChatGPT paid plans, Codex, the OpenAI API as gpt-6-astra, Azure, and AWS Bedrock.
How much does GPT-6 Astra cost in the API?#
$10 per million input tokens, $1 per million cached input tokens, and $50 per million output tokens, with cache writes at $12.50. Above 272K tokens of context the rates rise to $20 input and $75 output.
What is the difference between GPT-6 Astra, Sol, and Luna?#
Astra is the max-capability tier, Sol ($2/$10) is the coding-and-agentic default, and Luna ($0.10/$0.50) is the high-volume workhorse. All three share a 1,050,000-token context window. Sol and Luna shipped on September 22, 2026.
Can I use GPT-6 Astra in OpenCode?#
Yes, through the openai provider with your own key (openai/gpt-6-astra) or through OpenCode Zen (opencode/gpt-6-astra). It is not included in the OpenCode Go plan today.
Why does Astra refuse some security tasks?#
OpenAI rated Astra Critical for cybersecurity capability, so the public version refuses advanced offensive tasks like writing exploit proofs of concept. Secure code review and patching are allowed; less restricted access runs through OpenAI's Daybreak program.
Sources#
| Source | URL |
|---|---|
| GPT-6 Astra announcement (OpenAI), fetched September 28, 2026 | https://openai.com/index/gpt-6-astra/ |
| GPT-6 Astra API model page (OpenAI), fetched September 28, 2026 | https://developers.openai.com/api/docs/models/gpt-6-astra |
| GPT-6 Sol API model page (OpenAI), fetched September 28, 2026 | https://developers.openai.com/api/docs/models/gpt-6-sol |
| GPT-6 Astra system card (OpenAI) | https://deploymentsafety.openai.com/gpt-6-astra |
| Artificial Analysis Coding Agent Index v1.5, fetched September 28, 2026 | https://artificialanalysis.ai/agents/coding |
| models.dev registry (OpenCode model ids), fetched September 28, 2026 | https://models.dev/api.json |
| OpenCode docs, fetched September 28, 2026 | https://opencode.ai/docs/ |
Last updated: September 28, 2026
Continue Reading#
- GPT-6 Sol vs Claude Opus 5.5 vs Grok 4.7: The September 2026 Price War - the cheaper GPT-6 tiers against their direct competitors
- OpenAI DevDay 2026: What to Expect - confirmed schedule, credible reports, and labeled rumors for tomorrow's event
- How GPT-6 Astra Builds a Real E-Commerce Site - Astra driving Codex through a full storefront build
- OpenAI Says It Can't Rule Out Critical Cyber Capability for Astra - the August disclosure that set up the launch safeguards
- Claude Opus 5.5 Release Guide - the Anthropic model most teams will benchmark Astra against
Get the next comparison like this in your inbox
One email a week on News and the rest of the AI dev stack. Free.
Read next on AI coding tools
GPT-6 Sol vs Claude Opus 5.5 vs Grok 4.7: The September 2026 Price War
Three frontier launches in 48 hours repriced the agentic workhorse tier: Grok 4.7 at $2/$6 (Sep 21), Claude Opus 5.5 at $4/$20 with $0.20 cache reads (Sep 22), and GPT-6 Sol at $2/$10 with Luna at $0.10/$0.50 (Sep 22). Same-day-verified rates, honest benchmark attribution, and a decision guide.
9 min readOpenAI Says It Can't Rule Out Critical Cyber Capability for Astra, a First for the Preparedness Framework
On August 7 OpenAI disclosed that preliminary evaluations of its upcoming Astra model show strong enough agentic coding and cybersecurity performance that the company cannot rule out the Critical threshold under its Preparedness Framework. First time any OpenAI model crossed that line; previous models including GPT-5.6 Sol were assessed High. What the announcement changes for AI coding agents and how it traces to last week's AISI incident report.
7 min readHow GPT-6 Astra Builds a Real E-Commerce Site: Product Shots, Pages, and an AI Video Hero
GPT-6 Astra built a complete yoga clothing store from one prompt in the latest Developers Digest video - generated product photography, click-through product pages, and a custom AI video hero, wired through Codex and the Higgsfield CLI. Here is the verified workflow from the docs, plus how Genjutsu re-versions a single video for every market.
8 min readNew here? Start with
Technical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.








