Build Interactive 3D Worlds With GPT-6 & Blender
Briefing · Wednesday, September 23, 2026

Good morning. It's Wednesday, September 23, and we're covering the most aggressive front-page day in months: Anthropic launched Claude Opus 5.5, and about an hour and a half later OpenAI answered with GPT-6 Sol and Luna. Between them, they turned the one-line ledger of frontier pricing upside down.
The Claude Opus 5.5 thread sits at 1,558 points with 963 comments, and the GPT-6 Sol and Luna thread at 1,546 with 741. Simon Willison was already live-pricing both within hours. Here is the signal, sourced.
In today's brief:
THE BIG ONE
Anthropic announced Claude Opus 5.5 as the first model in its new Claude 5.5 family, and positioned it as the first release of the "pace the frontier" era, tested before launch by external evaluators including METR and Frontier Design. The headline is a gap collapse: it performs at the level of Claude Fable 5.1 on most work while costing 40% less to run than Opus 5. On the agentic-coding leaderboards it sets new highs - Terminal-Bench 4.0 at 66.4% against Fable 5.1's 55.8% and Opus 5's 52.3% - and on GDPval-AA v2.1 knowledge work it scores 1846 Elo, ahead of Fable 5.1's 1735.
The efficiency numbers are doing the heavy lifting on cost-per-task. Per-million-token pricing drops 20% to $4 input and $20 output, but the more consequential cut is cache reads, down 60% to $0.20 - the line that dominates agentic and coding workloads where over 90% of input tokens are typically cached. Output is more than 30% faster than Opus 5, and the system card claims the best scores of any Claude on the automated behavioral audit, with prompt-injection resistance that matches or beats Opus 5 in every setting tested. Early-tester stories lean into long, unattended runs: one completed a 680,000-line code migration in under a day; another ran a six-repo engineering task overnight for over 18 hours without going off-task.
Simon Willison ran it through his pelican test and flagged the day's one real caveat: at "max" effort, Opus 5.5 over-thinks to the point of failure, hitting its 128,000-token output ceiling while still reasoning about an SVG and refusing to return - twice, at roughly $2.56 and nearly 20 minutes each. He calls max "effectively useless" and recommends the default effort levels, and he has already made Opus 5.5 his default in Claude Code. Anthropic says Sonnet 5.5 and Haiku 5.5 follow in the coming weeks.
Why it matters: the most expensive token tier just got 20-60% cheaper while its capability gap to the flagship narrowed, which reshapes the cost-per-task math for every agent deployment - the same ledger our Fable 5 cost-per-task analysis tracks against the newer pricing comparison.
PLATFORMS
About an hour and a half after Anthropic's post, OpenAI introduced GPT-6 Sol and GPT-6 Luna, and the pricing is the story: both land at half the price of their GPT-5.6 equivalents. Sol, the coding-and-agentic flagship, goes from GPT-5.6 Sol's $4/$20 to $2 input and $10 output per million tokens with cached input at $0.20; Luna, the focused high-volume workhorse, drops from $0.20/$1.20 to $0.10 input and $0.50 output with cached input at $0.01. Both share a 1,050,000-token context window with 128,000 max output tokens (Sol, Luna), and both are live in the API today under gpt-6-sol and gpt-6-luna.
Willison frames Luna as one of the cheapest models OpenAI has ever released - beaten only by the far weaker GPT-4.1 Nano and GPT-5 Nano - and notes the timing pressure behind the drop: GPT-5.6 has a scheduled 25% price increase for November, so GPT-6 is priced against the promotional rates. The result is a shelf where Grok 4.7's aggressive $2/$6 from two days ago now reads as equals on Sol's input price, and where a capable Luna at $0.10/$0.50 becomes a plausible default for high-volume extraction and classification. Willison is already running GPT-6 Sol as his default in Codex and switched his Datasette agent demo over to GPT-6 Luna.
The HN thread is busy arguing whether the price is suppression of open-weights competition or just cost engineering on a more efficient model (742 comments), and early benchmark comparisons against Gemini 3.8 Flash and the MiMo v2.6 line are doing real work in the comments - one poster goes so far as to call Luna "insane value from a closed source model."
Why it matters: when both frontier labs cut the two tiers developers actually bill by half inside the same hour, rotation logic and per-task economics change everywhere agents run - the exact trade we broke down for model routing and in the fleet-economics post.
RESEARCH
The most striking research note of the day comes from cryptology, not a lab blog: Frode Weierud of Crypto Cellar Research authenticated that OpenAI's GPT-6 Astra broke the German Army Enigma message MVUEH of 10 July 1941, which had resisted all attempts since 2005 (387 comments). On 15 September, Carter Leffer asked the model to try the unbroken Enigma messages on the site. Astra picked MVUEH as most promising, suspected a link to the adjacent message SIPVX, settled on the repeated place name ROSENOW as a crib, and - after writing its own Enigma simulator and Bombe in Python and C++, plus Python for the traffic analysis - produced the correct key and plaintext.
Weierud's notes on how it got there are the point. The MVUEH key turns out to differ at the wheel order itself (253 instead of 512), the left-hand rotor turns over at the 72nd letter - a rare event known to complicate breaks - and the original ciphertext transcription contained errors. The model also behaved "like a very professional cryptanalyst and archive researcher": it traced the received corpus and independently cited concrete Bundesarchiv radio-message volumes (RS 3-3/20a and RS 3-3/63b) that match the archive's own catalogue. "What it has achieved in two days," Weierud writes, "would take a human researcher weeks or even months."
It's a hard-to-fake result - ciphertext and plaintext either fit the Enigma wiring or they don't - and it follows the pattern of agents doing long-horizon autonomous work: pick a target, form a hypothesis, build the tooling, iterate until verified. The breadth is where it gets uncomfortable, because the same autonomy applied to cryptanalysis is the frontier capability that has Anthropic and OpenAI both gating cyber work behind fallback models, and Astra's own family members ship with safeguards modeled on Fable 5.1.
Why it matters: a frontier model turned a 21-year-old open research puzzle into a two-day autonomous job, which is measured progress on the long-horizon agent claims - and another data point on the growing asymmetry between model capability and verification tooling we keep coming back to.
PLATFORMS
Nathan Lambert turned his prepared testimony for a Congressional briefing on open-weight models into the definitive state-of-the-union read on the open-model landscape (91 points on HN), and the numbers settle the "are open weights competitive" question in one paragraph. Chinese open-weight models are the clear leaders, and have been since about July 2025: Alibaba's Qwen alone is mentioned in roughly 30% of recent AI research papers against Llama's ~21%, and cumulative Hugging Face downloads put China at a lead of about 1.6 billion. On the Artificial Analysis Intelligence Index, the top model from an American lab sits in the teens, behind 15 Chinese models.
The capability gap to the closed frontier is now measured in months, not years: Lambert estimates Chinese open weights sit 2-5 months behind the American closed frontier, while American open weights trail OpenAI and Anthropic by more like 6-9 months. His estimate that fully preventing distillation would only widen the Chinese gap by 1-2 months is the counter to the "everything is a DeepSeek clone" take. Since September 2025, OpenRouter usage of open models grew from roughly 1 trillion tokens a week to about 80 trillion, with Chinese models taking over 80% of it - and the coding-agent lane is even more concentrated, with open models carrying the bulk of inference.
The framing for developers is practical rather than patriotic: open weights have crossed the viability threshold in exactly the agentic-coding lane this brief covers daily, from MiMo v2.6's MIT-licensed 1T flagship to the open-weight Jev remakes and the pricing pressure they put on every API shelf. Lambert's bottom line - "open models are going to be the substrate for everyone else in the world outside of the few true frontier AI labs" - is the strategic read under today's price war.
Why it matters: when a $0.10/$0.50 closed model meets open weights that are months behind instead of years, the default for high-volume agentic work is about to get a lot more contested across every budget tier - our open-weights economics analysis and local coding LLM roundup are the decision frames.
SECURITY
WordPress published a security advisory describing an unauthenticated path traversal that can lead to conditional remote code execution (193 points on HN, 100 comments). It is the kind of finding that draws immediate attention given the platform's installed base: an unauthenticated traversal chaining toward code execution is exactly the class of bug that turns a patch backlog into a front-page incident. As with most conditional RCEs, applying the advisory's fix plus updating plugins and themes is the practical resolution, and sites running older versions should treat it as a patch-now signal.
The security top of the thread is rounded out by Trail of Bits' SAML: A fractal of bad design (245 points), a well-argued case for why the protocol keeps breaking despite decades of audits, and by 404 Media reporting that the ShinyHunters group claims a breach of FBI-related services with data on employees and applicants (612 points), including a verified sample of 5,000 records.
Why it matters: the WordPress finding is the practical one for the fleet - an unauthenticated path toward RCE in the most deployed CMS in the world is a patch-now signal, and the SAML essay is the durable read on why SSO keeps leaking.
TOOLS WORTH A LOOK
WHAT ELSE IS HAPPENING
FROM THE SITE
Agent-Native Apps Need Shared Actions, Not UI Puppeteering - BuilderIO's Agent-Native framework, the top trending repo across all languages, and the case for one action layer shared by UI, agent, HTTP, MCP, A2A, and CLI. The architecture frame for the "agent-native app" claim that today's model releases keep making stronger.
Every link above goes to a primary source or our sourced coverage. Tomorrow's brief lands when the news does - subscribe to get it by email.
The daily brief, delivered. Free, unsubscribe anytime.