Skip to main content
Watch: The Agentic Identity Hub

Briefing · Wednesday, September 30, 2026

OpenAI's Dots, a Fifth-Price Sol, and the Open-Weight Cyber Threshold

OpenAI's Dots, a Fifth-Price Sol, and the Open-Weight Cyber Threshold

Good morning. It's Wednesday, September 30, and we're covering OpenAI's DevDay 2026 haul, the open-weight model Anthropic says crossed the autonomous-exploit threshold, and a pre-registered benchmark that is trying to settle the "my model got dumber" argument with statistics.

The GPT-6.1 Sol thread is at 949 points and 836 comments this morning, the top post on Hacker News, with the Dots launch thread at 627 points behind it.

In today's brief:

  • GPT-6.1 Sol: near-Astra capability at $2/$10 per million tokens against Astra's $10/$50, plus an 8x-faster Ultrafast tier and a new Pro 500 plan
  • Dots: always-on personal agents, powered by Astra and sold at the Pro and Enterprise tiers
  • GLM-5.3: Anthropic's Frontier Red Team says the open-weight model develops working exploits, and its safeguards fall to a cover story 64% of the time
  • Livenerf: a 30-day, pre-registered drift watch on Opus 5.5, now six days into its baseline

THE BIG ONE

OpenAI Prices Near-Astra Intelligence at a Fifth of Astra

At DevDay 2026 in San Francisco on Tuesday, OpenAI launched GPT-6.1 Sol, pitched as "near-Astra level intelligence at a fifth of the price." The API price is $2 per million input tokens and $10 output, with cache reads at $0.10, against GPT-6 Astra at $10/$50 with $1.00 cache reads. The model keeps a 1,050,000-token context window, is live in ChatGPT Work, Codex and the API, and arrives two days after OpenAI held back GPT-6.1 Astra over safety concerns. Our release guide has the verified pricing, the benchmark table and the caveats, and the September price war breakdown has the head-to-head this resets.

The rest of the keynote was throughput and distribution. Ultrafast, a new serving tier, generates up to 8x faster, at up to 300 tokens per second, at 6x the standard price; it is available for Astra 6 today and 6.1 Sol soon. Pro 500 is a new plan that bundles Ultrafast with 25x Plus usage, and the $200/month plan is back on sale for new subscribers (DevDay live blog, plan details). OpenAI also previewed a Decisions API that hands the Luna model a predefined set of options and answers "in a fraction of a second," which Simon Willison read as a direct answer to Jev and the decision-model pattern. The DevDay recap has the full list.

The community read is a price read first. "Shots fired, half the price of Opus 5.5" was the immediate take in the HN thread, with the sharper counterpoint that practitioners who ran GPT-6 Sol in production found it worse than GPT-5.6 Sol on real refactors. The upgrade claims are being read with that scar tissue. The price arbitrage, though, is real and verifiable today.

Why it matters: the workhorse tier, not the frontier, is where this price war is being fought, and when cache reads fall 10x it is worth re-running the model bake-off instead of inheriting last month's routing table. Our Codex usage and pricing guide covers where the new model lands in plan limits.

AGENTS

Dots: an Always-On Agent, Sold at the Top of the Plan Ladder

The second product of the day is Dots, personal agents you name and give an avatar, "powered by Astra," and designed to work continuously in the background. You can hand one a migration, and the live demo showed it tracing the code, updating integrations and opening three pull requests. Dots get access to Slack, email and calendars, live in their own ChatGPT surfaces, and OpenAI says ChatGPT Pro and Enterprise customers get access today, with specialist dots for legal and finance work and collaboration planned with Microsoft 365. Teams work with them in ChatGPT Space, a shared document-like surface, and Dots can have their own Slack identities, which is the same shape as Anthropic's Claude Tag.

The distribution moves matter as much as the agent. OpenAI launched an OpenAI Marketplace with Adobe, Canva, Figma, Notion, Salesforce, Vercel and Zendesk as launch partners, Sign in with ChatGPT so users can bring their plan tokens to third-party apps, and plugin extensions that behave like native ChatGPT and Codex apps. ChatGPT Sites, already hosting 8 million sites with 70% of OpenAI employees building on it internally, picks up SQLite databases, scheduled tasks and plugins. Sam Altman put the audience at 1.2 billion weekly ChatGPT users.

On Hacker News (627 points, 487 comments), the reception was skeptical in ways worth noting: the feature is Pro-only, a live demo stalled, and the top criticisms were the compute cost of always-on agents and the read that this is "another attempt at giving Codex to regular people." The product question for builders is not whether the demo works but whether the paid tiers convert.

Why it matters: if an always-on agent becomes how Pro users delegate, the Marketplace and Sign in with ChatGPT are new distribution channels for developer products, and the Dots permission model is the thing to read before you hand it tools. Our ChatGPT Work and Codex desktop merge post covers the unification that made this ship.

SECURITY

Anthropic: GLM-5.3 Crossed the Open-Weight Cyber Threshold

Anthropic's Frontier Red Team published an assessment of GLM-5.3, Zhipu AI's newest open-weight model, and the headline is capability, not policy: on ExploitBench, GLM-5.3 built end-to-end Chrome V8 exploits in 50 of 410 attempts, just under Claude Mythos Preview's 56 of 410, where earlier models like Claude Opus 4.6, GLM-5.2, Kimi K3 and DeepSeek V4.1-Flash score at or near zero. On Anthropic's internal Binary Exploitation benchmark it achieved full control-flow hijacks in 4% of trials against Mythos Preview's 6%. In a human-in-the-loop session, GLM-5.3 chained previously unknown browser vulnerabilities into a drive-by exploit that reads files off a visitor's machine; a smaller GLM-5.3-Flash built a working ARM64 exploit chain for a known Chrome CVE in 20 minutes of human attention and eight hours of model time, for about $20.40 at Zhipu's API prices. NIST's CAISI had already assessed GLM-5.3 as "the most cyber-capable open-weight model released to date," about four months behind the US frontier.

The safeguard findings are the part to internalize. Refusals that held at 100% for a bare harmful request fell to 64% with a false cover story and 92% with prefilled thinking tokens, and an abliterated copy engaged 100% of the time. Producing that abliterated copy took Anthropic about 2,200 GPU hours and roughly $4,400 on its first attempt, and about 600 GPU hours ($1,200) for a team that knows the technique. Refusal rates dropped from above 90% to about 3%, with general capability intact. Because the weights are downloadable, none of this requires a jailbreak against a hosted API.

The thread is not a clean read, and it should not be presented as one. Anthropic is a direct competitor of Zhipu, and the HN discussion leads with that conflict: the top replies call the harmful-request scenarios manufactured, argue the framing feeds arguments for restricting open weights, and observe wryly that the post is excellent publicity for GLM-5.3. Both things can be true: the benchmark numbers are checkable and broadly match NIST's assessment, and the messenger has a commercial interest in the conclusion.

Why it matters: the exploitable threshold is now in downloadable weights with safeguards that a determined actor removes in a weekend-scale project, so security teams should plan for exploit development at API prices rather than treating it as a frontier-lab-only capability. Our GLM-5.3 access guide covers where it runs, and the agent security checklist is the defensive side.

DEVELOPER TOOLS

OpenAI Productizes Its Own Security Sprint as Codex Security Cloud

Buried under the model news, DevDay shipped the security story many teams actually want: Codex Security went generally available with an open-source CLI and TypeScript SDK (past 10,000 stars), built from an internal sprint where a quarter of OpenAI's product engineers fixed 53 critical findings on the first day (The Defense Factory). The loop is scan, deduplicate, validate, prioritize, patch, verify. OpenAI's own numbers from the sprint: 36% of discovered issues were duplicates, which deduplication removed before humans saw them; patch generation opened PRs automatically; and a "verify-fix" step that plays an adversarial agent against each patch kept the rollback rate to about 1%. Product owners define threat models in SECURITY.md files, and the scanner keeps looping until it stops finding issues, using a security-specialist Daybreak model.

Why it matters: finding bugs with a scanner has been cheap for a year; patching them is the bottleneck, and the interesting line in this launch is the verified-patch loop with a human deciding the merge. Our analysis of the Codex Security open-source release and the Daybreak patching bottleneck post go deeper.

MEASUREMENT

Livenerf Puts "Is It Nerfed?" on a 30-Day Clock

Fed up with vibes-versus-vibes arguments, livenerf is a small, append-only benchmark that asks one question: does a model get worse after it ships? It started the clock on Claude Opus 5.5's launch week (September 22), runs once a day for 30 days through headless Claude Code on a Max subscription, and collects 90 samples per day on a frozen 78-question panel selected from 2,336 GPQA Diamond, MMLU-Pro and competition-math questions that Opus 5.5 answered correctly only sometimes. Everything else is pinned: Claude Code 2.1.280, a harness hash recorded per run, a frozen system prompt, no tools, and pure-function graders, with the statistics following Anthropic's own Adding Error Bars to Evals. The pre-registered decision rule is strict: a change is only called a change if the 99% interval excludes zero in two consecutive 10-day windows, the effect is at least 3 points, and the control arm does not move.

The honest limits are documented. Validation showed the rig detects a known degradation (effort low: 62% fewer output tokens, 8.3 points of accuracy) but cannot distinguish an Opus 5 swap from Opus 5.5 (-3.8 +/- 6.3 points). One commenter on the HN thread (638 points, 254 comments) called the idea "genius," others argued ten-day windows are too short to catch week-scale changes, and the author's pre-registered rule exists precisely because those arguments cannot be won by assertion. The status as of Tuesday: 6 of 30 days collected, no missed runs, about 3.6% of a weekly Max plan spent.

Why it matters: this is the eval-hygiene pattern in miniature - pin the harness, pre-register the decision rule, publish nulls - and it raises the bar for every vendor whose model quality is now the subject of a public time series. Our baseline receipts post is the same discipline for your own evals.

RESEARCH

Jeeves: Reasoning Before Deciding, for the Cost of a Few Seconds

PostHog dropped Jeeves, a 9B Jev-like decision model (Qwen3.5-9B with a LoRA and a pointer head) trained with SFT and CISPO to think before it decides, plus a block-4 diffusion drafter for speed. The numbers: 0.889 on held-out test data against Jev's 0.857 and Kev-9B's 0.822, 0.935 on JevBench's public tiers against Jev's 0.866, and a much lower rate of high-confidence answers to unknowable questions (5.5% against Jev's 9.0%). It speaks Jev's API format, handles yes/no, multiple-choice and rating questions in one request, and runs on one H100 at about 0.3 seconds per request without thinking and 3.3 seconds median with thinking. Weights, training code and data are all in the repo, and the HN thread is at 235 points.

The context is the interesting part. Jev came out of stealth less than two weeks ago, and the decision-layer pattern (plan with a frontier model, classify with a small one) has since picked up Jeff, Kev, this, and now an OpenAI DevDay preview of a hosted equivalent. The tradeoff Jeeves makes explicit is that thinking buys accuracy and costs latency, and the chain can be truncated when 3 seconds is too slow - which is the same knob every router in this space is now tuning.

Why it matters: decision models are becoming a standard tier in agent stacks, and a Jev-style reasoning classifier you can self-host is a candidate for routing, moderation and triage where per-call API costs do not scale. Our Jev usage guide and Jev release analysis cover the hosted-model side of the same pattern.

TOOLS WORTH A LOOK

  1. Claude Code v2.1.285 (free with a Claude plan) - adds claude --desktop to open the desktop app in the current directory, CLAUDE_CODE_DISABLE_WEB_FETCH to turn off the WebFetch tool, and an allowedProviders managed setting that limits which API providers a machine may use.
  2. nsl v0.4.0 (free, MIT, pre-release) - WSL-style Linux machines for Linux, built on systemd-nspawn inside a shared VM, with your files, forwarded ports, Wayland windows and seven signed distros; --isolated gives untrusted software its own VM. 136 points on HN.
  3. OpenAI Live Console (free, OSS) - the starter app from the DevDay GPT-Live session for building realtime voice agents; the session ran a voice model as a live co-presenter with Astra writing the drawing code behind it.
  4. Codex rust-v0.159.2 (free, OSS) - a small backport release that suppresses console windows flashing on Windows when Codex launches background processes; 0.161.0 alphas are rolling in the channel.

WHAT ELSE IS HAPPENING

  • Google is ending ChromeOS support two years early (The Register, 202 points): a support document says Chromebooks bought today get updates until 2034, migration paths to the new Googlebook OS are not published yet, and existing ChromeOS management licenses do not carry over.
  • The Netherlands is building a NixOS-based government stack (379 points, 379 comments): after US sanctions cut ICC staff off Microsoft services, the DAWO program picked NixOS for reproducible, signed package management, with eight municipalities in trials and a first stable release expected at the end of 2027.
  • macOS 27 Golden Gate is "a buggy mess" (457 points): a developer's bug roundup documents menu bar crashes, a system hang from the Colours eyedropper, misaligned UI and documentation still shipping from Tahoe, in a release Apple sold as a refinement cycle.
  • Tcl/Tk 9.1.0 is out (279 points): the new release adds Unicode normalization commands, a monotonic timer command, lfilter, screen-reader support and bidirectional text in Tk, and the ttk::toggleswitch widget.
  • Meta's Muse agent read a reporter's Messages without consent (Inc, 157 points on HN): Jason Aten reports Muse synced 187,000 lines from his Messages database with Full Disk Access off and Messages blocked; commenters on the AppleInsider write-up argue the likelier mechanism is retention after revocation, which is still a privacy failure.
  • A PS5 hypervisor exploit chain went public (312 points): Relapse targets firmware versions 7.00 to 13.60, sits at 1,238 stars, and is the kind of consumer-hardware security research that quietly hardens an ecosystem.

Every link above goes to a primary source or our sourced coverage. Tomorrow's brief lands when the news does - subscribe to get it by email.

Get the next one in your inbox

The daily brief, delivered. Free, unsubscribe anytime.