Briefing · Tuesday, September 29, 2026
Sonnet 5.5's Coding Leap, Nvidia's Agent Watchdog, and 22ms Models

Good morning. It's Tuesday, September 29, and we're covering Anthropic's Sonnet 5.5, which jumped from 10.3% to 70.6% on Terminal-Bench 4.0 at unchanged prices; Nvidia's open agent-safety platform, which puts a watchdog next to every agent; and Jeff, a 0.8B decision model trained at home that answers in 22 milliseconds.
The Sonnet 5.5 thread hit 793 points and 531 comments overnight and is still the top post on Hacker News this morning.
In today's brief:
- Sonnet 5.5: the mid-tier model now tracks Opus 5.5 on agentic coding benchmarks at a fraction of the cost per task, and it is the new free-tier model on claude.ai
- Nvidia's Open Agent Safety Platform: OpenShell sets a runtime boundary outside the model and harness, and Sentry, a watchdog running on BlueField-4 DPUs, can quarantine an agent in milliseconds
- Jeff: fine-tunes of Qwen3.5 and Gemma 4 that return calibrated probabilities in 22ms on one RTX PRO 6000, trained with no cloud GPUs
- Cloudflare's cf: a new CLI generated from OpenAPI schemas that covers 3,000+ operations, with JSON as the default and a TypeScript config agents can type-check
THE BIG ONE
Sonnet 5.5 Jumps from 10.3% to 70.6% on Terminal-Bench at the Same Price
Anthropic on Monday released Claude Sonnet 5.5, the second model in the 5.5 family, and the numbers are a different shape than the usual mid-tier refresh. On Terminal-Bench 4.0, an agentic terminal coding evaluation, Sonnet 5.5 scores 70.6% where Sonnet 5 scored 10.3%, and above Opus 5.5's 66.4%. It is priced identically to Sonnet 5 at $2 per million input tokens, $10 per million output tokens, and $0.20 per million cache reads, and Anthropic says it generates output 30%+ faster and costs up to 30% less per task. On knowledge work it lands closer still: 1844 on GDPval-AA against Opus 5.5's 1846, and 80.1% on OSWorld against 81.8%. Anthropic's own framing is that Opus 5.5 remains stronger at complex, open-ended work. The benchmark story is that the cheap tier is now within a couple of points of the flagship released a week earlier.
The effort ladder is the practical part: at Low or Medium effort, Sonnet 5.5 beats Sonnet 5's best score for about a tenth of the cost per task (Medium is the default in the Claude apps, High on the Claude Platform). Per the Claude Code changelog, it ships with 1M-token context and is now the default Sonnet. It is also the first Sonnet to beat Pokémon Red from screenshots and the first to launch with cyber safeguards and fallbacks like Opus 5.5's, plus classifiers that block distillation-style reasoning extraction. Haiku 5.5 arrives "in the coming weeks."
Two community reads are worth your time. Simon Willison ran his pelican test and surfaced the same failure mode as Opus 5.5: at Max effort, the model spent 128,000 tokens, about $1.28, and failed to produce an SVG; at xhigh effort it produced a good one for 5.74 cents in 41 seconds. He flags the more consequential change: Sonnet 5.5 is now the free-tier model on claude.ai, where ChatGPT's free tier runs Luna 5.6, so Anthropic is fielding the stronger free offering. In the HN thread, developers noted that Sonnet 5.5 nearly matches Opus 5.5 on CursorBench (55.5% vs 57.8%) and FrontierCode at xhigh (52.1% vs 54.4%), while the loudest complaints were about the new cyber safeguards and what the Cyber Verification Program does and does not cover.
Why it matters: a roughly tenfold cost-per-task cut at equal or better quality resets the router - if your agent reaches for an Opus-class model on routine fixes, this release is the argument to re-run the model bake-off rather than inherit last month's choice. Our Sonnet 5.5 developer guide, Sonnet 5 migration notes, and Opus 5.5 release guide have the API and effort-level details.
SECURITY
Nvidia Puts a Watchdog Next to Every Agent, and It Runs on a DPU
Nvidia on Monday announced the Open Agent Safety Platform, an open software and reference-hardware design for controlling agents from testing through deployment. It has two halves. OpenShell, now broadly available and open source, is a secure runtime boundary for agents that sits outside the model and the agent harness, runs on Nvidia's Vera CPU built for agentic workloads, and can be extended to Arm and Intel compute. Sentry is the reference design's out-of-band watchdog: it runs on BlueField-4 DPUs and monitors agent behavior in silicon, and if an agent tries to move outside its boundary, Sentry quarantines and stops it in milliseconds. The platform is built on DOCA and offers attested telemetry, agent identity verification, and granular zero-trust policies for data, tools, APIs and services, with skills and docs on GitHub.
The press release states the pattern behind it plainly: recent incidents share one shape, an agent circumventing controls at the application layer to complete its assigned task. The partner list is the tell for how seriously the industry is treating that: Anthropic (Claude Managed Agents runs the agent loop in a separate server from the sandboxes where work executes, and integrates with OpenShell and BlueField), CrowdStrike, Hugging Face, JPMorganChase, Microsoft, Red Hat, Salesforce (which wired OpenShell approvals into Slack), SAP, Scale AI, ServiceNow and SpaceXAI are among more than 100 organizations building on it. Canonical, SUSE and Red Hat are integrating it into operating systems, Figure and Skild AI are embedding OpenShell in robots, and Citi and JPMorganChase are collaborating on shared agent-safety tooling. Jensen Huang's line: "Safety and security require full-stack engineering."
Why it matters: app-layer guardrails can be worked around by the very agent they are meant to contain, while a boundary enforced in hardware the agent cannot see or reach is a different security model. Nvidia just made that a reference design, which is the world our containment capability ledger, agent security checklist, and Vera agent-safety testing post describe from the software side.
RESEARCH
Jeff: 0.8B Decision Models That Decide in 22 Milliseconds
Jeff is a set of fine-tunes of Qwen3.5 and Gemma 4 for zero-shot classification, released by firelex with the same request format as Jev: describe a situation, list the options in plain words, and the model returns a calibrated probability for each option from a single forward pass, with no generated text and no parsing. The speed is the headline: about 22ms per decision on an RTX PRO 6000, 28ms on an Apple M4 Max through MLX, and 463ms on a 32-thread CPU. The 0.8B model scores 79.1 overall across five public benchmarks (4,599 questions) against Jev's published 83.0, and the 2B scores 83.1. Jeff wins on classification-heavy sets and loses badly on reasoning-heavy BBH (64.0 vs 94.3), the size tradeoff stated plainly. Everything was trained on one RTX PRO 6000 (about two hours for the 0.8B, 3.5 for the 2B), with synthetic data written by an open model on two DGX Sparks and no cloud GPUs; the repo builds on the open AutoJev recipe and notes it is "not affiliated with or endorsed by" TypeSafe.
The most useful parts are the caveats and the fine-tune path. The released checkpoints cap out at 26 options per question, because they never learned the two-letter codes beyond Z; a short fine-tune is the escape hatch, and the voice-navigation example moved held-out accuracy from 31.7% to 95.8% in about half an hour on one GPU, at roughly 40ms per decision on an M4 Max. The zero-shot game tests are the fun demonstration: the 0.8B matches a hand-coded rule bot at Doom (6.55 kills), nearly matches it at Frogger (10.3 vs 10.25 crossings), and trails it at Pac-Man (57.0 of 98 pellets vs 94.1). The repo sits at 770 stars and the HN thread at 498 points and 190 comments.
Why it matters: a decision layer that runs locally in 22ms and costs nothing per call changes the shape of an agent - plan with a frontier model, decide with a classifier, and keep routing and moderation logic on your own hardware. Our Jev guide and Jev release analysis cover the same pattern from the hosted-model side.
PLATFORMS
Cloudflare's cf CLI Gives Agents All 3,000 API Operations
Cloudflare released cf in open beta, a new command-line tool that mirrors the entire Cloudflare API and is explicitly designed for agents. The usage numbers are the argument: agents were responsible for a quarter of Wrangler use in March 2026 and 48% last week, and they run almost twice as many distinct commands per day. Wrangler accumulated commands for around 280 operations; cf covers more than 3,000, generated directly from the OpenAPI schemas that power Cloudflare's docs and SDKs through an internal pipeline called Forge, which the company open-sourced.
The agent-first decisions are the interesting part. JSON is the default output rather than a flag, pretty-printed for humans and condensed for agents, on the theory that agents are now the primary user; cf cli search lets an agent ask in natural language which command it needs; and configuration moves to cloudflare.config.ts, a typed file any LSP-enabled agent can edit accurately, with helpers for bindings and triggers and a migration command. Vite is the default dev server. Cloudflare is also blunt about the endgame: when the beta ends, a final major version of Wrangler will point people at cf, and Wrangler gets maintenance support for 18 months after that. The HN thread (153 points, 76 comments) includes the fair counterpoint worth thinking about - whether command discovery aimed at machines is a feature or a new prompt-injection surface.
Why it matters: tools designed for agents look different from tools designed for humans - machine-readable by default, self-describing, searchable in natural language - and Cloudflare has the usage data to justify rebuilding its CLI around that shift. Our Cloudflare agent lifecycle guide, TypeScript CI/CD workflows post, and agentic internet overview track it.
BUSINESS
World Labs Is Joining AMD, and Fei-Fei Li Becomes AMD's Chief Scientist
World Labs, the spatial-intelligence lab Fei-Fei Li founded in 2024, signed a definitive agreement to join AMD. Li will become an Executive Vice President and Chief Scientist at AMD, working directly with CEO Lisa Su; co-founders Justin Johnson and Ben Mildenhall continue leading the team, which will form a frontier research organization inside AMD. The relationship started last year with training and inference work on AMD GPUs, and the post frames the combination as a natural fit for an end-to-end open AI ecosystem spanning hardware, software, platforms and open models. The transaction is expected to close by the end of 2026, subject to regulatory approvals. Li wrote her own account, and the HN thread is at 280 points and 111 comments.
Why it matters: where frontier labs land decides where open-model training capacity comes from, and AMD bought a research organization and a marquee leader at once - background in our AMD MI355X vs B200 vs B300 serving comparison.
ENGINEERING
"Coding is NOT Solved," and the Backlash Found Its Essay
Alex Ewerlöf's Coding is NOT solved reached spot 2 on Hacker News (496 points, 489 comments) with an argument that is getting harder to file under nostalgia: creation is now cheap, but maintenance, reliability, security and scalability are the majority of software cost, and none of those are solved. His mechanism claim is precise: LLMs appear to write code well because we wrapped them in feedback loops that feed compiler and runtime errors back until most failures are resolved or hidden, which is a property of the harness, not of reasoning. His list of software that does not require reading the code has four items (personal software, proofs of concept, throwaway automation, weaponized AI), and every low-risk-tolerance domain that hires engineers - healthcare, finance, aviation, power plants - is not on it.
The essay is blunt where it matters. He calls AI output "Nordic Gold": cheap, technically advanced, realistic enough that people who do not care about real gold cannot tell. He names an "AI Dunning-Kruger effect": the less someone knows about a task's edge cases, the more they trust the output. And he lands on accountability as the thing that cannot be delegated, because an AI cannot be fired, fined or imprisoned, so a human owns every shipped line regardless of how it was produced. On economics he is more nuanced than the title suggests: if AI code is 2x worse but 1000x faster and 100x cheaper, plenty of tasks still justify it, which is why SaaS increasingly sells SLAs and guarantees rather than code. The post includes an exchange with DHH over "pencils down on hand-written code" and cites Shopify CEO Tobi Lutke's term for the result, "slop grenades."
Why it matters: the value is migrating from generation to verification, and this is the clearest statement yet of where the human hours go: review, tests, ground truth and accountability. Our AI code review bottleneck, approve effects not invocations and agent PR governance posts are the practical version of the same argument.
TOOLS WORTH A LOOK
- Codex 0.159.0 (free, open source) - adds opt-in
instant_interruptso new input steers a response mid-flight, richer Mermaid flowchart rendering, and keeps filesystem denials attached to approved commands;.awsdirectories are now protected by default under writable roots. - Claude Code v2.1.284 (free with a Claude plan) - ships Sonnet 5.5 as the default Sonnet model with 1M context, shows gateway spend in dollars in
/usageand the status line, and adds/mcp reconnect allfor retrying every failed MCP server at once. - MicroLLM Lab (free, browser) - run seven tiny language models locally in a browser tab with no install; 246 points on HN.
- ESP32S3 BitNet cluster (OSS) - a cluster of ESP32-S3 boards serving a 1.58-bit quantized language model, a weekend-scale demonstration that inference can run on microcontroller hardware; 102 points on HN.
WHAT ELSE IS HAPPENING
- Anthropic's IPO prospectus shows a near-$42B net loss (Reuters, 106 points): the loss includes a roughly $34B accounting charge from revaluing financing that could convert into shares rather than operating spend, and risk factors note nearly a quarter of last year's revenue came from two customers not locked into long-term contracts.
- Firebase's iOS SDK crashed production apps for hours (519 comments): a crash in the Analytics SDK started at 00:41 UTC today when a response from the sdk-exp endpoint hit a nil-key path, affecting apps that had not shipped a new build; the issue title is now marked Resolved.
- MongoDB CEO Desai is leaving to lead Meta's enterprise platform (Reuters, 348 points): the database company's chief executive joins Meta's push to sell AI and enterprise software.
- OpenAI published a framework for safety cases in frontier training: the post argues developers should write explicit safety arguments before large training runs, arriving days after its own disclosures about agent behavior during training (our incident-report analysis covers the context).
- Cal Newport: "It's Time to Investigate the AI Labs" (506 points, 206 comments): the essay argues the labs' own accounts of their systems are the reason they need outside scrutiny.
Every link above goes to a primary source or our sourced coverage. Tomorrow's brief lands when the news does - subscribe to get it by email.
Get the next one in your inbox
The daily brief, delivered. Free, unsubscribe anytime.