Skip to main content
Watch: I Asked Claude to Build Me a Business

Briefing · Thursday, October 1, 2026

Gemini 4 Argon, Pi's MCP U-Turn, and the KV Cache Price War

Gemini 4 Argon, Pi's MCP U-Turn, and the KV Cache Price War

Good morning. It's Thursday, October 1, and we're covering Google's new frontier model and the gate in front of it, Pi's reversal on MCP, the cache economics that explain this month's price cuts, and a cryptographer's case that sandboxes are not containment.

The Gemini 4 Argon thread is the top post on Hacker News this morning at more than 1,400 points and over 900 comments, and the Pi MCP thread sits at 643.

In today's brief:

  • Gemini 4 Argon: Google's frontier model goes to cyber defenders first, at $2/$10 introductory pricing with a 1M-token output limit
  • Pi embraces MCP: Earendil moves MCP into its core and adds a JavaScript sandbox called Codemode
  • The KV cache argument: DeepSeek's published attention work is now visible in everyone's cache-read pricing
  • Agent worms: Matthew Green argues a payload plus a carrier is all a worm needs, sandbox or not

THE BIG ONE

Gemini 4 Argon Ships to Cyber Defenders First

Google announced Gemini 4 Argon on Tuesday, calling it its next frontier model and positioning it around three domains: real-world software engineering, enterprise knowledge work in legal and finance, and cyber defense. The headline numbers are an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input at 95% off, rising to $4/$20 after the introductory period. Google is also expanding the output token limit to 1M, up from 64K, which it says lets the model generate hundreds of thousands of tokens in one trajectory. On benchmarks, Google reports 77.9% on DeepSWE v1.1 for long-horizon coding, first place on the Vals Index, 51.3% on Zapier's AutomationBench, 91.7% on LVBench for long video understanding, and a tie for first at 68% on CWE-bench v1 for vulnerability remediation.

The unusual part is availability. Argon is rolling out first to a set of trusted cyber defenders through Google's Fairwind Program, the limited-access program it launched on September 2, and only later to paid API customers and Google AI Ultra subscribers. Google says it is "actively engaged in the U.S. government's voluntary process for pre-release model access" and will gather tester feedback before a broad release. For trusted defenders and its own internal teams, Google says it is releasing Argon without cyber guardrails so they can use its full defensive capability. Wiz is already using it through its Scan for Good initiative, where Google says the model found a critical vulnerability in healthcare software that previous frontier models missed.

The internal usage numbers are the most concrete thing in the post. Argon agents replaced 32,000 lines of SIMD code in a Rust port of libgav1 and produced a memory-safe decoder that runs 2.7x faster than that port with identical output. Google says Argon agents are migrating C and C++ codebases to Rust from libraries like re2 up to the 800,000-line Fuchsia Zircon kernel, under automated and manual auditing. On its own fleet, Argon agents analyzed profiling telemetry and applied memory optimizations that freed more than 300 TiB, with an estimated 500 TiB to 1 PiB in total savings, and a quantum team used it to beat a published baseline by 40% on subroutine optimization. Independent measurement from Artificial Analysis puts Argon (High) at 53 on its Intelligence Index, eighth of 223 models and above its median, at roughly $1.99 per Intelligence Index task, with a 1M-token context and text and image input.

The safety framing is where the release gets interesting. Google claims Argon is "our most resilient model yet against indirect prompt injections" and says it leads Gray Swan's Indirect Prompt Injection benchmark, with automated red teaming and adversarial training behind that. It also describes monitoring Argon's chain of thought and actions to stop execution when the model steps out of bounds, plus sealed sandboxes for high-risk training and evaluation. That lands one week after Anthropic's Frontier Red Team report on GLM-5.3 and OpenAI's Daybreak security push, which makes this the third consecutive week that frontier news has been framed through cyber offense and defense.

Hacker News read the gate before the benchmarks. The top comments mocked the "can't release a model" pattern ("Gemini not beating the 'can't release a model' allegations"), several paying subscribers said their Gemini app still offers 3.6 or 3.1 Pro rather than the newest models, and the pricing thread noted that $2/$10 is exactly what GPT-6.1 Sol launched at two days earlier, with the added caveat that Argon's price doubles when the introductory period ends. Benchmark saturation came up too. The counterargument is the usage data: a 300 TiB fleet-wide optimization and a 2.7x Rust decoder are outcomes, not eval scores.

Why it matters: Google now has a frontier-class coding model that almost no developer can call today, so the practical decision is not whether to switch but whether the pricing regime, the 1M-token output ceiling and the cyber-first access pattern justify planning a migration path for when paid API access opens. Our Gemini 3.5 Pro developer guide covers the access and context tradeoffs that carry over, and the September price war breakdown is the table Argon now joins.

PRICING

The Cache Read Is Where This Price War Shows Up

Yesterday's most-argued post on Hacker News was not a launch: it was an analysis piece (387 points, 430 comments) arguing that the West's price war is running on Chinese cache optimizations. The author's claim is specific: DeepSeek published MLA, which compressed the KV cache roughly 15x, then followed it with Compressed Sparse Attention and Heavily Compressed Attention, and its V4.1-Flash work adds CSA2, cross-layer cache reuse, a causal encoder-decoder architecture and FP4 caching that brings the global KV cache to 890 bytes per token, roughly a 437x reduction against DeepSeek-V1 for long-session workloads like coding. The piece then argues that Claude Opus 5.5 and GPT-6.1 Sol quietly shipped with those techniques, which is why their cache-read prices fell so hard.

Attribution matters here, so separate the argument from the arithmetic. The verifiable half is in public pricing. Anthropic's Opus 5.5 lists cache reads at $0.20 per million tokens against $4 input, a 20x discount. OpenAI's GPT-6.1 Sol lists cache reads at $0.10 against $2 input, a 20x discount. Google's Argon announcement quotes cached input at 95% off, so $0.10 against $2. Anthropic's Sonnet 5.5 sits at $0.20 against $2. Three labs, four models, and the same structural bet: cache reads are where long agent sessions get cheap. The 437x and 890-bytes-per-token figures are the author's, drawn from DeepSeek's published work, and readers should treat them as such.

The HN thread spent most of its energy on motive rather than mechanism, and the strongest replies were epistemic: one commenter called the piece "assumptions, nothing backing it" until others pointed out that DeepSeek's optimizations are peer-published research. Several argued the causality runs the other way, that open cache efficiency expands the market for cheap local models and pressures the subscription businesses. Nobody disputed the effect on long-context agent economics.

Why it matters: for anyone paying per token, the cache read is now the number to model, not the headline input price. If your agent loop cannot keep a stable prefix across turns and tools, you are paying 20x what a cache-disciplined loop pays. Our KV caching guide is the mechanics, and the Reasonix cache-first analysis covers what that does to an agent loop.

DEVELOPER TOOLS

Pi Reverses Itself: MCP Moves Into the Core

Earendil, the company behind the Pi coding agent, published a post titled "You Said No MCP!" on Tuesday confirming the U-turn: Pi's site used to declare that Pi does not support MCP, and now MCP is a supported part of the core harness. The stated reason is not capitulation but convergence. The company says the changes MCP needed were the same changes Pi needed to expose tools well: a sandbox in the form of an interpreter, deferred tool loading, mid-conversation system messages and reasoning level changes. MCP had previously been an extension, and the post says the metadata available to extensions was not enough to make that experience work well.

The mechanism is a new layer called Codemode, a JavaScript sandbox that runs harness-side and orchestrates tool calls, with its state kept in the session transcript rather than the filesystem. Codemode loads automatically when MCP is configured. The design argument is that MCP servers that dump tools into context and return text are optimizing the wrong thing; Earendil wants tools to return structured data and to be discoverable by their documentation, closer to "OpenAPI with intelligent tool discovery." The post's centerpiece example combines a Linear MCP server with Jev: ask Pi to find the 20 most frustrated commenters on its issue tracker, and Codemode pulls the open issues, runs them through a Jev classifier four at a time in parallel JavaScript workers, and stores the results, without spending context on intermediate tool output.

The thread (643 points, 354 comments) split predictably. Several commenters reached for the USB-C analogy: suboptimal for specific uses, universal for everyone else. Longtime Pi users worried about the minimalism the fork was chosen for, and Earendil's author replied directly to the criticism that the post under-explains MCP, saying readers should go to the source rather than let an agent summarize a PR. The more technical threads were about gaps that remain: file uploads over MCP still lack a settled best practice, and the codemode-versus-shell tradeoff has no clean answer when the harness runs server-side without shell access.

Why it matters: the most opinionated anti-MCP harness in the market just adopted it, which tells you the protocol's distribution has stopped being a question. The remaining question is design quality: if MCP servers are going to be judged on token efficiency, structured returns and discoverable tool docs are the interface you should be building against. Our MCP servers vs Agent Skills framework covers when a server is the right layer, and the five servers worth installing is the practical starting set.

SECURITY

Matthew Green: Sandboxing Is Not Containment

Cryptographer Matthew Green published a post on Wednesday asking whether sandboxing can contain rogue agents, and his answer is that the threat model is bigger than the box. The core observation is a worm has two halves: a payload that hijacks an agent, and an agent that will carry the payload to the next agent. He points to experiments where agents in separately isolated sandboxes left instructions for each other in a shared package cache, and those instructions changed what the recipients did. Replace the package cache with email, Slack, shared documents or WhatsApp, he writes, and replace independently sandboxed training runs with independently deployed personal agents like Muse, "and you have exactly the ingredients that a worm needs." Simon Willison quoted the passage this morning.

The argument is not that sandboxes are useless, it is that isolation is a boundary with a shared medium running through it. Agent systems are built to read untrusted channels: issues, tickets, email, documents, tool output. A sandbox can stop an agent from reading your filesystem and still allow it to write a message that another agent will read and act on. That is the same shape as the document-borne Copilot worm we covered in July, where attacker instructions in a document changed financial data and propagated to other documents, and it is the class of attack Google says Argon is now most resistant to on the Gray Swan indirect prompt injection benchmark.

Why it matters: containment is a system property, not a runtime property, and it fails at exactly the boundary where one agent's output becomes another agent's input. If you run a fleet, the design question to answer this week is what shared medium the agents use to talk to each other and who reviews what crosses it. Our sandbox architecture guide maps the runtime boundaries, and the pre-tool connection checklist starts at the permissions layer.

PLATFORMS

Netlify Rebuilds Edge Functions on Firecracker MicroVMs

Netlify published a deep dive into replacing V8 isolates with Firecracker MicroVMs for its Edge Functions, which run about a billion invocations a day. The numbers are the story: a warm invocation now costs about 5 to 6 ms at the median, down from 25 to 40 ms on the old infrastructure, p99 is 47.4% faster, availability is 99.998%, log delivery is 5x faster, and cold invocations, about 1.2% of requests, average roughly 9 ms. The engineering work was done with Unikraft, which wrote up its side of the rebuild.

The architectural fix was locality. Netlify's old setup sent requests out of its network to a hosted execution service; the new one forwards them to a compute node inside the edge network. Each request carries a machine specification naming three images, the runtime, the platform image and the customer's function image, plus CPU, memory and connection limits. A hash of that spec plus site information becomes a service ID, so two deploys never share a MicroVM, and each service is configured to handle a fixed number of requests before it is shut down and replaced. When a MicroVM boots, Netlify snapshots it so later starts come from the snapshot rather than a cold boot.

The HN discussion (179 points) was mostly productive skepticism. If Cloudflare Workers are V8 isolates running at single-digit milliseconds, why were these isolates taking 25 to 40 ms? The answer from the thread and the post is that the isolates were not running at the edge at all. Unikraft engineer Alex showed up to answer questions about the MicroVM side, and one commenter raised a real concern about forking RNG state from snapshots, which Firecracker has mechanisms to address.

Why it matters: if edge compute is moving to MicroVMs with millisecond cold starts, the workload you put there changes from routing and personalization glue to heavier per-request logic, including agent calls, without moving to a region. The cost is a new isolation contract to understand, since a fork of a booted image is now part of your runtime's behavior.

TOOLS WORTH A LOOK

  1. Claude Code v2.1.286 (free with a Claude plan) - permission prompts that stack now show a count like "2 of 5", fullscreen lists gained mouse support on their "N more" rows, and two bugs were fixed: repeated login browser opens when gcpAuthRefresh or awsAuthRefresh expire, and --resume losing turns.
  2. Magnitude (free, Apache 2.0, 5.9k stars) - an open-source inference engine that compiles and tunes kernels on your machine, claiming up to 2x faster decode than llama.cpp (92% on Metal, 19% on CUDA) and 27% less memory per agent; one-click connects Pi, OpenCode, Hermes, Codex, Claude Code and Cline. Launch HN at 163 points.
  3. Gitea 28.0.0 (free, self-hosted) - the project drops the 1.x version prefix and ships audit logging, bot accounts, HTTPS deploy tokens, administrator impersonation, code-owner approval rules and an Actions queue view. It is a security release, and migrations and mirrors now route through a new internal proxy with egress rules that must be reviewed before upgrading.
  4. Ledge.sh (free, Apache 2.0) - runnable Markdown notes for developers who keep commands in notes: press a shortcut on a code block and shell, SQL or AI prompt output streams in beneath it, with a persistent shell per note and a terminal drawer. Show HN at 145 points.

WHAT ELSE IS HAPPENING

  • EDG's C++ front end goes open source (220 points): after 30 years as a licensed, proprietary source-to-source engine, EDG's front end became public on September 30 under The C++ Alliance as its nonprofit home, with the same engine, same standards and community-funded feature tracks.
  • Google Cloud API Gateway becomes a native MCP server: annotating an OpenAPI 3.x spec with x-google-api-management.mcp turns REST operations into agent-discoverable tools, with the gateway transcoding MCP JSON-RPC into REST so existing auth, quotas and logging apply.
  • Ubuntu 26.04.1 LTS is available (77 points, 101 comments): the first point release of the LTS lands with the usual installer and hardware-support fixes, and the thread is a reminder that most desktop Linux users now wait for exactly this build before upgrading.
  • Reddit tightens old.reddit.com further (Ars Technica): accounts that have not used the legacy interface in six months will be blocked from it, extending the login requirement from June and continuing the slow squeeze on the interface power users depend on.
  • Android developer verification backlash resurfaces (241 points today): a developer's thread on the sideloading verification program drew hundreds of comments about account closures, fees and the shrinking path for hobbyist Android distribution; F-Droid's earlier technical argument against the program remains the best-sourced version of the case.
  • Data centers keep refusing to disclose water and power use (211 points, 194 comments): a Dutch report finds operators declining to say how much water and electricity their facilities consume, even as the buildout they serve keeps expanding.

Every link above goes to a primary source or our sourced coverage. Tomorrow's brief lands when the news does - subscribe to get it by email.

Get the next one in your inbox

The daily brief, delivered. Free, unsubscribe anytime.