Skip to main content
Watch: I Asked Claude to Build Me a Business

Briefing · Friday, October 2, 2026

Pi 1.0, Cloudflare Opens Clef, and Git 3.0's SHA-256 Bill

Pi 1.0, Cloudflare Opens Clef, and Git 3.0's SHA-256 Bill

Good morning. It's Friday, October 2, and we're covering Pi's 1.0 release and the durable harness built beside it, Cloudflare's open decision models and the RL service attached, DeepSeek's new desktop harness, and the case that Git 3.0's hash upgrade is a costly mistake.

The Pi 1.0 thread is the top post on Hacker News this morning at 1,297 points and 413 comments, and Pi Durable adds another 385 points of its own.

In today's brief:

  • Pi 1.0: Earendil's minimal coding agent ships native MCP through Codemode, deferred tool loading, and Anthropic cache warming, alongside a new experimental framework called Pi Durable
  • Clef: Cloudflare open-sources two Jev-compatible decision models on Workers AI at $0.24 and $0.09 per million input tokens, with an RL fine-tuning service
  • DeepSeek Harness: an MIT-licensed desktop harness where every capability, including the agent loop itself, is a plugin
  • Git 3.0's SHA-256 default: the GitHub co-founder's argument that two hash formats will bifurcate every forge, submodule, and tool

THE BIG ONE

Pi 1.0 Arrives, With a Durable Harness Beside It

Earendil shipped Pi 1.0 on Thursday, calling it "a hardened, minimal, extensible agent harness that you can make your own." The release folds in the changes the team says survived months of testing: Codemode, which makes MCP and non-LLM models like Jev and image models native to the harness, extension support for virtual models, deferred tool loading, cache warming for Anthropic models, mid-conversation system messages, a new TUI theme, and fullscreen mode by default. Pi is MIT licensed, and Earendil says hundreds of thousands of people use it every week.

The more interesting ship is Pi Durable, an experimental package for long-running agents that does not replace the coding agent. It opens a harness over a storage backend (memory, SQLite, or JSONL, with SQLite and JSONL using no Node APIs so they can run on Bun or inside a Cloudflare Durable Object), stores a checkpoint before every step, and resumes interrupted tasks when a new process opens the same storage. A requestId makes submissions exactly-once, conversations can fork at any transcript position, and one harness runs many conversations concurrently.

The design treats everything as durable state. Tools, hooks, compaction, and application documents all ride the same task and commit machinery, and any number of clients can attach to a conversation, watch it stream, and steer it. Earendil says the source is about 15,000 lines without tests, small enough that an agent can read the whole harness. The HN discussion split between users who want the minimalism preserved and users who want the new capabilities, with several noting that this week's MCP reversal is now simply part of the core.

Why it matters: durable execution, crash recovery, and multi-client steering are becoming table-stakes properties of agent harnesses, and Earendil shipped them as a separate framework rather than compromising the minimal product. If you build on an agent runtime, the questions to ask are the ones Pi Durable answers: what survives a crash, and what does a fork share. Our architecture deep dive and run modes guide cover the Pi lineage, and the fx vs pi comparison maps the minimal-agent spectrum.

PLATFORMS

Cloudflare Opens Clef and Attaches an RL Fine-Tuning Service

Cloudflare released Clef and Clef-flash, two decision models trained in-house and hosted on Workers AI, with Apache 2.0 weights on Hugging Face. Both are fully Jev-API compatible, and both add something Jev does not have today: a vision encoder and a 64k context window against Jev's 32k. Workers AI lists Clef at $0.24 per million input tokens and Clef-flash at $0.09.

The architecture is the pitch. Clef uses Qwen backbones (Qwen3.8-27B for Clef, Qwen3.5-9B for Clef-flash), runs a prefill-only pass, then scores every valid schema choice in parallel rather than generating text token by token, so it returns structured, calibrated decisions instead of prose. Cloudflare's numbers put Clef on top of the Jev Decision Index, with 98.47 on BFCL case exact, 91.93 on API-Bank, and 82.95 on the home appliances benchmark against Jev's 52.27, while Clef-flash posts a 38.8 ms median decision against Jev's 524.1 ms. Cloudflare's own threat intelligence team uses it to classify domains, where it says fetching, rendering, and classifying a site took 2.2 seconds against 4.7 seconds for gpt-oss-120b on the same workflow.

The fine-tuning plan is the product play. Cloudflare is starting with a forward-deployed engineering service to post-train Clef on customer data, then plans a self-serve pipeline built from AI Gateway (to capture request data), Workers AI rollouts, Containers as RL sandboxes, a new Trainer component, and redeployment back onto Workers AI. The default guarantee is that Cloudflare does not read, store, or train on requests or responses unless a customer opts into fine-tuning.

Why it matters: decision models are becoming a distinct layer under agent loops, and the economics are the argument. When a classification returns in tens of milliseconds for a fraction of an LLM call, routing deterministic decisions to a small model and reserving the frontier model for open-ended work is a straightforward cost win. Our Clef benchmark and pricing breakdown has the full table, and the Jev guide explains what System One models changed.

DEVELOPER TOOLS

DeepSeek Ships a Desktop Harness Where Everything Is a Plugin

DeepSeek entered public preview with DeepSeek Harness, an MIT-licensed agent harness available as desktop apps for macOS and Windows and as a web UI you launch with npx @deepseek-ai/dsh web. The architecture comes from Cordis, an "everything is a plugin" framework DeepSeek credits in its developer docs: tools, skills, the interface, and even the agent loop itself are plugins.

The official plugin list shows how far that goes: the agent loop, subagents, and a terminal, with agent teams, auto approval review, scheduled tasks, and voice input all marked experimental. A "creator mode" builds plugins through chat, and the demo has the agent write, install, and verify a floating Pomodoro timer plugin end to end. There is a trajectory view for inspecting tool calls, timings, and runtime state, and the screenshots show the harness running DeepSeek-V4.1-Flash.

DeepSeek is not first to a harness, but it is the first major model vendor to ship one as open source, desktop-native, and fully plugin-extensible. The HN thread (240 points, 119 comments) focused on the architecture more than the model, with commenters comparing it to Pi, opencode, and Codex and debating whether a plugin-first harness is easier to extend or easier to break.

Why it matters: the harness is where model vendors now compete for developers, and DeepSeek is betting that open source, MIT, and a plugin API win more usage than a closed first-party experience. If your workflow runs through an agent harness, expect the plugin surface, not the model, to be the switching cost. Our DeepSeek V4 developer guide covers the model line, and the budget coding agents breakdown is the cost comparison.

ECOSYSTEM

Figma Whitelists MCP Clients, and Pi Is Left Out

Figma has restricted its remote MCP server, the only one that grants agents edit access to Figma documents, to a whitelist of clients, and Pi is not on it. A Figma staff member announced the restriction, and the HN thread (180 points, 101 comments) filled in the mechanics: the local desktop MCP that ships with Figma remains open to any agent but is read-only, while write access requires the remote server and vendor approval.

The thread is a catalog of friction. The founder of opencode said on Threads that getting Figma's MCP set up took eight months of email and that his team did not want to keep spending time on "a simple mcp server." Multiple commenters said they bypass the gate by declaring a whitelisted client name in their OAuth configuration, and one Pi user reported that Pi's new OAuth client-name field plus "Codex" was enough to connect. Figma's own agent runs on Figma credits, and its MCP docs say the current free access is temporary with usage-based pricing to follow.

The defense of an allowlist is not crazy: MCP OAuth flows are hard to secure, callbacks are often localhost, and client identity is one of the few handles a vendor has. But the effect is that platform vendors, not agent builders, decide which harnesses can write to their data. Commenters pointed to Penpot's MCP and Paper as alternatives, and to computer use as the fallback that always works and always costs tokens.

Why it matters: MCP is becoming a distribution channel, and an allowlist is a governance decision that determines which agents can act on your design files. If your workflow depends on a vendor's MCP server, the fallback ladder is worth mapping now: first-party server, third-party bridge, or computer use. Our MCP servers vs Agent Skills breakdown covers where the protocol fits.

INFRASTRUCTURE

Scott Chacon: Git 3.0's SHA-256 Default Will Be a Costly Mistake

Git 3.0 will change the default content hash from SHA-1 to SHA-256, and GitButler co-founder and GitHub co-founder Scott Chacon argues in a long post that it will cost the ecosystem far more than it protects. The immediate problem is that the two formats cannot mix: a repository initialized with one cannot push to a host initialized with the other, submodules must match their parent, tools that assume a 40-character hash need updates, existing URLs containing SHA-1 hashes stop resolving after conversion, and every existing signature breaks because the objects it signed are rehashed.

Chacon's security case is that the threat model does not justify the migration. SHA-1's documented collision attacks (SHAttered in 2017, SHA-1 is a Shambles in 2020) matter for manufactured collisions, not accidental ones, and he puts the accidental-collision birthday bound at roughly 1.4 septillion files in one project. A second-preimage attack, the shape that would let someone replace a file you already trust, remains impractical against even MD5, let alone SHA-1. Real supply-chain attacks, he writes, come from compromising maintainers and package registries, which is "maybe a billion times simpler, cheaper and more likely to succeed" than a GPU-funded collision. He quotes Linus Torvalds from 2005: "The real security is in distribution."

His alternative is to sign an independent tree hash inside the objects that already get signed, the approach Colin Walters' git-evtag has used since 2015. Chacon built a proof of concept and measured it: Chromium's 35 GB, 2.1 million-file tree checksums in 5 seconds on an M5 Mac, the Linux tree in 257 ms, the Git project in 17 ms. His post also cites a talk by Emily Shaffer about Google's preparation for the migration, and says Google may simply override the default internally to keep new repositories on SHA-1 for as long as possible. The HN thread (382 points, 362 comments) split between agreement that the migration cost is under-discussed and the position that Git's inertia is not a reason to keep a broken primitive.

Why it matters: if Git 3.0 ships this default as planned, every forge, submodule, CI script, and hash-length assumption becomes compatibility work, and the alternative Chacon proposes is additive rather than breaking. Watch whether the project keeps the default or adopts a dual-hash path before you plan any repository migration.

DATA

Turbopuffer Declares the Vector Database Dead

Turbopuffer published the first post in a series about v3, a storage rewrite that removes the ANN vector index from the center of the architecture and makes it "just another" secondary index. The company says 100% of CI now passes on v3, that it will publish benchmarks at turbopuffer.com/v3 as it grinds toward performance parity, and that the change lays the foundation for SQL-style queries and aggregations on top of the same engine. The service currently hosts more than a trillion documents, handles 10M+ writes per second, and serves 25k+ queries per second.

The post is unusually specific about why the old design had to go. Keying everything by the ANN address causes storage amplification, because multi-vector representations duplicate document contents across vectors; write amplification, because rebalancing clusters moves full documents and their inverted indexes; and limited vectorization, because every query plan is stuck at the ANN index's 100 to 200-document cluster size while engines like DuckDB and ClickHouse process blocks of thousands. The team's own history is the proof: reworking full-text postings into fixed blocks of about 256 made the index 10x smaller and queries up to 20x faster, and that was only possible because postings do not have to follow cluster boundaries.

The HN discussion (331 points, 88 comments) mostly appreciated a vendor publishing a real architecture post rather than a launch. The customer list makes the stakes clear: Cursor and Notion for vector search, and Linear, which uses turbopuffer as the syncing engine behind its delta sync read path, a non-search workload the old storage layout did not fit.

Why it matters: the "vector database" category is consolidating into general search and storage engines, and the vector index is being demoted from the architecture to one access path among several. If you are choosing retrieval infrastructure for agents, the question is shifting from vector recall benchmarks to whether one engine can serve filtering, full-text, aggregation, and sync without a second system. Our vector database comparison still maps the field as it stands.

TOOLS WORTH A LOOK

  1. Claude Code v2.1.287 (free with a Claude plan) - adds Claude Mods, plugins that can modify deeper harness behavior, plus a built-in "You should know" mod where a side agent flags things you may have missed. Also adds URL prompts from MCP servers on the 2025-11-25 protocol and fixes a dangerous rm losing its always-ask safeguard when output is redirected.
  2. Codex CLI 0.160.0 (free CLI, model usage billed) - the agent command center gets history pagination, fullscreen sessions gain X11 middle-click paste, projectless sessions pick up workspace defaults, and Guardian reviews can pull earlier user instructions and handoff context.
  3. SvelteKit 3 (free, open source) - configuration moves into vite.config.ts, the $lib alias becomes the standard subpath import #lib, environment variables and service workers are simplified, and npx sv migrate sveltekit-3 automates most of the upgrade. Remote functions remain behind an experimental flag.
  4. Janus (free, open source) - a single Go binary that runs GGUF models over Vulkan on AMD, Intel, and Nvidia GPUs, no CUDA required. Show HN at 75 points.

WHAT ELSE IS HAPPENING

  • OpenAI partners with Synopsys on chip design (178 points): the multi-year agreement has OpenAI licensing Synopsys EDA tools to build GPT-Synopsys, a specialized model that operates chip design workflows directly, with a revenue-sharing framework and early engagements underway.
  • The FTC is looking into AI product risks (204 points): CNBC reports the commission has opened inquiries into OpenAI, Anthropic, and other AI companies, a signal that regulator attention is moving from model releases to deployed products.
  • Debian ships a massive kernel security update (311 points): DSA-6528-1 moves the trixie kernel to 6.12.111-1 and fixes hundreds of CVEs spanning privilege escalation, denial of service, and information leaks. If you run Debian stable anywhere, this is the upgrade.
  • Micron warns on memory supply (356 points): the CEO says memory supply will be substantially tighter in 2027 and 2028 than this year, a forward indicator for GPU and inference costs.
  • Cloudflare K2 brings serverless event streams (240 points): a durable, globally distributed stream primitive for Workers, aimed at event-driven and agent workloads that need replayable ordering without running Kafka.
  • Opus 5.5 finds a lost dodo record (149 points): a historian used the model to locate a previously unknown eyewitness account of a dodo, a nice case study in using a frontier model for archival research rather than code.

Every link above goes to a primary source or our sourced coverage. Tomorrow's brief lands when the news does - subscribe to get it by email.

Get the next one in your inbox

The daily brief, delivered. Free, unsubscribe anytime.