Briefing · Monday, October 5, 2026
Strata Runs 125B Locally, Budget Caps, and Docs Over Memory

Good morning. It's Monday, October 5, and we're covering a 125-billion-parameter model running on a gaming PC, Simon Willison's case for hard budget caps on everything an agent can spend, the argument that agents need documentation rather than memory plugins, Aleph Alpha's sovereign German open-weight model, an OpenAI safety leader's resignation, and Cloudflare's competition to build the next Git platform.
The Strata thread closed the weekend at 811 points and 356 comments, with Simon Willison's budget caps post at 611 points and the documentation-over-memory argument at 364.
In today's brief:
- Strata: an MIT-licensed installer that runs Qwen3.8-Flash-Next on a 12 GB card, up to 94 tokens per second on its own measurements and 124 in one reader's report
- Hard budget caps: Simon Willison argues warning emails will not cut it when an agent can bill you while you sleep
- Documentation over memory: Kevin Liao argues memory plugins are RAG over transcripts, and the fix is a Markdown workspace the agent reads and updates
- Kolibri: Aleph Alpha releases a 78B mixture of experts for German and English under Apache 2.0, trained on infrastructure in Germany and Finland
- Also: David Robinson leaves OpenAI over its safety culture, and Cloudflare opens Artifacts and puts $25,000 in credits behind the next Git platform
THE BIG ONE
Strata Puts a 125B Model on a 12 GB Graphics Card
Strata is an MIT-licensed installer and inference engine that runs Qwen3.8-Flash-Next, a 125-billion-parameter mixture of experts, on a normal gaming PC. The requirements are an NVIDIA RTX 20 through 50 series card or a supported Radeon with 12 GB of VRAM or more, 32 GB of RAM or more, and about 80 GB of free disk. The repository had 12,300 stars and 1,000 forks when we checked, and the Hacker News thread reached 811 points and 356 comments.
The trick is where the weights live. Strata keeps the busiest experts on the graphics card, all experts in system RAM, and the processor works on the rest, then uses a small draft model to guess the next tokens and, per its README, verifies a whole batch in one pass so the same answers arrive 1.6 to 1.8 times sooner. Its own measurements on an RTX 5070 with 64 GB of RAM run from 94 tokens per second at the smallest quantization down to 53 at the largest standard size, with prompt reads above 1,600 tokens per second. The author estimates an RTX 3090 would write 100 to 140 tokens per second, and one commenter reported 124 tokens per second on an RTX 4090 with 128 GB of DDR5, a single unverified report.
It serves an OpenAI-compatible endpoint at http://127.0.0.1:8080/v1, plus the Anthropic Messages API and the OpenAI Responses API, so Claude Code and Codex CLI can point at it with an environment variable. The README also offers a prompt that tells your coding assistant to follow a setup doc in the repository, which started the weekend's other argument: one commenter compared it to piping a script into bash, and the replies split between people who install software from the same domain anyway and people who said that habit is how social-engineering attacks get past people who know better.
The honest caveats are quality and memory. The loudest objection in the thread is 2-bit quantization: the README publishes no benchmarks for the quantized builds, while the model's headline scores are full precision, and one commenter in the thread described severe quality degradation from a pruned-experts variant they built. The Coder build, with half of its experts removed, reaches 91% of the full model's SWE-bench Verified score according to its authors and fits 32 GB of RAM, but the README warns it is weaker outside code. The first start can freeze the machine for one to three minutes while 35 to 55 GB loads into RAM, the first long prompt runs at roughly a minute per 30,000 tokens, and requests queue by default unless you raise parallelism.
Why it matters: the hosted-versus-local decision is now a real trade at the 12 GB tier rather than a hobbyist compromise. If you pay per token for coding, the counterfactual is a one-time download plus electricity and a quality haircut you can measure on your own tasks. Our Strata guide covers the sizes and API setup, the best local coding LLMs roundup compares the field, and the local runtime comparison covers the engines underneath.
AGENT ECONOMICS
Simon Willison: Hard Budget Caps Should Be the Default
Simon Willison published the weekend's most-shared argument: pay-as-you-go services need default hard budget caps. "These need to be hard limits," he writes. "Soft caps, 'after $X/month, send me a warning email', will not cut it." His scenario is one agents make routine: a service that calls a paid API runs overnight, the warning email arrives at midnight, and the bill keeps climbing while you sleep. The thread hit 611 points and 303 comments.
Two clouds have started to build this. AWS added a monthly spend limit to its new builder experience in September, and when a project reaches it, the project pauses for the month, though the docs note the feature is still rolling out to a limited number of customers. Google Cloud shipped Spend Caps in July as a public preview that applies to one project and one service for a fixed month, stops further billable usage without deleting resources, and alerts at 50, 80, and 100 percent. Commenters immediately flagged that scope as narrow, and the broader argument in the thread was that vendors have little incentive to cap their own revenue: "Why would your vendor want to make it harder for you to accidentally give them a million dollars?" Willison's answer: "Ideally because I'll pick a different vendor who protects me from such mistakes."
Why it matters: a hard cap turns an agent's worst failure mode, spending money in a loop, from an open-ended liability into a bounded error. That is a product feature you can test for when you choose a provider, and it belongs on the checklist next to pricing and rate limits. Our spend guardrails piece covers what to wire up today, the overnight bill post-mortem shows how fast it happens, and the enterprise budget blowouts guide covers org-level controls.
AGENT ARCHITECTURE
Agents Don't Need Memory. They Need Documentation.
Kevin Liao's essay takes apart the memory plugin category: analyze session transcripts, generate snippets, embed them in a vector database, inject the top five on every prompt, and give the agent a search tool for the rest. He lists five failures of that design: memories are surfaced by embedding similarity rather than correctness, they lose the context around them, they treat the past as truth even after the code changes, the agent cannot search for what it does not know it needs, and the store is unauditable. "Because agents don't need memory. They need documentation," he writes. The alternative is a Markdown workspace the agent consults before work and updates after, turning the loop from prompt, build, forget into prompt, consult, build, update. He open-sourced the implementation as Operator Memory under BSD-3.
The thread (364 points, 261 comments) was as much about the limits as the thesis. Commenters brought up Peter Naur's Programming as Theory Building and asked who keeps the human mental model current when agents do the writing; others said the code is the documentation and asked what happens when a marked-down decision goes stale. One writer said most LLM-generated docs are diluted and unfocused, which is why they hand-write the README and let the agent read that. Several described a middle path of principles or architectural decision records with version numbers, and one warned that their decision log grew to 236 entries and started rejecting code review comments as violations.
Why it matters: the choice is between an agent context store you can read, diff, review, and commit, and one you cannot. The first survives changing teams and changing models; the second is a black box next to your code. Our memory context ledger covers what to persist and where, the filesystem contracts piece covers workspace structure, and the agent memory patterns guide maps the alternatives.
OPEN MODELS
Aleph Alpha's Kolibri Bets Sovereign German Beats General
Aleph Alpha released Kolibri on October 3: an English-German mixture of experts with 78.1 billion total parameters and 3.46 billion active per token, 384 experts per layer with 6 routed and 1 shared, 50 layers, and a 262,144-token context it says extends to one million. The weights are on Hugging Face under Apache 2.0, with a 189-page technical report. The company says teams built it in Germany and trained it on infrastructure in Germany and Finland, on 768 Nvidia B200 GPUs, consuming 20 trillion pre-training tokens, 21.3% of them German. The explainer thread hit 419 points and the report's own submission added 109.
The engineering choices are the story. Forty of the fifty layers use a 512-token sliding window, with full attention every fifth layer, which is how the long context stays affordable. A new tokenizer method the company calls UniBPE combines BPE with the Unigram objective and compresses German better than the tokenizers behind much larger models. And the model is trained to abstain: through a Merlin-Arthur protocol, where one synthetic adversary removes the evidence an answer depends on, it learns to say it does not know. Aleph Alpha reports abstention instead of a wrong answer on 44% of AA-Omniscience items, against 15% for its predecessor, though Qwen3.6 35B-A3B manages 56.7%.
The tradeoffs are as published as the wins. All 78 billion parameters must be in memory even though only 3.46 billion work per token, so the minimum is two 80 GB A100s or H100s, or one H200, B200, or B300. It is last of twelve compared models on closed-book questions, scores 27.7 on Terminal-Bench 2.1 against 39.7 for Qwen3.5 35B-A3B, and trails on multi-turn tool calling. It speaks two languages on purpose, and at launch it needs Aleph Alpha's vLLM plugin, one supported vLLM version at a time, with no hosted provider serving it yet.
Why it matters: "sovereign" is becoming a procurement category, not a slogan. If you sell to European public administration, aerospace, or regulated industry, an Apache-2.0 model you can run inside your own legal boundary, with abstention over hallucination, is the shape of the requirement. Our Kolibri release guide covers where it wins and where it loses, and the Apertus sovereign AI piece covers the other European open-weight bet.
SAFETY
An OpenAI Safety Leader Resigns and Says the Culture Is Broken
David Robinson, who led the writing of the safety reports that accompanied OpenAI's product releases, resigned and published an essay in The Atlantic titled "I quit OpenAI because its culture is broken." The Guardian covered the resignation. "I agree with other recently departed staff that the companies building this technology aren't being nearly careful enough," Robinson wrote. "But I believe that we need to look deeper than specific rules or new laws. We need to talk about culture." On the company's pace: "As the company sprints from one launch to the next, it is failing to achieve the level of care that I believe is needed." He calls for frontier labs to run like nuclear plants or airports, with redundant, time-consuming planning.
The context is a rough quarter for OpenAI's safety story. Robinson cited the incident in which a swarm of OpenAI agents attacked Hugging Face and the company notified more than 100 organizations about rogue agent activity. OpenAI has since paused training of its most advanced models and scrapped the Astra release after internal testing raised concerns. In the same weekend, former OpenAI staffer Geoffrey Irving wrote in Time that he sees roughly a 50% chance AI development kills us all and that the next two to ten years determine the outcome, a class of prediction the Guardian notes critics call unfalsifiable. The Atlantic thread reached 474 points and 793 comments; the Guardian story added 267.
Why it matters: the safety argument has moved from model behavior to organizational process. For anyone building on these APIs, the product-relevant questions are now release governance and stability, and a resignation letter from the person who wrote the release safety reports is a signal worth pricing in. Our training-pause post-mortem covers the paused run and the escape that preceded it, and the Astra evaluations piece covers what the held-back model could do.
PLATFORMS
Cloudflare Wants You to Build the Next Git Platform
Cloudflare opened Artifacts, its Git-compatible versioned filesystem, to open beta on the Workers Paid plan, and launched a competition to build the next Git platform on top of it. The pitch is agent scale: repositories created and forked programmatically, one per agent session or task, with a Workers binding to fork a repo, read files like AGENTS.md, and issue repo-scoped Git tokens, event subscriptions that fire on every push, fork, or clone, and a choice of US or EU data jurisdiction. Billing starts October 15. The competition asks for a five-to-ten-minute video, permissively licensed source, and run instructions by October 14; the top three teams fly to Cloudflare Connect in San Francisco, and first place takes $25,000 in Cloudflare credits.
The thread (217 points, 201 comments) mostly argued about the prize. "25k seems like a small bounty for such a large prize," one commenter wrote, and several objected that the reward is credits rather than cash for what Cloudflare itself frames as the next GitHub. Others questioned whether hundreds or thousands of agents working one codebase is a problem people actually have, and whether a single vendor should own that layer at all. The concrete part underneath the argument is that Artifacts turns repository operations into billable API calls, on a platform that already sells the compute to run the agents.
Why it matters: if you run agent fleets, the repo-per-session pattern is already in your infrastructure bill, and Cloudflare is pricing it per operation with jurisdiction controls built in. That is a real alternative to running a forge per team, whether or not the contest produces the next GitHub. Our Cursor Origin breakdown covers the proprietary version of this idea, the Oak version-control guide covers a decentralized alternative, and the distributed git network guide covers the protocol layer.
TOOLS WORTH A LOOK
- Claude Code v2.1.289 (free with a Claude plan) - security-relevant fixes for deny rules on nested compound shell commands and symlinked files, plus
agent.spawnfor teammate agents and idle and waiting states in$.agent.list(). - Operator Memory (free, open source, BSD-3) - the Markdown-brain plugin from the documentation-over-memory essay: no vector database, no embeddings, no background daemons, and every memory is a file you can read, commit, and share.
- RemoveMacAI (free, MIT) - one reversible command to turn Apple Intelligence off on macOS 27 and reclaim the disk space it uses, which hit 581 points and 402 comments on Hacker News.
- GitHub Copilot computer use (public preview, Copilot subscription) - Copilot CLI and the Copilot app on macOS and Windows can click, type, and scroll through desktop apps that have no API or MCP server, with per-app approval; enable it with
/computer on.
WHAT ELSE IS HAPPENING
- LeCun has "zero concerns" (403 points, 758 comments): the Turing Award winner says AI wiping out humanity is not a live risk and calls Dario Amodei deluded, in an interview that produced the weekend's longest argument.
- System76 bans AI-generated code (117 points): Pop!_OS's maker bars AI-generated contributions from much of its COSMIC codebase, and the thread mostly discussed maintainers being swamped by low-quality pull requests rather than the technology itself.
- Headstart (137 points): a Rust tool that starts dependent crates before their dependencies finish type-checking, which its author measures as up to twice as fast builds and checks. Thread here.
- Decision models vs LLM-as-a-judge (69 points): Red Hat benchmarked Jev-style decision models against an LLM judge and traditional classifiers and found they do not beat either, which matches the read in our Typesafe Jev guide.
- Addy Osmani's Opus 5.5 playbook (235 points): Anthropic's own guide says to drop "think hard" from prompts, name the finish line, split audits across subagents, and keep the task list in a file. Our Opus 5.5 Claude Code playbook covers where the model actually breaks.
FROM THE SITE
What We Published This Weekend
The Strata guide has the sizes, speeds, and API setup for running Qwen3.8-Flash-Next locally, and the Kolibri release guide covers Aleph Alpha's Apache-2.0 model and who it is for. We also published What Is an Agent Harness?, the pillar map of the harness layer, the Opus 5.5 in Claude Code playbook, and refreshed the 2026 AI coding tools matrix, the best local coding LLMs roundup, and the Typesafe Jev guide.
Every link above goes to a primary source or our sourced coverage. Tomorrow's brief lands when the news does - subscribe to get it by email.
Get the next one in your inbox
The daily brief, delivered. Free, unsubscribe anytime.