How Bun Coordinated 64 Concurrent Claude Agents to Port 535K Lines of Zig to Rust

TL;DR
A deep dive into the agent orchestration behind the Bun Rust rewrite - the workflow architecture, adversarial review gates, what one human actually did, and the Zig vs Rust debate including Andrew Kelley's response.
Official Sources#
| Resource | Link |
|---|---|
| Bun Rust Rewrite Blog Post | bun.com/blog/bun-in-rust |
| Andrew Kelley Response | andrewkelley.me/post/my-thoughts-bun-rust-rewrite.html |
| Bun Unsafe Audit | bun.com/bun-unsafe-audit |
| GitHub Issue - Miri Checks | github.com/oven-sh/bun/issues/30719 |
| Claude Fable 5 Announcement | anthropic.com/news/claude-fable-5-mythos-5 |
Why This Is the Most Important Agent Case Study of 2026#
On July 8, Jarred Sumner published Rewriting Bun in Rust, the technical postmortem of porting the Bun JavaScript runtime from Zig to Rust. The headline numbers are wild: 535,496 lines of Zig, 6,502 commits, 11 days, one engineer.
But the headline you have probably seen repeated - "64 agents rewrote Bun" - is not quite what the post says. Here is the exact quote:
"At peak, we were running 4 of these workflows at once each in a separate worktree, each with 16 Claudes per workflow. About 64 Claudes at a time."
64 was the peak concurrency, not the team size. The actual unit of organization was what Sumner calls "about 50 dynamic workflows in Claude Code run continuously over the course of 11 days," using "a pre-release version of Claude Fable 5." That distinction matters, because the workflows - not the raw agent count - are the transferable lesson. This post breaks down the orchestration architecture, the verification gates, what the human actually did, and the Zig vs Rust debate that followed, including Andrew Kelley's response.
We covered the news itself in our earlier post. This is the deep dive for people who coordinate agent fleets.
The Numbers, Precisely Sourced#
Everything below is from the primary post:
- 535,496 lines of Zig (excluding comments), across 1,448 .zig files
- May 3 to May 14, 2026: start to merge, 11 days
- 6,502 commits (merges excluded), peaking at 695 commits per hour
- 5.9 billion uncached input tokens, 690 million output tokens, 72 billion cached input token reads
- "around $165,000 at API pricing"
- Roughly 50 dynamic workflows; peak of about 64 concurrent Claude instances
- Model: "a pre-release version of Claude Fable 5, a Mythos-class model"
Results, per the same post: a roughly 20% smaller binary on Linux and Windows, 2-5% faster overall with HTTP throughput up 2.8-4.8%, 128 memory bugs fixed, and the full test suite passing on all 6 platforms (Linux, macOS, Windows, each on x64 and arm64) before merge.
The Orchestration Architecture#
The fleet was not a swarm of agents with a shared goal. It was a hierarchy of pipelines with hard role separation.
The core unit: implementer plus adversarial reviewers#
Every workflow was built around one loop: an implementer writes code, then "2 or more adversarial reviewers per implementer" attack it. Sumner is explicit about the reviewer mandate:
"The reviewer's only job: find bugs & reasons why the code does not work."
This is the single most reusable pattern in the whole post. Reviewers are not collaborators. They are not asked to be balanced. They are prompted to be hostile, and a separate fixer agent applies their feedback. In the compiler-error phase this became a strict assembly line: "1 fixes 2 review 1 applies," with commits landing per crate.
Sharding: worktrees as the isolation boundary#
Parallelism was sharded across git worktrees, not just processes:
"I split it into just 4 workflow shards each with their own worktree (4 worktrees total), each running 16 claudes committing and pushing files."
Why worktrees? Because early on, agents sharing one checkout destroyed each other's work. From the post: about 2 minutes into looping the port over all 1,448 files, "one Claude ran git stash before committing. Another ran git stash pop. And then git reset HEAD --hard. They were stepping on each other!" Full per-agent worktrees were too expensive (Bun's repo is huge, and changes eventually need to compile together), so the compromise was 4 shards with 16 agents each, coordinated inside a shard by the workflow itself.
If you run multi-agent coding at any scale, this is the same lesson everyone hits: file-scope isolation is the first thing you design, not the last.
Phased pipelines, not one big prompt#
The rewrite was not "port Bun to Rust" as a single instruction. It was a sequence of distinct pipelines, each with its own verification signal:
- Preparation: generate a Zig-to-Rust porting guide, analyze lifetimes across struct fields, and run a trial on 3 files with 1 implementer and 2 reviewers before scaling.
- Mass translation: port all 1,448 files using the implementer/reviewer loop.
- Compile: per crate, run
cargo check, group errors by file, and run the fix/review/apply line until crates compile. - Smoke tests: loop over failing CLI subcommands until the binary behaves.
- Test suite: run batches of roughly 100 random test files, sharded across the 4 worktrees, until 100% passed in CI on all platforms.
Each phase has a machine-checkable exit condition. That is the quiet genius of the design: the agents never needed to judge their own success, because the compiler, the smoke tests, and 1.38 million expect() assertions did it for them.
Fix the process, not the output#
The post's most quotable engineering principle:
"fixing the process that generates the code instead of hand-fixing the code."
When agents produced bad output, Sumner edited the workflow, not the diff. One example from the post: Claude interpreted "let's get all the crates to compile" as "stub out the functions with compilation errors." The response was not to un-stub functions by hand; it was to change the workflow instructions so the failure mode could not recur across thousands of files. At 695 commits per hour, hand-fixing is not an option anyway. The workflow is the program; the agents are the runtime.
The Verification Gates#
The port survived because verification was independent of the thing being ported.
- A language-independent test suite. "Bun's own test suite is written in TypeScript which means it doesn't depend on the runtime's programming language." The Zig-era tests ran unchanged against the Rust binary: 1,386,826
expect()calls across 60,624 tests on Debian x64, with comparable counts on macOS and Windows, and "0 tests skipped or deleted" during the rewrite. - Adversarial review as a standing gate, not a final pass - two hostile reviewers on every change.
- CI on all 6 platforms as the merge condition, plus a manual audit: "I manually verified the tests were in fact running and not being skipped."
That last one deserves emphasis. Agents under pressure to make tests pass will sometimes make tests not run. Sumner's checklist assumed exactly that failure mode.
The gates were not perfect. The post owns "19 known regressions, each of which has been fixed," mostly from code that is "syntactically identical in both languages but semantically different" - a debug_assert! that erased side effects, off-by-one bounds checks Rust caught but Zig did not, a slice panic where Zig would truncate. The lesson: a million assertions catch a lot, but semantic gaps between languages slip through precisely because the code looks right.
What the Human Actually Did#
One engineer. Sumner's own description of his role during the 11 days:
"For most of those 11 days (and after), I monitored workflows - manually reading the outputs to check for issues and bugs"
Concretely, the human's job was: design the phased pipelines, watch outputs for false starts, edit workflow instructions when the process produced bad code, verify the tests were really running, review that "the adversarial code review agents were correctly catching discrepancies," handle infrastructure failures (the machine "ran out of disk space and crashed several times"), run manual local checks after CI went green, and press merge.
An HN commenter (yomismoaqui) put a name on this role: "coding agent herders," where "the test harnesses, linters, workflows, etc will be our herding dogs." That maps to what we see in every serious fleet deployment: the human moves up one level of abstraction, from writing code to writing and debugging the system that writes code.
One caveat from the HN thread worth carrying: commenter grandimam pointed out that this was not any engineer plus any codebase. Sumner had deep full-context knowledge of Bun (itself a reimplementation of Node, so correct behavior was known in advance) and an exhaustive test suite. The fleet amplified an expert; it did not replace one.
The Cost Debate#
At "around $165,000 at API pricing," the port was not cheap, and the HN thread litigated the comparison thoroughly. One commenter (jeremyloy_wt) ran the napkin math: a comparable human team effort at loaded Bay Area rates lands several times higher, before counting coordination overhead. Others (IshKebab) countered that cheaper engineering markets narrow the gap, and that the 11-day timeline, not the dollar figure, is the real advantage. Sumner's own framing in the post: "This Rust rewrite would've taken a team of engineers with full-context on the codebase a year of work."
There is also a disclosure worth stating plainly: Bun is part of Anthropic, the model was a pre-release Fable 5 that nobody outside Anthropic could use in May, and the post doubles as a Claude showcase. Several HN commenters (rvz, cube00) flagged exactly this. The orchestration patterns are real and reproducible; the specific cost and timeline came with insider model access.
The Zig vs Rust Debate, Fairly#
Bun's case#
The post's stated motivation is a specific bug class: mixing JavaScriptCore's garbage-collected values with Zig's manually managed memory produced recurring use-after-free, double-free, and leak-at-error-boundary bugs. Rust's borrow checker turns those into "compiler errors" instead of conventions "enforced through code review." The team reports 128 memory bugs fixed and instrumentable leaks eliminated.
Andrew Kelley's response#
Zig's creator responded on July 9 with My Thoughts on the Bun Rust Rewrite, and his argument deserves a fair reading:
- It was not about language features. "The main issue here had nothing to do with the language features of Zig vs Rust, and everything to do with the diverging value systems."
- The bugs reflect engineering practice, not Zig. He contrasts Bun with TigerBeetle, another large Zig codebase: "Quite simply they put in the time to find and eliminate the bugs."
- The performance claims are shaky. "Performance increase is attributed to LTO, which Zig has supported for all of Bun's existence." The post also does not report compilation speed, a metric where Zig typically wins.
Despite sharp words about Bun's engineering culture, Kelley closes on reconciliation: "I don't wish him any ill will. Even in the midst of my frustration, I am happy for him and his success." His post hit 784 points on its own HN thread, slightly outscoring the original.
The unsafe code question#
The strongest technical criticism of the port is about what "memory safe" means here. The Bun post itself discloses that "about 4% of Bun's Rust code sits inside an unsafe block" - roughly 13,000 unsafe keywords. When the port first merged to main in May, a GitHub issue reported that the codebase failed basic Miri checks and allowed undefined behavior in safe Rust, and HN commenters (dfabulich, lunar_mycroft) argued the merged state was far rougher than the announcement tone suggested. Simon Willison's counter in the thread: "that's what this whole post is about. It's about the process of going from that original state to something that's now shipping in production."
Both things are true. The May merge shipped known-rough code, and the July post documents two months of hardening, 19 fixed regressions included. If you cite this project as evidence that agent fleets produce production-ready code in 11 days, you are overclaiming; 11 days got to tests-green, and the path to production ran through June.
What to Steal for Your Own Fleet#
Patterns from this case study that transfer to normal-sized teams and codebases:
- Adversarial reviewers with a single hostile mandate. Do not ask review agents for feedback; ask them for reasons the code is broken. Separate the fixer from the reviewer.
- Machine-checkable exit conditions per phase. Compiler, smoke tests, then the full suite. Agents should never grade their own work.
- Worktree-level isolation. Shared checkouts fail fast and catastrophically. Budget disk space for shards.
- A verification oracle outside the blast radius. Bun's TypeScript test suite survived the rewrite untouched. Whatever you are migrating, your tests must not be part of what changes.
- Fix the workflow, never the diff. At fleet scale, hand-edits are a smell that your process is broken.
- Audit that tests actually ran. "Manually verified the tests were in fact running and not being skipped" belongs in every fleet operator's checklist.
- Pilot before you scale. Three files with one implementer and two reviewers came before 1,448 files with 64 concurrent agents.
FAQ#
Did 64 AI agents rewrite Bun in Rust?#
Not exactly as usually stated. The primary source says: "At peak, we were running 4 of these workflows at once each in a separate worktree, each with 16 Claudes per workflow. About 64 Claudes at a time." So 64 was peak concurrency across about 50 dynamic Claude Code workflows run over 11 days, using a pre-release version of Claude Fable 5, orchestrated by one engineer.
How was the work verified?#
Three gates: adversarial review agents on every change (two or more reviewers per implementer whose only job was finding bugs), phase-specific machine checks (cargo check per crate, then CLI smoke tests), and Bun's language-independent TypeScript test suite - over 1.38 million expect() assertions - passing in CI on all 6 platforms before merge, with a manual audit that tests were genuinely running.
What did the human do while agents wrote the code?#
Jarred Sumner designed the phased workflows, monitored outputs continuously ("manually reading the outputs to check for issues and bugs"), edited workflow instructions when agents produced bad patterns, verified the review agents were catching real discrepancies, handled machine crashes and disk exhaustion, and made the merge decision.
What is Andrew Kelley's counterargument?#
The Zig creator argues the rewrite "had nothing to do with the language features of Zig vs Rust" and everything to do with engineering values, pointing to TigerBeetle as a large Zig codebase without Bun's bug profile. He also notes the performance gains are attributed to LTO, which Zig has long supported, and that compilation speed went unreported.
Is the Rust port actually memory safe?#
Partially. About 4% of the Rust code is inside unsafe blocks (roughly 13,000 unsafe keywords), and the initially merged code failed Miri checks per a GitHub issue filed in May. The team reports 128 memory bugs fixed and instrumentable leaks eliminated, plus 19 known regressions from the rewrite, all since fixed. The safety story improved between the May merge and the July writeup.
How much did it cost and was it worth it?#
Around $165,000 at API pricing (5.9 billion uncached input tokens, 690 million output tokens), plus 11 days of one expert engineer. Comparable human-team estimates in the HN discussion ranged from a few hundred thousand dollars to a year of team time. The bigger caveat: the project used a pre-release model with insider access and an unusually strong test suite, so treat the timeline as an upper bound on what was possible in mid-2026, not a baseline.
Continue Reading#
- Cloudflare CI/CD as Workflows: TypeScript Pipelines, Agent Self-Healing, and the End of YAML Fatigue
- Cloudflare Wallets Gives Agents a Credit Card, an ID, and a Spending Cap
- TypeScript 7.0 Native Compiler: What Breaks, What Gets 10x Faster, and How to Migrate
Sources#
- Rewriting Bun in Rust - primary source, Jarred Sumner, July 8, 2026
- My Thoughts on the Bun Rust Rewrite - Andrew Kelley, July 9, 2026
- HN: Rewriting Bun in Rust (528 comments)
- HN: My thoughts on the Bun Rust rewrite (687 comments)
- GitHub issue: Miri checks and UB in safe Rust
- Bun unsafe audit
Get the next deep dive like this in your inbox
One email a week on AI Agents and the rest of the AI dev stack. Free.
Read next on AI coding tools
Bun Rewrites 535K Lines of Zig to Rust in 11 Days Using Claude
The Bun runtime completed an AI-assisted rewrite from Zig to Rust, fixing memory safety issues and improving performance. Here is what HN thinks and why it matters for LLM-assisted code migration.
6 min readWhat a Fleet of Claude Agents Actually Costs (July 2026 Math)
Claude Code parallel agents cost real money because every session draws from one quota - here is the July 2026 budgeting math, verified against live pricing.
10 min readWhy Skills Beat Prompts for Coding Agents in 2026
The coding-agent workflow is maturing past giant hand-written prompts. The winning pattern in 2026 is a control stack: project rules, reusable skills, bounded sub-agents, and deterministic tools around the model.
9 min readNew here? Start with
Technical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.








