AI AGENTS
382 items
376 posts, 2 tools, 4 guides
CodeNib's July paper argues that coding agents should stop rediscovering the same repo through grep and reads. Repository context is becoming compiled infrastructure.
Codex Computer History gives agents a rolling view of work across apps. Here is how it works, where it helps, and the privacy boundaries developers should understand.
Vercel Labs' fx installs as a 7.8 MiB Zig binary and runs as a shell-like CLI, a JSON script endpoint, an ACP server, or an embeddable WebAssembly module. Here is the verified setup path, the three auth routes, and the workflows each surface unlocks.
fx is Vercel Labs' experimental coding agent written in Zig: a roughly 6 MiB native binary built to be embedded anywhere from CI sandboxes to the browser. We read the source, the docs, and the launch thread so you can decide fast.
Two minimal coding agents are taking swings at the platform era, but they minimize opposite things: pi refuses features, fx shrinks the bytes. Placing both on the harness spectrum against Claude Code and raw shells shows who each bet is actually for.
Every Grok Bot works on a persistent cloud computer with browser and terminal access, and that single primitive explains everything else about the product. Here is why own-computer beats chat drafts, API integrations, and session-scoped agents.
Grok Bot ships controls over agents rather than controls by agents: approval gates, a chief-of-staff structure, and escalation learning stand in for a settings page. Here is how that oversight model works, and the control questions xAI has not answered publicly.
Grok Bot ships four primitives that compose - a text thread, its own cloud computer, a chief of staff over specialist Bots, and show-it-once routines - and deliberately nothing else. That restraint is the product: you message a coworker instead of configuring an automation platform.
Grok Bot's routines flip the automation playbook: do the job once while a Bot follows along, correct it in plain language, then let the Bot own the schedule. Here is how the mechanic works, where it fits, and how approval gates keep it safe.
How Herdr went from a solo project to 41,000 GitHub stars and Y Combinator: how agent-aware terminals work and the gap they fill that tmux does not.
The hands-on guide to running a fleet of coding agents on Herdr: verified install and config steps, three fleet patterns pulled from real projects, the extension ecosystem, and the gaps nobody advertises.
Herdr vs pi vs tmux compared against their own docs: which agent harness fits 2, 10, or 20 agents, and where Herdr genuinely loses.
Within weeks of going public, Herdr collected policy gates, OS-level agent surfaces, editor bridges, a plugin marketplace, and a YC acceptance letter. We measured the ecosystem layer to test what that velocity actually proves about where agent tooling lands next.
The MCP maintainers published an updated roadmap on August 22, 2026 with five priority areas, including progressive discovery for tool catalogs and standardized agent identity. Here is what changes for developers building MCP servers and agent platforms.
OpenClaw went from an unlisted repo created on November 24, 2025 to more than 100,000 stars in under two weeks, and stood at 387,250 stars as of August 23, 2026 - the near-vertical line WIRED described as a rocket launch. Here is how that chart happened, and where the curve stands now.
How a one-developer protest against bloated coding harnesses became a 95,000-star agent toolkit: pi's five-package architecture, branching JSONL session trees, four run modes, and the philosophy that refuses to build sub-agents, plan mode, or MCP.
The practical guide to earendil-works/pi: verified install and auth steps, all four run modes from TUI to SDK, JSONL session trees with branch, fork and resume, and the rough edges nobody advertises.
A decision-intent comparison of pi, Claude Code, OpenCode and Codex CLI as your main agentic coding harness in late 2026, with a verified capability matrix and pick-X-if verdicts.
Compression is the default answer to the agent bill, and a new three-model, eleven-method audit says the bill is the wrong place to look: quantized and pruned agents lose their head knowledge first, stay confidently wrong about what they lost, and hide subgroup preference flips behind flat bias scores. The same week, the serving side produced cost cuts that touch none of that. Our bet: cheapness comes from the cache before it comes from the weights.
A feedback-driven test-generation loop reported steady improvement. An audit found a single-reference oracle had inflated the measured gain by 9.46 to 14.85 points, independent resampling beat the evolution at equal budget, and a placebo arm erased the feedback benefit. The judge was never the only layer that lied - the reference underneath shares the disease. Independent verification is the only real verification.

Get Smarter About AI Dev
New tutorials, open-source projects, and deep dives on coding agents - delivered weekly.