
Codex CLI Worktrees Turn Agent Runs Into Durable Sessions
Codex CLI 0.154.0 adds experimental worktrees, inline answers, Windows daemon support, and approval hardening. The important shift is durable agent workspace control.
GPT-6 Builds Websites That Actually Look This Good...
The Archive
Deep dives into AI agents, coding tools, and building with LLMs.
982 articles921 topics

Codex CLI 0.154.0 adds experimental worktrees, inline answers, Windows daemon support, and approval hardening. The important shift is durable agent workspace control.

A new code-editing paper finds full-file generation beating iterative diff edits on Flutter/Dart tasks. The useful takeaway is not to abandon diffs, but to route by task locality.

A new position paper argues that AI coding-agent research is optimizing for solo autonomy while the real bottleneck is how developers steer, verify, and adapt agents in live work.

Anthropic says Claude worked largely autonomously for 11 days to formalize Fermat's Last Theorem in Lean. The developer lesson is less about one theorem and more about verified repo-scale research artifacts.

Qwen's Terminal-Universe paper argues that terminal-agent trajectories are more useful when you reconstruct the workspace behind them, then generate new verifiable tasks from that environment.

Google shipped agentic video understanding on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite: the model decides which frames, audio, and transcripts to inspect instead of swallowing video at a fixed frame rate. Verified numbers: up to 88% fewer tokens, up to 66% lower cost, and up to 7% better accuracy on video benchmarks. Here is what changed and where the agentic loop still leaks.

Cloudflare's new bot detection engine drops the keep-everyone-out wall for a continuously retraining model, disposable rules, and a memory of past attacks. The first component ships today as a toggle in Bot Management, and the design is an inversion of how every bot product has worked until now.

A changelog nobody reads is a story nobody heard. A coding agent reads your real git history and writes a two-speaker script, the ElevenLabs Text to Dialogue API turns it into a host-and-guest conversation, and ffmpeg stitches the episode. The complete one-hour build.

The Rime CLI streams natural-sounding text-to-speech straight from your terminal, so Claude Code, Codex, Devin, and OpenCode can end each step with a brief spoken summary plus a next-step question. Install, commands, flags, and the agent prompt pattern behind the demo.

Tencent's Hy4 preview ships 770B total parameters with 49B active under Apache 2.0 - a 1M-context text MoE with DeepSeek-style sparse attention, posted Terminal-Bench 85.4 and DeepSWE 64.3, and an OpenRouter price of $0.834/$2.501. Verified against the model card and the live OpenRouter page on August 31, 2026.

Give an agent one instruction and it obeys. Give it eight and it obeys all of them about five percent of the time, no matter which frontier model you bought. The phase transition is measured, the constraints also die in compaction and handoff notes, and in security-critical code the failure ships as infrastructure. The fix is not a better prompt. It is a smaller simultaneous budget and a side channel for the rules that must survive.

GitHub announced three Copilot changes with firm deadlines: Business and Enterprise seats go prepaid starting October 1, the cloud agent and chat surfaces converge into one agent-session experience by September 28, and Balanced becomes the default code review effort level. Here is what each means for your team's budget and workflows.
Showing 12 of 982 articles
Reading Paths
Start with the strongest evergreen pieces, then branch into the weekly notes, guides, and comparisons around the same problem.
Claude Code, Cursor, Codex, and the operating habits that make agentic coding useful in real repositories.
How to choose between orchestration frameworks, UI copilots, stateful agents, and production app patterns.
The practical layer for giving agents access to tools, files, services, and repeatable developer workflows.
Permissions, prompt injection, tool access, audit logs, and rollback patterns for agents that can touch real systems.
Model routing, AI SDKs, deploy platforms, and the infrastructure decisions behind reliable AI products.
How the site, app directory, public APIs, and content loops are built, wired, and improved in public.
Start Here
Series
Deep dives on AI agents, coding tools, and building with LLMs - delivered weekly. No spam.
Free forever. Unsubscribe anytime.

New tutorials, open-source projects, and deep dives on coding agents - delivered weekly.