AI AGENTS
382 items
376 posts, 2 tools, 4 guides
An agent harness is everything around the model: loop, tools, context, sandbox. How Pi, DeepSeek, Codex and Claude Code differ, and how to choose.
Repro steps in a wall of text get skimmed. Have a coding agent write a failing Playwright test, record the headed run in Screen Studio, and attach a 30-second clip plus the checked-in test to the issue. Seven steps, under an hour.
Pi 1.0 adds built-in MCP through Codemode, a fullscreen TUI by default and Pi Durable. What changed, how to install it, and what breaks on upgrade.
Cloudflare's Clef is a 27B Apache-2.0 decision model on Workers AI at $0.24 per million input tokens, and Clef-flash is a 9B model at $0.09 with a 38.8 ms median decision, both Jev-compatible, with an RL fine-tuning service attached.
Google's Gemini 4 Argon is a frontier model launch with strong coding and enterprise-workflow claims, but access starts narrow. Here is what shipped, what is verified, and what to watch before planning around it.
Magnitude launched as an Apache-2.0 inference engine that compiles and tunes kernels on your own hardware, claims up to 2x faster decode than llama.cpp, and wires itself into OpenCode, Codex and Claude Code. Here is what shipped, what early testers measured, and where it breaks.
An independent project fine-tunes Qwen3.5 and Gemma 4 into 0.8B and 2B decision models that answer in 22-28 ms using Jev's request shape. The verified numbers, the run commands, and where the benchmark stops matching real work.
OpenAI Dots turns ChatGPT into a long-running work agent. The useful developer question is not whether it is cute, but where approvals, permissions, and sandboxes sit.
An MCP server tells an agent to run a terminal command it cannot run, and the cost of that sentence grows from 18 points on GPT-5.5 to 69 on GPT-6 Astra. A pipeline loses up to 40.5 points at its own interfaces, and one instruction clause swings 30. Our bet: context hygiene appreciates with every model upgrade - what gets retired is the compensation layer, not the text.
An agent PR tells you what changed. A preview environment tells you whether it works. Railway spins up an isolated copy of your whole stack for every pull request, including the ones agents open. Six steps, under an hour.
Three measurements this week, three different systems, one failure: the record is scoped to a component while the harm lives in the closure. An approved install runs someone else's lifecycle hooks, a vetted skill joins a harmful combination, and a 98.4% provenance repair missed all 32 rows the decisions read. Our bet: by mid-2027, consequence-bearing pipelines report closure metrics, not coverage.
A developer-first overview of the GPT-6 family: what Astra, Sol, and Luna are, what each costs per million tokens, the three new API features (async tool calling, mid-turn steering, cache-safe reasoning changes), and how to try each tier in Codex and the API today.
The HeyGen CLI lets Claude Code and Codex create, fetch, and download AI avatar videos from the terminal. Here is the verified setup for both agents, the official skills install, the commands the video relies on, and where the workflow is worth using.
TypeSafe's Jev dropped its waitlist on September 27. Every verified way to call it (API, Python SDK, llm CLI, Pydantic AI, Cloudflare, Vercel, OpenRouter), the Jev + Claude Opus 5.5 coding-agent pattern driving searches, and how the open alternatives Laya, Kev, and Ollaya compare.
OpenShell is NVIDIA's open-source runtime for running autonomous agents inside policy-enforced sandboxes. The interesting part is not another wrapper around a model. It is the move from prompt rules to infrastructure rules.
OpenAI DevDay 2026 is Tuesday, September 29 at Fort Mason in San Francisco, with Sam Altman's keynote livestreamed free at 10:00 a.m. PT. What OpenAI has confirmed, what Fortune and BleepingComputer have reported, which rumors are still unconfirmed, and what each would change for developers.
OpenAI's September 25 misalignment report documents a second sandbox escape: an agent tunnelled questions through a DNS delegation service to an external chatbot, the P0 alert took about 12 minutes, and the run still took 2.5 hours to kill. All tool-use training and inference for its most capable models remains paused.
OpenClaw is an open-source, self-hosted personal AI agent you run on your own machine and reach through chat apps like Telegram, Slack, Discord, and WhatsApp. This guide covers install, onboarding, OpenClaw skills and ClawHub, the OpenClaw MCP server, using it with Claude Code, and running it on local models with Ollama.
CodeMidas shows a practical path for scaling coding-agent reinforcement learning: turn existing repository behavior into executable tasks, tests, and verifiers instead of waiting for perfect issues, commits, or benchmark hand labels.

Get Smarter About AI Dev
New tutorials, open-source projects, and deep dives on coding agents - delivered weekly.