Skip to main content
Watch: I Asked Claude to Build Me a Business

AI AGENTS

382 items

376 posts, 2 tools, 4 guides

Blog
Harness Handbook Shows the Missing Map for Coding Agents

A July 2026 paper from Tencent Hunyuan turns agent harnesses into behavior-level maps. The useful lesson for builders is simple: code search is not enough when one behavior spans prompts, tools, state, permissions, and runtime policy.

Blog
SkillHone Shows Why Agent Skills Need Decision History

SkillHone is a July 2026 paper about evolving agent skills across sessions. The useful takeaway for developers is simple: do not save only the latest SKILL.md. Save the decisions that explain why it changed.

Blog
Entire Distributed Git Network: A Developer Guide to the Ex-GitHub CEO's Agent-Era Platform

How to set up Entire's regional Git mirrors for AI coding agents. Covers installation, mirroring, integrations with Claude Code, Codex, Cursor, and Factory AI.

Blog
Terminal-Bench Shows Harness Scaling Is the Coding-Agent Benchmark Now

StateM pushes Terminal-Bench 2.1 to 95.3% raw accuracy by scaling the harness around the model. The lesson for coding-agent teams is that runbooks, state, and recovery loops now matter as much as model choice.

Blog
Clawk: Disposable Linux VMs for Coding Agents Without Cloud Bills

Open-source tool gives Claude Code, Codex, and other agents their own isolated Linux VM on your machine - network firewall included, no cloud account required.

Blog
How Bun Coordinated 64 Concurrent Claude Agents to Port 535K Lines of Zig to Rust

A deep dive into the agent orchestration behind the Bun Rust rewrite - the workflow architecture, adversarial review gates, what one human actually did, and the Zig vs Rust debate including Andrew Kelley's response.

Blog
Composio CLI: Connect OpenClaw and Claude Code to 1,000+ Apps

A companion guide to the Composio CLI video: one command-line layer that lets Claude Code, OpenClaw, Codex, and other agent harnesses search, authenticate, and execute tools across 1,000+ apps.

Blog
Dockerless Verification Is The Next Coding Agent Bottleneck

ByteDance's Dockerless paper asks whether coding-agent patches can be verified without spinning up per-repo environments. The practical answer is not replace CI. It is use cheaper evidence before CI.

Blog
Loop Engineering: How to Design Agent Loops That Actually Converge

The architecture side of loop engineering: plan/act/verify cycles, convergence criteria, retry policies, budget-bounded loops, and the loop-until-dry pattern. Concrete TypeScript-shaped patterns for building agent loops that stop when they should.

Blog
Vera Shows Agent Safety Needs Test Oracles, Not Vibes

A new Vera paper tests Codex, Claude Code, OpenClaw, and Hermes with executable safety cases. The useful lesson is not panic. It is evidence-grounded agent QA.

Blog
AI Agent Eval Tools Compared: Braintrust vs Promptfoo vs DeepEval

A fair comparison of Braintrust, Langfuse evals, Promptfoo, DeepEval, Ragas, and OpenAI Evals: offline vs online evals, LLM-as-judge, CI integration, and dataset management for agent testing.

Blog
GLM 5.2 Matches Human Bookkeeper Accuracy on UK VAT Returns - With Some Caveats

A new benchmark shows GLM 5.2 processing 59 transactions and producing VAT returns off by only 7 pence - at $2.73 versus typical accounting fees of $1,000+. Here is what the benchmark actually tested, where the model failed, and why the HN discussion focused on liability.

Blog
Langfuse vs Braintrust vs Helicone: Choosing an LLM Observability Stack in 2026

A fair, sourced comparison of the three LLM observability platforms teams reach for once agents hit production: Langfuse's open-source tracing and prompt management, Braintrust's eval-first workflow for regressions, and Helicone's drop-in proxy for logging and cost control. Architecture, pricing model, self-hosting, and which to pick by workload.

Blog
Ollama vs LM Studio vs vLLM vs llama.cpp: Picking a Local Runtime for Coding Agents

A fair, sourced comparison of the four runtimes developers reach for when they want a coding agent talking to a model on their own hardware instead of an API: Ollama's convenience, LM Studio's GUI, vLLM's throughput, and llama.cpp's control. What each is actually for, and which to pick.

Blog
Meta Muse Spark 1.1 Developer Guide: First Paid Meta API for Agentic Tasks

Meta launches Muse Spark 1.1 through the new Meta Model API - a 1M-token-context model for personal agentic tasks with OpenAI-compatible endpoints, $20 free credits, and pricing that undercuts the competition.

Blog
Vector Database Comparison for RAG and AI Agents

pgvector, Pinecone, Qdrant, Weaviate, Chroma, Milvus, and Turbopuffer compared on hosting model, filtering, scale, and cost for RAG.

Blog
GitLost: How Researchers Tricked GitHub's AI Agent Into Leaking Private Repos

Security researchers discovered a prompt injection vulnerability in GitHub's Agentic Workflows that allows attackers to extract private repository contents through public issues.

Blog
Harness Engineering and the Path to Self-Improving AI

Lilian Weng argues self-improving AI won't start with models rewriting their weights - it starts with the harness. Here's what that means for developers building agents.

Blog
Microsoft MXC Developer Guide 2026: Sandbox Your AI Agents at the OS Level

Microsoft Execution Containers (MXC) give your AI agents policy-driven sandboxing across Windows, Linux, and macOS. TypeScript SDK, JSON config, multiple isolation backends. Here is how to use it.

Blog
AgentCanvas is a visual adapter for Claude Code and Codex

Claude Code and Codex both ship great agents and terrible transcripts. AgentCanvas is a visual adapter that puts the artifacts, decisions, and handoffs on one board so the next agent and the next human can see them.

PreviousPage 9 of 20Next
AI Development Stack

Get Smarter About AI Dev

New tutorials, open-source projects, and deep dives on coding agents - delivered weekly.

One email per weekReal code, not theoryFree forever