Skip to main content
Watch: I Asked Claude to Build Me a Business

AI AGENTS

382 items

376 posts, 2 tools, 4 guides

Blog
Cloudflare Gateway Can Now Detect MCP Traffic on the Wire: Shadow MCP Gets a Network Boundary

Cloudflare Gateway now classifies MCP traffic by protocol headers instead of hostname heuristics, ships a shadow-MCP dashboard, and lets admins block any MCP connection that does not arrive through an approved portal. The 2026-07-28 stateless spec is what made it possible.

Blog
The ICML 2026 Agent Reproduction Audit: 23% of Examined Papers Had Falsified or Contested Claims

Hugging Face's open challenge used 1,200+ participants and their coding agents to attempt 2,226 ICML 2026 papers claim by claim. 51% had claims independently verified, 23% had a falsified or contested claim, and four documented falsifications include a spotlight theorem that fails after step 224.

Blog
The Judge Is Now a System You Design

LLM judges flip 25 to 71 percent of their verdicts under pushback and 62 to 91 percent under a trainable persuader, model rankings reverse across token budgets, and a deliberating jury of cheap open-weight models beats frontier single judges at 8 to 15 percent of the cost. The single-judge era is over. Here is the design spec that replaces it.

Blog
GitHub's AutoGPT Playbook: The Repo Instruction File Is Now an API for Other People's Agents

AutoGPT's founding AI engineer published the gates that keep an open source repo sane when agents submit the majority of pull requests: enforced PR templates, AGENTS.md placement, skills that fire on trigger phrases, a CLA as a human detector, and a commit-SHA rule that kills fake review resolutions. GitHub published the playbook August 12, and the details are sharper than the headline.

Blog
We Read DeepSeek Harness: What 453K Lines of Agent Runtime Actually Say

DeepSeek open-sourced its agent harness today. We cloned it and read the code: a 453K-line plugin runtime on a vendored Cordis fork, three patterns worth stealing, V4 line signals hiding in the model adapter, and a 3-line BENCHMARK.md from a lab that published zero eval claims.

Blog
Mendel Godel Machine: Why Self-Improving Coding Agents Need Lineage

A new August 2026 paper argues that coding agents improve faster when they compare attempts across tasks and lineages, not just retry one failed trajectory.

Blog
CLAUDE.md Files Never Stop Growing: A New Paper Names the Mechanism

A study of 247,694 instruction lifetimes in 1,867 repositories shows agentic prompt files grow +226% on average because the reasoning behind each rule decays. Comments encoding that reasoning remove 99.3% of the excess.

Blog
Dub Your Videos into Every Language: The ElevenLabs Dubbing Pipeline

Your best video speaks one language. A coding agent extracts your vocabulary, the ElevenLabs Dubbing API transcribes, translates, and re-voices the file into 90+ languages, keeping each speaker, the timing, and the background audio intact. The complete one-hour build, from repo to a folder of market-ready dubs.

Blog
The $44 Compiler: Persistent Projects Beat Persistent Agents

EvoX Genesis built a 250k-line Rust C compiler with DeepSeek V4 Flash for $44 in tokens by making the project the persistent thing and keeping agents finite-lived. The paper's three runs, the design that made them possible, and what it says about agent memory.

Blog
OpenAI Enterprise Signals: The Agentic Gap Is Now Measurable, and It Is Not About Models

OpenAI published real usage data from its enterprise customer base: Codex now drives 64% of enterprise output tokens, and the top 10% of firms generate 8.3x the tokens of typical ones. What the frontier gap says about agentic AI's spread beyond engineering.

Blog
Skill Files Are the New Supply Chain Attack Surface

Adversarial skill files - folders of instructions agents load dynamically - exploit a mainstream enterprise coding agent in 95.5 to 96.1 percent of runs, while the agent recognizes danger 1.99 percent of the time. The skill folder is now a measured attack surface, and the defense is admission engineering, not better prompts.

Blog
ACE vs ALTK-Evolve: How You Deliver Agent Memory Determines the Token Bill

ACE and IBM's ALTK-Evolve both turn agent trajectories into reusable lessons. The difference is delivery: one injects the whole playbook every step, the other calibrates. On AppWorld, calibration wins with the same accuracy at a fraction of the tokens.

Blog
Cactus Needle 2: The 14MB Agentic LLM That Runs on a Raspberry Pi 5

Cactus open-sourced Needle 2, a 45M-parameter agentic LLM in a single 14MB binary that runs a full tool-calling session in 28MB of RAM. 500 tok/s on a Raspberry Pi 5, ESP32-S3 class parts, Apache 2.0. Here is what the benchmarks actually show.

Blog
Deploy From Your Coding Agent: Wire Railway's MCP Server Into OpenCode

Your coding agent can write the code. With Railway's official MCP server it can ship it too: create the project, deploy the service, assign a domain, tweak variables, and read logs, all as tool calls. The complete one-hour build.

Blog
GitHub Copilot SDK for Java: Annotations, Virtual Threads, and BYOK for Enterprise Agent Harnesses

GitHub shipped a Java-native Copilot SDK (1.0.7-preview.1) with @CopilotTool annotations, virtual-thread support, Jakarta EE and Spring composition, and BYOK mode that works against any OpenAI-compatible endpoint with no Copilot subscription. Here is what changed and what it unlocks.

Blog
Stop Means Stop: New Paper Finds Agent Approval Gates and Cancellation Leak in Six Frameworks

A new arXiv paper probes six widely used open-source agent frameworks and finds the barrier semantics of approval gates, cancellation, and timeouts hold on none of them. A sibling branch can execute while the user is rejecting another one, and replay can double-execute. The fix is a verified external gate called SoundGate.

Blog
Vercel Sandbox Gets a Real Network Boundary: Why Egress Control Is the Missing Half of Agent Security

Vercel Sandbox now polices all outbound traffic on the host, outside the microVM, with SNI-based domain policies, CIDR rules, host-level credential injection, and a deny-all default. Here is why a network boundary is the half of agent isolation that VM escapes missed.

Blog
AgentChaos: Fault Injection Shows Agent Robustness Is a Systems Problem, Not a Model Problem

A new ASE 2026 framework injects server errors, truncated responses, and corrupted tool calls into live agent systems at the HTTP layer. Every system degrades, pass@1 drops up to 50 points, and the ranking stays the same no matter which LLM is behind it.

Blog
LivePlan: Monitoring and Corrective Steering for Coding Agents, Without the LLM Tax

A new arXiv paper builds a deterministic monitor on top of SWE-agent that watches long agent trajectories and only calls an advisor LLM when the run actually drifts. Resolution rates go up by up to 15.2 points at an extra $0.08 per instance, and the paper argues the expensive approach is re-planning from inside the loop.

Blog
Muse Glimmer 30B: Meta's Open-Weight Local Agent Model, Benchmarks, and Hardware Reality

Meta open-sourced Muse Glimmer, a 30B Apache 2.0 multimodal agent model that runs in a 24GB envelope at up to 233 tok/s. MCP Atlas 75.5, SWE-Bench Verified 76.0, 131K context. Here is what the numbers actually say.

PreviousPage 4 of 20Next
AI Development Stack

Get Smarter About AI Dev

New tutorials, open-source projects, and deep dives on coding agents - delivered weekly.

One email per weekReal code, not theoryFree forever