Skip to main content
Watch: I Asked Claude to Build Me a Business

AI AGENTS

382 items

376 posts, 2 tools, 4 guides

Blog
What If AI Was Free Tomorrow, at Exactly Today's Capabilities?

A thought experiment with the sci-fi removed: freeze the models at today's capability, drop the price to zero overnight, and work out what actually changes for a working developer. Less than you fear, more than you think, and not where you expect.

Blog
OpenAI Cuts GPT-5.6 Luna 80% and Terra 20%: The Cost-Per-Task Math for Agent Builders

Luna drops from $1/$6 to $0.20/$1.20 per million tokens, Terra from $2.50/$15 to $2/$12, and Sol gets a paid Fast mode. What the new floor means for agent economics, Codex quotas, and the competition.

Blog
MCP Apps vs Tool Calling vs Standalone UIs: Interactive Interfaces for Agent Tools Compared

MCP Apps shipped with the 2026-07-28 final spec - sandboxed interactive UIs for MCP servers. How they compare to standard tool calling and standalone web UIs, and when to use each approach.

Blog
Wiki Skills: The Missing Graph Layer in Agent Context

The Agent Skills spec gave agents progressive disclosure in three tiers - name, SKILL.md, bundled files. What it did not give them is a graph. Skills that link to each other, and say when to follow the link, let an agent navigate knowledge instead of front-loading it. Here is the argument, the measurements from our own 36-skill repo, and what to change.

Blog
AI Coding Agent Firewalls and Security Layers Compared 2026

Belay, Claude Code built-in guards, Codex CLI sandboxing, and MCP proxy patterns compared - how to protect your system from destructive commands, secret leaks, and prompt injection in AI coding agents.

Blog
AI Coding Agent Security Models Compared 2026: Permissions, Sandboxing, and Threat Models for Every Major Tool

How Claude Code, Cursor, Codex, GitHub Copilot, Aider, and Windsurf handle permissions, sandboxing, credential protection, and prompt injection. A structured comparison for engineering teams evaluating agent security.

Blog
Deep Research Agents Need Constraint Ledgers

AREX and the July deep-search papers point to the next useful research-agent primitive: a ledger of claims, constraints, failed paths, and unresolved questions that survives beyond the chat transcript.

Blog
Anthropic Removed 80% of Claude Code's System Prompt. Here Is What They Learned.

Anthropic cut 80% of Claude Code's system prompt for Opus 5 and Fable 5 with zero regression on coding evals. The post landed on HN with 197 points and 133 comments. Here is what the article says, what HN thinks, and what it means for your agent harness.

Blog
The New Rules of Context Engineering for Claude 5 Models: A Developer Guide

Anthropic removed over 80% of Claude Code's system prompt for Claude 5 models. Here is how the rules changed and what it means for your CLAUDE.md files, skills, and system prompts.

Blog
Self-Improving Agents in 5 Minutes: Reflect, Refine, Repeat

Agents that critique their own output, learn from mistakes, and get better over time - the three patterns that actually ship, from simple reflection loops to tree search and meta agents.

Blog
Replit Agent 4: Design-to-Full App with Parallel Agents and Infinite Canvas

Replit Agent 4 adds an infinite design canvas, parallel agents, and team collaboration to the prompt-to-app platform. Here is what changed, what it costs, and when to use it.

Blog
AI Agent Auth Platforms Compared: Arcade vs Composio vs Nango vs Stytch

A practical comparison of the four authentication platforms developers reach for when connecting AI agents to third-party APIs: Arcade, Composio, Nango, and Stytch. OAuth 2.1, MCP support, integration counts, and which to pick by workload.

Blog
DataFlow-Harness Shows Why Agents Need Editable Pipelines

The DataFlow-Harness paper is a useful reminder that coding agents should not just emit scripts. For data work, the durable artifact is an editable, validated pipeline.

Blog
SearchOS Shows Deep Research Agents Need Shared State

SearchOS turns web research from a growing chat transcript into shared state: frontier tasks, evidence graphs, coverage maps, and failure memory. That is the pattern serious deep-research agents need.

Blog
Cursor's SQLite Swarm Is a Test of Goal-Driven Software Engineering

Cursor's latest agent-swarm experiment rebuilt a SQLite-like database from documentation and passed a held-out conformance suite. The bigger story is the shift from assigning code tasks to specifying, measuring, and governing a goal.

Blog
SWE-Pruner Pro Makes Tool Output Pruning an Agent Runtime Problem

SWE-Pruner Pro points at a practical coding-agent design shift: do not only compress prompts outside the model. Teach the runtime to prune tool outputs before they become the next turn's context.

Blog
Resource2Skill Turns Tutorials Into Agent Skills

Microsoft's Resource2Skill paper points at the next agent-skills problem: converting videos, repos, articles, and reference artifacts into executable skills without losing provenance.

Blog
Securing AI Coding Agents: A Practical Threat Model for 2026

Prompt injection, sandbox escapes, and hallucinated dependencies are now documented, patched, CVE-numbered realities. Here is the threat model for agent-written code and the defenses worth adopting this week, ranked by effort.

Blog
Spare Mac for Claude Code: The Remote Control Setup Guide

A guide to setting up an isolated spare Mac that Claude Code can control remotely over SSH, Remote Control from your phone, and Tailscale.

Blog
LM Studio Bionic: A Local-First AI Agent for Open Models

LM Studio launches Bionic, a standalone agent harness for open models with local inference, voice input, and zero data retention cloud options.

PreviousPage 8 of 20Next
AI Development Stack

Get Smarter About AI Dev

New tutorials, open-source projects, and deep dives on coding agents - delivered weekly.

One email per weekReal code, not theoryFree forever