Skip to main content
Watch: I Asked Claude to Build Me a Business

LLM

32 items

25 posts, 7 tools

Blog
ACE vs ALTK-Evolve: How You Deliver Agent Memory Determines the Token Bill

ACE and IBM's ALTK-Evolve both turn agent trajectories into reusable lessons. The difference is delivery: one injects the whole playbook every step, the other calibrates. On AppWorld, calibration wins with the same accuracy at a fraction of the tokens.

Blog
DCAS: Why Fine-Tuned Coding Agents Fall Apart When You Switch Scaffolds

A Huawei-Queen's study finds open coding models fine-tuned under OpenHands degrade sharply under other scaffolds - SWE-Lego-Qwen3-32B drops from 52.6% to 8.4% Pass@1 on OpenCode. The fix: train planning as a model capability, not a scaffold artifact.

Blog
OpenJDK Bans AI-Generated Code: What the New Policy Means for Java Contributors

OpenJDK's interim policy bans AI-generated contributions in full or in part, while Oracle runs on AI-written code internally. What the policy actually says, how it compares to Rust and Debian, and what it means for Java contributors.

Blog
SkillSV: A Shapley Framework That Values the Lines Inside an Agent Skill

Automated skill optimizers write long SKILL.md files whose credit is a black box. SkillSV attributes value to rules, examples, and scripts inside a skill: pruning to 69% of tokens without significant loss on four benchmarks.

Blog
SIGIL Compiles Agent Skills into Harnesses: Prose Runs Skip 44% of Mandated Steps

A Michigan team measures prose SKILL.md files against compiled harnesses: agents execute only 56% of the steps their own skill mandates. SIGIL compiles skills into typed graph harnesses, hitting 86% compliance with 0.58x the tokens.

Blog
The Underground Relay Market for AI API Tokens: How Resellers Get 97% Off

An inside look at the gray-market relay economy that resells OpenAI, Anthropic, and Google API access at up to 97.8% off -- and what it means for developers building on AI APIs.

Blog
Debian Debates LLM Usage: Four Proposals, One Fork in the Road

Debian is voting on four proposals to regulate LLM-generated contributions - from an outright ban to full acceptance. The HN discussion reveals the fault lines in open source's biggest AI policy debate yet.

Blog
Running Gemma 4 26B at 5 Tokens/Sec on a 13-Year-Old Xeon With No GPU

A developer got Google's Gemma 4 26B running on 2013 Xeon hardware for under $300. The fix for a silent MoE bug is now upstream - here's what it means for local inference.

Blog
Mesh LLM: Run 235B Models Across Your Home Lab with iroh

A new distributed inference system pools GPU resources across multiple machines and exposes them through a single OpenAI-compatible API. No RDMA, no NVLink - just QUIC and your existing hardware.

Blog
Self-Hosted vs Managed AI Gateways: A Decision Guide

Self-host LiteLLM or Kong, or use a managed gateway like Portkey, OpenRouter, or Cloudflare AI Gateway? A factual breakdown of cost, control, and ops tradeoffs.

Blog
Program-as-Weights Turns Prompts Into Local Fuzzy Functions

The Program-as-Weights paper is a useful signal for developers: some LLM calls may move from per-request API prompts into compact local artifacts that behave like reusable fuzzy functions.

Tool
Langfuse

Open-source LLM engineering platform: tracing, evals, prompt management, and datasets. Self-hostable, OpenTelemetry-native, with 50+ framework integrations.

Blog
The One-Cent Attack: Prompt Injection Through Bank Transfer Memos

Security researchers showed a €0.02 bank transfer could compromise a banking AI assistant. Here is the exact attack chain - and what every developer building agents needs to do differently.

Blog
Mastra: Review and Setup Guide for TypeScript Agent Apps (2026)

A hands-on look at Mastra, the open source TypeScript framework for building production-ready AI agents and workflows -- with verified setup commands, honest tradeoffs, and current pricing.

Blog
OpenRouter in 2026: Review, Setup, and When Model Routing Pays

OpenRouter gives you one API key for 300+ models, automatic fallbacks, and intelligent provider routing. Here is what it actually costs, how to set it up in five minutes, and when you should skip it entirely.

Blog
LLM Routers Compared: LiteLLM vs Portkey vs OpenRouter in 2026

A practical comparison of LLM routing tools - LiteLLM, Portkey, and OpenRouter - covering cost management, fallbacks, caching, and when to use each for production AI applications.

Blog
KV Caching: A Practical Guide to Optimizing Transformer Inference

How KV caching speeds up LLM inference - the math, the code, the memory tradeoffs, and when it stops helping. Every dev running local models hits this wall.

Blog
Mercury 2 Developer Guide: Building With a Diffusion LLM in Production

A hands-on developer guide to Mercury 2 from Inception Labs. OpenAI-compatible API, reasoning levels, tool use, structured outputs, and when a diffusion LLM beats an autoregressive one in real apps.

Blog
Promptlock: Deterministic Prompt Versioning for LLM Apps

Promptlock gives every prompt a 12-char content-addressable id and a diff-able artifact, turning silent prompt drift into a reviewable change.

Tool
Browser Harness

Self-healing browser automation harness that lets LLMs complete any browser task. 5,000+ stars in under a week.

Page 1 of 2Next
AI Development Stack

Get Smarter About AI Dev

New tutorials, open-source projects, and deep dives on coding agents - delivered weekly.

One email per weekReal code, not theoryFree forever