Harness Engineering and the Path to Self-Improving AI

TL;DR
Lilian Weng argues self-improving AI won't start with models rewriting their weights - it starts with the harness.
Official Sources#
| Source | What it covers |
|---|---|
| Harness Engineering for Self-Improvement | Lilian Weng's essay (July 4, 2026) on why the harness - not the weights - is the near-term path to recursive self-improvement |
| ACE: Agentic Context Engineering | Treating context as an evolving playbook via a generator, reflector, and curator |
| ADAS: Automated Design of Agentic Systems | A meta-agent that programs new agent workflows in code and keeps an archive of solutions |
| Darwin Gödel Machine | Agents that empirically rewrite their own harness code and validate on SWE-bench |
| AlphaEvolve | Evolutionary search where a frozen LLM proposes diffs against marked code blocks |
| Frontis-MA1 | Open 35B AI4AI model trained for executable machine-learning engineering loops |
| OpenRSI | Released OpenMLE stack, tasks, traces, and Frontis-MA1 model links |
Lilian Weng's new essay makes a claim that cuts against most of the "the models will rewrite themselves" hype: recursive self-improvement is coming, but it won't start with weights. It starts with the harness - the software wrapped around a base model that decides how it thinks, what tools it calls, what it remembers, and how its work gets judged.
If you build agents for a living, this is the most useful framing of 2026 so far. The thing you already control - the scaffolding - is the same thing that improves first.
Last verified: September 12, 2026.
What a harness actually is#
Weng defines a harness as "the system surrounding a base model that orchestrates execution and decides how the model thinks and plans, calls tools and acts, perceives and manages context, stores artifacts, and evaluates results."
That is broader than "agent framework." It includes workflow design, evaluation, permission controls, and persistent state - the boring plumbing that determines whether a capable model produces reliable work or expensive slop. Every coding agent you use - Claude Code, Codex, OpenCode - is a harness. They've quietly converged on the same interface: file discovery, read and edit, shell execution, external context, artifact handling, backend jobs, and subagent delegation.
The three patterns that make it work#
Strip away the research names and modern harnesses share three moves.
Workflow as a goal-oriented loop. Plan → execute → observe/test → improve → iterate until the goal is met. The Codex agent loop is the canonical example, and it's the same shape whether the goal is "fix this bug" or "reproduce this paper."
The file system as persistent memory. Instead of dragging the whole workflow through the context window, the harness writes durable state to disk: experiment logs, code diffs, paper summaries, error traces, past rollout trajectories. This is how long-running agents survive context limits - and why token budget is a harness design problem, not just a billing line.
Subagents and backend jobs. The harness spawns parallel workers, but keeps the parallelism explicit and inspectable - outputs land as files and logs so the run can recover after an interruption. That "leave receipts" discipline is exactly what separates a real agent swarm from a demo.
How the harness starts improving itself#
Here's where it gets recursive. If the harness is code, and a coding agent can write code, then the agent can rewrite the harness. A whole research literature is now doing precisely that:
- Agentic Context Engineering (ACE) treats context as an evolving playbook. A generator produces trajectories, a reflector distills lessons, and a curator merges structured bullets - with IDs, deterministically - so the context grows without collapsing into mush.
- ADAS and AFlow automate workflow discovery itself: ADAS uses a meta-agent to program new agents in code and archive them; AFlow represents workflows as graphs and searches them with Monte Carlo Tree Search.
- STOP (Self-Taught Optimizer) recursively improves its own scaffolding and rediscovered tricks like genetic algorithms and prompt bandits on its own. The catch matters: it improved results on GPT-4 and degraded them on weaker models. Recursion needs a strong base.
- AlphaEvolve and the Darwin Gödel Machine go further - evolutionary pools of candidate programs, with the DGM rewriting its own agent codebase and matching handcrafted agents on SWE-bench Verified.
The pattern across all of them: self-improvement is a search problem, and the harness is the search space.
The evidence is still thin#
Weng is careful not to oversell it, and the benchmarks back her up. On PaperBench (replicate 20 ICML 2024 papers), the best models reach ~21% against ML PhDs. On MLE-bench (75 Kaggle competitions), the best setup hits bronze-medal level just 16.9% of the time. On RE-Bench, humans still score non-zero in 82% of open-ended ML research attempts. Autonomous research works in narrow, verifiable slices - not end to end.
What changed: OpenMLE makes the loop executable#
The July Hugging Face papers list added a useful update to this story: Frontis-MA1, a 35B open model trained for AI4AI in machine-learning engineering. The important part is not the phrase "recursive self-improvement." It is the choice of domain. Machine-learning engineering gives the agent a place to draft code, improve it, debug failures, recombine attempts, execute the result, and score the outcome with real task feedback.
That makes Frontis-MA1 closer to a harness-engineering proof point than a general intelligence claim. The authors release OpenMLE, a stack with verifiable task environments, operator learning, and long-horizon search. In their reported setup, Frontis-MA1 improves MLE-Bench Lite Medal Average from 39.39% to 60.61% over its base model when paired with OpenMLE-Evo, and reaches 71.21% with the larger OpenMLE-Evo-Max search setup. They also report transfer on NatureBench Lite: keeping the framework fixed, the trained model raises Match-SOTA from 50% to 70%; keeping the model fixed, OpenMLE-Evo raises it from 20% to 50%.
Read that result carefully. It does not show an agent autonomously rewriting its own weights until takeoff. It shows something more useful for developers: when the operators, execution environment, evaluation loop, and search policy are all made explicit, a smaller open model can compound experience inside a bounded engineering domain. That is the same thesis as AI4AI-Bench: self-improvement becomes falsifiable only when the task, scorer, and improvement loop are executable artifacts.
Google Trends is a good reality check here. On September 12, 2026, a US 90-day comparison of recursive self improvement AI, AI agent benchmark, MLE Bench, self improving AI, and AI coding agent showed the exact MLE and RSI terms as sparse, while AI coding agent remained the steadier demand cluster. So the durable reader framing is not "RSI is trending." It is "AI coding agents are turning self-improvement into benchmarked engineering loops."
Seven things standing in the way#
The heart of the essay is a sober list of why full recursive self-improvement isn't here yet.
The one that should worry builders most is weak evaluators. Self-improvement loops are only as good as the signal they optimize, and "research taste, novelty, and long-term scientific value are much harder to measure" than a passing test suite. Pair that with reward hacking - loops that game whatever signal you give them - and the design rule writes itself: your evaluator and your permission controls should sit outside the loop, on held-out tests and human review, or the agent will optimize the referee instead of the game.
The rest rhyme with anything you've shipped: context that degrades over long horizons, a training bias toward success that makes models bad at admitting failure, evolutionary loops that collapse to one solution, optimization that ignores maintainability and migration cost, and the human who needs to move up the stack without leaving the loop.
What this means if you're building agents#
You don't need a Darwin Gödel Machine to use any of this. The near-term, practical reading:
- Invest in the harness, not just the prompt. The loop, the file-backed memory, and the tool surface are where reliability actually lives.
- Make everything leave receipts. Logs, diffs, and trajectories on disk are what let an agent recover, and what let you evaluate whether it's improving.
- Keep the evaluator honest and external. Held-out tests and human review are the only defense against a loop that learns to cheat.
- Treat context as a curated artifact, not an ever-growing transcript. The ACE playbook idea - structured, deduplicated, ID'd entries - is something you can apply today with plain context engineering.
- Name the operators. Frontis-MA1's Draft, Improve, Debug, and Crossover loop is a good product-design hint: self-improving systems become inspectable when the edits they are allowed to make are typed, logged, and scored.
The takeaway is oddly empowering. The frontier of self-improving AI isn't locked inside a training run you can't touch. It's the scaffolding on your own machine - and harness engineering is a skill you can start compounding now.
The security version of the same argument is now visible in NVIDIA OpenShell. Once a harness can read files, run code, call APIs, and request credentials, prompt-level safety rules are not enough. The harness needs a runtime policy layer that can say which files, hosts, processes, and credentialed requests are actually allowed.
FAQ#
What is a harness in AI?#
A harness is the software system wrapping a base model that orchestrates how it plans, calls tools, manages context and memory, stores artifacts, and evaluates results. Coding agents like Claude Code and Codex are harnesses.
How is a harness different from an agent framework?#
An agent framework is one piece. A harness is broader - it also covers evaluation, permission controls, persistent state, and workflow design, all the machinery that turns a capable model into a reliable system.
Why does self-improvement start with the harness instead of the weights?#
Because the harness is code an agent can already read and rewrite, and its behavior can be validated empirically. Rewriting weights needs training infrastructure and reliable reward signals we largely don't have yet.
What's the biggest blocker to recursive self-improvement?#
Weak and fuzzy evaluators. Without fast, precise verifiers, a self-improvement loop has no honest signal to optimize - and tends to hack whatever proxy you hand it.
References#
- Lilian Weng, Harness Engineering for Self-Improvement, 2026
- ACE: Agentic Context Engineering
- ADAS: Automated Design of Agentic Systems
- AFlow: Automating Agentic Workflow Generation
- STOP: Self-Taught Optimizer
- AlphaEvolve
- Darwin Gödel Machine
- Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
- OpenRSI GitHub repository
- Hugging Face July 2026 Daily Papers
- PaperBench · RE-Bench · MLE-bench
Continue Reading#
Get the next deep dive like this in your inbox
One email a week on AI Agents and the rest of the AI dev stack. Free.
Read next on AI coding tools
Long-Running Agents Need Harnesses, Not Hope
A long-running coding agent is only useful if the environment around it can queue tasks, capture logs, checkpoint state, verify behavior, limit cost, and recover from failure.
9 min readHarness Engineering Makes Tokens a Systems Budget
OpenAI's harness engineering post and new token-use research point to the same lesson: agentic coding teams need token budgets, receipts, and eval loops, not vibes.
8 min readSelf-Improving AI Agents: Building Systems That Learn From Their Mistakes
AI agents that reflect on failures, accumulate skills, and get better with every session. Reflection patterns, memory architectures, skill extraction, and working code examples for building agents that actually learn.
13 min readNew here? Start with
Technical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.





