LLM / Agent / Training
AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement
Yizhe Chi, Wenyi Li, Deyao Hong, Xiaoqiu Wang, Mingju Gao, Kaisen Yang, Bingxiang He, Youjie Zheng, Calvin Xiao, Qinhuai Na
AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement
Authors: Yizhe Chi, Wenyi Li, Deyao Hong, Xiaoqiu Wang, Mingju Gao, Kaisen Yang, Bingxiang He, Youjie Zheng, Calvin Xiao, Qinhuai Na
arXiv ID: 2608.20318
Problem: Recursive self-improvement (RSI) asks whether an AI system can improve the process that produces AI systems. That process is the training algorithm: a better objective or update rule improves the compute-capability exchange rate for every subsequent run. No existing benchmark isolates whether an agent can design training algorithms - existing suites are won by collecting data or tuning hyperparameters, and none separates a change to how a run is executed from a change to how the model learns.
Key Methodology:
- 10 frozen research repositories spanning 10 training algorithm families
- In each task, an agent has 4 hours on one B300 to rewrite the training algorithm; its code is rerun from scratch for up to 12 hours and scored by a fixed evaluator hidden from the agent, against the repository's original algorithm under the same procedure
- All 10 tasks mapped onto one scale: 0 is an uninformative model, 0.1 is the algorithm the repository ships, 1.0 is the task optimum
- Task suite, evaluators, and every scored submission released for repeatable measurement
Key Results:
- Across 29 configurations of 6 systems on all 10 tasks, the mean score is 0.166; the best system reaches 0.250
- Most submissions never change how the model learns at all; the minority that do average 0.226 against 0.126 for the rest
- More reasoning effort mostly buys the willingness to change the algorithm, taking that minority from 8% of submissions to 64% and the mean score from 0.094 to 0.196
What it means for developers: Self-improvement claims now have a benchmark that cannot be gamed by data collection or hyperparameter tuning: the agent's output is a training algorithm re-run from scratch under a hidden evaluator, so the score measures whether the learning rule itself improved. Current systems close only a fifth of the distance between the shipped algorithm and the optimum, and most never touch the learning rule at all - while the reasoning-budget result shows extra compute mostly changes which actions agents attempt, not the quality of the outcome. For teams running agentic loops, the benchmark's discriminator is the practical takeaway: check whether the loop's output actually changes how the model learns, or just how it is run.
Paper: arXiv:2608.20318