Skip to main content
Watch: Claude Opus 5.5 Built an Entire 3D World

LLM / Agent / Training

AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement

Yizhe Chi, Wenyi Li, Deyao Hong, Xiaoqiu Wang, Mingju Gao, Kaisen Yang, Bingxiang He, Youjie Zheng, Calvin Xiao, Qinhuai Na

arXiv:2608.20318

AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement

Authors: Yizhe Chi, Wenyi Li, Deyao Hong, Xiaoqiu Wang, Mingju Gao, Kaisen Yang, Bingxiang He, Youjie Zheng, Calvin Xiao, Qinhuai Na

arXiv ID: 2608.20318

Problem: Recursive self-improvement (RSI) asks whether an AI system can improve the process that produces AI systems. That process is the training algorithm: a better objective or update rule improves the compute-capability exchange rate for every subsequent run. No existing benchmark isolates whether an agent can design training algorithms - existing suites are won by collecting data or tuning hyperparameters, and none separates a change to how a run is executed from a change to how the model learns.

Key Methodology:

  • 10 frozen research repositories spanning 10 training algorithm families
  • In each task, an agent has 4 hours on one B300 to rewrite the training algorithm; its code is rerun from scratch for up to 12 hours and scored by a fixed evaluator hidden from the agent, against the repository's original algorithm under the same procedure
  • All 10 tasks mapped onto one scale: 0 is an uninformative model, 0.1 is the algorithm the repository ships, 1.0 is the task optimum
  • Task suite, evaluators, and every scored submission released for repeatable measurement

Key Results:

  • Across 29 configurations of 6 systems on all 10 tasks, the mean score is 0.166; the best system reaches 0.250
  • Most submissions never change how the model learns at all; the minority that do average 0.226 against 0.126 for the rest
  • More reasoning effort mostly buys the willingness to change the algorithm, taking that minority from 8% of submissions to 64% and the mean score from 0.094 to 0.196

What it means for developers: Self-improvement claims now have a benchmark that cannot be gamed by data collection or hyperparameter tuning: the agent's output is a training algorithm re-run from scratch under a hidden evaluator, so the score measures whether the learning rule itself improved. Current systems close only a fifth of the distance between the shipped algorithm and the optimum, and most never touch the learning rule at all - while the reasoning-budget result shows extra compute mostly changes which actions agents attempt, not the quality of the outcome. For teams running agentic loops, the benchmark's discriminator is the practical takeaway: check whether the loop's output actually changes how the model learns, or just how it is run.

Paper: arXiv:2608.20318