Skip to main content
Watch: I Asked Claude to Build Me a Business

Harness / Training / Optimization

ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization

Sungho Park, Wonjoong Kim, Jue Zhang, Wook-Shin Han, Pengfei Gao, Chanyoung Park, Yongqiang Yao, Rao Fu, Elsie Nallipogu, Qingwei Lin, Victor Rühle · Microsoft; POSTECH

arXiv:2610.0090654 upvotes

ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization

Authors: Sungho Park, Wonjoong Kim, Jue Zhang, Wook-Shin Han, Pengfei Gao, Chanyoung Park, Yongqiang Yao, Rao Fu, Elsie Nallipogu, Qingwei Lin, Victor Rühle

arXiv ID: 2610.00906

Problem: Automated harness optimization improves agents by iteratively updating prompts, tool interfaces, and control logic from execution feedback. Existing methods optimize how the harness is updated while largely fixing which training scenarios generate that feedback. As the harness evolves, the scenarios most useful for further optimization change, so the training curriculum should adapt alongside the harness - a dimension earlier harness optimizers left fixed.

Key Methodology:

  • Formulates curriculum selection for harness optimization as an automated curriculum learning problem.
  • Models the evolving curriculum as a non-stationary bandit with dynamically instantiated optimization targets.
  • Abstracts recurring failures into reusable failure-pattern arms, estimates the potential learning progress from further targeting each pattern, and balances revisiting known weaknesses with exploring unseen scenarios for new ones.
  • Optimization outcomes continually update both the set of discovered failure patterns and their priorities, so the curriculum co-evolves with the harness.

Key Results:

  • On GAIA2 and Terminal-Bench 2.0, ActiveSaddler consistently discovers stronger harnesses, improving test Pass@1 by 4.4 and 7.5 percentage points over the same harness optimizer using a scenario order fixed before optimization.
  • Ablations show the gains depend on dynamically constructing optimization targets, estimating their evolving utility, and balancing continued optimization with new failure discovery.

Applied Context: The curriculum is now a first-class harness-optimization dimension: the failure set fed to an improvement loop should be re-ranked as the artifact changes, not frozen before optimization begins. A discovered-failure backlog with priorities is the cheap artifact to keep, and the bandit framing gives a principled way to spend limited evaluation budget between known weaknesses and unseen scenarios. Pairs with the same team's AutoSaddler result on offline trace-driven patches: the update rule and the training scenarios are each tunable, and the scenario side had been held fixed.

Paper: arXiv:2610.00906