ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization

πŸ“… 2026-09-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitation of existing agent training methods that rely on fixed scenarios and lack dynamic curriculum adjustment mechanisms aligned with evolving capabilities. To this end, we propose an automated curriculum learning framework that enables the co-evolution of training scenarios and agents. This framework introduces curriculum adaptability as a core dimension for the first time, employing a non-stationary multi-armed bandit model to abstract reusable failure modes, quantify learning progress, and dynamically update training priorities. Experimental results demonstrate that our method improves Pass@1 by 4.4 to 7.5 percentage points on benchmarks such as GAIA2, significantly enhancing overall agent performance.
πŸ“ Abstract
Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback. However, existing methods primarily optimize how the harness is updated while largely fixing which training scenarios generate the feedback that drives those updates. As the harness evolves, the scenarios most useful for further optimization can change, suggesting that the training curriculum itself should adapt alongside the harness. We formulate this missing dimension of harness optimization as an automated curriculum learning problem and introduce ActiveSaddler. ActiveSaddler models the evolving curriculum as a non-stationary bandit with dynamically instantiated optimization targets. It abstracts recurring failures into reusable failure-pattern arms, estimates the potential learning progress from further targeting each pattern, and adaptively balances revisiting known weaknesses with exploring unseen scenarios for new ones. Optimization outcomes continually update both the set of discovered failure patterns and their priorities, allowing the curriculum to co-evolve with the harness. Experiments on GAIA2 and Terminal-Bench 2.0 show that ActiveSaddler consistently discovers stronger harnesses, improving test Pass@1 by 4.4 and 7.5 percentage points over the same harness optimizer using a scenario order fixed before optimization, respectively. Ablations further show that these gains depend on dynamically constructing optimization targets, estimating their evolving utility, and balancing continued optimization with new failure discovery. Together, these results establish automated curriculum learning as a new crucial optimization dimension for harness optimization.
Problem

Research questions and friction points this paper is trying to address.

Automated harness optimization
Curriculum learning
LLM agents
Non-stationary bandit
Failure pattern
Innovation

Methods, ideas, or system contributions that make the work stand out.

Automated Curriculum Learning
Non-stationary Bandit
Harness Optimization
Failure Pattern Discovery
LLM Agents