🤖 AI Summary
This study addresses the unreliability of data generation and verification in reinforcement learning for terminal agents by systematically diagnosing three failure modes: invalid benchmarks, fragile frameworks, and reward misalignment. To optimize end-to-end training, we propose a meta-agent pipeline that leverages frontier models as meta-agents to generate tasks and verifiers, integrated with prompt engineering and context extension techniques. The evaluation establishes solvability calibration, verifier auditing, and infrastructure error accounting as core criteria. Experiments demonstrate that this approach improves baseline solvability rates by 5.6×. Furthermore, our analysis reveals rapid performance saturation in 9B-parameter models on specific tasks and highlights the significant suppressive effect of high-difficulty tasks on pass rates.
📝 Abstract
Using a frontier model like Claude Opus as a meta-agent to generate terminal tasks and verifiers for RL training is increasingly common. Yet a runnable Docker image and executable test suite do not guarantee a faithful end-to-end pipeline for terminal agent training. We present a meta-agent pipeline motivated by this gap, diagnosing three classes of failure: benchmark invalidity, harness brittleness, and reward misalignment. Prompt redesign and context extension raise baseline solvability 5.6 times, but a 9B model saturates at 81.3% mean pass@2 within 20 steps on Claude Opus-generated tasks. Adding hard tasks reduces mean pass@2 to 20.6% without changing the training configuration, a strong evidence that the solvability band is model-specific. These findings demonstrate that meta-agent reliability requires solvability-band calibration, verifier audits, and infrastructure error accounting as first-class evaluation criteria, not post-hoc diagnost.