π€ AI Summary
This study addresses the limitation of existing terminal agent training paradigms that neglect the alignment between trajectories and execution frameworks, thereby hindering co-evolutionary efficiency. We propose an alternating co-evolution framework alongside CoTrace, a novel harness-aware data recipe that decouples framework search from policy training through explicit trajectory routing governance to resolve cross-framework trajectory value discrepancies. By integrating supervised fine-tuning, online reinforcement learning, and component-level promotion decisions, our approach enables closed-loop optimization of automatically synthesized frameworks. This method substantially reduces computational costs while enhancing generalization. Notably, Qwen3.5-9B improves its score on Tmax tasks from 78 to 88 out of 90, and demonstrates superior out-of-distribution transferability on benchmarks such as Terminal-Bench.
π Abstract
Terminal-agent capability depends jointly on model weights and the runtime harness that formats prompts, binds tools, and handles error recovery. Existing harness-model co-evolution approaches improve both components, yet often treat trajectories produced during harness search as an undifferentiated replay buffer. This practice overlooks that a trajectory's value for model training depends on the harness under which it was generated. To systematically analyze this interface, we establish an alternating co-evolution framework that decouples harness search and policy training through component-wise promotion decisions. Within this framework, we introduce CoTrace, a harness-aware data recipe that explicitly governs trajectory routing, provenance matching, and curriculum refresh. Under CoTrace, recurring execution failures guide harness synthesis, while policy training is strictly conditioned on verified rollouts matched to the adopted runtime for supervised fine-tuning (SFT) or fresh online interactions for reinforcement learning (RL). On the Tmax promotion split, CoTrace advances Qwen3.5-9B from 78 to 88 solved tasks under supervised fine-tuning while an online reinforcement variant reaches 90. Specifically, a compact harness-matched corpus produces steady model gains at substantially lower compute than much larger corpora pooled across sibling harnesses. Furthermore, evaluations on Terminal-Bench 2.1 and SWE-bench Lite show that out-of-distribution transfer depends fundamentally on harness compatibility, where maintaining consistency between training and evaluation runtimes prevents procedural execution breakdowns observed under foreign scaffolds.