🤖 AI Summary
This study addresses the environment interaction latency caused by the "think-then-act" paradigm in online distillation for multi-turn agents. To this end, we propose ActFirst-OPD, a novel framework that pioneers an "act-then-reason" paradigm to decouple interaction from generation. Specifically, it directly executes actions via reference-conditioned inverse dynamics while asynchronously generating complete responses to acquire teacher supervision. Furthermore, trajectory deviation detection enables automatic switching to autonomous prediction, effectively overcoming the bottleneck where reasoning blocks action. Experiments demonstrate that this approach accelerates training on Qwen3 models by 1.8× to 4.9× while maintaining or surpassing baseline task success rates, thereby validating the feasibility of highly efficient agent training.
📝 Abstract
On-policy distillation (OPD) trains multi-turn language agents with dense teacher supervision on student-generated responses. However, standard think-then-act rollouts require lengthy reasoning before each short action, delaying environment transitions and experience collection. Generating actions directly reduces this delay but can degrade rollout quality. To address this, we propose ActFirst-OPD, an act-first, reason-later training framework that decouples environment interaction from full-response generation. The student infers and executes actions through reference-conditioned inverse dynamics using its current interaction context and a reference next observation, and switches to autonomous next-action prediction when the resulting transition deviates from the reference trajectory. From the collected interaction contexts, the student asynchronously generates full think-then-act responses for token-level teacher supervision. Experiments across 0.6B-, 1.7B-, and 4B-parameter Qwen3 students show that ActFirst-OPD achieves average wall-clock training speedups of $2.3\times$ on ALFWorld, $1.8\times$ on WebShop, and $4.9\times$ on ScienceWorld over Vanilla OPD. It matches or exceeds all compared OPD baselines in mean task success rate across eight of nine benchmark-model settings. These results demonstrate that reasoning need not block acting during multi-turn agent distillation.