Act First, Reason Later: Accelerating On-Policy Distillation for Multi-Turn Agents via Reference-Conditioned Inverse Dynamics

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the environment interaction latency caused by the "think-then-act" paradigm in online distillation for multi-turn agents. To this end, we propose ActFirst-OPD, a novel framework that pioneers an "act-then-reason" paradigm to decouple interaction from generation. Specifically, it directly executes actions via reference-conditioned inverse dynamics while asynchronously generating complete responses to acquire teacher supervision. Furthermore, trajectory deviation detection enables automatic switching to autonomous prediction, effectively overcoming the bottleneck where reasoning blocks action. Experiments demonstrate that this approach accelerates training on Qwen3 models by 1.8× to 4.9× while maintaining or surpassing baseline task success rates, thereby validating the feasibility of highly efficient agent training.
📝 Abstract
On-policy distillation (OPD) trains multi-turn language agents with dense teacher supervision on student-generated responses. However, standard think-then-act rollouts require lengthy reasoning before each short action, delaying environment transitions and experience collection. Generating actions directly reduces this delay but can degrade rollout quality. To address this, we propose ActFirst-OPD, an act-first, reason-later training framework that decouples environment interaction from full-response generation. The student infers and executes actions through reference-conditioned inverse dynamics using its current interaction context and a reference next observation, and switches to autonomous next-action prediction when the resulting transition deviates from the reference trajectory. From the collected interaction contexts, the student asynchronously generates full think-then-act responses for token-level teacher supervision. Experiments across 0.6B-, 1.7B-, and 4B-parameter Qwen3 students show that ActFirst-OPD achieves average wall-clock training speedups of $2.3\times$ on ALFWorld, $1.8\times$ on WebShop, and $4.9\times$ on ScienceWorld over Vanilla OPD. It matches or exceeds all compared OPD baselines in mean task success rate across eight of nine benchmark-model settings. These results demonstrate that reasoning need not block acting during multi-turn agent distillation.
Problem

Research questions and friction points this paper is trying to address.

on-policy distillation
multi-turn agents
training acceleration
experience collection
language agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Distillation
Inverse Dynamics
Multi-Turn Agents
Asynchronous Generation
Act-First Reason-Later
🔎 Similar Papers
No similar papers found.