DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of optimizing Vision-Language-Action (VLA) models for long-horizon tasks using offline data, where online interaction is prohibitively expensive. We propose a teacher-free, rollback-free sequence-level distillation framework that innovatively decomposes the reverse Kullback-Leibler divergence into block-level and future potential terms, enabling policy optimization via a one-step drift objective. Furthermore, our method learns a Q-function from offline demonstrations to serve as an evaluator, eliminating the need for additional teacher models. Experiments demonstrate that the proposed approach surpasses existing baselines in both simulated and real-world robotic scenarios, achieving task success rates comparable to those of multi-step teacher policies.
📝 Abstract
Vision-Language-Action (VLA) models increasingly rely on action experts that generate short action chunks under receding-horizon control. While chunk-level training is convenient across robot embodiments, it optimizes local action likelihood without explicitly accounting for long-horizon task success. Sequence-level reinforcement learning can address this limitation, but typically requires policy rollouts and closed-loop interaction, which are costly for real-robot manipulation. We introduce DriftOPD, a teacher-free, rollout-free framework for sequence-level on-policy distillation of continuous VLA action experts. We show that the sequence-level reverse Kullback-Leibler (KL) divergence decomposes into a chunk-level reverse-KL term and a future-potential term that captures the long-horizon effect of the current action. DriftOPD optimizes these two terms using a one-step drifting objective and a Q-function critic learned from offline demonstrations, respectively, enabling sequence-level optimization with only offline data and one-step action generation. Across multiple VLA architectures in simulation and real-world manipulation, DriftOPD generally outperforms existing one-step distillation baselines while achieving task success performance comparable to multi-step teacher policies. These results demonstrate that long-horizon behavior can be effectively distilled into one-step VLA action experts without online interaction or a separate teacher.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
sequence-level optimization
knowledge distillation
offline reinforcement learning
long-horizon tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reverse-KL Distillation
Vision-Language-Action Models
One-Step Policy
Sequence-Level Optimization
Offline Reinforcement Learning
🔎 Similar Papers
No similar papers found.