FastOPD: On-Policy Distillation for Lightweight VLA Deployment

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive computational overhead and real-time deployment challenges of large-scale Vision-Language-Action (VLA) models by proposing the FastOPD framework. This work introduces a novel online policy distillation technique that uniquely integrates flow matching with single-state supervision and a self-consistency objective to compress foundation VLA models into lightweight student models. Theoretically, this approach is proven to recover the ideal few-step teacher distribution. Extensive evaluations on the LIBERO benchmark demonstrate that the proposed framework retains 84% of baseline performance using only two inference steps while reducing latency by 78.1%. Furthermore, successful deployment on physical robots validates its practical low-latency capabilities in real-world scenarios.
📝 Abstract
Vision-Language-Action (VLA) foundation models have scaled rapidly to enhance manipulation performance and generalizability, but this scaling incurs high computational costs that render real-world deployment increasingly challenging. Existing approaches typically mitigate this issue by designing smaller architectures or reducing the iterative denoising steps in flow-based policies. In this work, we propose FastOPD, a foundation-to-lightweight VLA framework that enables the practical deployment of large-scale VLAs through efficient on-policy distillation. Specifically, FastOPD adapts a flow map for single-state teacher supervision and combines it with a self-consistency objective to construct a compact student that learns the teacher dynamics. Furthermore, we theoretically demonstrate that minimizing this objective allows the distilled student to recover a distribution on par with that induced by an ideal few-step teacher model. We evaluate FastOPD across diverse foundation policies in simulation and real-world experiments. On LIBERO, FastOPD retains 84% of the performance of $π_{0.5}$ with only two inference steps, reducing inference latency by 78.1% while outperforming existing few-step distillation baselines in average success rate. With LingBot-VLA as the teacher, FastOPD improves the single-step success rate over the base student by 15.9 percentage points on RoboTwin 2.0. We further demonstrate its applicability to a World Action Model (WAM) and deploy a compact student distilled from MolmoAct2 on a real robot.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
computational cost
real-world deployment
lightweight VLA
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Distillation
Vision-Language-Action
Flow Map
Self-Consistency Objective
Lightweight Deployment
🔎 Similar Papers
No similar papers found.