ESP: Energy-Score Policy for One-Step Multimodal Action Generation

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high inference latency caused by iterative sampling in diffusion-based vision-language-action (VLA) policies. To overcome this limitation, we propose Teacher-Free Energy Score Prediction (ESP), a novel method that introduces the strictly proper energy score as an objective function for the first time. ESP directly learns multimodal action distributions and generates action chunks through a single forward pass, eliminating the need for teacher model guidance. Experimental results demonstrate that ESP maintains high success rates across both simulated and real-world tasks while significantly reducing action generation latency. By enabling efficient one-step inference, this work provides a practical solution for the real-time deployment of VLA policies.
📝 Abstract
Generative action models based on diffusion and flow matching have been increasingly adopted in vision-language-action (VLA) policies for their ability to capture diverse behaviors, including multiple valid action sequences under the same observation and instruction. Their iterative sampling procedures, however, require repeated network evaluations to generate each action chunk, increasing inference latency in closed-loop control. We propose ESP (Energy-Score Policy), a teacher-free approach that maps policy context and noise directly to an action chunk in a single network evaluation. ESP trains the action head with the energy score rather than mean squared error. Whereas squared-error regression targets the conditional mean, the energy score is strictly proper: its expected value is uniquely minimized by the target distribution. This provides a principled objective for learning multimodal action distributions without iterative sampling, with exact recovery at the population optimum when the model can represent the target distribution. Experiments on both simulation and real-world manipulation tasks demonstrate competitive task success with substantially lower action-generation latency than the flow matching baseline. These results support direct distributional learning as an efficient alternative to iterative generative robot policies.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
generative action models
inference latency
multimodal action generation
closed-loop control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Energy-Score Policy
One-Step Action Generation
Multimodal Action Distribution
Vision-Language-Action (VLA)
Strictly Proper Scoring Rule
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Lilika Makabe
Microsoft Research Asia - Tokyo; Graduate School of Information Science and Technology, the University of Osaka
Heecheol Kim
Heecheol Kim
Microsoft Research Asia - Tokyo
RoboticsEmbodied AI
Yasuyuki Matsushita
Yasuyuki Matsushita
Microsoft Research Asia Tokyo
Embodied AIComputer Vision