World-Calibrated Proposal-to-Action Flow for Vision-Language-Action Models

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work proposes ProAct to address the lack of local continuity in action generation for Vision-Language-Action (VLA) policies and the limitation of world models that solely modulate dynamics while neglecting the generation origin. The framework employs a Proposal Expert to preserve motion continuity and a World Expert to predict scene evolution. It innovatively renders the generation source predictable by integrating motion anchoring with a future-compatibility assumption, achieving anisotropic source calibration via flow matching and low-rank geometric constraints under a condition-number budget. Compared to π0.5, ProAct demonstrates superior performance across both simulated and real-world tasks, reducing denoising steps by 50%, decreasing inference latency by 25.8%, and increasing throughput by 34.8%.
📝 Abstract
Flow-based Vision-Language-Action (VLA) policies generate action chunks by transporting samples from a task-agnostic isotropic Gaussian source. As this source is conditioned on neither recent execution nor predicted future evolution, (i) it discards the local continuity established by recently executed motion. (ii) Even when predictive world representations are introduced, they often only condition the transport dynamics rather than determine where generation starts, how far it may deviate, or along which action directions it may expand. Building on this observation, we introduce ProAct, a world-calibrated proposal-to-action framework that makes the generative source itself predictable. (i) To preserve motion continuity, a lightweight Proposal Expert converts recent actions into a scene-aware hypothesis via one motion-anchored endpoint flow-matching step, initializing generation near the demonstrated action manifold. (ii) To jointly capture intended scene evolution and proposal-future compatibility, a prospective World Expert treats the hypothesis as a soft motion prior while predicting the task-consistent latent future. (iii) From this compatibility, the model calibrates a proposal-centered anisotropic source, where a bounded per-step extent controls the allowed deviation and a trace-normalized low-rank geometry under a condition-number budget allocates refinement over coupled translation, rotation, and gripper directions. Compared with $π_{0.5}$, ProAct improves performance across simulation and real-world tasks while reducing denoising steps by 50%, inference latency by up to 25.8%, and increasing throughput by up to 34.8%.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
Flow matching
Motion continuity
World representation
Action generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action Models
Flow Matching
World-Calibrated Proposal
Anisotropic Generative Source
Motion Continuity
🔎 Similar Papers