🤖 AI Summary
This work addresses the limitations of fixed-rate action chunking in robotic manipulation, including constrained temporal resolution, restricted prediction horizons, and precision redundancy. To overcome these challenges, we propose Vela, a model that represents robot behavior as continuous trajectories by integrating spline parameterization with motion-dependent adaptive temporal support. This approach transcends fixed output budgets to achieve an optimal balance between long-horizon coverage and local precision, while providing a unified shared interface for heterogeneous embodiments. Extensive multi-embodiment pretraining experiments based on a vision-language-action architecture demonstrate that Vela achieves superior performance across the LIBERO-X and EBench benchmarks, as well as in real-world cooking tasks. These results validate the effectiveness of continuous action representations for scalable and precise robot learning.
📝 Abstract
Most vision-language-action models represent future motion as fixed-rate action chunks, tying temporal resolution and prediction horizon to a fixed output budget. This pointwise representation wastes capacity on highly correlated neighboring actions, leaves temporal continuity and smoothness to be learned implicitly, and forces a tradeoff between long-horizon coverage and the local precision required for contact-rich manipulation. To address these limitations, we introduce Vela, a vision-language-action foundation model that represents future robot behavior as continuous trajectories. Vela combines a compact spline-based action representation with motion-dependent temporal support and a shared action interface for heterogeneous embodiments, allowing a fixed output budget to adapt its temporal resolution across motions. We pretrain Vela on large-scale multi-embodiment robot data and evaluate it on LIBERO-X, EBench, and two real-world long-horizon tasks, egg-cake cooking and potato shredding, obtaining promising results across simulation and physical manipulation. These results highlight the potential of continuous action representations as a foundation for future embodied foundation models. Project page and more results: https://clementine24.github.io/Vela/ .