🤖 AI Summary
Vision-Language-Action (VLA) models face critical challenges in real-world embodied deployment, including poor generalization, low action precision, and the persistent simulation-to-reality gap. Method: We conduct a systematic literature review and technical analysis of VLA research across four dimensions—model architecture, training data, pretraining and post-training methodologies, and evaluation protocols—establishing the first multi-dimensional analytical framework tailored to embodied manipulation tasks. Contribution/Results: Our analysis identifies key evolutionary trends and empirically validated training paradigms in state-of-the-art VLA models. We distill fundamental bottlenecks hindering real-world applicability and propose concrete, theoretically grounded yet engineering-practical directions for future work. This yields a clear, actionable technical roadmap for VLA model design, evaluation, and deployment in robotic systems.
📝 Abstract
Embodied intelligence systems, which enhance agent capabilities through continuous environment interactions, have garnered significant attention from both academia and industry. Vision-Language-Action models, inspired by advancements in large foundation models, serve as universal robotic control frameworks that substantially improve agent-environment interaction capabilities in embodied intelligence systems. This expansion has broadened application scenarios for embodied AI robots. This survey comprehensively reviews VLA models for embodied manipulation. Firstly, it chronicles the developmental trajectory of VLA architectures. Subsequently, we conduct a detailed analysis of current research across 5 critical dimensions: VLA model structures, training datasets, pre-training methods, post-training methods, and model evaluation. Finally, we synthesize key challenges in VLA development and real-world deployment, while outlining promising future research directions.