Survey of Vision-Language-Action Models for Embodied Manipulation

📅 2025-08-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Vision-Language-Action (VLA) models face critical challenges in real-world embodied deployment, including poor generalization, low action precision, and the persistent simulation-to-reality gap. Method: We conduct a systematic literature review and technical analysis of VLA research across four dimensions—model architecture, training data, pretraining and post-training methodologies, and evaluation protocols—establishing the first multi-dimensional analytical framework tailored to embodied manipulation tasks. Contribution/Results: Our analysis identifies key evolutionary trends and empirically validated training paradigms in state-of-the-art VLA models. We distill fundamental bottlenecks hindering real-world applicability and propose concrete, theoretically grounded yet engineering-practical directions for future work. This yields a clear, actionable technical roadmap for VLA model design, evaluation, and deployment in robotic systems.

Technology Category

Computer Vision: Large Vision ModelsIntelligent Robots: Embodied AIMachine Learning: Large Multimodal Models (LMMs)

Application Category

User Modeling, Personalization and Recommendation: User modeling and simulation for interactive and conversational systemsSearch and Retrieval-Augmented AI: Large language models for searchSystems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applications
📝 Abstract
Embodied intelligence systems, which enhance agent capabilities through continuous environment interactions, have garnered significant attention from both academia and industry. Vision-Language-Action models, inspired by advancements in large foundation models, serve as universal robotic control frameworks that substantially improve agent-environment interaction capabilities in embodied intelligence systems. This expansion has broadened application scenarios for embodied AI robots. This survey comprehensively reviews VLA models for embodied manipulation. Firstly, it chronicles the developmental trajectory of VLA architectures. Subsequently, we conduct a detailed analysis of current research across 5 critical dimensions: VLA model structures, training datasets, pre-training methods, post-training methods, and model evaluation. Finally, we synthesize key challenges in VLA development and real-world deployment, while outlining promising future research directions.
Problem

Research questions and friction points this paper is trying to address.

Surveying Vision-Language-Action models for embodied robotic manipulation
Analyzing VLA architectures, training methods, and evaluation techniques
Identifying challenges and future directions for real-world VLA deployment
Innovation

Methods, ideas, or system contributions that make the work stand out.

VLA models as universal robotic control frameworks
Comprehensive analysis across five critical dimensions
Addressing key challenges in VLA development deployment
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Haoran Li
Institute of Automation, Chinese Academy of Sciences, Beijing 100190
Y
Yuhui Chen
Institute of Automation, Chinese Academy of Sciences, Beijing 100190
W
Wenbo Cui
Institute of Automation, Chinese Academy of Sciences, Beijing 100190
W
Weiheng Liu
Institute of Automation, Chinese Academy of Sciences, Beijing 100190
K
Kai Liu
Institute of Automation, Chinese Academy of Sciences, Beijing 100190
M
Mingcai Zhou
Institute of Automation, Chinese Academy of Sciences, Beijing 100190
Z
Zhengtao Zhang
Institute of Automation, Chinese Academy of Sciences, Beijing 100190
Dongbin Zhao
Dongbin Zhao
Institute of Automation, Chinese Academy of Sciences
Deep Reinforcement LearningAdaptive Dynamic ProgrammingGame AISmart drivingrobotics