Score
Designs, implements, and evaluates models that fuse visual and linguistic inputs to produce action outputs or control policies for embodied or simulated agents, encompassing perception modules, multimodal fusion, language grounding, and action/planning components. This work covers architecture and training (supervised, reinforcement, imitation), the interfaces between vision, language, and control, and analyses of behavior alignment, robustness, generalization, and sim‑to‑real transfer.
This paper addresses the core challenge of how Vision-Language-Action (VLA) models support language-conditioned robotic tasks in embodied intelligence. Methodologically, it introduces the first systematic, panoramic survey framework, proposing a three-dimensional taxonomy—“Component Design–Low-level Action Policies–High-level Task Planning”—that unifies VLA modeling, embodied control, task decomposition, simulation integration, and cross-benchmark evaluation. Key contributions include: (1) the first explicit characterization of three principal VLA technical paradigms; (2) a comprehensive survey of multimodal datasets, embodied simulation platforms, and standardized evaluation benchmarks; and (3) a structured knowledge graph that identifies critical open challenges—including scalable architecture design, world model integration, and real-world deployment—and outlines promising future research directions.
The rapid advancement of Vision-Language-Action (VLA) models in embodied intelligence lacks systematic organization and critical synthesis. Method: We conduct a comprehensive literature review, taxonomic analysis, and technical roadmap construction to identify five core research frontiers—representation, execution, generalization, safety, and data & evaluation—and propose the first structured “module–milestone–challenge” analytical framework aligned with the evolutionary trajectory of general-purpose embodied agents. Contribution/Results: This work delivers an authoritative, pedagogically instructive, and research-forward survey that has become a de facto standard reference in embodied AI. It has catalyzed community consensus on evaluation benchmarks and data curation standards, and enables dynamic knowledge accumulation via a continuously updated online platform.
This paper addresses the unified modeling of Vision-Language-Action (VLA) systems to achieve deep integration of perception, language understanding, and embodied action. Methodologically, it synthesizes vision-language models (VLMs), hierarchical action control, neuro-symbolic planning, and parameter-efficient fine-tuning, while introducing real-time inference acceleration and multimodal action representations. The contribution comprises: (1) a systematic survey of 80+ works from the past three years; (2) a novel five-pillar framework covering conceptual foundations, architectural evolution, application domains, core challenges, and evaluation paradigms; and (3) the first VLA-specific five-dimensional evaluation standard, clarifying the progression from VLMs to general-purpose embodied agents. Results demonstrate broad applicability across six domains—including humanoid robotics and autonomous driving—with significant improvements in cross-task generalization and real-time control robustness, establishing a scalable foundational framework for embodied intelligence.
Vision-Language-Action (VLA) models face critical challenges in real-world embodied deployment, including poor generalization, low action precision, and the persistent simulation-to-reality gap. Method: We conduct a systematic literature review and technical analysis of VLA research across four dimensions—model architecture, training data, pretraining and post-training methodologies, and evaluation protocols—establishing the first multi-dimensional analytical framework tailored to embodied manipulation tasks. Contribution/Results: Our analysis identifies key evolutionary trends and empirically validated training paradigms in state-of-the-art VLA models. We distill fundamental bottlenecks hindering real-world applicability and propose concrete, theoretically grounded yet engineering-practical directions for future work. This yields a clear, actionable technical roadmap for VLA model design, evaluation, and deployment in robotic systems.
This work systematically investigates the integration paradigms of foundation models (FMs) in embodied robotics, focusing on complex instruction understanding and dexterous manipulation under dynamic environments. We quantitatively compare three paradigms—end-to-end vision-language-action (VLA) models, modular pipelines combining vision-language models (VLMs) with multimodal large language models (MLLMs)—on instruction grounding and object manipulation tasks, providing the first zero-shot and few-shot generalization evaluation across these settings. Results show that VLAs achieve superior manipulation transfer but suffer from low data efficiency; modular approaches demonstrate greater robustness in instruction grounding, with VLMs attaining 78.3% zero-shot accuracy; few-shot fine-tuning boosts VLA manipulation success by 41.6%. The study distills design principles for real-world embodied agents and identifies scalability—particularly in bridging perception, reasoning, and action—as a critical challenge.
This work investigates the feasibility of large language models (LLMs) as end-to-end embodied controllers, addressing the core challenge of directly mapping continuous perceptual inputs (e.g., state vectors) to continuous action outputs. Methodologically, we propose a novel paradigm integrating symbolic reasoning with embodied perception–action closed-loop learning: an initial control policy is generated via textual prompting, then iteratively refined using sensory-motor feedback from Gymnasium/MuJoCo simulation environments, leveraging in-context learning and iterative prompt engineering. Our key contribution is the first demonstration of LLM-driven, gradient-free, iterative controller learning—without explicit policy networks or reinforcement learning updates. Experiments across multiple canonical control benchmarks achieve optimal or near-optimal performance, empirically validating LLMs’ efficacy and generalization capacity as universal embodied controllers.
To overcome the limitations of end-to-end vision-language-action systems, this work proposes Guava—an efficient and general-purpose embodied manipulation interface framework. Built upon three core design principles—iterative perception-reasoning-action loops, semantic action abstraction, and multimodal observation—Guava decouples high-level language reasoning from external perception, planning, and control modules. By distilling a 4B-parameter open-source model using only 2K simulated trajectories, Guava achieves performance on par with state-of-the-art closed-source models in both simulation and real-world environments. It demonstrates strong generalization to unseen objects, novel instructions, and long-horizon tasks, thereby validating its model-agnostic nature and broad applicability.
Existing vision-language-action (VLA) models are typically confined to single tasks—such as navigation or manipulation—and lack generalizability across domains. This work proposes OneVLA, a unified architecture that jointly models navigation and manipulation within a single framework for the first time. Its key innovations include a task-agnostic shared action head, a multi-stage progressive training strategy, a cross-task dataset construction methodology, and a Chain-of-Thought fine-tuning mechanism. Experimental results demonstrate that OneVLA substantially outperforms both specialized single-task models and existing cross-task approaches in both simulated and real-world environments, achieving state-of-the-art overall performance.
Existing vision-language-action models are largely confined to reactive policies and lack explicit modeling of how the physical world evolves under interventions. This work proposes a new paradigm—World Action Models (WAMs)—formally defining its framework for the first time to unify environment dynamics prediction and action generation by modeling the joint distribution over future states and actions. We establish a structured taxonomy distinguishing cascaded and joint architectures, clarifying boundaries with related concepts. Leveraging diverse data sources—including robot teleoperation, human demonstrations, simulation, and in-the-wild first-person videos—we design an evaluation protocol encompassing visual fidelity, physical commonsense, and action plausibility. The study systematically surveys the current landscape, reveals key architectural trade-offs, and identifies open challenges and promising directions for future research.