vision-language-action models

Designs, implements, and evaluates models that fuse visual and linguistic inputs to produce action outputs or control policies for embodied or simulated agents, encompassing perception modules, multimodal fusion, language grounding, and action/planning components. This work covers architecture and training (supervised, reinforcement, imitation), the interfaces between vision, language, and control, and analyses of behavior alignment, robustness, generalization, and sim‑to‑real transfer.

vision-language-actionmodels

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.96
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$229K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges

Dec 12, 2025
CX
Chao Xu
🏛️ IROOTECH TECHNOLOGY | King's College London | Hong Kong Polytechnic University | Technische Universität Darmstadt | University of Agder | Imperial College London

The rapid advancement of Vision-Language-Action (VLA) models in embodied intelligence lacks systematic organization and critical synthesis. Method: We conduct a comprehensive literature review, taxonomic analysis, and technical roadmap construction to identify five core research frontiers—representation, execution, generalization, safety, and data & evaluation—and propose the first structured “module–milestone–challenge” analytical framework aligned with the evolutionary trajectory of general-purpose embodied agents. Contribution/Results: This work delivers an authoritative, pedagogically instructive, and research-forward survey that has become a de facto standard reference in embodied AI. It has catalyzed community consensus on evaluation benchmarks and data curation standards, and enables dynamic knowledge accumulation via a continuously updated online platform.

Analyzing five core challenges in representation, execution, generalization, safety, and evaluationProviding a structured guide for researchers to understand and advance embodied intelligenceSurveying Vision-Language-Action models' architecture, evolution, and research challenges

Vision-Language-Action Models: Concepts, Progress, Applications and Challenges

May 07, 2025
RS
Ranjan Sapkota
🏛️ Cornell University | The Hong Kong University of Science and Technology | University of the Peloponnese

This paper addresses the unified modeling of Vision-Language-Action (VLA) systems to achieve deep integration of perception, language understanding, and embodied action. Methodologically, it synthesizes vision-language models (VLMs), hierarchical action control, neuro-symbolic planning, and parameter-efficient fine-tuning, while introducing real-time inference acceleration and multimodal action representations. The contribution comprises: (1) a systematic survey of 80+ works from the past three years; (2) a novel five-pillar framework covering conceptual foundations, architectural evolution, application domains, core challenges, and evaluation paradigms; and (3) the first VLA-specific five-dimensional evaluation standard, clarifying the progression from VLMs to general-purpose embodied agents. Results demonstrate broad applicability across six domains—including humanoid robotics and autonomous driving—with significant improvements in cross-task generalization and real-time control robustness, establishing a scalable foundational framework for embodied intelligence.

Addressing challenges in real-time control and task generalizationExploring applications in robotics and autonomous systemsUnifying perception, language, and action in AI systems

Survey of Vision-Language-Action Models for Embodied Manipulation

Aug 20, 2025
HL
Haoran Li
🏛️ Institute of Automation | Chinese Academy of Sciences

Vision-Language-Action (VLA) models face critical challenges in real-world embodied deployment, including poor generalization, low action precision, and the persistent simulation-to-reality gap. Method: We conduct a systematic literature review and technical analysis of VLA research across four dimensions—model architecture, training data, pretraining and post-training methodologies, and evaluation protocols—establishing the first multi-dimensional analytical framework tailored to embodied manipulation tasks. Contribution/Results: Our analysis identifies key evolutionary trends and empirically validated training paradigms in state-of-the-art VLA models. We distill fundamental bottlenecks hindering real-world applicability and propose concrete, theoretically grounded yet engineering-practical directions for future work. This yields a clear, actionable technical roadmap for VLA model design, evaluation, and deployment in robotic systems.

Analyzing VLA architectures, training methods, and evaluation techniquesIdentifying challenges and future directions for real-world VLA deploymentSurveying Vision-Language-Action models for embodied robotic manipulation

From Grounding to Manipulation: Case Studies of Foundation Model Integration in Embodied Robotic Systems

May 21, 2025
XS
Xiuchao Sui
🏛️ A*STAR | Singapore University of Technology and Design

This work systematically investigates the integration paradigms of foundation models (FMs) in embodied robotics, focusing on complex instruction understanding and dexterous manipulation under dynamic environments. We quantitatively compare three paradigms—end-to-end vision-language-action (VLA) models, modular pipelines combining vision-language models (VLMs) with multimodal large language models (MLLMs)—on instruction grounding and object manipulation tasks, providing the first zero-shot and few-shot generalization evaluation across these settings. Results show that VLAs achieve superior manipulation transfer but suffer from low data efficiency; modular approaches demonstrate greater robustness in instruction grounding, with VLMs attaining 78.3% zero-shot accuracy; few-shot fine-tuning boosts VLA manipulation success by 41.6%. The study distills design principles for real-world embodied agents and identifies scalability—particularly in bridging perception, reasoning, and action—as a critical challenge.

Assessing generalization and data efficiency in zero-shot and few-shot settingsComparing VLA, VLM, and LLM paradigms for complex instruction followingExploring FM integration strategies for robotic language-action bridging

LLMs for sensory-motor control: Combining in-context and iterative learning

Jun 05, 2025
JT
Jônata Tyska Carvalho
🏛️ Federal University of Santa Catarina | Institute of Cognitive Sciences and Technologies

This work investigates the feasibility of large language models (LLMs) as end-to-end embodied controllers, addressing the core challenge of directly mapping continuous perceptual inputs (e.g., state vectors) to continuous action outputs. Methodologically, we propose a novel paradigm integrating symbolic reasoning with embodied perception–action closed-loop learning: an initial control policy is generated via textual prompting, then iteratively refined using sensory-motor feedback from Gymnasium/MuJoCo simulation environments, leveraging in-context learning and iterative prompt engineering. Our key contribution is the first demonstration of LLM-driven, gradient-free, iterative controller learning—without explicit policy networks or reinforcement learning updates. Experiments across multiple canonical control benchmarks achieve optimal or near-optimal performance, empirically validating LLMs’ efficacy and generalization capacity as universal embodied controllers.

Enabling LLMs to control agents via observation-action mappingIteratively refining control strategies using feedback and dataValidating method on classic and MuJoCo control tasks

Latest Papers

What's happening recently
View more

To overcome the limitations of end-to-end vision-language-action systems, this work proposes Guava—an efficient and general-purpose embodied manipulation interface framework. Built upon three core design principles—iterative perception-reasoning-action loops, semantic action abstraction, and multimodal observation—Guava decouples high-level language reasoning from external perception, planning, and control modules. By distilling a 4B-parameter open-source model using only 2K simulated trajectories, Guava achieves performance on par with state-of-the-art closed-source models in both simulation and real-world environments. It demonstrates strong generalization to unseen objects, novel instructions, and long-horizon tasks, thereby validating its model-agnostic nature and broad applicability.

embodied manipulationgeneralizationharness design

Existing vision-language-action (VLA) models are typically confined to single tasks—such as navigation or manipulation—and lack generalizability across domains. This work proposes OneVLA, a unified architecture that jointly models navigation and manipulation within a single framework for the first time. Its key innovations include a task-agnostic shared action head, a multi-stage progressive training strategy, a cross-task dataset construction methodology, and a Chain-of-Thought fine-tuning mechanism. Experimental results demonstrate that OneVLA substantially outperforms both specialized single-task models and existing cross-task approaches in both simulated and real-world environments, achieving state-of-the-art overall performance.

embodied intelligencegeneral-purpose roboticsmanipulation

Existing vision-language-action models are largely confined to reactive policies and lack explicit modeling of how the physical world evolves under interventions. This work proposes a new paradigm—World Action Models (WAMs)—formally defining its framework for the first time to unify environment dynamics prediction and action generation by modeling the joint distribution over future states and actions. We establish a structured taxonomy distinguishing cascaded and joint architectures, clarifying boundaries with related concepts. Leveraging diverse data sources—including robot teleoperation, human demonstrations, simulation, and in-the-wild first-person videos—we design an evaluation protocol encompassing visual fidelity, physical commonsense, and action plausibility. The study systematically surveys the current landscape, reveals key architectural trade-offs, and identifies open challenges and promising directions for future research.

Embodied AIPredictive State ModelingVision-Language-Action Models

Hot Scholars

OS

Oleg Sautenkov

Skoltech
RoboticsComputer VisionUAVSwarm of Drones
AL

Artem Lykov

PhD student, Skolkovo Institute of Science and Technology
RoboticsAICognitive roboticsVLA
MA

Muhammad Ahsan Mustafa

Msc Student
Aerial RoboticsAgile DronesModel Predictive ControlReinforcement Learning
YY

Yasheerah Yaqoot

MSc Student, Skolkovo Institute of Science and Technology
Aerial RoboticsUAV Path PlanningUAV NavigationVLM/LLM
XY

Xumin Yu

Tencent Hunyuan
computer vision