BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing 3D vision-language-action (VLA) models suffer from limitations in data efficiency, generalization, and historical memory modeling, hindering their performance in data-scarce, open-world, and memory-dependent manipulation tasks. This work proposes the first unified spatiotemporal memory architecture built upon BridgeVLA, which explicitly models persistent spatial context and temporal interaction history to enable reasoning over observation sequences while preserving the original framework’s efficiency and generalization capabilities. By integrating a pretrained vision-language model, multi-view point cloud projection, intermediate heatmap prediction, and a spatiotemporal memory network, the method constructs an end-to-end memory-augmented 3D operating system. It achieves state-of-the-art performance on two memory-intensive benchmarks and demonstrates strong cross-task, cross-environment, and cross-platform scalability in both dual-arm simulation and real-robot experiments.
📝 Abstract
Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as a promising paradigm for 3D robot manipulation. However, existing 3D VLA methods remain data-hungry, exhibit limited generalization under distribution shifts, and lack explicit memory of past observations. These limitations hinder their application to data-scarce, open-world, and memory-dependent manipulation scenarios. Our previous work, BridgeVLA, improves data efficiency and generalization by preserving the input--output alignment of a pre-trained VLM during 3D action learning: raw point clouds are projected into multi-view images, and intermediate heatmaps are predicted before generating robot actions. In this work, we develop BridgeVLA++ by equipping BridgeVLA with a unified spatio-temporal memory architecture that models persistent spatial context and temporal interaction history. The resulting memory-augmented framework can reason over observation histories while preserving BridgeVLA's data efficiency and generalization capabilities. Extensive experiments show that our framework achieves strong performance on spatial manipulation tasks while exhibiting robust generalization. BridgeVLA++ further achieves state-of-the-art performance on two challenging memory-dependent manipulation benchmarks without sacrificing the data efficiency and generalization of the original BridgeVLA. In addition, BridgeVLA++ performs effectively in bimanual manipulation settings and is validated on an additional real-world robotic platform, demonstrating its scalability across tasks, environments, and robotic platforms. These results establish BridgeVLA++ as a unified 3D vision-language-action framework that simultaneously supports data-efficient learning, robust generalization, and effective memory-aware robot manipulation. Project website: https://bridgevla-plus.github.io/.
Problem

Research questions and friction points this paper is trying to address.

data efficiency
generalization
memory-augmented
3D manipulation
vision-language-action
Innovation

Methods, ideas, or system contributions that make the work stand out.

memory-augmented
vision-language-action
data-efficient
3D manipulation
generalization
Peiyan Li
Peiyan Li
Ludwig-Maximilians-Universität München
data mininggraph mining
Y
Yuze Zhu
New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences, Beijing, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China
Y
Yixiang Chen
New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences, Beijing, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China
Q
Qisen Ma
New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences, Beijing, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China
Yuan Xu
Yuan Xu
Associate Professor, Cumming School of Medicine, University of Caglary
Health Data MethodsEpidemiologyHealth Services Research
J
Jiabing Yang
New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences, Beijing, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China
H
He Guan
FiveAges, Beijing, China
Yan Huang
Yan Huang
Institute of Automation, Chinese Academy of Sciences
computer visiondeep learningmultimodal learning
Hongtao Wu
Hongtao Wu
Bytedance Research
roboticsrobot visionrobot learning
Xiao Ma
Xiao Ma
ByteDance Seed
Robot LearningReinforcement LearningRobotics
Tao Kong
Tao Kong
ByteDance Research
Robot Foundation ModelRobot LearningComputer Vision
Liang Wang
Liang Wang
National Lab of Pattern Recognition
Computer VisionPattern RecognitionMachine Learning
Tieniu Tan
Tieniu Tan
Institute of Automation, Chinese Academy of Sciences