MaskVLA: Visual Masking Against Trajectory Overfitting of Vision-Language-Action Model

📅 2026-09-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决VLA模型在有限数据集上轨迹过拟合问题,提出MaskVLA方法,通过随机遮掩部分主摄像头信息,引导模型学习更精细、任务相关的视觉特征,提高其泛化性能。
📝 Abstract
Vision-Language-Action (VLA) models integrate vision-language understanding with executable robot actions, enabling end-to-end learning for robot control. However, our empirical analysis reveals that existing models exhibit severe trajectory overfitting when finetuned on limited datasets. To guide the model in effectively utilizing wrist camera information, we propose MaskVLA, a masking-based fine-tuning strategy. By randomly masking a small portion of the main camera's visual information, the model is guided to autonomously learn more fine-grained, task-relevant, and effective visual features. This process leads to the emergence of robust policies, thereby enhancing the model's capability to tackle complex manipulation tasks and improving its generalization performance. Our method has been comprehensively evaluated on RoboTwin 2.0, achieving an average success rate improvement of 23.2% and 16.8% compared to $π_0$ and OpenVLA-OFT, respectively. Furthermore, experiments on real-world ALOHA robots also demonstrate the effectiveness of our approach.
Problem

Research questions and friction points this paper is trying to address.

trajectory overfitting
Vision-Language-Action model
limited datasets
Innovation

Methods, ideas, or system contributions that make the work stand out.

Masking-based Fine-tuning
Trajectory Overfitting
Vision-Language-Action Model
Robust Policies
Generalization Performance
Y
Yuxuan Jiang
FNii-Shenzhen, The Chinese University of Hong Kong, Shenzhen, China
J
Jiaying Huang
FNii-Shenzhen, The Chinese University of Hong Kong, Shenzhen, China; School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen, China
G
Ge Wang
FNii-Shenzhen, The Chinese University of Hong Kong, Shenzhen, China; School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen, China; Ising AI
S
Shenhao Yan
FNii-Shenzhen, The Chinese University of Hong Kong, Shenzhen, China; Ising AI
J
Jiahao Yang
FNii-Shenzhen, The Chinese University of Hong Kong, Shenzhen, China; Ising AI; University of Glasgow
C
Chengsi Yao
FNii-Shenzhen, The Chinese University of Hong Kong, Shenzhen, China; Ising AI; School of Automation, Southeast University
Qi Liu
Qi Liu
University of Hong Kong, Reka AI
Natural Language ProcessingMachine LearningArtificial General Intelligence
Q
Qing Zhao
FNii-Shenzhen, The Chinese University of Hong Kong, Shenzhen, China; Ising AI
Shuguang Cui
Shuguang Cui
Distinguished Presidential Chair Professor, School of Science and Engineering, CUHKSZ
AI+NetworkingWireless Communications
Y
Yiming Zhao
Ising AI
Yatong Han
Yatong Han
Chinese University of Hong Kong(SZ)
Embodied AINeuro AILiquid Neural NetworksComputational Biology
Zhen Li
Zhen Li
The Chinese University of Hong Kong
Computer VisionGenerative Models