GALA: Geometry-Aware Latent Action Modeling for Vision-Language-Action Model Pretraining across Embodiments

📅 2026-09-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决多实体数据集中动作空间异质性问题,提出GALA框架,结合视觉和几何潜在动作建模,增强跨实体泛化能力。
📝 Abstract
Learning large-scale vision-language-action (VLA) models from multi-embodiment datasets remains challenging due to heterogeneous action spaces across end effectors. Although latent action models (LAMs) can learn embodiment-agnostic action representations from diverse video data, existing image-based LAMs often fail to capture fine-grained end-effector articulation, particularly finger-level geometric changes in human and dexterous robot hands. To address this limitation, we propose GALA, a Geometry-Aware Latent-Action modeling framework that augments image-based latent actions with 3D end-effector geometric motion. However, naively incorporating point clouds yields fine-grained action representations with limited shared semantics, hindering cross-embodiment pretraining. To address this issue, we introduce the Unified End-effector Motion Representation (UEMR), which preserves fine-grained motion information while improving the cross-embodiment generalizability of latent actions. Building upon UEMR, GALA combines visual latent actions that capture scene-level dynamics with geometric latent actions that capture shared fine-grained end-effector articulation, providing effective supervision for VLA pretraining from multi-embodiment data, including action-free ego-centric human videos. Experiments on fine-grained motion probing, cross-embodiment retrieval, and downstream VLA evaluation demonstrate GALA's effectiveness in modeling generalizable fine-grained motions across embodiments, achieving 68.3% RoboCasa-GR1 success rate and 75.5% real-world success rate. Code, appendix, and demos are available at https://puzhenyuan.github.io/GALA-website/.
Problem

Research questions and friction points this paper is trying to address.

vision-language-action
latent action models
end-effector articulation
geometry-aware
cross-embodiment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Geometry-Aware Latent Action
Unified End-effector Motion Representation (UEMR)
Cross-embodiment Generalizability
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yichen Liu
Institute for Interdisciplinary Information Sciences, Tsinghua University, China
P
Puzhen Yuan
Institute for Interdisciplinary Information Sciences, Tsinghua University, China
Xiang Zhu
Xiang Zhu
Institute for Interdisciplinary Information Sciences, Tsinghua University
robotics
Yanjiang Guo
Yanjiang Guo
Tsinghua University
Embodied AIGenerative Model
Jianyu Chen
Jianyu Chen
Assistant Professor, Tsinghua University
AIRobotics