Disentangling Spurious Correlations in Vision-Language-Action Models via Predicting Domain-Invariant Latent Lookahead

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of Vision-Language-Action (VLA) models to visual distribution shifts caused by their reliance on spurious correlations. To mitigate this issue, we propose the DILL framework, which employs a task-domain dual encoder and domain-invariant latent lookahead prediction. By integrating contrastive learning with Gaussian decoupling regularization, DILL effectively disentangles task-relevant structures from domain-specific visual variations, thereby facilitating robust policy learning. Experimental results demonstrate that DILL achieves a success rate of 69.1% on the LIBERO-Plus benchmark, outperforming baseline methods by 11.4 percentage points. Furthermore, its generalization capability is validated through real-world robotic manipulation tasks.
📝 Abstract
Vision-Language-Action (VLA) models remain brittle under visual distribution shifts, often relying on spurious correlations tied to domain-specific factors rather than task-relevant structure. We propose Domain-Invariant Latent Lookahead (DILL), a representation-learning framework that mitigates shortcut learning in VLA policies. Our key idea is to supervise policies with domain-invariant future latents learned from domain-transformed trajectory data. A Task-Domain Encoder is trained with contrastive objectives and Gaussian disentanglement regularization to separate task-relevant structure from domain-specific visual variation. The learned encoder then provides future latents for VLA policy learning through lookahead prediction and domain disentanglement, encouraging the policy to focus on task-relevant structure rather than incidental visual factors. Counterfactual task-view evaluations show that DILL reduces shortcut reliance, while LIBERO-Plus evaluations demonstrate improved visual robustness, with 69.1% average success, 11.4 percentage points above the strongest baseline. Real-world manipulation experiments further support DILL's applicability beyond controlled simulation. Complementary latent-space diagnostics show that these behavioral gains are accompanied by representations that better preserve task-consistent structure while suppressing domain-specific variation. Our project page is available at https://dill-vla.github.io/.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
spurious correlations
visual distribution shifts
shortcut learning
domain-invariant representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action Models
Domain-Invariant Latent Lookahead
Disentangled Representation Learning
Spurious Correlations
Contrastive Learning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Junghyun Kim
Junghyun Kim
Ph.D. Candidate, Seoul National University
Embodied AIMachine LearningRobot LearningVisual Grounding
N
Ngseo Kim
Seoul National University, Seoul, Korea
C
ChungWoo Lee
Seoul National University, Seoul, Korea
S
Seoyeon Lee
Seoul National University, Seoul, Korea
W
Woo-Jeong Baek
Hyundai Motors, Korea
A
Adam Zhou
OpenMind, San Francisco, CA, USA
Chip Huyen
Chip Huyen
OpenMind, San Francisco, CA, USA
J
Jun-Ki Lee
Seoul National University, Seoul, Korea
Gi-Cheon Kang
Gi-Cheon Kang
Ajou University, Korea
Byoung-Tak Zhang
Byoung-Tak Zhang
Professor of Computer Science, Cognitive Science, and Brain Science, Seoul National University
Machine LearningArtificial IntelligenceCognitive Science