🤖 AI Summary
This study addresses the vulnerability of Vision-Language-Action (VLA) models to visual distribution shifts caused by their reliance on spurious correlations. To mitigate this issue, we propose the DILL framework, which employs a task-domain dual encoder and domain-invariant latent lookahead prediction. By integrating contrastive learning with Gaussian decoupling regularization, DILL effectively disentangles task-relevant structures from domain-specific visual variations, thereby facilitating robust policy learning. Experimental results demonstrate that DILL achieves a success rate of 69.1% on the LIBERO-Plus benchmark, outperforming baseline methods by 11.4 percentage points. Furthermore, its generalization capability is validated through real-world robotic manipulation tasks.
📝 Abstract
Vision-Language-Action (VLA) models remain brittle under visual distribution shifts, often relying on spurious correlations tied to domain-specific factors rather than task-relevant structure. We propose Domain-Invariant Latent Lookahead (DILL), a representation-learning framework that mitigates shortcut learning in VLA policies. Our key idea is to supervise policies with domain-invariant future latents learned from domain-transformed trajectory data. A Task-Domain Encoder is trained with contrastive objectives and Gaussian disentanglement regularization to separate task-relevant structure from domain-specific visual variation. The learned encoder then provides future latents for VLA policy learning through lookahead prediction and domain disentanglement, encouraging the policy to focus on task-relevant structure rather than incidental visual factors. Counterfactual task-view evaluations show that DILL reduces shortcut reliance, while LIBERO-Plus evaluations demonstrate improved visual robustness, with 69.1% average success, 11.4 percentage points above the strongest baseline. Real-world manipulation experiments further support DILL's applicability beyond controlled simulation. Complementary latent-space diagnostics show that these behavioral gains are accompanied by representations that better preserve task-consistent structure while suppressing domain-specific variation. Our project page is available at https://dill-vla.github.io/.