Disentangling Spurious Correlations in Vision-Language-Action Models via Predicting Domain-Invariant Latent Lookahead
This study addresses the vulnerability of Vision-Language-Action (VLA) models to visual distribution shifts caused by their reliance on spurious correlations. To mitigate this issue, we propose the DILL framework, which employs a task-domain dual encoder and domain-invariant latent lookahead prediction. By integrating contrastive learning with Gaussian decoupling regularization, DILL effectively disentangles task-relevant structures from domain-specific visual variations, thereby facilitating robust policy learning. Experimental results demonstrate that DILL achieves a success rate of 69.1% on the LIBERO-Plus benchmark, outperforming baseline methods by 11.4 percentage points. Furthermore, its generalization capability is validated through real-world robotic manipulation tasks.