MixVLA: Adaptive Mixing of Non-Invariant Information for Generalizable Vision-Language-Action Models

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited zero-shot generalization of Vision-Language-Action (VLA) models under out-of-distribution (OOD) conditions by proposing the MixVLA framework. This framework introduces an Adaptive Mixing of non-Invariant information (AMI) technique that regularizes environment-specific variables by fusing invariant and non-invariant features while preserving complementary predictive cues. The core innovation lies in achieving model-agnostic generalization enhancement without requiring additional OOD data or architectural modifications, thereby effectively balancing robustness with policy expressiveness. Experimental results demonstrate that the proposed method significantly improves zero-shot robustness on the LIBERO benchmark and real-world robotic tasks while maintaining strong in-domain performance.
📝 Abstract
Vision-Language-Action (VLA) models have achieved remarkable advances in robotic manipulation, yet their zero-shot generalization under out-of-distribution (OOD) conditions remains limited. These models often entangle task-relevant invariant structure with environment-specific non-invariant factors, causing policies to rely on spurious appearance cues during action prediction. In this work, we propose \textbf{MixVLA}, a model-agnostic training framework that improves the generalization of VLA models without requiring additional OOD data or architectural modifications. The key component of MixVLA is \textbf{Adaptive Mixing of Non-Invariant Information (AMI)}. AMI stochastically mixes non-invariant representations to regularize distribution-specific variability while preserving complementary predictive cues. The mixed non-invariant features are then fused with invariant representations for final action prediction, resulting in improved robustness without sacrificing policy expressiveness. Extensive experiments across challenging manipulation settings, including LIBERO, LIBERO-Plus, the RoboTwin perturbation suite, and real-world tasks, demonstrate that MixVLA improves overall zero-shot robustness while retaining strong in-domain performance.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
zero-shot generalization
out-of-distribution
spurious correlations
robotic manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action Models
Adaptive Mixing
Zero-shot Generalization
Out-of-Distribution Robustness
Non-Invariant Information
🔎 Similar Papers
No similar papers found.