🤖 AI Summary
This study addresses the failure of compositional generalization in Vision-Language-Action (VLA) models, which tend to rely on visual shortcuts due to insufficient diversity in training data. To overcome this limitation, we propose ReGuide, a plug-and-play wrapper that requires neither retraining nor modifications to the backbone network. By leveraging semantic and geometric rebinding mechanisms, ReGuide guides the executor toward supporting configurations to recover local skills, thereby effectively resolving the challenge of recombining unseen elements. Extensive evaluations in both simulation and real-world experiments demonstrate that our method improves task success rates by up to 75% under compositional shifts, while preserving performance on standard tasks without degradation.
📝 Abstract
While Vision-Language-Action (VLA) models enable flexible action generation, their generalization across diverse environmental elements, including manipulated objects, destinations, and backgrounds, is limited by the lack of diversity in robotic training data. Trained end-to-end on such data, VLAs tend to exploit visual shortcuts, associating actions with task-irrelevant visual features rather than the intended task semantics. These shortcuts block recomposition of elements already seen by the policy, that is, compositional generalization. Existing approaches mitigate such entanglement through task-relevant perception or targeted data diversification, but offer no explicit mechanism for unseen recomposition and require backbone-specific modifications with retraining. We observe that under such recomposition, VLAs often fail at global grounding while retaining local manipulation skills that recover near the correct target in familiar configurations. Therefore, we propose Referential Guidance (ReGuide), a training-free wrapper that, given object poses from a grounding module, combines semantic and geometric rebinding to guide the end-effector into demonstration-supported configurations of the instructed referent, where the frozen policy can resume execution. Experiments in simulation across multiple VLA backbones as well as on a real robot show that ReGuide improves success rates under compositional shifts by up to 56.8 and 75.0 percentage points, respectively, while preserving standard-task performance.