PerturBot: Breaking Shortcut Priors in Vision-Language-Action Models with Perturbative Training

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Vision-Language-Action (VLA) models are prone to exploiting modality shortcuts, thereby neglecting genuine task-relevant evidence. This work proposes the GroundingFscore metric to precisely diagnose such shortcut reliance while decoupling training from evaluation. Furthermore, it introduces perturbation-based training strategies—including wrist-view perturbation, instruction enrichment, and failure trajectory relabeling—to disrupt shortcut priors and compel the model to attend to task-critical information. These methods effectively strengthen the VLA's reliance on authentic task evidence, significantly enhancing policy generalization and scaling capabilities. Ultimately, this approach ensures that robotic decision-making is grounded in verifiable task evidence rather than spurious correlations.
📝 Abstract
A vision--language--action (VLA) policy can complete complex tasks while ignoring the evidence that should determine its actions. An object held near the wrist camera can displace the instructed target. Language and action show the same pattern: a familiar noun can trigger the operation it was paired with in training even after the verb changes, and a gripper that closed on nothing may lift anyway. We call these dependencies modality shortcuts: regularities in successful demonstrations make visual, lexical, or motor cues sufficient to predict expert actions without the task evidence needed for the underlying decision. More demonstrations of the same kind can raise task success while leaving these shortcuts intact. We propose Perturbot which makes task-relevant evidence easier to use and shortcuts insufficient on their own: it applies task-preserving wrist-view perturbations, enriches instructions with decision-relevant captions, and adds random and failed trajectory segments relabeled with the behavior they contain. It complements scaling by changing what is scaled, and leaves inference unchanged. Moreover, we propose GroundingFscore, an offline score that diagnoses how severely a policy relies on modality shortcuts. Task success rate shows whether a policy improves, while GroundingFscore reveals whether the policy scales healthily, relying on task evidence rather than shortcuts. Together, Perturbot and GroundingFscore provide a training-and-evaluation framework for disentangling VLA decisions from shortcut priors while preserving responsiveness to task-relevant evidence.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action Models
Modality Shortcuts
Shortcut Priors
Robotic Policy Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action Models
Shortcut Priors
Perturbative Training
GroundingFscore
Modality Shortcuts
🔎 Similar Papers