Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of Vision-Language-Action (VLA) models to visual shortcuts, which cause them to disregard instructions and struggle to generalize to unseen skill compositions. To overcome this limitation, this work proposes the CRAFT framework, which leverages counterfactual data augmentation to eliminate spurious visual correlations. Furthermore, it introduces a novel skill-representation-based cross-execution supervision transfer mechanism that facilitates effective skill decoupling and reliable reuse. Extensive evaluations across multiple simulation benchmarks and real-world robotic platforms demonstrate that the proposed approach significantly improves execution success rates for undemonstrated skill compositions while preserving high performance on demonstrated tasks. Consequently, this method effectively resolves the combinatorial generalization challenge inherent in VLA models.
📝 Abstract
Vision-language-action (VLA) models often struggle to generalize to skill combinations absent from their fine-tuning demonstrations, even when every constituent skill has been demonstrated. We focus on a vision shortcut as one failure mode: during fine-tuning, visual observations can serve as a proxy for the instruction, so a policy may execute a demonstrated combination associated with similar observations rather than the instructed combination. This motivates training with counterfactual pairs formed by holding a demonstration observation fixed while changing the instruction to specify an undemonstrated combination. These pairs, however, lack corresponding demonstrated action targets. Crucially, the currently required skill has already been demonstrated, but actions from those executions cannot serve as direct targets because the same skill can require different actions across observations. We propose CRAFT, which transfers supervision from demonstrated executions of the required skill to counterfactual pairs using skill representations that can be reused across executions of the same skill. Across three VLA models and two simulation benchmarks, CRAFT improves success on undemonstrated combinations while maintaining high success on demonstrated ones; it also improves compositional generalization on a real robot. Project website: https://taegeunyang.github.io/craft/
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action Models
Compositional Generalization
Skill Alignment
Visual Shortcut
Counterfactual Pairs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Compositional Generalization
Vision-Language-Action Models
Skill Alignment
Counterfactual Training
Visual Shortcut
💼 Related Jobs
No related jobs found.
T
Taegeun Yang
Korea Advanced Institute of Science and Technology (KAIST)
Y
Youngju Na
Korea Advanced Institute of Science and Technology (KAIST)
Y
Yoonki Cho
Korea Advanced Institute of Science and Technology (KAIST)
Sung-Eui Yoon
Sung-Eui Yoon
Professor of Dept. of Computer Science, KAIST
GraphicsVisionRobotics