🤖 AI Summary
This study addresses the insufficient task conditioning in Vision-Language-Action (VLA) models operating in cluttered scenes due to their reliance on text prompts alone. We propose an enhanced VLA framework that integrates electromyography (EMG) signals with visual segmentation annotations, innovatively introducing non-linguistic modalities for continuous conditioning to overcome the limitations of conventional text-only instructions. Specifically, eight-channel EMG envelopes are concatenated with proprioceptive vectors, while image segmentation annotations are injected to supplement multimodal task information. Experimental results demonstrate that the proposed model significantly outperforms baseline methods in cluttered out-of-distribution scenarios, validating the potential of beyond-language task conditioning for robotic manipulation.
📝 Abstract
Vision-Language-Action (VLA) models rely strongly on language for describing task information, despite having multimodal inputs. We hypothesize that other modalities in the state space may present opportunities for supplemental task conditioning, which may be particularly relevant in cluttered or otherwise ambiguous scenes. We introduce two tuned models to test this hypothesis: (1) an electrophysiology-conditioned VLA (EC-VLA) that incorporates 8-channel electromyography envelopes as continuous conditioning input concatenated to the proprioceptive vector, and (2) a visually-annotated VLA (VA-VLA) that incorporates visual segmentation annotations to the image inputs. On a cube-selection task evaluated across three participants, EC-VLA matches a language-prompted baseline in uncluttered, in-distribution conditions and substantially outperforms it in cluttered, out-of-distribution scenes. Similarly, VA-VLA shows modest improvements over a language-prompted baseline in in-distribution scenes with substantial improvement in cluttered, out-of-distribution trials. Together, these results provide strong evidence for the potential benefit of task-conditioning beyond language.