🤖 AI Summary
This study addresses the absence of visual token influence modulation mechanisms in Vision-Language-Action (VLA) policies by proposing a Dual-Space Intent-Aware Visual Decay module. This method introduces a novel "anchor-then-decay" paradigm that estimates relevance anchors by fusing task intent with visual evidence, applying dual weighting and state-persistent decay at both the entrance and interior of the backbone network. This achieves complementary regulation across dual spaces without requiring additional grounding supervision. Experimental results demonstrate that the proposed module attains an average success rate of 98.0% on the LIBERO benchmark, elevates the zero-shot score to 72.6, and significantly enhances the model's robustness against interference.
📝 Abstract
Vision-language-action (VLA) policies typically feed dense visual patch tokens into a language-action backbone, preserving scene context but offering no explicit mechanism to regulate how strongly different visual tokens influence policy computation. We introduce DIVA, a Dual-Space Intent-Aware Visual Attenuation module with an anchor-then-attenuate design. DIVA combines high-level task intent with low-level visual evidence to estimate patch-wise relevance anchors, then applies them in two complementary spaces: it reweights projected visual tokens before backbone entry and persistently attenuates low-relevance visual states within the backbone. DIVA preserves the full visual token sequence and requires no external grounding supervision. On LIBERO, DIVA improves OpenVLA-OFT from 96.6% to 98.0% average success and raises its zero-shot LIBERO-Plus score from 69.6 to 72.6. Real-world experiments further show consistent gains under task-irrelevant visual perturbations, supporting the robustness of intent-aware visual attenuation beyond simulation.