🤖 AI Summary
This study addresses the high inference costs of Vision-Language-Action (VLA) models and the underutilization of text-vision synergistic information in existing caching mechanisms by proposing TVCache, a training-free acceleration framework. Specifically, TVCache eliminates redundant computations through attention head filtering and leverages entropy difference to guide layer reuse selection, thereby enhancing cache stability while optimizing visual grounding and cache resource allocation. Experimental results demonstrate that TVCache improves task success rates by 14.5% across multiple benchmarks while reducing FLOPs by 2.45×, achieving a significant balance between performance and efficiency.
📝 Abstract
Vision-Language-Action (VLA) models enable generalizable robotic control but remain computationally expensive. Token caching provides a training-free, plug-and-play acceleration alternative. However, existing VLA caching does not fully exploit a key inductive bias of VLA models: text-vision synergy, wherein textual semantics guide the precise visual grounding of task-relevant regions. In particular, existing designs insufficiently account for head-wise reliability in attention aggregation and layer-wise stability in cache reuse. To address this, we propose Text-Vision Synergistic Token Caching (TVCache), a training-free framework for efficient VLA inference. TVCache filters attention heads based on text-vision information focus to improve task-relevant and physically consistent visual grounding. Concurrently, we introduce a reuse-layer selection mechanism guided by text-vision entropy differences to avoid caching unstable representations and improve cache resource allocation. Extensive experiments across four representative VLA models, two simulation benchmarks, and real-world robotic tasks demonstrate the effectiveness and generality of TVCache. At matched token-retention ratios, TVCache consistently improves task success over existing VLA caching with comparable computational cost. On OpenVLA-OFT, it improves average success by up to 14.5 percentage points over VLA-Cache at 12.5% retention while reducing FLOPs by 2.45x relative to full-token inference.