🤖 AI Summary
This study addresses the premature loss of critical visual tokens caused by fixed-layer pruning in Vision-Language-Action (VLA) models by proposing SAPrune, a training-free dynamic pruning framework. To our knowledge, this work is the first to introduce a stage-aware mechanism based on action-visual attention evolution, which dynamically determines the optimal pruning layer by analyzing attention patterns over a calibration set. Furthermore, a dual-path rule is designed to jointly preserve essential tokens and contextual information. Experimental results demonstrate that SAPrune can prune 87.5% of visual tokens on benchmarks such as LIBERO, achieving up to 1.718× inference acceleration while maintaining high task success rates. These findings indicate that the proposed method effectively balances inference efficiency and performance for VLA models.
📝 Abstract
Visual token pruning is an effective way to accelerate vision-language models and is especially useful for vision-language-action (VLA) inference, where many visual tokens must be processed before predicting robot actions. Existing pruning methods usually estimate which tokens can be pruned based on attention scores or feature diversity, retaining tokens that are either highly attended or visually different from others. However, most of them use fixed pruning schedules, such as pruning once at a preset layer or pruning at uniformly spaced layers. Such schedules can be risky for VLA models, because the model may not know which visual regions matter for the action in early layers. Tokens that look unimportant at first may become useful after the model combines visual observations with the language instruction. In this work, we propose SAPrune, a training-free visual token pruning framework for efficient VLA inference. Instead of pruning at fixed layers, SAPrune uses a small calibration set to observe how action-to-visual attention changes across layers, and chooses pruning layers only after the attention pattern becomes more reliable. At each selected layer, SAPrune applies a dual-path pruning rule: one path protects strongly attended visual tokens from pruning, while the other prevents useful surrounding context from being discarded. Experiments on LIBERO, SIMPLER, and real-world robotic tasks show that SAPrune prunes 87.5% of visual tokens and achieves up to 1.718x inference speedup while maintaining competitive task success rates.