Beyond Token Importance: Preserving Spatial Scaffolds for Efficient Vision-Language-Action Inference

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing Vision-Language-Action (VLA) pruning methods, which constrain robotic manipulation success rates by neglecting spatial information. We propose GeoScaffold, a tuning-free visual token pruning framework that transcends reliance on semantic importance alone by revealing a strong correlation between spatial coverage radius and task success. Through spatial region partitioning, task-relevance budget allocation, and farthest point sampling, GeoScaffold effectively preserves the spatial structure of visual scenes. Evaluated on the LIBERO benchmark, the proposed method achieves a 93.2% success rate while retaining only 20% of visual tokens, alongside a 1.78× prefill acceleration.
📝 Abstract
Existing VLA pruning strategies primarily select individual visual tokens according to task-level semantic relevance, while overlooking the spatial information required for robotic manipulation. To examine this limitation, we construct a simple Stride baseline that uniformly samples tokens along the flattened one-dimensional visual sequence, representing a purely geometric pruning strategy. Surprisingly, Stride outperforms semantic pruning and random pruning at certain pruning ratios, but collapses when the token budget is only slightly reduced. We characterize this phenomenon through the spatial coverage radius, defined as the largest spatial blind spot induced by the retained token set after pruning. Our analysis reveals a strong correlation between the spatial structure of retained tokens and task success, suggesting that reliable VLA pruning requires preserving not only task-relevant tokens but also the spatial scaffold of the scene. Motivated by this diagnosis, we propose GeoScaffold, a training-free visual token pruning method that partitions each image into spatial regions, allocates inter-region token budgets using task-relevance weights, and selects intra-region scaffold tokens via farthest point sampling to reduce the local coverage radius. On pi 0.5 and LIBERO, GeoScaffold retains only 20% of visual tokens while preserving a 93.2% average success rate, and achieves a 1.78 times prefill speedup over the unpruned baseline.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
visual token pruning
spatial scaffold
robotic manipulation
spatial coverage
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action (VLA)
Visual Token Pruning
Spatial Scaffold
Training-free
Farthest Point Sampling
🔎 Similar Papers