🤖 AI Summary
This study addresses the high computational overhead of global attention in feed-forward 3D vision models and the tendency of existing sparse methods to select redundant regions. We propose a training-free sparse attention framework featuring a value-aware block selection mechanism that integrates query-key correlation with neighborhood value contrast to precisely eliminate redundancy. Additionally, an execution-aware cross-layer memory mechanism is introduced to dynamically track unfulfilled reconstruction demands and enable competitive optimization across layers. Experiments demonstrate that our method significantly improves pose estimation accuracy and reconstruction quality on benchmarks such as 7Scenes. Furthermore, it achieves a 2.29× inference speedup over dense attention models, establishing a superior balance between computational efficiency and performance.
📝 Abstract
Feed-forward 3D vision models such as VGGT have achieved remarkable progress, unifying camera estimation and dense scene reconstruction in a single pass. However, their quadratic global attention makes long image sequences expensive, while existing sparse methods may favor highly attended yet value-redundant regions. To address these limitations, we introduce VASC, a training-free sparse attention method combining value-aware block selection and execution-aware cross-layer memory. Our value-aware block selection integrates pooled query--key relevance with neighboring value contrast, reducing redundancy while preserving query-relevant and distinctive content. Cross-layer memory tracks unserved demand across layers and updates this state according to actual execution, enabling previously underserved blocks to compete under a fixed computation budget. Experiments on 7Scenes and NeuralRGB-D with VGGT and $π^3$ demonstrate improved pose estimation and reconstruction quality compared with FasterVGGT, together with up to $2.29\times$ faster inference than dense VGGT. Code is available at https://github.com/kosakayamahoo-design/VASC.