π€ AI Summary
This study addresses the latency bottleneck in the global attention mechanism of the VGGT model caused by long sequences, proposing a hardware-aware block-clustered attention method. By incorporating hash hyperplane calibration and threshold error compensation techniques, the proposed approach effectively reduces clustering errors while optimizing memory access patterns and computational overhead, thereby accelerating 3D scene reconstruction. Experimental results demonstrate that this method achieves a 2.1β2.9Γ speedup for the global attention module and a 1.8β2.6Γ acceleration for the overall backbone network, with a performance degradation of less than 5%. These findings indicate that the proposed technique successfully establishes a favorable balance between computational efficiency and reconstruction accuracy.
π Abstract
The Visual Geometry Grounded Transformer (VGGT) marks a significant leap forward in 3D scene reconstruction, as it is the first model that directly infers all key 3D attributes (camera poses, depths, and dense geometry) jointly in one pass. However, this joint inference mechanism requires global attention layers with extremely long sequences that causes a significant latency bottleneck. In this paper, we propose blockwise clustered attention (BC attention) to accelerate the global attention layers in VGGT. By limiting the clustering within HW-friendly neighborhood blocks, BC attention reduces the computation overhead of query clustering as well as the costly data movement between on- and off-chip memory. This enables BC attention to scale to long sequences and deliver practical latency improvements on GPUs. Moreover, we introduce a hashing hyperplane calibration method and a threshold-based error compensation method to reduce clustering errors efficiently, which is a bottleneck in the current clustered attention mechanism. Overall, our experiments on GPU demonstrate that calibrated BC attention accelerates the global attention layers by 2.10-2.63$\times$ and the whole backbone by 1.77-2.35$\times$ with negligible loss (1%) for large scenes. With a small performance loss (< 5%), calibrated BC attention further achieves a 2.26-2.87$\times$ latency improvement on the global attention layers and a 1.90-2.55$\times$ improvement on the backbone.