Does the VGGT Family Need All Its Layers?

šŸ“… 2026-09-29
šŸ“ˆ Citations: 0
✨ Influential: 0
šŸ“„ PDF
šŸ¤– AI Summary
This study addresses layer redundancy in feedforward geometry models by exploring the minimal number of layers required to preserve camera poses and dense 3D structures. Methodologically, it employs Centered Kernel Alignment (CKA) as a proxy metric to accelerate the search, revealing a distinctive early- and late-layer dual redundancy pattern in VGGT-series models that reduces pruning complexity from O(L⁓) to O(L²). Furthermore, a retraining-free closed-form linear calibration strategy is proposed to recover accuracy. Experiments demonstrate that this approach reduces aggregator parameters by 44% while maintaining performance, requiring only 100 scenes for calibration and generalizing effectively to unseen datasets.
šŸ“ Abstract
Which layers of a feed-forward geometry model are needed to preserve both camera poses and dense 3D structure? We study layer redundancy in VGGT, $Ļ€^3$, and VGGT-$Ī©$: 3,018 pruned configurations, scored on seven camera-pose and dense-geometry metrics across four indoor and outdoor datasets. Four findings follow: (i) Removable layers cluster in two redundancy regions: a dominant early region and a narrower late one, while deletions spanning the intervening layers are consistently more disruptive. This recurring pattern holds across models, datasets, and metrics, and contrasts with the middle-to-late redundancy commonly reported in the literature. (ii) Within these regions, we observe that the joint degradation from deleting two intervals is approximately the sum of their individual degradations, reducing the number of model evaluations for pruning search from $O(L^4)$ to $O(L^2)$, where $L$ is the aggregator depth. (iii) We find that CKA provides a cheaper representation-based proxy for interval degradation, offering a practical trade-off between pruning quality and calibration cost. (iv) Closed-form linear calibration recovers accuracy after pruning without end-to-end retraining. A least-squares analysis shows that using a shared map for special and patch tokens generally incurs excess reconstruction loss, motivating token-aware recovery. Recovery maps fitted on just 100 calibration scenes generalize to held-out scenes and unseen datasets. The resulting models reduce aggregator parameters by up to 44% while maintaining accuracy comparable to their intact counterparts. Code and experimental results will be available at our project page: https://xian-bei.github.io/vggt-family-layer-redundancy/
Problem

Research questions and friction points this paper is trying to address.

layer redundancy
feed-forward geometry model
camera pose estimation
dense 3D reconstruction
model pruning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Layer Pruning
Feed-forward Geometry Models
Complexity Reduction
CKA Proxy
Linear Calibration
šŸ”Ž Similar Papers
No similar papers found.