🤖 AI Summary
This study addresses the prohibitive computational costs of feed-forward 3D reconstruction models and the absence of efficient neural pseudo-label generation systems by proposing a lightweight framework based on knowledge distillation. The proposed method compresses a large Transformer teacher model into a compact student model and integrates a high-speed dense pixel point map pipeline with SE(3) pose estimation, enabling low-cost multi-view reconstruction and dataset generation. Remarkably, it achieves a 9.4× model compression and a 7× inference speedup while requiring only 1.6% of the original training compute. Furthermore, this work presents a high-throughput pseudo-labeling alternative to COLMAP that delivers superior accuracy in specific scenarios alongside a 980× throughput improvement, effectively overcoming failure cases inherent to conventional approaches.
📝 Abstract
Feed-forward 3D reconstruction models have achieved impressive performance by scaling model and dataset size, but their cost excludes most research groups and precludes edge deployment. Additionally, generating 3D supervision without sensors still relies on slow, unreliable Structure-from-Motion, as the community lacks a COLMAP-like system for neural 3D pseudo-label generation. We present OTT3R (RGB-Only Tiny Transformer for 3D Reconstruction), a knowledge distillation framework that addresses both problems on a single workstation equipped with 2 GPUs. Distilling $π^3$ (959M parameters) into a 102M-parameter student yields 9.4$\times$ compression and up to 7$\times$ faster inference, trained at 1.6% of VGGT's training compute. An integrated pseudo-label pipeline offers a reliable, high-throughput alternative to COLMAP, generating dense per-pixel point maps and SE(3) camera poses for a 667K-image corpus in 3.5 hours on two commodity GPUs and succeeding on every sequence we tested, including those where COLMAP fails. The general student tracks the teacher on in-distribution monocular depth and, zero-shot, outperforms COLMAP on 7-Scenes and on DTU completion, but it does not replace the teacher on out-of-distribution multi-view geometry. The deployable artifact is the domain-specialized student: after specialization at 0.2% compute, it is 4$\times$ more accurate than COLMAP on 7-Scenes at 980$\times$ throughput, with near-teacher completion. Code is available at https://github.com/TheFourthKaramazov/OTT3R