Review of Feed-forward 3D Reconstruction: From DUSt3R to VGGT

📅 2025-07-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Traditional SfM/MVS methods suffer from complex pipelines, high computational cost, and poor robustness in textureless regions. To address these limitations, this paper systematically reviews feedforward, single-pass 3D reconstruction techniques—particularly end-to-end models such as DUSt3R—and proposes a Transformer-based paradigm for joint pose and geometry estimation, eliminating iterative optimization and unifying multi-view geometry and camera pose estimation. The method supports arbitrary numbers of input images, ensuring strong generalization and flexibility. By integrating correspondence modeling, multi-view feature fusion, and joint regression, it achieves superior efficiency and robustness over both classical approaches and learning-based methods (e.g., MVSNet) on standard benchmarks. We further survey prevailing architectural frameworks, benchmark datasets, and evaluation metrics, and identify key future directions—including dynamic scene modeling and model scalability.

Technology Category

Computer Vision: 3D Computer VisionIntelligent Robots: Multimodal Perception & Sensor FusionMachine Learning: Multi-instance/Multi-view Learning

Application Category

Graph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphsUser Modeling, Personalization and Recommendation: Federated recommendation systems and personalizationSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAG
📝 Abstract
3D reconstruction, which aims to recover the dense three-dimensional structure of a scene, is a cornerstone technology for numerous applications, including augmented/virtual reality, autonomous driving, and robotics. While traditional pipelines like Structure from Motion (SfM) and Multi-View Stereo (MVS) achieve high precision through iterative optimization, they are limited by complex workflows, high computational cost, and poor robustness in challenging scenarios like texture-less regions. Recently, deep learning has catalyzed a paradigm shift in 3D reconstruction. A new family of models, exemplified by DUSt3R, has pioneered a feed-forward approach. These models employ a unified deep network to jointly infer camera poses and dense geometry directly from an Unconstrained set of images in a single forward pass. This survey provides a systematic review of this emerging domain. We begin by dissecting the technical framework of these feed-forward models, including their Transformer-based correspondence modeling, joint pose and geometry regression mechanisms, and strategies for scaling from two-view to multi-view scenarios. To highlight the disruptive nature of this new paradigm, we contrast it with both traditional pipelines and earlier learning-based methods like MVSNet. Furthermore, we provide an overview of relevant datasets and evaluation metrics. Finally, we discuss the technology's broad application prospects and identify key future challenges and opportunities, such as model accuracy and scalability, and handling dynamic scenes.
Problem

Research questions and friction points this paper is trying to address.

Overcoming limitations of traditional 3D reconstruction methods like SfM/MVS
Exploring feed-forward deep learning models for joint pose-geometry inference
Addressing challenges in accuracy, scalability, and dynamic scene handling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Feed-forward deep network for 3D reconstruction
Joint camera pose and geometry regression
Transformer-based correspondence modeling
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
W
Wei Zhang
School of Artificial Intelligence, Optics and Electronics (iOPEN), Northwestern Polytechnical University, Xi’an 710072, China
Y
Yihang Wu
School of Artificial Intelligence, Optics and Electronics (iOPEN), Northwestern Polytechnical University, Xi’an 710072, China
S
Songhua Li
School of Artificial Intelligence, Optics and Electronics (iOPEN), Northwestern Polytechnical University, Xi’an 710072, China
W
Wenjie Ma
School of Artificial Intelligence, Optics and Electronics (iOPEN), Northwestern Polytechnical University, Xi’an 710072, China
X
Xin Ma
School of Artificial Intelligence, Optics and Electronics (iOPEN), Northwestern Polytechnical University, Xi’an 710072, China
Q
Qiang Li
School of Artificial Intelligence, Optics and Electronics (iOPEN), Northwestern Polytechnical University, Xi’an 710072, China
Q
Qi Wang
School of Artificial Intelligence, Optics and Electronics (iOPEN), Northwestern Polytechnical University, Xi’an 710072, China