LiteMVS: Efficient Multi-View Stereo with Foundation Distillation and Expert Aggregation

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of traditional multi-view stereo methods, which struggle in textureless or repetitive regions, and monocular depth models, which lack geometric constraints necessary for high-quality 3D and temporal reconstruction. The authors propose LiteMVS, a novel approach that integrates plane-sweeping geometric reasoning with monocular semantic and structural priors distilled from vision foundation models to construct a semantics-enhanced cost volume. A mixture-of-experts (MoE) mechanism is further introduced to enable adaptive geometric aggregation. Evaluated on ScanNetv2 and 7-Scenes benchmarks, LiteMVS achieves high-accuracy depth estimation and efficient 3D reconstruction, establishing a robust geometric foundation for 4D representation learning.
📝 Abstract
Real-time 3D perception is crucial for robotics, augmented reality, and embodied intelligence applications. Existing multi-view stereo (MVS) methods primarily rely on geometric correspondences, which often fail in textureless or repetitive regions, while monocular depth models leverage strong image-level priors but lack robust multi-view geometric constraints. More importantly, in robotics and embodied manipulation scenarios, high-quality 3D geometry is not only essential for static reconstruction, but also serves as a critical foundation for learning temporally consistent 4D representations. To obtain visual representations with stronger structural awareness and greater potential for spatiotemporal extension, we present LiteMVS, a lightweight multi-view depth estimation model that integrates plane-sweep geometric reasoning with strong monocular semantic and structural priors. The central idea of LiteMVS is to efficiently inject high-level monocular knowledge, obtained from lightweight segmentation models and large-scale vision foundation models, into a multi-view stereo framework. In particular, LiteMVS enriches the cost volume with semantic descriptors and employs a Mixture-of-Experts (MoE) formulation to enable adaptive geometric aggregation across depth hypotheses. Moreover, geometric priors distilled from vision foundation models further strengthen monocular guidance without increasing inference cost. Through this design, LiteMVS not only improves depth estimation and 3D reconstruction quality in static scenes, but also provides a more reliable geometric foundation for subsequent temporal modeling and 4D representation learning. Experiments on ScanNetv2 and 7-Scenes demonstrate that LiteMVS achieves high-quality depth prediction and 3D reconstruction while maintaining competitive efficiency.
Problem

Research questions and friction points this paper is trying to address.

multi-view stereo
depth estimation
3D reconstruction
geometric constraints
4D representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-View Stereo
Foundation Model Distillation
Mixture-of-Experts
Semantic-Aware Cost Volume
Lightweight 3D Reconstruction
🔎 Similar Papers
No similar papers found.
T
Tianbao Zhang
Shanghai Key Laboratory of Intelligent Sensing and Recognition, Shanghai Jiao Tong University
Z
Zeyu Liu
Shanghai Key Laboratory of Intelligent Sensing and Recognition, Shanghai Jiao Tong University
S
Shuyu Wu
Shanghai Key Laboratory of Intelligent Sensing and Recognition, Shanghai Jiao Tong University
Fanxing Li
Fanxing Li
Thomas M. Clausi Distinguished Professor in Chemical Engineering, North Carolina State University
energyCO2 capture and utilizationchemical loopingcatalysis
Z
Zhaoxin Fan
Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, School of Artificial Intelligence, Beihang University
W
Wenjun Wu
Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, School of Artificial Intelligence, Beihang University
Danping Zou
Danping Zou
Professor, Shanghai Jiao Tong University
Visual SLAMRobotic VisionVision-based navigation