🤖 AI Summary
This study addresses the drift problem arising from the exclusion of multi-view geometry in optimization when integrating feedforward 3D models with online SLAM. To overcome this, we propose directly converting feedforward predictions into persistent factor graph constraints. The method employs a dual-stream architecture operating at high and low frequencies to maintain tracking while refreshing measurements. For the first time, this enables multi-view priors to participate actively in dense bundle adjustment rather than serving merely as post-hoc alignment, thereby achieving unified optimization of poses and depth. Experimental results demonstrate that the proposed approach significantly improves both trajectory estimation and reconstruction accuracy across multiple benchmarks, reducing the ATE RMSE to 0.002 meters on the uncalibrated Replica dataset.
📝 Abstract
Feed-forward 3D models provide strong multi-view geometric priors, while on- line simultaneous localization and mapping (SLAM) relies mainly on local mea- surements and can accumulate drift over long sequences. Existing attempts to combine the two typically treat feed-forward predictions as an external geomet- ric state that is aligned or fused with the online estimate after the fact, which keeps broader multi-view evidence outside the optimizer that refines the SLAM state. We present F2SLAM, which instead converts feed-forward geometry di- rectly into optimization-native target-weight measurements attached to a persis- tent dense factor graph. A high-frequency stream maintains local tracking con- straints and graph connectivity, while a low-frequency stream uses wider multi- view context to selectively refresh existing measurements after a state-consistency check. Both streams constrain the same poses, inverse depths, and optional cam- era intrinsics through a single dense bundle adjustment. Experiments on multiple benchmarks demonstrate consistently strong trajectory estimation and improved dense reconstruction in both calibrated and uncalibrated settings. Notably, the uncalibrated configuration reduces the average ATE RMSE from 0.030 m for the strongest feed-forward baseline to 0.002 m on the Replica dataset.