robust initialization

Designs and implements algorithms that compute a robust coarse geometric and camera-pose initialization for image-based 3D reconstruction and tracking pipelines; this includes methods to recover global scene topology and to bootstrap structure-from-motion or 3D Gaussian-splatting models, producing stable dense-tracking initializations even from degraded inputs such as motion-blurred images.

robustinitialization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.12
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$201K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study investigates the interplay between initialization and densification strategies in 3D Gaussian splatting and their impact on reconstruction quality. Addressing the challenge that existing methods struggle to effectively leverage densely initialized point clouds, we establish the first systematic benchmark to evaluate combinations of diverse initialization sources—including LiDAR scans, multi-view stereo, monocular depth estimates, and sparse structure-from-motion (SfM)—with representative densification protocols. Our experiments reveal that current densification mechanisms yield limited performance gains when initialized with dense point clouds and often fail to surpass reconstructions based on sparse SfM initialization. To foster further research in this direction, the benchmark will be publicly released.

3D Gaussian Splatting3D reconstructiondensification

Better Pose Initialization for Fast and Robust 2D/3D Pelvis Registration

Mar 10, 2025
YS
Yehyun Suh
🏛️ Vanderbilt University | Vanderbilt University Medical Center

Poor initial pose estimation often leads to failure in 2D/3D pelvic registration. To address this, this paper proposes a data-driven coarse pose initialization method that, for the first time, integrates a deep learning–based initialization network into an optimization-based registration pipeline. The method jointly models 2D projection geometry constraints and 3D shape priors, and is trained end-to-end with a conventional optimizer (L-BFGS). Evaluated on multicenter clinical data, the approach achieves a registration success rate of 98.2%, with mean localization error <1.3 mm and orientation error <1.1°, while requiring only 0.8 seconds per inference. It significantly improves robustness, accuracy, and real-time performance—meeting stringent intraoperative requirements.

Enhances computational efficiency in pose estimationEnsures robust registration under extreme pose variationsImproves 2D/3D pelvis registration accuracy

This work addresses the challenge of jointly reconstructing dense dynamic scenes and estimating camera poses from multiple freely moving cameras, overcoming limitations of monocular inputs or reliance on pre-calibrated rigid camera arrays. The authors propose a two-stage optimization framework: the first stage constructs a spatiotemporal connectivity graph that extends visual SLAM to multi-camera settings by integrating temporal continuity and spatial overlap to achieve consistent scale and robust tracking; the second stage jointly optimizes dense depth and camera poses using wide-baseline optical flow. Key innovations include the spatiotemporal graph structure and a wide-baseline initialization strategy, which significantly enhance robustness in low-overlap scenarios. The study also introduces MultiCamRobolab, the first real-world multi-camera dataset with motion-capture ground truth. Experiments demonstrate superior performance over existing feedforward models on both synthetic and real data, with reduced memory consumption.

camera pose estimationdense dynamic scene reconstructionfreely moving cameras

Self-Supervised Geometry-Guided Initialization for Robust Monocular Visual Odometry

Jun 03, 2024
TK
Takayuki Kanai
🏛️ Toyota Motor Corporation | Toyota Research Institute

Monocular visual odometry suffers from insufficient robustness under large motions, dynamic scenes, and challenging illumination conditions. This paper proposes a geometry-guided self-supervised initialization method that, for the first time, integrates a frozen pre-trained monocular depth model (e.g., MiDaS) as a geometric prior to initialize depth and pose for dense bundle adjustment—without fine-tuning the SLAM backbone network. Leveraging self-supervised learning and explicit geometric constraints, the approach effectively suppresses motion blur artifacts and mitigates interference from dynamic objects. Evaluated on KITTI and the challenging DDAD benchmarks, the method achieves significant improvements in pose accuracy, particularly enabling stable tracking under rapid ego-motion and outdoor dynamic scenarios. Results demonstrate both the efficacy of incorporating frozen depth priors directly into the optimization pipeline and the strong generalization capability of the proposed framework.

Addresses failures in monocular visual odometry under poor conditionsEnhances dense bundle adjustment initialization using self-supervised priorsImproves robustness against large motion and object dynamics

On-the-fly Reconstruction for Large-Scale Novel View Synthesis from Unposed Images

Jun 05, 2025
AM
Andréas Meuleman
🏛️ Inria | Université Côte d’Azur | TU Wien

Existing methods for novel view synthesis from large-scale unstructured image sequences suffer from high computational latency, memory bottlenecks, and SLAM failure under wide-baseline or large-scale scenarios—particularly during camera pose estimation and 3D Gaussian optimization. This paper introduces the first real-time, online dynamic Gaussian radiance field framework that enables simultaneous capture and reconstruction. Our method jointly optimizes camera poses and the radiance field via learning-based fast initial pose estimation and a GPU-efficient miniature bundle adjustment. We propose a novel dynamic incremental Gaussian primitive generation scheme coupled with anchor-point clustering and offloading to alleviate memory and computational constraints. Furthermore, direct Gaussian sampling and progressive anchor storage enhance rendering efficiency. Evaluated across diverse datasets, our approach achieves “reconstruction-on-the-fly,” significantly outperforming offline methods in processing speed while matching state-of-the-art rendering quality—fully supporting dense, wide-baseline, and ultra-large-scale scenes.

Efficient pose estimation and 3DGS optimization for wide-baseline capturesFast on-the-fly reconstruction for large-scale novel view synthesisScalable radiance field construction for handling large scenes incrementally

Latest Papers

What's happening recently
View more

This work addresses the failure of conventional Structure-from-Motion (SfM) methods under extreme motion blur by proposing PRISM3D, a unified framework that, for the first time, leverages deep dense tracking for robust initialization of 3D Gaussian splatting. Integrating a physics-based imaging model grounded in Bézier trajectories with a Markov Chain Monte Carlo (MCMC)-based probabilistic densification strategy, PRISM3D enables physically consistent blind 3D reconstruction. Furthermore, the authors introduce PRISM3D-E, the first multimodal reconstruction framework fusing RGB and event camera data, along with a dedicated benchmark dataset. Experiments demonstrate that the proposed approach significantly outperforms existing methods under extreme blur conditions, with both its single-modality and multimodal variants achieving state-of-the-art performance.

3D scene reconstructionblind inverse problemextreme degradation

This work addresses the challenge of efficiently and robustly reconstructing high-fidelity geometric structures in dense monocular SLAM without relying on appearance modeling. To this end, the authors propose GeoGS, a purely geometry-driven 3D Gaussian Splatting representation that discards appearance parameters and retains only spatial geometric information. The method introduces a local plane-guided Gaussian initialization strategy to accelerate convergence and incorporates a global map alignment mechanism during loop closure to effectively prevent map tearing. Experimental results demonstrate that GeoGS consistently outperforms existing approaches on both synthetic and real-world datasets, achieving state-of-the-art performance in terms of online mapping efficiency and geometric reconstruction accuracy.

3D Gaussian Splattingdense visual SLAMgeometry-only reconstruction

This work addresses the challenge of balancing reconstruction accuracy and efficiency in event-based 3D reconstruction without initial pose estimates. We propose EvTrajGS, a novel framework that integrates continuous-time trajectory modeling with 3D Gaussian Splatting, enabling temporally consistent, high-fidelity scene reconstruction through joint optimization of camera poses and geometry. The method introduces a temporally coupled pose representation and an adaptive event sampling strategy with loss reweighting, which mitigates drift accumulation while avoiding the high computational overhead of traditional SLAM pipelines. Experiments on both synthetic and real-world datasets demonstrate substantial improvements over state-of-the-art approaches, achieving a 3.8 dB gain in PSNR, a 0.1 increase in SSIM, and over 40% reduction in ATE RMSE, all while maintaining computational efficiency.

3D reconstructionevent camerasGaussian Splatting

Hot Scholars

XC

Xieyuanli Chen

Associate Professor, NUDT, China
RoboticsSLAMLocalizationLiDAR Perception
WW

Wenping Wang

Texas A&M University
Computer GraphicsGeometric Computing
HC

Hongyang Chen

SUN YAT-SEN UNIVERSITY
SDNCloud ComputingMicroserviceAIOps
XL

Xiaobing Li

University of Wisconsin-Madison SUNY College of Optometry
saccade attention decision making