Score
Designs and implements algorithms and systems that estimate sparse camera or sensor poses and 3D landmarks from visual inputs while explicitly modeling measurement and state uncertainty, producing per-state and per-landmark uncertainty estimates. Analyzes and propagates those uncertainties consistently through localization and mapping using probabilistic estimation (filtering, smoothing or factor-graph formulations) and engineers sparse representations and computational optimizations to meet limited-compute, near-real-time requirements.
Existing SLAM methods struggle to provide provable inclusion guarantees for pose and map uncertainties in safety-critical scenarios. This work proposes the first provably correct uncertainty quantification framework tailored for 3D–3D landmark-based SLAM, delivering uncertainty sets in polyhedral form with deterministic inclusion guarantees through three core modules: forward mapping, backward pose tracking, and pose composition. The approach offers, for the first time, theoretically rigorous guarantees over the entire SLAM pipeline while integrating conformal prediction to enable data-driven probabilistic calibration—thus balancing strong theoretical assurances with practical utility. Extensive simulations and real-world experiments validate the method’s effectiveness, and the implementation is publicly released.
To address the high uncertainty in monocular visual localization and mapping during spacecraft proximity operations, this paper proposes an autonomous environmental modeling and self-localization method integrating active perception with factor graph smoothing-based SLAM. We introduce information-entropy-driven active perception into the spacecraft SLAM framework for the first time, enabling joint optimization of observation planning and state estimation through online camera pose optimization. The method models the joint posterior distribution of the state trajectory and map landmarks using a factor graph, and dynamically selects observation strategies that maximize information gain based on information-theoretic metrics. Numerical simulations demonstrate that, compared to passive perception, the proposed approach significantly improves pose and map accuracy, accelerating state uncertainty convergence by approximately 40%. These results validate the effectiveness and engineering applicability of the closed-loop “perception–planning–estimation” coordination mechanism.
In visual localization, devices struggle to accurately match landmarks and estimate their positions when wireless signals are unavailable or unreliable and nearby landmarks exhibit high visual similarity. This paper proposes a purely vision-aided localization method leveraging multi-source distance measurements and geometric constraints. First, landmarks are modeled as a marked Poisson point process (PPP); we theoretically prove that only three noise-free distance measurements suffice to uniquely identify the landmark set in 2D. Second, we construct a joint distribution model for key random variables and derive a closed-form analytical expression for the landmark-set identification probability under noisy measurements. Experiments demonstrate that, even when landmarks are visually indistinguishable, the method achieves robust identification and precise localization with high confidence.
This work addresses the challenge in factor graph–based state estimation where local optimization methods often converge to suboptimal solutions, while existing globally optimal approaches based on convex relaxation are hindered by modeling complexity and high computational cost. To overcome these limitations, the paper proposes an efficient method for globally optimal estimation that automatically constructs a semidefinite programming (SDP)–based convex relaxation within general factor graphs. The approach innovatively leverages the Bayes tree structure from the GTSAM framework together with chordal sparsity to decompose and solve the SDP problem efficiently. Experimental results on 3D pose-graph SLAM and 2D localization benchmarks demonstrate that the proposed method achieves global optimality while significantly outperforming conventional local solvers in both scalability and computational efficiency.
This paper addresses the relative pose estimation problem for three calibrated cameras given only four correspondences across all views. To overcome limitations of conventional methods—namely, their reliance on more correspondences or insufficient robustness—we propose a novel strategy that approximates a fifth correspondence using the centroid of the four observed points. We further introduce the first joint three-view pose estimation framework integrating a 4-point affine fundamental matrix solver, a standard 5-point relative pose solver, and a P3P solver. Geometric modeling enhances robustness against noise and outliers, while local optimization refines accuracy. Evaluated on real-world datasets, our method achieves state-of-the-art performance: the centroid-based strategy significantly outperforms pure affine approaches, striking a superior balance among accuracy, robustness, and computational efficiency, with straightforward implementation.
This work addresses the lack of spatial uncertainty modeling in YOLO-Pose for keypoint localization. The authors propose a lightweight, post-hoc probabilistic extension that introduces an additional probability head to predict input-dependent 2×2 covariance matrices, enabling calibrated bivariate Gaussian or Student-t distributions over original keypoints. This is the first approach to equip YOLO-Pose with keypoint-level predictive distributions. A novel evaluation protocol is introduced, combining distribution calibration diagnostics with Average Keypoint Precision (AKP). Experiments on COCO demonstrate that the method effectively supports reliability-based keypoint ranking, with the Student-t formulation yielding more accurate residual distribution fitting. Furthermore, in an aircraft visual landing task, the calibrated covariance enables uncertainty-aware pose estimation and sensor fusion.
为解决快速摄像机运动和复杂动态环境下的3DGS-SLAM问题,提出SCOUT-SLAM框架,通过共享基础网络估计双不确定性,提高相机跟踪精度和静态场景重建质量。
该研究解决了3D场景图中不确定性表示和传播问题,通过引入概率场景图(PSG)及高斯层次图(HGG),实现了实时感知与精确定位。
This work addresses the degraded navigation performance of vision-only mobile robots in unknown environments caused by motion blur, low texture, and illumination changes, as well as the insufficient perception-action coupling and inconsistent uncertainty propagation in existing approaches. To this end, we propose UNSEEN—a unified uncertainty-aware co-navigation framework that leverages only a forward-facing monocular camera. Within a receding horizon, UNSEEN tightly couples localization, mapping, and planning modules, achieving, for the first time, end-to-end joint optimization of all three components. A consistent uncertainty propagation mechanism enhances both robustness and estimation accuracy. The system operates in real time at 6 Hz, supporting sparse mapping, pose and uncertainty estimation, and trajectory planning. Evaluated in both simulated and real-world unknown environments, it achieves 100% task success rate, reduces absolute translation error by 9.8%, and improves state estimation accuracy by up to 45%.
This work addresses the sensitivity of 3D Gaussian Splatting (3DGS) pose optimization to initial camera poses and reconstructed geometry, which often leads to reprojection distortions and unstable optimization due to uncertainties in pose priors and geometry. For the first time, this study systematically analyzes these two sources of uncertainty and proposes a relocalization framework that requires neither retraining nor additional supervision. The method explicitly models the pose distribution via Monte Carlo pose sampling and jointly handles uncertainty through a PnP optimization guided by the Fisher Information Matrix. Evaluated on multiple indoor and outdoor benchmarks, the approach significantly improves localization accuracy and demonstrates enhanced optimization stability and robustness under pose and depth noise.