🤖 AI Summary
To address the lack of rigorous uncertainty quantification in visual and robotic pose estimation, this paper introduces the first differentiable-rendering-based framework for pose uncertainty quantification. Our method linearizes the rendering process via small perturbations on the pose manifold, enabling derivation of a rendering-aware Cramér–Rao lower bound (CRLB)—the first systematic integration of differentiable rendering into CRLB theory. The resulting closed-form lower bound on camera pose covariance aligns with classical bundle adjustment uncertainty estimates. The framework natively supports multi-camera systems, enabling Fisher information fusion without keypoint correspondence—facilitating cooperative perception and novel-view synthesis. Its core innovation lies in the deep coupling of geometry-aware differentiable rendering with statistical lower-bound theory, providing interpretable and verifiable uncertainty guarantees for learning-based dense pose estimation.
📝 Abstract
Pose estimation is essential for many applications within computer vision and robotics. Despite its uses, few works provide rigorous uncertainty quantification for poses under dense or learned models. We derive a closed-form lower bound on the covariance of camera pose estimates by treating a differentiable renderer as a measurement function. Linearizing image formation with respect to a small pose perturbation on the manifold yields a render-aware Cramér-Rao bound. Our approach reduces to classical bundle-adjustment uncertainty, ensuring continuity with vision theory. It also naturally extends to multi-agent settings by fusing Fisher information across cameras. Our statistical formulation has downstream applications for tasks such as cooperative perception and novel view synthesis without requiring explicit keypoint correspondences.