Score
Designs, implements, and evaluates algorithms and systems that recover an object’s spatial pose and related parameters—rotation, translation, scale, and 3D landmark coordinates—from visual inputs (single images or video) using approaches such as pose regression networks, pose representation learning, tracking, and self-supervision. Builds pose-and-scale alignment and tracking pipelines, defines and computes pose evaluation metrics, and analyzes robustness and temporal consistency of pose estimates for use as inputs to downstream alignment, compositing, or control pipelines.
This work addresses the bottleneck in general object pose estimation—its reliance on hard-to-acquire CAD models—by investigating the feasibility of substituting them with image-reconstructed 3D models. To this end, we introduce the first pose-estimation-oriented 3D reconstruction quality benchmark, built upon the YCB-V dataset and featuring calibrated reconstructions aligned with ground-truth poses. The benchmark integrates geometric reconstruction pipelines (COLMAP for SfM/MVS), learning-based methods (PixelNeRF, iMAP), and the BOP evaluation framework. Key contributions are: (1) the first reconstruction quality benchmark explicitly designed for pose estimation; (2) empirical evidence that traditional geometric methods outperform learning-based approaches in the accuracy–speed trade-off; (3) discovery that standard reconstruction metrics (e.g., Chamfer distance) exhibit weak correlation with pose estimation accuracy; and (4) demonstration that most image-reconstructed models support high-accuracy pose estimation, albeit systematically underperforming CAD models. Code and benchmark are publicly released.
To address the unreliable pose graph construction and motion synchronization challenges in multi-view point cloud registration, this paper proposes an end-to-end absolute pose estimation paradigm. First, matching distance is introduced as a principled reliability metric for pose graph construction, replacing handcrafted loss functions with direct global pose regression. Second, the method jointly optimizes feature interaction and structural awareness by integrating local geometric distribution modeling with adaptive attention mechanisms. Fully data-driven, it eliminates iterative optimization and post-processing. Evaluated on diverse indoor and outdoor datasets, the approach achieves a 12.7% improvement in pose graph construction accuracy and reduces overall registration error by 21.3%, demonstrating significantly enhanced robustness and cross-scene generalization capability.
Existing category-level pose estimation methods for articulated objects with known categories but unknown geometries suffer from high annotation costs, reliance on multi-view inputs, or strong supervision. This paper proposes the first unsupervised, single-frame point cloud-based category-level pose estimation framework. It introduces a self-supervised paradigm that jointly optimizes canonical global pose and joint states through object-level and part-level co-alignment. Key technical components include a point-cloud autoencoder for canonical reconstruction, differentiable part segmentation, hierarchical pose disentanglement, and joint-aware geometric consistency constraints. Contributions are threefold: (1) the first object-part co-alignment mechanism for articulated pose estimation; (2) the first large-scale real-scene articulated object benchmark dataset; and (3) competitive performance against state-of-the-art supervised methods on multiple benchmarks, with strong generalization and robustness validated on the new dataset.
This work addresses the lack of provable correctness guarantees in visual pose estimation for safety-critical applications by proposing a certifiable pose estimation algorithm that integrates physics-driven geometric modeling with learning-based methods. The core innovation lies in introducing a Geometric Generative Model (GGM) combined with neural network reachability analysis to construct a multi-stage, certification-aware pipeline capable of delivering verifiable pose estimates and object detection under conditions ranging from unoccluded to complex occlusion scenarios. The approach is validated on both synthetic and real-world imagery—including event camera data—for planar objects such as traffic signs. Experimental results demonstrate that the estimated poses rigorously satisfy pre-specified certification error bounds, thereby achieving, for the first time, provably robust pose perception suitable for safety-critical deployment.
This work addresses the unreliability of existing keypoint-based pose estimation methods, which often neglect geometric constraints inherent to object shape, leading to undetectable failures. To overcome this limitation, the authors propose a failure detection approach that does not rely on keypoint uncertainty estimates. Instead, it leverages handcrafted geometric consistency features—specifically pairwise keypoint distances, reprojection errors, and consistency between rendered and observed masks—to characterize spatial relationships among 2D keypoints. These features are fed into a logistic regression classifier to determine whether a given pose estimate has failed. Experimental results demonstrate that the proposed method significantly outperforms existing confidence-based failure detection schemes, such as conformal keypoint prediction, offering superior reliability and practical utility in real-world applications.
This work addresses the challenge of simultaneously achieving high accuracy, robustness, and loop-closure capability in two-frame pose optimization by proposing a unified framework that integrates geometric and photometric information. For the first time, dense geometric feature descriptors are incorporated into differential photometric optimization, replacing conventional photometric residuals with descriptor-based residuals to enable subpixel-level pose estimation in descriptor space. By synergistically combining the strengths of both geometric and photometric paradigms, this approach explores a novel trajectory for pose optimization grounded in descriptor similarity. Experimental results demonstrate a significant improvement in tracking accuracy; however, overall performance remains slightly inferior to reprojection error–based methods, with the primary bottleneck identified as the relatively flat landscape of descriptor similarity, which limits optimization efficacy.
This work addresses the computationally demanding yet critical nonlinear least-squares problem in pose estimation for real-time computer vision. By introducing a suitable parametrization of rotations, the problem is reformulated as a system of polynomial equations. The authors propose a novel class of resultant solvers based on Sylvester matrices that enable efficient closed-form solutions. This approach substantially reduces computational complexity while preserving high numerical accuracy. Experimental results demonstrate that the method outperforms state-of-the-art techniques in terms of runtime on both 3D–3D and 3D–2D pose estimation tasks, offering a practical solution for time-sensitive applications.
This work addresses the challenge of estimating 6D poses of unseen CAD-model objects from RGB images without fine-tuning. To this end, we propose a lightweight network that directly matches correspondences between a query image and a set of reference images through an innovative multi-view reference feature fusion strategy. Our approach drastically reduces the number of required reference images while maintaining high pose estimation accuracy, thereby significantly lowering both storage and computational overhead. Evaluated on the seven core datasets of the BOP Challenge, the method achieves performance comparable to state-of-the-art approaches using fewer reference views and a smaller model footprint, while substantially reducing memory consumption and inference time.