RefPose: Leveraging Reference Geometric Correspondences for Accurate 6D Pose Estimation of Unseen Objects

📅 2025-05-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the challenge of 6D pose estimation for unseen objects in monocular RGB images—where no prior CAD models are available—this paper proposes a reference-image-driven zero-shot pose estimation method. The approach introduces three key contributions: (1) a dynamic geometric correspondence modeling mechanism that operates without predefined CAD templates; (2) a render-and-compare iterative optimization framework, integrating template rendering with correlation-volume-guided attention to achieve geometrically consistent pose prediction and refinement; and (3) explicit modeling of 3D structural constraints between reference and target images via cross-image geometric correspondence prediction. Evaluated on the BOP benchmark, the method achieves state-of-the-art performance, striking an optimal trade-off between accuracy and inference speed while significantly enhancing generalization to previously unseen objects.

Technology Category

Computer Vision: 3D Computer VisionIntelligent Robots: State EstimationMachine Learning: Calibration & Uncertainty Quantification

Application Category

Graph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphsUser Modeling, Personalization and Recommendation: On-Device user modeling, personalization, and recommendationSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
Estimating the 6D pose of unseen objects from monocular RGB images remains a challenging problem, especially due to the lack of prior object-specific knowledge. To tackle this issue, we propose RefPose, an innovative approach to object pose estimation that leverages a reference image and geometric correspondence as guidance. RefPose first predicts an initial pose by using object templates to render the reference image and establish the geometric correspondence needed for the refinement stage. During the refinement stage, RefPose estimates the geometric correspondence of the query based on the generated references and iteratively refines the pose through a render-and-compare approach. To enhance this estimation, we introduce a correlation volume-guided attention mechanism that effectively captures correlations between the query and reference images. Unlike traditional methods that depend on pre-defined object models, RefPose dynamically adapts to new object shapes by leveraging a reference image and geometric correspondence. This results in robust performance across previously unseen objects. Extensive evaluation on the BOP benchmark datasets shows that RefPose achieves state-of-the-art results while maintaining a competitive runtime.
Problem

Research questions and friction points this paper is trying to address.

Estimating 6D pose of unseen objects from RGB images
Leveraging reference images and geometric correspondences for pose refinement
Dynamic adaptation to new object shapes without pre-defined models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses reference image and geometric correspondence
Employs correlation volume-guided attention mechanism
Dynamically adapts to new object shapes