🤖 AI Summary
Existing RGB-based object pose estimation methods under sparse-view settings suffer from limited robustness to occlusion and poor generalization across object categories. This paper proposes a generic single-modality RGB pose estimation framework. Its core innovation lies in adopting the 3D bounding box’s eight corners as an intermediate semantic representation and introducing a reference-driven 2D corner synthesis network, which robustly establishes 2D–3D correspondences from merely one or few reference views. Pose estimation is then solved via Perspective-n-Point (PnP). The method requires no depth input, CAD models, or category-specific priors, significantly enhancing occlusion robustness and few-shot generalization. It achieves state-of-the-art performance on YCB-Video and Occluded-LINEMOD, reducing average rotation and translation errors by 18.7%–32.4%, and demonstrates markedly improved cross-object transfer capability.
📝 Abstract
This paper presents a generalizable RGB-based approach for object pose estimation, specifically designed to address challenges in sparse-view settings. While existing methods can estimate the poses of unseen objects, their generalization ability remains limited in scenarios involving occlusions and sparse reference views, restricting their real-world applicability. To overcome these limitations, we introduce corner points of the object bounding box as an intermediate representation of the object pose. The 3D object corners can be reliably recovered from sparse input views, while the 2D corner points in the target view are estimated through a novel reference-based point synthesizer, which works well even in scenarios involving occlusions. As object semantic points, object corners naturally establish 2D-3D correspondences for object pose estimation with a PnP algorithm. Extensive experiments on the YCB-Video and Occluded-LINEMOD datasets show that our approach outperforms state-of-the-art methods, highlighting the effectiveness of the proposed representation and significantly enhancing the generalization capabilities of object pose estimation, which is crucial for real-world applications.