BoxDreamer: Dreaming Box Corners for Generalizable Object Pose Estimation

📅 2025-04-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing RGB-based object pose estimation methods under sparse-view settings suffer from limited robustness to occlusion and poor generalization across object categories. This paper proposes a generic single-modality RGB pose estimation framework. Its core innovation lies in adopting the 3D bounding box’s eight corners as an intermediate semantic representation and introducing a reference-driven 2D corner synthesis network, which robustly establishes 2D–3D correspondences from merely one or few reference views. Pose estimation is then solved via Perspective-n-Point (PnP). The method requires no depth input, CAD models, or category-specific priors, significantly enhancing occlusion robustness and few-shot generalization. It achieves state-of-the-art performance on YCB-Video and Occluded-LINEMOD, reducing average rotation and translation errors by 18.7%–32.4%, and demonstrates markedly improved cross-object transfer capability.

Technology Category

Computer Vision: Biometrics, Face, Gesture & PoseIntelligent Robots: State EstimationReasoning under Uncertainty: Relational Probabilistic Models

Application Category

Graph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphsUser Modeling, Personalization and Recommendation: On-Device user modeling, personalization, and recommendationWeb Mining and Content Analysis: Robustness and generalizability of Web mining methods
📝 Abstract
This paper presents a generalizable RGB-based approach for object pose estimation, specifically designed to address challenges in sparse-view settings. While existing methods can estimate the poses of unseen objects, their generalization ability remains limited in scenarios involving occlusions and sparse reference views, restricting their real-world applicability. To overcome these limitations, we introduce corner points of the object bounding box as an intermediate representation of the object pose. The 3D object corners can be reliably recovered from sparse input views, while the 2D corner points in the target view are estimated through a novel reference-based point synthesizer, which works well even in scenarios involving occlusions. As object semantic points, object corners naturally establish 2D-3D correspondences for object pose estimation with a PnP algorithm. Extensive experiments on the YCB-Video and Occluded-LINEMOD datasets show that our approach outperforms state-of-the-art methods, highlighting the effectiveness of the proposed representation and significantly enhancing the generalization capabilities of object pose estimation, which is crucial for real-world applications.
Problem

Research questions and friction points this paper is trying to address.

Generalizable RGB-based object pose estimation in sparse-view settings
Overcoming occlusion and sparse-view limitations in pose estimation
Using bounding box corners as intermediate pose representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses bounding box corners as pose representation
Novel reference-based point synthesizer for 2D corners
Enhances generalization with 2D-3D corner correspondences
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yuanhong Yu
State Key Lab of CAD & CG, Zhejiang University; Ant Group
X
Xingyi He He
State Key Lab of CAD & CG, Zhejiang University
C
Chen Zhao
EPFL
Junhao Yu
Junhao Yu
University of Science and Technology of China
LLMMachine Learning
J
Jiaqi Yang
Northwestern Polytechnical University
R
Ruizhen Hu
Shenzhen University
Yujun Shen
Yujun Shen
Ant Group
Generative ModelingComputer VisionDeep Learning
X
Xing Zhu
Ant Group
Xiaowei Zhou
Xiaowei Zhou
Professor of Computer Science, Zhejiang University
Computer VisionComputer Graphics
Sida Peng
Sida Peng
Zhejiang University
Computer VisionComputer Graphics