GenCOPE: Syn2Real Generalized Category-Level Object Pose Estimation for Robotic Picking

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reliance of category-level object pose estimation on extensive real-world data and its limited generalization to novel categories. To this end, it proposes an end-to-end pose regression framework trained exclusively on synthetic data yet deployable in real-world scenarios. Methodologically, the approach enforces 2D/3D semantic consistency constraints and employs cross-modal dense fusion to learn domain-invariant representations, effectively bridging the synthetic-to-real domain gap. Furthermore, by leveraging global features within a lightweight and efficient architecture, it achieves a minimalist Syn2Real generalization paradigm. Experimental results demonstrate that the proposed method exhibits superior generalization performance across the REAL275 and Wild6D benchmarks, as well as in real-world robotic grasping tasks. The source code has been made publicly available.
📝 Abstract
Category-level object pose estimation (COPE), capable of generalizing to intra-class unknown objects, has become a core technique for robotic 3D scene understanding. However, existing COPE methods still require labor-intensive recollection of real-world training data for novel object categories, which limits their scalability in practical applications. This paper aims to achieve synthetic-to-real (Syn2Real) generalized COPE, where a model is trained solely on rendered synthetic data and directly generalized to real-world deployments. The central challenge lies in the significant domain gap between synthetic and real-world data, particularly in texture appearance. To address this, we aim to enhance domain generalization by learning domain-invariant representations that capture semantic commonalities among objects within the same category. We introduce 2D and 3D semantic consistency constraints to reduce the sensitivity of feature encoders to domain-specific features. In addition, we propose an end-to-end pose regression framework that performs 2D-3D cross consistency learning, leveraging dense cross-modality fusion to further refine pose estimation. Since simplicity and effectiveness are essential for real-world robotic deployment, our model operates exclusively on global features, yielding a highly lightweight and efficient architecture. Extensive experiments on the REAL275 and Wild6D benchmarks, as well as real-world robotic manipulation scenes, show superior Syn2Real generalization performance of our paradigm. Code and demos are released at https://paperreview99.github.io/GenCOPE/.
Problem

Research questions and friction points this paper is trying to address.

Category-level object pose estimation
Synthetic-to-real generalization
Domain gap
Robotic picking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Syn2Real Generalization
Category-Level Pose Estimation
Semantic Consistency
Cross-Modality Fusion
Lightweight Architecture
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.