RYOPO: Bringing End-to-End Category-Level Object Pose Estimation into Real Time

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the bottleneck of existing category-level pose estimation methods that rely on external segmentation and per-object cropping, hindering real-time deployment. We propose an end-to-end, query-based RGB-D set prediction framework that jointly performs detection, segmentation, and 9DoF pose estimation. The method employs a shared scene encoder to eliminate redundant computation and designs a query-conditioned geometric pathway to associate 3D point clouds with feature representations. Furthermore, pose-conditioned cross-attention and iterative residual refinement mechanisms are introduced to optimize explicit pose states without requiring CAD priors or independent segmentation modules. The proposed architecture significantly outperforms existing joint detection approaches on the NOCS dataset, achieves accuracy comparable to two-stage methods on the REAL275 benchmark, and enables real-time inference at 31.8 FPS on an RTX A6000 GPU.
📝 Abstract
Category-level object pose estimation predicts the rotation, translation, and metric size of unseen instances within known categories. Many accurate RGB-D methods rely on external instance segmentation and crop-based pose estimation, introducing separate stages and object-dependent processing costs that hinder real-time inference. To bring accurate pose estimation into real time, we present \ours{}, an end-to-end trainable query-based RGB-D set predictor. It jointly detects and segments objects and estimates their \mbox{9-DoF} poses without explicit CAD-derived shape priors or a separately trained instance segmentor. Shared image and scene encoding avoids repeated per-object crop encoding. A query-conditioned geometry pathway associates observed 3D points and RGB features with object queries and incorporates shared scene context. Object-centric refinement uses the resulting point descriptors to update an explicit pose state through pose-conditioned cross-attention and recurrent residual corrections. On NOCS, \ours{} substantially improves on published RGB-D joint detection and pose estimation results. It achieves competitive performance compared with two-stage methods under all-object evaluation on REAL275 and HouseCat6D, while enabling real-time full-frame pose estimation at $31.8$ FPS on an RTX~A6000. Project page: https://yopo-series.github.io/RYOPO-project-page/.
Problem

Research questions and friction points this paper is trying to address.

Category-level object pose estimation
Real-time inference
End-to-end
RGB-D
9-DoF pose
Innovation

Methods, ideas, or system contributions that make the work stand out.

End-to-End Pose Estimation
Query-Based Set Predictor
Category-Level Object Pose
Shared Scene Encoding
Real-Time RGB-D
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hakjin Lee
PIT IN Co., Anyang-si, Gyeonggi-do, Republic of Korea
Junghoon Seo
Junghoon Seo
PIT IN Corp.
Machine LearningComputer VisionRobot Vision
J
Jaehoon Sim
PIT IN Co., Anyang-si, Gyeonggi-do, Republic of Korea