🤖 AI Summary
This study addresses the bottleneck of existing category-level pose estimation methods that rely on external segmentation and per-object cropping, hindering real-time deployment. We propose an end-to-end, query-based RGB-D set prediction framework that jointly performs detection, segmentation, and 9DoF pose estimation. The method employs a shared scene encoder to eliminate redundant computation and designs a query-conditioned geometric pathway to associate 3D point clouds with feature representations. Furthermore, pose-conditioned cross-attention and iterative residual refinement mechanisms are introduced to optimize explicit pose states without requiring CAD priors or independent segmentation modules. The proposed architecture significantly outperforms existing joint detection approaches on the NOCS dataset, achieves accuracy comparable to two-stage methods on the REAL275 benchmark, and enables real-time inference at 31.8 FPS on an RTX A6000 GPU.
📝 Abstract
Category-level object pose estimation predicts the rotation, translation, and metric size of unseen instances within known categories. Many accurate RGB-D methods rely on external instance segmentation and crop-based pose estimation, introducing separate stages and object-dependent processing costs that hinder real-time inference. To bring accurate pose estimation into real time, we present \ours{}, an end-to-end trainable query-based RGB-D set predictor. It jointly detects and segments objects and estimates their \mbox{9-DoF} poses without explicit CAD-derived shape priors or a separately trained instance segmentor. Shared image and scene encoding avoids repeated per-object crop encoding. A query-conditioned geometry pathway associates observed 3D points and RGB features with object queries and incorporates shared scene context. Object-centric refinement uses the resulting point descriptors to update an explicit pose state through pose-conditioned cross-attention and recurrent residual corrections. On NOCS, \ours{} substantially improves on published RGB-D joint detection and pose estimation results. It achieves competitive performance compared with two-stage methods under all-object evaluation on REAL275 and HouseCat6D, while enabling real-time full-frame pose estimation at $31.8$ FPS on an RTX~A6000. Project page: https://yopo-series.github.io/RYOPO-project-page/.