LEGAU: Learning Semantic Gaussian Priors for Scalable Category-level Pose Estimation

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the under-constrained problem caused by geometric incompleteness in monocular RGB-D category-level 6D pose estimation. To this end, it proposes a unified multimodal framework that incorporates reconstruction priors. Methodologically, a Transformer module is employed to jointly predict Normalized Object Coordinate Space (NOCS) correspondences, pose and size parameters, and semantic Gaussian fields. Furthermore, text-embedding-conditioned decoding is introduced, leveraging shape reconstruction as a categorical prior to guide local reasoning and achieve coupled learning of pose and shape. Experimental results demonstrate that the proposed method achieves superior performance on both synthetic and real-world benchmarks, yielding a 22% improvement in the SOPE metric while exhibiting strong cross-category generalization capabilities.
📝 Abstract
Category-level 6D pose estimation from a single RGB-D observation is inherently under-constrained, since partial visible geometry must be interpreted together with a canonical object structure before a stable pose can be determined. We present LEGAU, a unified framework that jointly predicts NOCS correspondence, object pose and size, and a canonical Semantic Gaussian Field. Rather than treating reconstruction as a detached auxiliary task, LEGAU uses the Gaussian field as a category-conditioned structural prior that participates in multimodal feature fusion and provides global guidance for local pose reasoning. Conditioned on a categorical text embedding, LEGAU processes RGB-D observations through a transformer-based fusion module that integrates visual, geometric, and category-level cues, decoding the NOCS map, pose and size information and the Gaussian-based object representation. Extensive experiments on synthetic and real-world benchmarks show that this coupled pose-shape formulation achieves strong performance in a single-model multi-category setting, with up to 22\% on SOPE and competitive transfer to real-world data. These results highlight the benefit of jointly learning canonical correspondence, object shape, and pose alignment within a unified representation.
Problem

Research questions and friction points this paper is trying to address.

Category-level 6D pose estimation
RGB-D observation
under-constrained problem
Innovation

Methods, ideas, or system contributions that make the work stand out.

Category-level Pose Estimation
Semantic Gaussian Priors
NOCS Correspondence
Multimodal Feature Fusion
Unified Representation
🔎 Similar Papers
No similar papers found.