🤖 AI Summary
This work addresses the challenge of extracting multi-granular, hierarchical 3D scene groupings from 2D foundation models without relying on semantic labels or a fixed vocabulary. To this end, it proposes transforming 2D affinity cues into hierarchical supervision signals and embedding them into a unified feature field within Lorentzian hyperbolic space. The method uniquely integrates the Dasgupta hierarchical clustering objective with hyperbolic geometry to construct a hyperspherical affinity field capable of representing groupings at multiple granularities. A lowest common ancestor ordering constraint is further introduced to enhance hierarchical consistency. This framework enables end-to-end, unsupervised 3D hierarchical grouping, accurately recovering object- and part-level structures across multiple scales while effectively fusing 2D priors with 3D geometric representations.
📝 Abstract
Hierarchical 3D grouping aims to recover scene groups across multiple granularities, from fine object parts to complete objects, without relying on semantic labels or a fixed vocabulary. The main challenge is to transform 2D foundation-model cues into coherent hierarchy supervision and embed that hierarchy in a 3D representation. We propose H2G, a hyperbolic affinity field for hierarchical 3D grouping. Our method derives semantically organized tree supervision by interpreting foundation-model affinities through Dasgupta's objective for similarity-based hierarchical clustering. This supervision is distilled into a single Lorentz hyperbolic feature field, whose geometry is well suited for tree-like branching structures. A hierarchy-aware objective aligns the field with fine-level assignments, coarse object structure, compact feature clusters, and LCA (Lowest Common Ancestor) ordering. This formulation represents multiple grouping levels in one feature space, enabling semantic hierarchical grouping grounded in 2D foundation-model knowledge.