construct part-aware 3d scene graphs

Designs and implements methods that detect objects and their constituent parts in 3D scenes and assemble part-aware, hierarchical scene graphs encoding part–whole, spatial, and functional relations. Builds affordance-grounding modules that predict part- and object-level affordances, support open-vocabulary labels and zero-shot generalization, and separate semantic and geometric grounding to reduce heavy end-to-end data requirements.

constructpart-aware3dscene

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.21
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

FunGraph: Functionality Aware 3D Scene Graphs for Language-Prompted Scene Interaction

Mar 10, 2025
DR
Dennis Rotondi
🏛️ University of Stuttgart | University of Bonn

Existing 3D scene graph methods rely on object-level, coarse-grained representations, limiting their applicability to functional robot–environment interaction. This work proposes a fine-grained, function-oriented 3D scene graph that explicitly models functionally manipulable parts—such as door handles and light switches—as first-class nodes, enabling a semantic shift from object-level to function-level reasoning. Methodologically, we synthesize multi-source 3D data to generate 2D functional part annotations, train a part-level detector, and integrate it into standard 3D scene graph construction; we further enhance functional grounding via vision-language alignment and task-driven affordance localization. Experiments demonstrate state-of-the-art performance in functional part segmentation and significantly improved accuracy and robustness in mapping natural language instructions to executable robot actions in real-world settings.

Augment 3D scene graphs using 2D data for improved affordance grounding.Detect and store affordance-relevant object parts for functional interaction.Develop a fine-resolution 3D scene graph for robot-environment interaction.

TB-HSU: Hierarchical 3D Scene Understanding with Contextual Affordances

Dec 07, 2024
WX
Wenting Xu
🏛️ The University of Sydney

Task-oriented 3D scene understanding requires joint modeling of spatial hierarchy (room → region → object) and functional affordances, yet existing methods lack explicit coupling between them. Method: We propose the 3D Hierarchical Scene Graph (3DHSG) framework, which jointly learns room classification, region segmentation, and region-/object-level affordance prediction from point clouds and semantic labels via a Transformer-based multi-task model. Contribution/Results: (1) We introduce the first 3DHSG benchmark dataset with fine-grained region- and object-level affordance annotations; (2) we establish a hierarchical 3DHSG generation paradigm that unifies spatial structure and functional semantics; (3) our method achieves significant improvements over state-of-the-art methods across multiple metrics. We publicly release both code and dataset to advance functional and structured 3D scene understanding.

Develops a hierarchical 3D scene graph modelImproves 3D scene understanding using transformer-based modelsIntegrates functional affordance with spatial context

This work addresses the problem of language-guided, precise localization of object affordances in 3D embodied environments—particularly under partial occlusion, multi-view observation, and arbitrary object rotations that induce incomplete sensory input. We formally introduce a novel task: language–vision–interaction joint-driven 3D affordance grounding. To support this, we present AGPIL, the first multimodal dataset covering full-view, occluded, and multi-rotation scenarios with fine-grained affordance annotations. Methodologically, we propose LMAffordance3D, a language-guided multimodal 3D affordance grounding network integrating vision-language models (VLMs), point-cloud encoders, cross-modal attention, and multi-view geometric alignment to jointly embed and spatially ground 2D images, 3D point clouds, and natural-language instructions. Evaluated on AGPIL, LMAffordance3D outperforms all baselines by +12.6 mAP and demonstrates strong generalization to unseen objects, viewpoints, and instruction combinations.

Address partial observations in 3D space due to occlusion or rotationDevelop multi-modal network for language-guided 3D affordance groundingGround 3D object affordance using language, vision, and interactions

GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance Grounding

Nov 29, 2024
YS
Yawen Shao
🏛️ University of Science and Technology of China | Northeastern University

Open-vocabulary 3D object affordance localization aims to precisely localize functional regions on 3D objects that enable actions specified by arbitrary natural language instructions. Existing methods suffer from limited semantic priors and struggle to model geometric invariance and implicit interaction intent. To address this, we propose a geometry-intent collaborative reasoning paradigm: (i) an implicit invariant geometric representation that disentangles intrinsic structural properties; (ii) an interaction intent analogy mechanism enabling human-like functional reasoning across objects and instructions; and (iii) PIADv2—the largest open-vocabulary 3D affordance understanding dataset to date. Our method integrates point-cloud–image joint representation learning, geometry-guided intent distillation, and contrastive analogy reasoning. It achieves significant improvements over state-of-the-art methods on open-vocabulary affordance localization, demonstrating strong generalization to unseen objects and instructions. Code and dataset are publicly released.

Addressing limited semantic space in existing methods through collaborative inferenceLeveraging invariant geometries and interaction intentions for affordance knowledgeOpen-Vocabulary 3D object affordance grounding with arbitrary instructions

Latest Papers

What's happening recently
View more

Existing 3D scene graph generation methods are object-centric and struggle to model part-level details and multi-level relationships, limiting fine-grained scene understanding. This work proposes the first open-vocabulary, part-aware unified 3D scene graph framework that jointly represents objects, interactable parts, spatial and functional relationships, and affordances. By integrating object-part knowledge-guided detection, part-aware 3D feature fusion, geometry-prior-initialized relation modeling, and joint optimization with large language models, our approach enables efficient and accurate relational reasoning. We introduce a new benchmark, UniGraph3D, on which our method achieves state-of-the-art performance and significantly enhances perception for a variety of robotic tasks.

3D scene graphopen-vocabularypart-aware

This work addresses the limitation of existing vision systems in modeling the structured hierarchical dependencies among scenes, objects, parts, and functions, which hinders interactive semantic understanding. It introduces, for the first time, a hierarchical scene parsing task that explicitly constructs a “scene → object → part → function” hierarchy and proposes a unified generative framework based on vision-language models. Key contributions include the formal definition of this new task, the design of structure-completion pseudo-labels and a curriculum learning strategy, and the creation of SceneParser-Bench—a large-scale benchmark with tailored evaluation metrics. Experiments demonstrate that the proposed method significantly outperforms current multimodal large language models and perception-based composition approaches on SceneParser-Bench, while also exhibiting strong generalization and practical utility on downstream tasks such as COCO and AGD20K.

affordance predictionhierarchical scene parsinginteraction-oriented understanding

This work addresses the challenge of precisely localizing fine-grained functional regions—such as handles and buttons—in 3D scenes using zero-shot vision-language systems, where such regions are often small, visually ambiguous, and repetitive. To this end, the authors propose AFFORDMEM, a novel framework that introduces, for the first time, a two-level memory mechanism requiring no model fine-tuning. A cross-scene, category-level memory guides a frozen vision-language model to attend to manipulable sub-regions, while an intra-scene spatial memory leverages structured scene graphs to resolve spatial referring relationships. Relying solely on a reusable RGB memory bank and 3D spatial modeling—without any annotations or training on target scenes—the method achieves AP50 scores on SceneFun3D that surpass existing zero-shot approaches by 3.23 and 3.7 points, respectively. Ablation studies confirm the complementary benefits of the two memory mechanisms.

3D scene understandingcross-scene memoryfunctional affordance grounding

Existing approaches struggle to accurately localize multifunctional regions in cluttered real-world scenes according to task instructions and lack benchmarks supporting complex mappings such as one-to-many task-to-region correspondences. This work proposes a task-conditioned, scene-level functional affordance localization framework and introduces the first real-scene affordance benchmark dataset encompassing both single- and multi-region instruction mappings. To enable efficient and high-quality annotation, we design the A2A-AffordGen pipeline, which integrates large language model filtering, interactive segmentation, mask refinement, and human verification. Models trained on this dataset significantly outperform baseline methods—including general-purpose segmentation models, vision-language models, and affordance distillation approaches—in both task-level localization accuracy and spatial priors for downstream manipulation tasks.

affordance groundingfunctional regionsinstruction-region correspondence

This work addresses the challenge of accurately localizing functional regions in cluttered 3D scenes that support specific actions described by natural language instructions—a task where existing methods often fail due to missed regions, improper granularity, or visual distractions. The authors propose a decoupled framework that first generates high-recall, fine-grained interaction region proposals without relying on object or part names, leveraging learnable affordance prompts and multi-level visual features (achieving 77.5% recall at IoU=0.25). Subsequently, a “think-before-respond” structured reasoning mechanism, termed VPAR, aligns language instructions with region selection through group relative policy optimization (GRPO) augmented with a 3D overlap reward. Evaluated on the SceneFun3D validation set, the method achieves 10.69% AP50 and 25.46% AP25, with GRPO-trained VPAR attaining a region selection accuracy of 72.1%.

3D groundingaffordancecluttered scenes

Hot Scholars

JT

Jin Tian

Mohamed bin Zayed University of Artificial Intelligence
artificial intelligencemachine learningcausal inference
YZ

Yue Zhou

Associate Professor, East China Normal University
Remote Sensing Vision-Language ModelOriented Object Detection
LX

Liang Xie

Wuhan University of Technology
Time Series ForecastingCross-modal Learning
XH

Xiaoshuai Hao

Beijing Academy of Artificial Intelligence,BAAI
vision and language