Score
Designs and implements methods that detect objects and their constituent parts in 3D scenes and assemble part-aware, hierarchical scene graphs encoding part–whole, spatial, and functional relations. Builds affordance-grounding modules that predict part- and object-level affordances, support open-vocabulary labels and zero-shot generalization, and separate semantic and geometric grounding to reduce heavy end-to-end data requirements.
Existing 3D scene graph methods rely on object-level, coarse-grained representations, limiting their applicability to functional robot–environment interaction. This work proposes a fine-grained, function-oriented 3D scene graph that explicitly models functionally manipulable parts—such as door handles and light switches—as first-class nodes, enabling a semantic shift from object-level to function-level reasoning. Methodologically, we synthesize multi-source 3D data to generate 2D functional part annotations, train a part-level detector, and integrate it into standard 3D scene graph construction; we further enhance functional grounding via vision-language alignment and task-driven affordance localization. Experiments demonstrate state-of-the-art performance in functional part segmentation and significantly improved accuracy and robustness in mapping natural language instructions to executable robot actions in real-world settings.
Task-oriented 3D scene understanding requires joint modeling of spatial hierarchy (room → region → object) and functional affordances, yet existing methods lack explicit coupling between them. Method: We propose the 3D Hierarchical Scene Graph (3DHSG) framework, which jointly learns room classification, region segmentation, and region-/object-level affordance prediction from point clouds and semantic labels via a Transformer-based multi-task model. Contribution/Results: (1) We introduce the first 3DHSG benchmark dataset with fine-grained region- and object-level affordance annotations; (2) we establish a hierarchical 3DHSG generation paradigm that unifies spatial structure and functional semantics; (3) our method achieves significant improvements over state-of-the-art methods across multiple metrics. We publicly release both code and dataset to advance functional and structured 3D scene understanding.
This work addresses the problem of language-guided, precise localization of object affordances in 3D embodied environments—particularly under partial occlusion, multi-view observation, and arbitrary object rotations that induce incomplete sensory input. We formally introduce a novel task: language–vision–interaction joint-driven 3D affordance grounding. To support this, we present AGPIL, the first multimodal dataset covering full-view, occluded, and multi-rotation scenarios with fine-grained affordance annotations. Methodologically, we propose LMAffordance3D, a language-guided multimodal 3D affordance grounding network integrating vision-language models (VLMs), point-cloud encoders, cross-modal attention, and multi-view geometric alignment to jointly embed and spatially ground 2D images, 3D point clouds, and natural-language instructions. Evaluated on AGPIL, LMAffordance3D outperforms all baselines by +12.6 mAP and demonstrates strong generalization to unseen objects, viewpoints, and instruction combinations.
Open-vocabulary 3D object affordance localization aims to precisely localize functional regions on 3D objects that enable actions specified by arbitrary natural language instructions. Existing methods suffer from limited semantic priors and struggle to model geometric invariance and implicit interaction intent. To address this, we propose a geometry-intent collaborative reasoning paradigm: (i) an implicit invariant geometric representation that disentangles intrinsic structural properties; (ii) an interaction intent analogy mechanism enabling human-like functional reasoning across objects and instructions; and (iii) PIADv2—the largest open-vocabulary 3D affordance understanding dataset to date. Our method integrates point-cloud–image joint representation learning, geometry-guided intent distillation, and contrastive analogy reasoning. It achieves significant improvements over state-of-the-art methods on open-vocabulary affordance localization, demonstrating strong generalization to unseen objects and instructions. Code and dataset are publicly released.
Existing 3D scene graph generation methods are object-centric and struggle to model part-level details and multi-level relationships, limiting fine-grained scene understanding. This work proposes the first open-vocabulary, part-aware unified 3D scene graph framework that jointly represents objects, interactable parts, spatial and functional relationships, and affordances. By integrating object-part knowledge-guided detection, part-aware 3D feature fusion, geometry-prior-initialized relation modeling, and joint optimization with large language models, our approach enables efficient and accurate relational reasoning. We introduce a new benchmark, UniGraph3D, on which our method achieves state-of-the-art performance and significantly enhances perception for a variety of robotic tasks.
This work addresses the limitation of existing vision systems in modeling the structured hierarchical dependencies among scenes, objects, parts, and functions, which hinders interactive semantic understanding. It introduces, for the first time, a hierarchical scene parsing task that explicitly constructs a “scene → object → part → function” hierarchy and proposes a unified generative framework based on vision-language models. Key contributions include the formal definition of this new task, the design of structure-completion pseudo-labels and a curriculum learning strategy, and the creation of SceneParser-Bench—a large-scale benchmark with tailored evaluation metrics. Experiments demonstrate that the proposed method significantly outperforms current multimodal large language models and perception-based composition approaches on SceneParser-Bench, while also exhibiting strong generalization and practical utility on downstream tasks such as COCO and AGD20K.
This work addresses the challenge of precisely localizing fine-grained functional regions—such as handles and buttons—in 3D scenes using zero-shot vision-language systems, where such regions are often small, visually ambiguous, and repetitive. To this end, the authors propose AFFORDMEM, a novel framework that introduces, for the first time, a two-level memory mechanism requiring no model fine-tuning. A cross-scene, category-level memory guides a frozen vision-language model to attend to manipulable sub-regions, while an intra-scene spatial memory leverages structured scene graphs to resolve spatial referring relationships. Relying solely on a reusable RGB memory bank and 3D spatial modeling—without any annotations or training on target scenes—the method achieves AP50 scores on SceneFun3D that surpass existing zero-shot approaches by 3.23 and 3.7 points, respectively. Ablation studies confirm the complementary benefits of the two memory mechanisms.
Existing approaches struggle to accurately localize multifunctional regions in cluttered real-world scenes according to task instructions and lack benchmarks supporting complex mappings such as one-to-many task-to-region correspondences. This work proposes a task-conditioned, scene-level functional affordance localization framework and introduces the first real-scene affordance benchmark dataset encompassing both single- and multi-region instruction mappings. To enable efficient and high-quality annotation, we design the A2A-AffordGen pipeline, which integrates large language model filtering, interactive segmentation, mask refinement, and human verification. Models trained on this dataset significantly outperform baseline methods—including general-purpose segmentation models, vision-language models, and affordance distillation approaches—in both task-level localization accuracy and spatial priors for downstream manipulation tasks.
This work addresses the challenge of accurately localizing functional regions in cluttered 3D scenes that support specific actions described by natural language instructions—a task where existing methods often fail due to missed regions, improper granularity, or visual distractions. The authors propose a decoupled framework that first generates high-recall, fine-grained interaction region proposals without relying on object or part names, leveraging learnable affordance prompts and multi-level visual features (achieving 77.5% recall at IoU=0.25). Subsequently, a “think-before-respond” structured reasoning mechanism, termed VPAR, aligns language instructions with region selection through group relative policy optimization (GRPO) augmented with a 3D overlap reward. Evaluated on the SceneFun3D validation set, the method achieves 10.69% AP50 and 25.46% AP25, with GRPO-trained VPAR attaining a region selection accuracy of 72.1%.