Score
Designs and implements structured scene representations (scene graphs) that encode objects and their semantic and spatial relationships. Builds analyzers and constraint reasoners that infer, enforce, or repair relationships such as support, contact, proximity, and occlusion, translate those constraints into actionable parameters, and detect or resolve inconsistencies for downstream processing.
Industrial CAD models often lack semantic, spatial, and functional information, limiting their utility in robotic simulation and high-level scene understanding. This work proposes an offline method that leverages large vision-language models (LVLMs)—introduced for the first time into CAD environments—to automatically generate structured 3D scene graphs that explicitly model manipulable objects and their functional relationships. By effectively integrating semantic parsing with functional reasoning, the approach achieves high-precision semantic annotation and relationship recognition on industrial structures such as piping systems. Both qualitative and quantitative evaluations demonstrate its effectiveness. The associated code and dataset have been made publicly available.
Existing video graph representations neglect fine-grained action semantics—such as location, tools employed, and object functionality—leading to inadequate modeling of human–context relationships. To address this, we propose the Situational Scene Graph (SSG), the first video graph representation that explicitly incorporates semantic role–value structures to jointly model human–object interactions and their associated attributes. We formulate a novel task, “Situational Scene Graph Generation,” construct the first annotated SSG dataset, and design InComNet—a multi-stage, interaction-complementary network integrating semantic role labeling, graph-structured reasoning, and joint multi-task learning. Experiments demonstrate that our approach significantly outperforms baselines on predicate classification, semantic role–value classification, and situational reasoning tasks. These results validate SSG as an effective, generalizable, structured representation for human-centered contextual understanding in videos.
Existing 3D scene graph methods rely on object-level, coarse-grained representations, limiting their applicability to functional robot–environment interaction. This work proposes a fine-grained, function-oriented 3D scene graph that explicitly models functionally manipulable parts—such as door handles and light switches—as first-class nodes, enabling a semantic shift from object-level to function-level reasoning. Methodologically, we synthesize multi-source 3D data to generate 2D functional part annotations, train a part-level detector, and integrate it into standard 3D scene graph construction; we further enhance functional grounding via vision-language alignment and task-driven affordance localization. Experiments demonstrate state-of-the-art performance in functional part segmentation and significantly improved accuracy and robustness in mapping natural language instructions to executable robot actions in real-world settings.
Current 3D semantic scene graph prediction methods rely on graph neural networks but suffer from insufficient discriminability and representational capacity in object and relational feature encoding. To address this, we propose a decoupled representation learning framework: first, a highly discriminative object feature encoder is designed, integrating geometric-semantic multimodal fusion; second, an object-centric contrastive pre-training strategy is introduced to explicitly decouple object representation learning from graph structure prediction. Crucially, our method requires no architectural modifications to downstream graph inference modules and can be seamlessly integrated as a plug-in enhancement. Evaluated on the 3DSSG benchmark, our approach significantly outperforms state-of-the-art methods, achieving consistent improvements in both object classification and relationship prediction—the two core evaluation metrics—thereby validating the effectiveness of decoupled representation learning for 3D scene semantic understanding.
This work addresses the challenge of simultaneously achieving structural precision, semantic interpretability, and identity controllability in existing 3D/4D scene representations. We propose “Scene Language”—a unified 3D/4D scene representation framework that integrates executable program structures, natural-language semantic tokens, and visual identity embeddings. To our knowledge, this is the first method enabling zero-shot cross-modal reasoning: without fine-tuning, it directly synthesizes structured scene programs from pretrained language models and vision encoders, while explicitly modeling hierarchical relationships to support fine-grained editing. The representation is renderer-agnostic, interfacing seamlessly with traditional, neural, and hybrid renderers to produce high-fidelity images. Experiments demonstrate significant improvements over baselines—including scene graphs—on complex scene generation tasks, achieving breakthroughs in fidelity, controllability, and editability.
This work addresses the challenge of modeling component assembly relationships in 2D scenes under few-shot settings without semantic labels. The authors propose an end-to-end scene graph generation method that integrates geometric feature extraction with structural reasoning. Specifically, Faster R-CNN is employed to extract geometric representations of components, and a Transformer architecture constructs an initial adjacency matrix. To refine relational inference, the approach incorporates an attention-based Graph Convolutional Network (aGCN) with a message-passing mechanism. Notably, the model operates without semantic supervision and achieves accurate recovery of ground-truth assembly relationships using only a minimal number of training samples. Experimental results on a toy vehicle dataset demonstrate the method’s effectiveness, significantly advancing the capability to model assembly relationships in semantically unlabeled scenarios.
This work addresses the fragmented state of 3D scene graph research, hindered by the absence of a unified definition, construction pipeline, and evaluation protocol, which impedes method comparison and real-world deployment. The paper presents the first formal definition and a cohesive theoretical framework for 3D scene graphs, systematically analyzing key modeling choices—including node and edge attributes, hierarchical structure, dynamic modeling, and functional awareness—and clarifying the terminology and mainstream methodologies that map perceptual data to scene graphs. Through a comprehensive literature review, taxonomic analysis, and technical comparison, it delineates the core components and evolutionary trajectories of the field, identifies critical research gaps in geometry–semantics integration, relational reasoning, dynamic modeling, and task-driven evaluation, and introduces an accompanying knowledge website to provide a clear roadmap for algorithm development, benchmarking, and standardized collaboration.
This work addresses the limitations of existing 3D scene graph methods, which are constrained by predefined relationship categories and struggle to capture open-ended semantics and causal connections. To overcome this, the authors propose a novel framework that integrates vision-language models (VLMs) with large language models (LLMs) to construct a hierarchical forest of 3D semantic scene graphs. The VLM extracts instance-level nodes and geometry-aware relationships, while the LLM performs high-level reasoning to generate abstract concepts and open-vocabulary semantic associations. This approach transcends the confines of closed relationship sets, substantially enhancing the semantic depth and expressiveness of scene representations. Experiments on uHumans2 and ScanNet demonstrate improved accuracy in relationship generation, and real-world deployment on a Spot robot successfully enables open-vocabulary object retrieval in physical environments.
This work addresses the challenge of generating scene graphs with hierarchical structure and strong dependencies from natural language by proposing a dependency-aware, hierarchy-constrained discrete diffusion model. It introduces, for the first time, dependency-aware mechanisms and hierarchical constraints into a discrete diffusion framework, decoupling structural and semantic modeling to separately handle conditional dependencies among objects, edges, and relationships during both forward and reverse diffusion processes. The approach further enables text-aligned sampling without requiring additional training. Experimental results demonstrate that the proposed model outperforms existing continuous and discrete graph generation methods on standard scene graph benchmarks, achieving superior performance in both graph structure and layout metrics. When applied to downstream image generation tasks, it significantly improves compositional alignment quality in multi-object scenes.