scene graph reasoning

Designs and implements structured scene representations (scene graphs) that encode objects and their semantic and spatial relationships. Builds analyzers and constraint reasoners that infer, enforce, or repair relationships such as support, contact, proximity, and occlusion, translate those constraints into actionable parameters, and detect or resolve inconsistencies for downstream processing.

scenegraphreasoning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.19
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$212K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Industrial CAD models often lack semantic, spatial, and functional information, limiting their utility in robotic simulation and high-level scene understanding. This work proposes an offline method that leverages large vision-language models (LVLMs)—introduced for the first time into CAD environments—to automatically generate structured 3D scene graphs that explicitly model manipulable objects and their functional relationships. By effectively integrating semantic parsing with functional reasoning, the approach achieves high-precision semantic annotation and relationship recognition on industrial structures such as piping systems. Both qualitative and quantitative evaluations demonstrate its effectiveness. The associated code and dataset have been made publicly available.

CADfunctional informationindustrial environment

Situational Scene Graph for Structured Human-centric Situation Understanding

Oct 30, 2024
CS
Chinthani Sugandhika
🏛️ Nanyang Technological University | Agency for Science, Technology and Research

Existing video graph representations neglect fine-grained action semantics—such as location, tools employed, and object functionality—leading to inadequate modeling of human–context relationships. To address this, we propose the Situational Scene Graph (SSG), the first video graph representation that explicitly incorporates semantic role–value structures to jointly model human–object interactions and their associated attributes. We formulate a novel task, “Situational Scene Graph Generation,” construct the first annotated SSG dataset, and design InComNet—a multi-stage, interaction-complementary network integrating semantic role labeling, graph-structured reasoning, and joint multi-task learning. Experiments demonstrate that our approach significantly outperforms baselines on predicate classification, semantic role–value classification, and situational reasoning tasks. These results validate SSG as an effective, generalizable, structured representation for human-centered contextual understanding in videos.

Action DetailsSemantic InformationVideo Understanding

FunGraph: Functionality Aware 3D Scene Graphs for Language-Prompted Scene Interaction

Mar 10, 2025
DR
Dennis Rotondi
🏛️ University of Stuttgart | University of Bonn

Existing 3D scene graph methods rely on object-level, coarse-grained representations, limiting their applicability to functional robot–environment interaction. This work proposes a fine-grained, function-oriented 3D scene graph that explicitly models functionally manipulable parts—such as door handles and light switches—as first-class nodes, enabling a semantic shift from object-level to function-level reasoning. Methodologically, we synthesize multi-source 3D data to generate 2D functional part annotations, train a part-level detector, and integrate it into standard 3D scene graph construction; we further enhance functional grounding via vision-language alignment and task-driven affordance localization. Experiments demonstrate state-of-the-art performance in functional part segmentation and significantly improved accuracy and robustness in mapping natural language instructions to executable robot actions in real-world settings.

Augment 3D scene graphs using 2D data for improved affordance grounding.Detect and store affordance-relevant object parts for functional interaction.Develop a fine-resolution 3D scene graph for robot-environment interaction.

Current 3D semantic scene graph prediction methods rely on graph neural networks but suffer from insufficient discriminability and representational capacity in object and relational feature encoding. To address this, we propose a decoupled representation learning framework: first, a highly discriminative object feature encoder is designed, integrating geometric-semantic multimodal fusion; second, an object-centric contrastive pre-training strategy is introduced to explicitly decouple object representation learning from graph structure prediction. Crucially, our method requires no architectural modifications to downstream graph inference modules and can be seamlessly integrated as a plug-in enhancement. Evaluated on the 3DSSG benchmark, our approach significantly outperforms state-of-the-art methods, achieving consistent improvements in both object classification and relationship prediction—the two core evaluation metrics—thereby validating the effectiveness of decoupled representation learning for 3D scene semantic understanding.

Decoupling object representation learning from relationship predictionEnhancing object feature discriminative capability for 3D scene graphsIntegrating geometric and semantic features for relationship prediction

The Scene Language: Representing Scenes with Programs, Words, and Embeddings

Oct 22, 2024
YZ
Yunzhi Zhang
🏛️ Stanford University | UC Berkeley

This work addresses the challenge of simultaneously achieving structural precision, semantic interpretability, and identity controllability in existing 3D/4D scene representations. We propose “Scene Language”—a unified 3D/4D scene representation framework that integrates executable program structures, natural-language semantic tokens, and visual identity embeddings. To our knowledge, this is the first method enabling zero-shot cross-modal reasoning: without fine-tuning, it directly synthesizes structured scene programs from pretrained language models and vision encoders, while explicitly modeling hierarchical relationships to support fine-grained editing. The representation is renderer-agnostic, interfacing seamlessly with traditional, neural, and hybrid renderers to produce high-fidelity images. Experiments demonstrate significant improvements over baselines—including scene graphs—on complex scene generation tasks, achieving breakthroughs in fidelity, controllability, and editability.

Enabling high-quality 3D and 4D scene generation and editingInferring scene representation from text or image inputsRepresenting visual scenes with structure, semantics, and identity

Latest Papers

What's happening recently
View more

This work addresses the challenge of modeling component assembly relationships in 2D scenes under few-shot settings without semantic labels. The authors propose an end-to-end scene graph generation method that integrates geometric feature extraction with structural reasoning. Specifically, Faster R-CNN is employed to extract geometric representations of components, and a Transformer architecture constructs an initial adjacency matrix. To refine relational inference, the approach incorporates an attention-based Graph Convolutional Network (aGCN) with a message-passing mechanism. Notably, the model operates without semantic supervision and achieves accurate recovery of ground-truth assembly relationships using only a minimal number of training samples. Experimental results on a toy vehicle dataset demonstrate the method’s effectiveness, significantly advancing the capability to model assembly relationships in semantically unlabeled scenarios.

Assembly RelationshipComponent AssemblyGeometric Representation

This work addresses the fragmented state of 3D scene graph research, hindered by the absence of a unified definition, construction pipeline, and evaluation protocol, which impedes method comparison and real-world deployment. The paper presents the first formal definition and a cohesive theoretical framework for 3D scene graphs, systematically analyzing key modeling choices—including node and edge attributes, hierarchical structure, dynamic modeling, and functional awareness—and clarifying the terminology and mainstream methodologies that map perceptual data to scene graphs. Through a comprehensive literature review, taxonomic analysis, and technical comparison, it delineates the core components and evolutionary trajectories of the field, identifies critical research gaps in geometry–semantics integration, relational reasoning, dynamic modeling, and task-driven evaluation, and introduces an accompanying knowledge website to provide a clear roadmap for algorithm development, benchmarking, and standardized collaboration.

3D Scene Graphsevaluation protocolssemantic representation

This work addresses the limitations of existing 3D scene graph methods, which are constrained by predefined relationship categories and struggle to capture open-ended semantics and causal connections. To overcome this, the authors propose a novel framework that integrates vision-language models (VLMs) with large language models (LLMs) to construct a hierarchical forest of 3D semantic scene graphs. The VLM extracts instance-level nodes and geometry-aware relationships, while the LLM performs high-level reasoning to generate abstract concepts and open-vocabulary semantic associations. This approach transcends the confines of closed relationship sets, substantially enhancing the semantic depth and expressiveness of scene representations. Experiments on uHumans2 and ScanNet demonstrate improved accuracy in relationship generation, and real-world deployment on a Spot robot successfully enables open-vocabulary object retrieval in physical environments.

3D scene graphsfoundation modelsrobotics

This work addresses the challenge of generating scene graphs with hierarchical structure and strong dependencies from natural language by proposing a dependency-aware, hierarchy-constrained discrete diffusion model. It introduces, for the first time, dependency-aware mechanisms and hierarchical constraints into a discrete diffusion framework, decoupling structural and semantic modeling to separately handle conditional dependencies among objects, edges, and relationships during both forward and reverse diffusion processes. The approach further enables text-aligned sampling without requiring additional training. Experimental results demonstrate that the proposed model outperforms existing continuous and discrete graph generation methods on standard scene graph benchmarks, achieving superior performance in both graph structure and layout metrics. When applied to downstream image generation tasks, it significantly improves compositional alignment quality in multi-object scenes.

compositional scene understandingdiscrete diffusionhierarchical dependencies

Hot Scholars

XZ

Xiatian Zhu

University of Surrey
Machine LearningComputer Vision
FT

Federico Tombari

Google, TU Munich
Computer VisionMachine Learning3D Perception
MP

Marc Pollefeys

Professor of Computer Science, ETH Zurich, and Director Spatial AI Lab, Microsoft
Computer VisionComputer GraphicsRoboticsMachine Learning
HJ

Haian Jin

Cornell University
Computer VisionSpicy Food
XL

Xiaodan Liang

Professor of Computer Science, Sun Yat-sen University, MBZUAI, CMU, NUS
Computer visionEmbodied AIMachine learning