Score
Designs and builds 3D scene-graph representations in which nodes are 3D entities and edges encode spatial, semantic, and relational links among them, producing a relational graph that captures scene structure. Implements methods to infer and label edges from multi-view or 3D inputs (including using vision–language or other reasoning), and to enforce geometric plausibility and reduce or learn graph connectivity without manual edge annotations.
Current 3D semantic scene graph prediction methods rely on graph neural networks but suffer from insufficient discriminability and representational capacity in object and relational feature encoding. To address this, we propose a decoupled representation learning framework: first, a highly discriminative object feature encoder is designed, integrating geometric-semantic multimodal fusion; second, an object-centric contrastive pre-training strategy is introduced to explicitly decouple object representation learning from graph structure prediction. Crucially, our method requires no architectural modifications to downstream graph inference modules and can be seamlessly integrated as a plug-in enhancement. Evaluated on the 3DSSG benchmark, our approach significantly outperforms state-of-the-art methods, achieving consistent improvements in both object classification and relationship prediction—the two core evaluation metrics—thereby validating the effectiveness of decoupled representation learning for 3D scene semantic understanding.
Existing 3D scene graph prediction methods predominantly adopt object-centric graph neural networks (GNNs), which struggle to capture high-order relational dependencies among entities. Method: This paper proposes a relation-centric inference paradigm: it transforms the original object-centric graph into a line graph—where relations become nodes and higher-order associations become edges—and designs a novel edge-centric line graph neural network. A link-guided mechanism is introduced to suppress noisy relations, enabling progressive fusion of relation-level contextual information into object-level understanding. The framework is modular and compatible with arbitrary baseline methods. Key components include line graph construction, object-aware feature fusion, link prediction, and multi-granularity message passing. Results: On the 3DSSG benchmark, our method consistently outperforms two strong baselines across all metrics, validating both the effectiveness and generalizability of the relation-to-object reasoning paradigm.
Addressing the challenge of robust 3D semantic scene graph generation from multi-view RGB images without 3D ground-truth annotations, this paper proposes an end-to-end framework. First, geometric priors are obtained via multi-view depth estimation and pseudo-point cloud reconstruction. Second, semantic masks guide cross-view feature aggregation to suppress background noise. Third, topological relationships among nodes and edges within one-hop neighborhoods are explicitly modeled, and a statistical-prior-based confidence rescaling mechanism jointly optimizes object, predicate, and relation predictions. Crucially, the method operates entirely without 3D supervision. Experiments demonstrate significant improvements in both accuracy and structural stability of 3D scene graphs, outperforming state-of-the-art unsupervised and weakly supervised approaches on mainstream benchmarks.
This work addresses the fragmented state of 3D scene graph research, hindered by the absence of a unified definition, construction pipeline, and evaluation protocol, which impedes method comparison and real-world deployment. The paper presents the first formal definition and a cohesive theoretical framework for 3D scene graphs, systematically analyzing key modeling choices—including node and edge attributes, hierarchical structure, dynamic modeling, and functional awareness—and clarifying the terminology and mainstream methodologies that map perceptual data to scene graphs. Through a comprehensive literature review, taxonomic analysis, and technical comparison, it delineates the core components and evolutionary trajectories of the field, identifies critical research gaps in geometry–semantics integration, relational reasoning, dynamic modeling, and task-driven evaluation, and introduces an accompanying knowledge website to provide a clear roadmap for algorithm development, benchmarking, and standardized collaboration.
This work addresses the limitations of existing 3D scene graph methods, which are constrained by predefined relationship categories and struggle to capture open-ended semantics and causal connections. To overcome this, the authors propose a novel framework that integrates vision-language models (VLMs) with large language models (LLMs) to construct a hierarchical forest of 3D semantic scene graphs. The VLM extracts instance-level nodes and geometry-aware relationships, while the LLM performs high-level reasoning to generate abstract concepts and open-vocabulary semantic associations. This approach transcends the confines of closed relationship sets, substantially enhancing the semantic depth and expressiveness of scene representations. Experiments on uHumans2 and ScanNet demonstrate improved accuracy in relationship generation, and real-world deployment on a Spot robot successfully enables open-vocabulary object retrieval in physical environments.
This work addresses the limitation of existing open-vocabulary 3D scene understanding methods, which often rely on context-free semantic representations and neglect the critical role of inter-object relationships in semantic refinement. To overcome this, the authors propose a relationship-aware 3D scene graph construction approach that requires no manual relation annotations. The method leverages vision-language reasoning to infer object relationships and employs multi-view geometric constraints to eliminate implausible connections. Furthermore, an adaptive gated dual-stream graph attention network is introduced to disentangle and fuse geometric and semantic features effectively. Hierarchical contrastive learning is incorporated to enhance both instance-level consistency and category-level discriminability. Evaluated on multiple benchmarks—including ScanNetV2, ScanNet200, ScanNet++, and Replica—the proposed approach significantly improves open-vocabulary 3D understanding performance and generalization capability.
This work addresses the challenge of modeling component assembly relationships in 2D scenes under few-shot settings without semantic labels. The authors propose an end-to-end scene graph generation method that integrates geometric feature extraction with structural reasoning. Specifically, Faster R-CNN is employed to extract geometric representations of components, and a Transformer architecture constructs an initial adjacency matrix. To refine relational inference, the approach incorporates an attention-based Graph Convolutional Network (aGCN) with a message-passing mechanism. Notably, the model operates without semantic supervision and achieves accurate recovery of ground-truth assembly relationships using only a minimal number of training samples. Experimental results on a toy vehicle dataset demonstrate the method’s effectiveness, significantly advancing the capability to model assembly relationships in semantically unlabeled scenarios.
Existing methods struggle to simultaneously model contextual relationships among objects and shape diversity, often resulting in distorted 3D scene layouts. To address this limitation, this work proposes a compositional 3D scene generation framework grounded in semantic scene graphs. The approach first constructs a semantic scene graph from an RGB image sequence, then employs a graph neural network enhanced with a cross-verified feature attention mechanism to predict scene structure. Furthermore, a graph variational autoencoder (Graph-VAE) is designed to jointly generate object shapes and layouts by integrating shape and layout priors. Evaluated on the 3RScan/3DSSG and SG-FRONT datasets, the method significantly outperforms existing approaches, generating semantically consistent and structurally plausible 3D scenes even in complex indoor environments and under strong constraints, thereby enabling high-quality personalized mixed reality content creation.
Existing methods for 3D scene graph generation are hindered by scarce annotated data and susceptibility to object priors, making it challenging to design effective self-supervised pretraining tasks. This work proposes a topological layout learning framework that, for the first time, formulates predicate-aware topological layout reconstruction as a self-supervised objective. By modeling spatial priors conditioned on anchor points and leveraging graph neural networks for topological-geometric reasoning, the approach recovers the global structure of subgraphs. To preserve semantic fidelity, the method incorporates structure-aware multi-view augmentation and enhances relational representations through self-distillation. Evaluated on the 3DSSG dataset, the proposed framework significantly outperforms current state-of-the-art baselines, demonstrating its effectiveness and robustness.