Score
Designs and implements systems that construct and iteratively refine scene graphs by using large language models to predict, validate, and update relational edges among objects. These methods combine geometry-initialized relational priors with LLM-based graph refinement to reduce spurious relation predictions and produce more efficient, accurate relational graph representations.
Existing 3D scene graph prediction methods predominantly adopt object-centric graph neural networks (GNNs), which struggle to capture high-order relational dependencies among entities. Method: This paper proposes a relation-centric inference paradigm: it transforms the original object-centric graph into a line graph—where relations become nodes and higher-order associations become edges—and designs a novel edge-centric line graph neural network. A link-guided mechanism is introduced to suppress noisy relations, enabling progressive fusion of relation-level contextual information into object-level understanding. The framework is modular and compatible with arbitrary baseline methods. Key components include line graph construction, object-aware feature fusion, link prediction, and multi-granularity message passing. Results: On the 3DSSG benchmark, our method consistently outperforms two strong baselines across all metrics, validating both the effectiveness and generalizability of the relation-to-object reasoning paradigm.
This work investigates large language models’ (LLMs) capacity to understand and generate scene graphs under complex narrative inputs. To this end, we introduce TSG Bench—the first bidirectional text-to-scene-graph benchmark tailored for LLMs—comprising a dual-task evaluation framework for scene graph understanding and generation. We propose structured prompting strategies and fine-grained triplet-level metrics (entity/attribute/relation) to assess structural fidelity. Extensive zero-shot and few-shot evaluations across 11 state-of-the-art LLMs reveal strong understanding performance (82.4% average F1), but critically deficient generation capability (39.7% average F1), especially in temporal decomposition of multi-event narratives. Our core contributions are threefold: (1) the first comprehensive bidirectional text↔scene-graph evaluation benchmark; (2) empirical identification of a fundamental bottleneck in LLMs’ structured visual-semantic generation; and (3) provision of a standardized diagnostic toolkit to advance controllable, faithful scene graph generation research.
该研究解决了3D场景图中不确定性表示和传播问题,通过引入概率场景图(PSG)及高斯层次图(HGG),实现了实时感知与精确定位。
To address the limited relational reasoning capability of embodied agents in 3D scenes—hindering natural language interaction—this paper proposes a learnable 3D semantic graph representation framework that explicitly models objects alongside their spatial, functional, and semantic relations. This structured graph is injected as input into large language models (LLMs) to enable vision-language joint reasoning. It marks the first effort to deeply integrate structured 3D scene graphs with LLMs, overcoming conventional limitations of coordinate- or point-cloud–based relational modeling. The methodology comprises four components: 3D scene graph construction, graph neural network–based representation learning, LLM instruction fine-tuning, and multimodal alignment prompt engineering. Evaluated on six benchmarks—including ScanRefer and ScanQA—the approach achieves average improvements of 5.2%–9.8% across referring localization, visual question answering, and description generation tasks, significantly outperforming state-of-the-art methods.
Existing scene graph generation models struggle to learn reliable visual commonsense under sparse annotations, leading to significantly degraded performance on rare relationships. This work proposes a model-agnostic, semantics-guided knowledge refinement framework that automatically uncovers spatial, functional, and qualitative relational patterns from training data during inference and dynamically corrects predictions through declarative commonsense reasoning. Requiring neither handcrafted rules nor model retraining, the method is readily transferable across datasets and architectures, and represents the first effective integration of structured visual commonsense reasoning into purely learning-driven scene graph generation pipelines. It consistently outperforms strong baselines across three standard benchmarks, demonstrating the critical role of visual commonsense in enhancing the robustness and accuracy of scene graph generation.
Existing 3D perception methods often rely on object-centric modeling or extensive scene-specific training, hindering unified and efficient open-vocabulary reasoning. This work proposes a training-free, unified framework that constructs a hierarchical 3D scene representation by distilling language-aligned Gaussian splats, refines geometry through Gaussian pruning, and aggregates multi-view 2D features via language-guided alignment to produce precise 3D object embeddings. Building upon this representation, the method constructs an open-vocabulary 3D semantic scene graph that jointly models hierarchical semantics and intra- and inter-object relationships, enabling unified reasoning across segmentation, retrieval, and relational understanding. Experiments demonstrate that the approach is both efficient and scalable across multiple tasks.
This work addresses the limitation of existing open-vocabulary 3D scene understanding methods, which often rely on context-free semantic representations and neglect the critical role of inter-object relationships in semantic refinement. To overcome this, the authors propose a relationship-aware 3D scene graph construction approach that requires no manual relation annotations. The method leverages vision-language reasoning to infer object relationships and employs multi-view geometric constraints to eliminate implausible connections. Furthermore, an adaptive gated dual-stream graph attention network is introduced to disentangle and fuse geometric and semantic features effectively. Hierarchical contrastive learning is incorporated to enhance both instance-level consistency and category-level discriminability. Evaluated on multiple benchmarks—including ScanNetV2, ScanNet200, ScanNet++, and Replica—the proposed approach significantly improves open-vocabulary 3D understanding performance and generalization capability.
This work addresses the fragmented state of 3D scene graph research, hindered by the absence of a unified definition, construction pipeline, and evaluation protocol, which impedes method comparison and real-world deployment. The paper presents the first formal definition and a cohesive theoretical framework for 3D scene graphs, systematically analyzing key modeling choices—including node and edge attributes, hierarchical structure, dynamic modeling, and functional awareness—and clarifying the terminology and mainstream methodologies that map perceptual data to scene graphs. Through a comprehensive literature review, taxonomic analysis, and technical comparison, it delineates the core components and evolutionary trajectories of the field, identifies critical research gaps in geometry–semantics integration, relational reasoning, dynamic modeling, and task-driven evaluation, and introduces an accompanying knowledge website to provide a clear roadmap for algorithm development, benchmarking, and standardized collaboration.
本文探讨了如何利用视觉表示来增强图神经网络的推理和学习能力,提出了视觉与图结合的新方法,并归纳了三个研究方向。
Existing 3D scene graph generation methods are object-centric and struggle to model part-level details and multi-level relationships, limiting fine-grained scene understanding. This work proposes the first open-vocabulary, part-aware unified 3D scene graph framework that jointly represents objects, interactable parts, spatial and functional relationships, and affordances. By integrating object-part knowledge-guided detection, part-aware 3D feature fusion, geometry-prior-initialized relation modeling, and joint optimization with large language models, our approach enables efficient and accurate relational reasoning. We introduce a new benchmark, UniGraph3D, on which our method achieves state-of-the-art performance and significantly enhances perception for a variety of robotic tasks.