llm-based scene-graph refinement

Designs and implements systems that construct and iteratively refine scene graphs by using large language models to predict, validate, and update relational edges among objects. These methods combine geometry-initialized relational priors with LLM-based graph refinement to reduce spurious relation predictions and produce more efficient, accurate relational graph representations.

llm-basedscene-graphrefinement

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.5
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Edge-Centric Relational Reasoning for 3D Scene Graph Prediction

Nov 19, 2025
YM
Yanni Ma
🏛️ Sun Yat-sen University | East China Normal University | University of Amsterdam

Existing 3D scene graph prediction methods predominantly adopt object-centric graph neural networks (GNNs), which struggle to capture high-order relational dependencies among entities. Method: This paper proposes a relation-centric inference paradigm: it transforms the original object-centric graph into a line graph—where relations become nodes and higher-order associations become edges—and designs a novel edge-centric line graph neural network. A link-guided mechanism is introduced to suppress noisy relations, enabling progressive fusion of relation-level contextual information into object-level understanding. The framework is modular and compatible with arbitrary baseline methods. Key components include line graph construction, object-aware feature fusion, link prediction, and multi-granularity message passing. Results: On the 3DSSG benchmark, our method consistently outperforms two strong baselines across all metrics, validating both the effectiveness and generalizability of the relation-to-object reasoning paradigm.

Current approaches struggle to capture high-order relational dependenciesExisting methods restrict relation representations to pairwise object contextObject-centric graph networks limit accurate 3D scene graph prediction

This work investigates large language models’ (LLMs) capacity to understand and generate scene graphs under complex narrative inputs. To this end, we introduce TSG Bench—the first bidirectional text-to-scene-graph benchmark tailored for LLMs—comprising a dual-task evaluation framework for scene graph understanding and generation. We propose structured prompting strategies and fine-grained triplet-level metrics (entity/attribute/relation) to assess structural fidelity. Extensive zero-shot and few-shot evaluations across 11 state-of-the-art LLMs reveal strong understanding performance (82.4% average F1), but critically deficient generation capability (39.7% average F1), especially in temporal decomposition of multi-event narratives. Our core contributions are threefold: (1) the first comprehensive bidirectional text↔scene-graph evaluation benchmark; (2) empirical identification of a fundamental bottleneck in LLMs’ structured visual-semantic generation; and (3) provision of a standardized diagnostic toolkit to advance controllable, faithful scene graph generation research.

Assess LLMs' ability to understand scene graphsEvaluate LLMs' capability to generate scene graphsIdentify limitations in complex narrative decomposition

3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding

Dec 24, 2024
TZ
Tatiana Zemskova
🏛️ Moscow Institute of Physics and Technology

To address the limited relational reasoning capability of embodied agents in 3D scenes—hindering natural language interaction—this paper proposes a learnable 3D semantic graph representation framework that explicitly models objects alongside their spatial, functional, and semantic relations. This structured graph is injected as input into large language models (LLMs) to enable vision-language joint reasoning. It marks the first effort to deeply integrate structured 3D scene graphs with LLMs, overcoming conventional limitations of coordinate- or point-cloud–based relational modeling. The methodology comprises four components: 3D scene graph construction, graph neural network–based representation learning, LLM instruction fine-tuning, and multimodal alignment prompt engineering. Evaluated on six benchmarks—including ScanRefer and ScanQA—the approach achieves average improvements of 5.2%–9.8% across referring localization, visual question answering, and description generation tasks, significantly outperforming state-of-the-art methods.

3D Spatial UnderstandingNatural Language ProcessingObject Relationships

Existing scene graph generation models struggle to learn reliable visual commonsense under sparse annotations, leading to significantly degraded performance on rare relationships. This work proposes a model-agnostic, semantics-guided knowledge refinement framework that automatically uncovers spatial, functional, and qualitative relational patterns from training data during inference and dynamically corrects predictions through declarative commonsense reasoning. Requiring neither handcrafted rules nor model retraining, the method is readily transferable across datasets and architectures, and represents the first effective integration of structured visual commonsense reasoning into purely learning-driven scene graph generation pipelines. It consistently outperforms strong baselines across three standard benchmarks, demonstrating the critical role of visual commonsense in enhancing the robustness and accuracy of scene graph generation.

Annotation SparsityKnowledge RefinementRelational Regularities

Latest Papers

What's happening recently
View more

Existing 3D perception methods often rely on object-centric modeling or extensive scene-specific training, hindering unified and efficient open-vocabulary reasoning. This work proposes a training-free, unified framework that constructs a hierarchical 3D scene representation by distilling language-aligned Gaussian splats, refines geometry through Gaussian pruning, and aggregates multi-view 2D features via language-guided alignment to produce precise 3D object embeddings. Building upon this representation, the method constructs an open-vocabulary 3D semantic scene graph that jointly models hierarchical semantics and intra- and inter-object relationships, enabling unified reasoning across segmentation, retrieval, and relational understanding. Experiments demonstrate that the approach is both efficient and scalable across multiple tasks.

3D perceptionopen-vocabularyrelational reasoning

This work addresses the limitation of existing open-vocabulary 3D scene understanding methods, which often rely on context-free semantic representations and neglect the critical role of inter-object relationships in semantic refinement. To overcome this, the authors propose a relationship-aware 3D scene graph construction approach that requires no manual relation annotations. The method leverages vision-language reasoning to infer object relationships and employs multi-view geometric constraints to eliminate implausible connections. Furthermore, an adaptive gated dual-stream graph attention network is introduced to disentangle and fuse geometric and semantic features effectively. Hierarchical contrastive learning is incorporated to enhance both instance-level consistency and category-level discriminability. Evaluated on multiple benchmarks—including ScanNetV2, ScanNet200, ScanNet++, and Replica—the proposed approach significantly improves open-vocabulary 3D understanding performance and generalization capability.

3D scene graphcontextual refinementobject relationships

This work addresses the fragmented state of 3D scene graph research, hindered by the absence of a unified definition, construction pipeline, and evaluation protocol, which impedes method comparison and real-world deployment. The paper presents the first formal definition and a cohesive theoretical framework for 3D scene graphs, systematically analyzing key modeling choices—including node and edge attributes, hierarchical structure, dynamic modeling, and functional awareness—and clarifying the terminology and mainstream methodologies that map perceptual data to scene graphs. Through a comprehensive literature review, taxonomic analysis, and technical comparison, it delineates the core components and evolutionary trajectories of the field, identifies critical research gaps in geometry–semantics integration, relational reasoning, dynamic modeling, and task-driven evaluation, and introduces an accompanying knowledge website to provide a clear roadmap for algorithm development, benchmarking, and standardized collaboration.

3D Scene Graphsevaluation protocolssemantic representation

Existing 3D scene graph generation methods are object-centric and struggle to model part-level details and multi-level relationships, limiting fine-grained scene understanding. This work proposes the first open-vocabulary, part-aware unified 3D scene graph framework that jointly represents objects, interactable parts, spatial and functional relationships, and affordances. By integrating object-part knowledge-guided detection, part-aware 3D feature fusion, geometry-prior-initialized relation modeling, and joint optimization with large language models, our approach enables efficient and accurate relational reasoning. We introduce a new benchmark, UniGraph3D, on which our method achieves state-of-the-art performance and significantly enhances perception for a variety of robotic tasks.

3D scene graphopen-vocabularypart-aware

Hot Scholars

MR

Martin R. Oswald

University of Amsterdam
3D Computer VisionRepresentation LearningApplied Machine LearningOptimization