diagram graph embedding

Designs and builds graph-structured representations (knowledge graphs) of diagrams and sketches and methods to map those graphs into vector embedding spaces; this includes learning joint sketch–diagram embeddings for alignment, grounding visual elements to graph nodes, and augmenting the knowledge graphs to improve representation, retrieval, synthesis, or reasoning.

diagramgraphembedding

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.36
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of generating high-quality scientific figures from incomplete hand-drawn sketches, where existing methods struggle to jointly preserve semantic content and topological structure. The authors propose a lightweight retrieval-augmented framework that, for the first time, integrates knowledge graphs with multi-granularity sketch variants to establish a structure-aware retrieval mechanism. By representing chart semantics via a knowledge graph, synthesizing multi-level simplified sketch variants, and training a shared embedding model, the approach achieves joint structural-semantic alignment between sketches and reference figures within a unified embedding space, further guided by visual priors during generation. Evaluated on DiagramBank and FigureBench, the method achieves F1 scores of 0.848 and 0.802, respectively, a VLM score of 7.170, and reduces single-sample inference latency to 35.48 seconds.

diagram synthesisretrieval-augmented generationscientific diagram generation

This work addresses the lack of interpretability and structural awareness in knowledge graph node embeddings by proposing a plug-and-play, general-purpose graph node representation method. The approach jointly models local proximity and long-range structural correlations through a multi-source subspace comprising hop-based topology, label overlap, Markov transition probabilities, and Recursive Spectral Bisection (RSB) clustering indices. It introduces, for the first time, a ground-truth-guided loss function that jointly estimates Jaccard similarity and label overlap. A multivariate stochastic gradient descent framework is designed to jointly optimize subspace weights. Experiments on multiple benchmark datasets demonstrate significant improvements in node similarity retrieval accuracy. The resulting embeddings are directly applicable to downstream tasks without fine-tuning, and each subspace’s contribution is both interpretable and ablatable.

Complex SimilarityKnowledge GraphNode Representation

Existing graph embedding methods overly rely on explicit topological structures, limiting their ability to capture implicit semantic and transitive relationships in sparse graphs. To address this, we propose a novel knowledge-augmented graph embedding paradigm: for the first time, we integrate a decay-based logical inference mechanism for knowledge completion into the graph preprocessing stage, enabling automatic discovery and strengthening of implicit edges and dynamic topology reconstruction. Subsequently, we combine GraphSAGE and Node2Vec to generate high-quality node representations. Our approach significantly improves the geometric structure and semantic coherence of the embedding space. Extensive experiments on multiple benchmark datasets demonstrate substantial gains in node classification and link prediction performance. These results empirically validate the critical importance of modeling implicit knowledge for graph representation learning.

Addresses implicit knowledge gaps in graph embeddings from sparse datasetsProposes knowledge completion to uncover latent semantics before embedding generationTransforms graph topology and representation quality through inference functions

Subgraph matching is critical for applications such as knowledge graph question answering and molecular design, yet existing neural approaches only explore fragmented regions of the graph matching network design space. This paper presents the first systematic exploration of this unified design space, identifying three core architectural axes: inter-graph interaction mechanisms (attention vs. soft permutation), node/edge alignment strategies, and scoring network architectures. Based on these, we propose a composable and scalable neural graph matching framework and uncover synergistic gains arising from cross-dimensional design choices. Empirical evaluation across multiple subgraph matching benchmarks shows that our optimal configuration significantly outperforms state-of-the-art methods. Furthermore, we distill generalizable design principles—yielding transferable architectural insights and practical guidelines for neural graph matching. (136 words)

Exploring neural graph representations for subgraph matching tasksIdentifying optimal combinations of neural representation and interaction methodsUnifying design space for graph matching networks across applications

This paper addresses the challenge of structured semantic understanding in visual narratives (e.g., comics). We propose a hierarchical multimodal knowledge graph framework that decomposes narratives into three granular levels: story arcs, event segments, and panels—unifying semantic and spatiotemporal modeling across levels. Our key innovation is a novel multi-granularity alignment mechanism enabling panel-level visual–textual coupling and cross-level symbolic reasoning. The framework constructs multimodal graphs, fuses them hierarchically, and adapts to a manually annotated Manga109 subset. Evaluated on four tasks—action retrieval, dialogue tracking, character mapping, and panel temporal reconstruction—it achieves high precision and recall. Experiments demonstrate significant advantages in interpretability, consistent multimodal representation, and cross-task generalization.

Enabling interpretable reasoning for narrative understanding tasksIntegrating multimodal elements across panel, event, and macro-event levelsModeling hierarchical semantic structures in visual narratives

Latest Papers

What's happening recently
View more

Existing approaches treat graphs merely as external knowledge sources fed into large language models, overlooking their structural role in organizing reasoning processes. This work proposes, for the first time, leveraging graphs as internal reasoning scaffolds by employing schematic mind maps to guide models through multi-hop structured reasoning. The method integrates supervised fine-tuning with KL divergence-based knowledge distillation to train a student model to construct reasoning trajectories aligned with visual graph structures. Experimental results demonstrate that, even when direct answer cues are removed, this graph-guided strategy substantially outperforms flattened textual graph representations, enhancing answer quality without compromising reasoning efficiency. These findings validate the efficacy of graphs as intrinsic tools for structuring and guiding complex reasoning in language models.

graph scaffoldslarge language modelsmulti-hop question answering

This work explores how graph structures can enhance large language models (LLMs) to mitigate hallucinations, strengthen complex reasoning, and improve comprehension of structured data. To this end, the authors propose a systematic graph-augmented LLM paradigm that reduces factual errors through dynamic knowledge injection, introduces graph-based prompting strategies—such as Graph-of-Thought—to reinforce reasoning chains, and integrates graph neural networks with knowledge graphs to model structured information from domains like e-commerce, source code, and relational databases. Experimental results demonstrate that the proposed approach significantly lowers hallucination rates and achieves notable performance gains across diverse structured reasoning tasks, while also opening new avenues for graph-based sparse architectures and brain-inspired memory systems.

GraphsHallucinationLarge Language Models

Existing visual graph recognition methods are often confined to specific tasks and lack generalizability and cross-scenario transferability. This work proposes GraSP, an end-to-end framework based on subgraph prediction that jointly models graph structure and visual features to enable unified recognition of diverse graph types and rendering styles. GraSP achieves cross-task transfer without task-specific fine-tuning, representing the first general-purpose and transferable approach for visual graph recognition. Evaluated on multiple synthetic benchmarks and a real-world application, GraSP demonstrates exceptional generalization and adaptability, advancing the field toward a unified paradigm for graph recognition.

graph recognitionsubgraph predictiontransferability

This study addresses the lack of systematic understanding regarding the applicability, timing, and methodology of integrating graph structures with large language models (LLMs) across diverse scenarios. It proposes a structured analytical framework that categorizes existing graph-LLM approaches according to task objectives, graph types, and integration strategies—encompassing techniques such as prompt engineering, data augmentation, joint training, and agent-based architectures—and covering multiple graph modalities including knowledge graphs and causal graphs. Through cross-domain evaluation, the work identifies optimal practices and boundary conditions for various fusion schemes in tasks like reasoning and retrieval, offering researchers a principled guideline for selecting appropriate methods based on task requirements, data characteristics, and reasoning complexity.

generative AIGraph-LLM integrationreasoning

Hot Scholars

CB

Carlos Bobed

Assistant Professor at University of Zaragoza, Spain
Semantic WebOntologiesKnowledge GraphsNLP
AS

Akrati Saxena

LIACS, Leiden University, Netherlands
Social Network AnalysisComplex NetworksMachine LearningSocial Computing
LS

Luciano Serafini

Head of Data and Knowledge Management Research Unit, Fondazione Bruno Kessler, Trento, Italy
Knowledge representationArtificial intelligenceSemantic WebApplied Ontology
AB

Alessandro B. Melchiorre

PostDoc at the Johannes Kepler University Linz, Austria
Recommender SystemMachine learning ExplainabilityMusic Recommender SystemBias and Fairness in
JQ

Jianzhong Qi

School of Computing and Information Systems, The University of Melbourne
Spatio-temporal data managementKnowledge basesMachine learning