retrieval-augmented diagram completion

Designs and builds systems that retrieve semantically and topologically compatible reference diagrams and use them to infer and synthesize missing content, connectivity/topology, and visual attributes to complete partially observed or underspecified diagrams. Analyzes retrieval strategies, reference selection and alignment, and the priors and fusion methods that guide downstream diagram generation and rendering.

retrieval-augmenteddiagramcompletion

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of generating high-quality scientific figures from incomplete hand-drawn sketches, where existing methods struggle to jointly preserve semantic content and topological structure. The authors propose a lightweight retrieval-augmented framework that, for the first time, integrates knowledge graphs with multi-granularity sketch variants to establish a structure-aware retrieval mechanism. By representing chart semantics via a knowledge graph, synthesizing multi-level simplified sketch variants, and training a shared embedding model, the approach achieves joint structural-semantic alignment between sketches and reference figures within a unified embedding space, further guided by visual priors during generation. Evaluated on DiagramBank and FigureBench, the method achieves F1 scores of 0.848 and 0.802, respectively, a VLM score of 7.170, and reduces single-sample inference latency to 35.48 seconds.

diagram synthesisretrieval-augmented generationscientific diagram generation

Vision-language models (VLMs) suffer from low accuracy and poor generalization in business document chart understanding due to inherent visual recognition limitations. Method: We propose a text-only paradigm for chart structure understanding—bypassing image-based analysis entirely and instead parsing structured metadata (e.g., shapes, connectors) directly from editable source files (XLSX/PPTX/DOCX) at the Office Open XML (OOXML) level, then feeding this structured input to large language models (LLMs) for relational reasoning and question answering. Contribution/Results: By eliminating VLMs’ visual bottlenecks and leveraging fine-grained XML parsing with structure-aware prompt engineering, our approach achieves high-precision semantic parsing. On system design document QA tasks, it significantly outperforms VLM baselines. Robust cross-format generalization is validated across PPTX, DOCX, and XLSX, demonstrating strong adaptability to real-world business scenarios. This work establishes a new, interpretable, cost-effective, and high-accuracy pathway for document intelligence.

Bypasses visual recognition limitations of Vision-Language ModelsEnhances diagram understanding with text-driven methodUtilizes source file metadata for accurate structure analysis

This study addresses a critical gap in the theoretical understanding of high-dimensional structured data by introducing a novel framework that integrates sparse representation with geometric deep learning. The proposed method leverages intrinsic manifold structures to enforce consistency across heterogeneous data modalities, significantly improving robustness under noise and missing observations. Through rigorous theoretical analysis, the authors establish convergence guarantees and sample complexity bounds that scale favorably with ambient dimensionality. Extensive experiments on benchmark datasets demonstrate state-of-the-art performance in tasks ranging from semi-supervised classification to cross-modal retrieval, outperforming existing approaches by substantial margins. The work further provides actionable insights into the interplay between data geometry, sparsity, and generalization, offering a principled foundation for future research in scalable and interpretable representation learning.

dataset curationdocument contextfigure corpora

This study addresses the absence of benchmarks and inherent modeling challenges in chart topology extraction using vision-language models (VLMs) by proposing a systematic solution. Methodologically, it introduces a novel symbolic generation technique to construct a large-scale, precisely aligned benchmark for chart topology. Furthermore, a structured reasoning framework is designed to decouple the complex extraction process into two subtasks—node enumeration and source-conditioned edge prediction—for supervised training. Experimental results demonstrate that this approach substantially enhances the topological understanding capabilities of open-source VLMs, achieving state-of-the-art average edge F1 scores on both internal and external real-world benchmarks.

benchmarkdiagram topology extractiondiagram-to-graph

This work addresses the challenge faced by visually impaired users in accessing node-link diagrams commonly distributed as bitmap images, a task for which existing assistive technologies are ill-suited due to their reliance on structured data rather than visual input. The paper presents the first lightweight deep learning approach for semantic segmentation of such diagram images, training a compact model on a large-scale synthetic dataset to achieve pixel-level parsing. The proposed method attains over 93% pixel accuracy on synthetic data and demonstrates strong performance both quantitatively and qualitatively. By enabling precise extraction of diagram semantics directly from rasterized images, this approach establishes a viable foundation for non-visual interaction and effectively bridges a critical gap in accessibility technology for bitmap-based graphical content.

accessibilityassistive technologybitmap images

Latest Papers

What's happening recently
View more

This study addresses the challenge of transferring figure-generation expertise from a single paper to the first figure of a new paper in scientific chart generation. To this end, it proposes the SIVIA-RSI framework, which enables effective cross-paper reuse of plotting skills through source-text anchoring. Methodologically, this work pioneers linking critique feedback to source paragraphs, constructing a persistent skill library with constrained editing, and decoupling candidate competition from skill acceptance mechanisms. The evaluation integrates comprehensive candidate assessment, source-grounding analysis, and automated preference selection. Experimental results demonstrate that the optimal candidate achieves an accuracy of 87.5%, outperforming the baseline at 83.3%. Furthermore, the findings reveal the limitation that local improvements alone cannot guarantee consistent transferability across papers.

diagram generationreusable skillsscientific diagrams

This study addresses the challenge that novice designers often struggle to extract deep concepts and visual motifs from reference images, frequently falling into superficial imitation. To overcome this limitation, this work proposes ReVision, an AI-assisted design tool grounded in multimodal parsing techniques. ReVision achieves, for the first time, the decoupling of conceptual interpretation from visual motifs, enabling cross-space recombination and generative rendering to produce divergent visual variants. By bridging the gap left by existing tools in deep semantic association and expressive diversity control, ReVision effectively mitigates cognitive fixation induced by reference imagery. Consequently, it significantly enhances the divergency of design exploration, empowering novice designers to move beyond surface-level replication toward more innovative ideation.

conceptual interpretationdesign fixationdivergent exploration

Existing vision-language model evaluation benchmarks inadequately address the distinctive characteristics of engineering drawings, such as dense layouts, specialized symbols, and cross-references between text and graphics. This work introduces the first open dataset and benchmark specifically designed for complex engineering drawings, proposing two tasks: structured parts-list extraction and free-form visual question answering. The study employs zero-shot learning, chain-of-thought prompting, and LLM-as-judge calibration for systematic evaluation. Experimental results show that state-of-the-art models achieve part-recognition recall rates of 0.61–0.87 but exhibit poor performance in descriptive token F1 scores (0.03–0.18) and generally underperform on visual question answering, revealing significant deficiencies in technical semantic understanding and factual reasoning. Furthermore, the findings demonstrate that conventional overlap-based metrics substantially underestimate models’ capacity for technical description.

benchmarkdatasetengineering diagrams

This study investigates whether structured text can substitute for vision in chart reasoning, explicitly disentangling errors arising from missing representational information versus insufficient solver capability. We propose a diagnostic protocol that contrasts original images, model-extracted text, and source structures, integrating edge interventions with token efficiency analysis to precisely localize structural bottlenecks during the textualization process. Experimental results demonstrate that answer-relevant topology predicts accuracy more effectively than global topology. Furthermore, source structure achieves 87% accuracy, whereas both direct visual processing and learned textual representations fall below 30%. Notably, editing a single critical edge suffices to reduce accuracy to zero. This work establishes a fine-grained evaluation framework for multimodal chart reasoning.

diagram reasoninginformation fidelitystructural bottlenecks

Hot Scholars

JS

Jean-Simon Pacaud Lemay

Macquarie University, Senior Lecturer/Associate Professor, DECRA Fellow
MathematicsCategory TheoryDifferential Categories
XL

Xiaodan Liang

Professor of Computer Science, Sun Yat-sen University, MBZUAI, CMU, NUS
Computer visionEmbodied AIMachine learning
SC

Shuhang Chen

Zhejiang University
deep learningcomputer vision
YZ

Yifan Zhu

Beijing University of Posts and Telecommunications
PEFT of LLMsGraph RAGGraph mining
TT

Thanh Tran

Senior Applied Scientist- Amazon
Simulation AgentResponsible AIRecommender SystemsNLP applications