Score
Design and build knowledge-graph representations of charts that capture chart entities (e.g., data series, axes, groups, and annotations) at multiple levels of granularity and organize them hierarchically. These graphs explicitly encode inter-chart relations and the analytical structure of charts (such as measures, aggregations, and mappings between visual elements and data) to enable retrieval, cross-chart linking, and structured analysis.
Existing cross-chart retrieval-augmented generation (RAG) benchmarks are limited by high lexical overlap between queries and evidence and inconsistent reasoning chains, hindering their ability to support complex multimodal analysis. This work proposes ChartWalker, a framework that constructs chart-oriented hierarchical knowledge graphs and employs a structure-aware sampling algorithm to explicitly control query difficulty and granularity, thereby synthesizing high-quality question-answer pairs with multi-hop reasoning paths. Leveraging this approach, we introduce ChartWalker-Bench, the first comprehensive cross-chart RAG benchmark, which exposes significant performance bottlenecks in current state-of-the-art methods. To facilitate future research, we also release ChartWalker-Agent, an open-source agent baseline designed for this challenging task.
This paper addresses the challenge of cross-scale modeling in data spaces by proposing a Multi-level Graph (MLG) structure that enables multi-granularity data abstraction—from local to global. We formalize topological contraction and expansion operations to establish an incremental, invertible graph transformation algebra, unifying the representation of both structured and unstructured data. Unlike conventional single-layer graph models, our approach achieves the first semantic-preserving hierarchical compression and expansion of data. Experiments on a real-world dream report dataset demonstrate a 42% improvement in cross-granularity exploratory efficiency and significantly enhanced semantic coherence. This work introduces a scalable, interpretable, and dynamically evolvable foundational representation paradigm for data spaces.
Existing multimodal large models (MLLMs) suffer from poor interpretability in complex chart understanding and reasoning, and lack comprehensive, fine-grained evaluation benchmarks. Method: We introduce ChartX—the first benchmark covering 18 chart types, 7 reasoning tasks, and 22 academic domains—and propose ChartVLM, a dedicated chart foundation model. ChartVLM innovatively integrates chart-structure-aware visual encoding, multimodal collaborative representation learning, and task-adaptive instruction tuning to enhance interpretability in pattern recognition. Contribution/Results: On ChartX, ChartVLM significantly outperforms mainstream MLLMs and matches the performance of GPT-4V. Both the open-source code and ChartX dataset have been widely adopted by the research community. This work bridges two critical gaps in chart understanding: the absence of a systematic, domain-diverse evaluation framework and the lack of specialized, interpretable modeling architectures.
To address the challenge of fine-grained classification in scientific chart accessibility, this paper proposes a coarse-to-fine curriculum learning framework grounded in inter-class similarity, emulating human cognitive progression through dynamically constructed, incremental training sequences. Methodologically, it integrates hierarchical category modeling, feature-space similarity measurement, and fine-tuning of deep classification networks to enable controllable, stage-wise training. Its key innovation lies in explicitly modeling inter-class similarity as the principled basis for curriculum design—establishing, to the best of our knowledge, the first generalizable curriculum learning paradigm for chart classification. Evaluated on the ICPR 2022 CHART-Infographics dataset, the method substantially outperforms prior state-of-the-art approaches, achieving significant gains in fine-grained classification accuracy. This work advances scientific chart understanding and contributes a novel, principled approach to improving accessibility for visually impaired users.
Large language models (LLMs) struggle to effectively comprehend graph-structured data due to their inherent sequence-based architecture and lack of native graph-aware representations. Method: This paper introduces *graph laws*—statistically derived, topologically parameterized features that are interpretable as natural language descriptions—establishing a novel paradigm for representing graphs as LLM-compatible inputs. We systematically construct a multi-dimensional graph law framework spanning macro/micro scales, low/high orders, and static/dynamic properties, integrating graph-theoretic analysis, multi-scale observational modeling, and natural language alignment techniques, while establishing semantic mappings to downstream graph tasks and retrieval-augmented generation (RAG) scenarios. Results: Experiments demonstrate that graph laws substantially mitigate LLM hallucination, overcome context-length limitations, and enable end-to-end graph reasoning. The approach achieves strong generalization across diverse domains, including molecular design, recommender systems, and protein structure modeling.
Existing visualization design knowledge bases (e.g., Draco) suffer from incomplete training corpora and insufficient coverage of design variants, hindering systematic evaluation of design trade-offs—resulting in low feature coverage and suboptimal recommendation accuracy. To address this, we propose a data augmentation method grounded in design permutation and identification of under-evaluated features, enabling automated generation of high-quality chart contrast pairs. We further introduce a scalable, multi-strategy annotation framework coupled with model-fitting-driven feature importance analysis to dynamically update and optimize knowledge base feature weights. Experimentally, we construct an expanded corpus comprising thousands of novel chart pairs and validate our approach within the Draco system: feature coverage increases by 32%, and chart recommendation accuracy improves significantly. This work marks the first systematic knowledge enhancement targeting the design trade-off space.
Current vision-language models (VLMs) lack the capability to perform joint reasoning across multiple semantically related charts. Method: We introduce InterChart, the first diagnostic benchmark for multi-chart reasoning, comprising synthetic aligned chart sets and real-world chart pairs. It features a three-tiered task hierarchy—entity inference, trend correlation, and multi-step abstract reasoning—to systematically evaluate VLMs’ semantic integration across 2–3 topically or structurally related charts. We propose a hierarchical evaluation framework and a “decompose-distribute” chart information processing mechanism to explicitly model cross-chart reasoning paths. Contribution/Results: Experiments reveal that state-of-the-art open- and closed-source VLMs suffer significant performance degradation as chart complexity increases; visual decomposition notably improves reasoning accuracy. InterChart is the first benchmark to uncover systematic limitations of VLMs in collaborative multi-chart understanding, providing an interpretable, scalable diagnostic tool for complex multimodal visual reasoning.
This study addresses the lack of consistent boundary definitions for chart types—such as Gantt charts—which complicates alignment in design space construction, grammar generation, and perceptual research. To resolve this, the authors propose a functional definition approach that distinguishes essential from variable features of chart types, establishing a boundary reasoning framework. Through conceptual analysis, design space modeling, and case studies, they validate the method and uncover hidden structures like feature entanglement, thereby making scope selection explicit. The approach yields a shared vocabulary and analytical tools for boundary analysis, demonstrated across Gantt charts, radar charts, and table maps. By clarifying how chart type definitions influence generalizability, this work offers a novel theoretical perspective for visualization research.
Industrial standards pose significant challenges for knowledge modeling due to their broad technical scope, intricate regulatory logic, and heterogeneous structure—including tables, constraints, exceptions, and numerical computations. To address this, we propose a hierarchical propositionalization-based joint modeling approach that, for the first time, unifies conditional logic, numerical rules, and tabular semantics within an ontology-enhanced knowledge graph (KG), substantially improving semantic expressivity for nested constraints and scoping relationships. Integrating large language model–driven triple extraction, structured table recognition, and ontology alignment, we construct a KG-RAG framework supporting multi-hop question answering and toxic clause detection. Evaluated on a novel, multi-type benchmark dataset, our method consistently outperforms existing KG-RAG approaches, demonstrating superior effectiveness and advancement in enhancing logical inferability and intelligent management of industrial standard knowledge.