Score
Designs and implements methods and pipelines to construct token or entity co-occurrence graphs/networks from corpora, including scalable algorithms for fast graph construction, sliding-window or temporal aggregation, and efficient indexing/lookup. Computes and interprets centrality measures and other topological metrics to rank nodes, identify hubs, brokers, and communities, and to characterize temporal and structural properties of the resulting network.
This paper addresses the fundamental challenge that large language models (LLMs) cannot natively process graph-structured data. To bridge this gap, we propose LLM4graph—a unified framework introducing the first systematic taxonomy of Graph2Text and Graph2Token paradigms, and identifying four core challenges in graph-to-text/token conversion. Methodologically, LLM4graph integrates graph encoding, structured serialization, prompt engineering, and LLM adaptation techniques to achieve multi-granular, semantics-preserving graph representation transformation. Key contributions include: (1) a principled LLM selection guideline tailored to problem characteristics, hardware constraints, and task requirements; (2) a comprehensive survey of over 100 works in graph–LLM integration; and (3) distillation of five critical future research directions. This work establishes a theoretical foundation, technical methodology, and practical paradigm for the emerging intersection of graph learning and large language modeling.
To address high query latency, poor scalability, and weak multi-source coordination in interactive exploration of large-scale heterogeneous knowledge graphs, this paper proposes a unified federated query system supporting document, relational, and graph data models. Our approach introduces three key contributions: (1) a semi-join decomposition technique that effectively curbs exponential blowup of intermediate results in path queries; (2) the first integration of query plan folding, semantic caching, and adaptive query planning—collectively enhancing execution efficiency and resource utilization; and (3) a novel distributed join algorithm enabling cross-modal collaborative analysis. Experimental evaluation demonstrates near-linear scalability across three dimensions—data volume, concurrent user count, and number of compute nodes—while sustaining sub-100ms end-to-end query latency.
To address the challenge of processing multi-source, heterogeneous data in social networks, this paper proposes a unified batch-stream-graph analytics framework built upon the Hadoop-Spark ecosystem. The method systematically integrates Hive (for SQL-based batch processing), HBase (for low-latency key-value lookups), and GraphX (for scalable graph computation) under a single Spark execution layer. It supports three core analytical tasks: user influence assessment, high-frequency term statistics, and community relationship mining. Leveraging HDFS for distributed storage, YARN for resource orchestration, and multi-language APIs, the framework achieves loosely coupled integration of computation and storage. End-to-end experiments on real-world social datasets demonstrate that the hybrid architecture accelerates complex relational analysis by 1.8–3.2× compared to single-component baselines, significantly improving both processing efficiency and system flexibility.
This work systematically evaluates the capabilities and limitations of large language models (LLMs) in graph pattern understanding—particularly graph pattern mining—a task largely unexplored for LLMs. Method: We introduce the first dedicated benchmark comprising three task categories—terminology comprehension, topological description, and autonomous discovery—covering synthetic and real-world graph datasets, seven mainstream LLMs, and eleven subtasks. Our zero-shot and few-shot evaluation framework integrates structured graph representations with natural language prompts and adopts a multidimensional assessment paradigm that jointly measures descriptive understanding and generative capability. Contribution/Results: Experiments reveal, for the first time, that LLMs possess preliminary graph pattern understanding ability (with O1-mini achieving best overall performance), exhibiting reasoning pathways fundamentally distinct from traditional algorithms. Performance improves substantially when input formats align with pretraining knowledge. This work establishes foundational benchmarks, methodological frameworks, and empirical evidence for interdisciplinary research at the intersection of graph AI and foundation models.
Existing LLM-based graph analysis benchmarks rely on direct structural reasoning, limiting scalability to large graphs; in contrast, human experts routinely solve such tasks programmatically using graph libraries (e.g., NetworkX, PyTorch Geometric). Method: We propose ProGraph—the first programming-centric benchmark for graph analysis—comprising three expert-level task categories, multi-scale real-world graphs, and six mainstream graph libraries. We further introduce LLMS4Graph, a high-quality dataset featuring authoritative documentation and automatically generated code. To enhance API comprehension and code generation, we employ documentation-augmented retrieval and fine-tune open-source LLMs on this data. Contribution/Results: Our approach yields 11–32% absolute accuracy gains on ProGraph, with the best model achieving 36% accuracy. All components—including the ProGraph benchmark, LLMS4Graph dataset, and enhanced models—are publicly released to advance LLMs’ programmatic understanding of structured graph data.
This work addresses the need for unified and efficient knowledge provisioning in large language models by proposing a novel architecture that integrates relational and property graph data models. The approach leverages record addresses from log files as immutable reference values in place of traditional foreign keys, enabling efficient graph-style link traversal instead of costly join queries while natively supporting triple-based knowledge representation. The resulting unified knowledge service framework combines the structural rigor of relational models with the flexible associative capabilities of graph models, significantly enhancing knowledge retrieval efficiency and effectively supporting knowledge integration and invocation in generative AI systems.
This work addresses the high computational cost of existing pretraining data selection methods, which often rely on auxiliary models or labeled data. The authors propose WebGraphMix, a novel framework that leverages the host-level web graph topology derived from Common Crawl to guide data mixing without requiring additional training or annotations. By employing unsupervised centrality scores, WebGraphMix adaptively balances the proportion of core and peripheral documents, revealing their complementary learning value in the web graph—a signal orthogonal to existing content-quality metrics. Evaluated on the DataComp-LM benchmark, a simple 1:1 mixing strategy achieves an average score of 41.4% on 400M and 1B language models, outperforming uniform sampling (39.8%). Further gains are realized by combining this approach with content-quality scoring, yielding a performance of 43.8%.
This work addresses the challenges of high construction costs and difficulties in language grounding for natural language query interfaces over enterprise private knowledge graphs, particularly underperforming in short-query and schema-paraphrasing scenarios. The authors propose KG2Cypher, a novel data-centric Text-to-Cypher framework that automatically generates high-quality text–Cypher pairs from the knowledge graph itself. The approach integrates candidate-aware supervised fine-tuning (SFT) data generation, category-conditional schema prompting, entity retrieval, and efficient LoRA-based fine-tuning, complemented by an LLM-driven automatic evaluation pipeline and human validation. Evaluated on Korean enterprise use cases, the method achieves execution F1 scores of 0.950 and 0.920 for broadcast program and company queries, respectively, and attains 95.2% exact match accuracy, 99.9% execution rate, and 0.964 F1 on an 11-way classification task.
This work addresses data contamination caused by irreversible entity merging and ontology misclassification based on name fragments in knowledge graph construction. The authors propose a “review-before-linking” mechanism featuring an identity-ladder strategy—leveraging identifiers, names, and type scopes—to enable controlled deduplication, alongside anchor-evidence constraints that govern multi-class ontology label assignment. This approach corrects the evidential asymmetry arising when names are treated as instance labels rather than type assertions. Integrated into a system combining automated merging, evidence validation, and a human review queue, the method was evaluated on a knowledge graph comprising 537,157 entities and 2,198,567 relations. It reduced role assignment errors from 36 to zero, requiring only 775 manual decisions to resolve 48,403 merge proposals, thereby significantly mitigating risks of over-merging and misclassification.