co-occurrence network analysis

Designs and implements methods and pipelines to construct token or entity co-occurrence graphs/networks from corpora, including scalable algorithms for fast graph construction, sliding-window or temporal aggregation, and efficient indexing/lookup. Computes and interprets centrality measures and other topological metrics to rank nodes, identify hubs, brokers, and communities, and to characterize temporal and structural properties of the resulting network.

co-occurrencenetworkanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Graph2text or Graph2token: A Perspective of Large Language Models for Graph Learning

Jan 02, 2025
SY
Shuo Yu
🏛️ Dalian University of Technology | Technical University of Darmstadt | RMIT University

This paper addresses the fundamental challenge that large language models (LLMs) cannot natively process graph-structured data. To bridge this gap, we propose LLM4graph—a unified framework introducing the first systematic taxonomy of Graph2Text and Graph2Token paradigms, and identifying four core challenges in graph-to-text/token conversion. Methodologically, LLM4graph integrates graph encoding, structured serialization, prompt engineering, and LLM adaptation techniques to achieve multi-granular, semantics-preserving graph representation transformation. Key contributions include: (1) a principled LLM selection guideline tailored to problem characteristics, hardware constraints, and task requirements; (2) a comprehensive survey of over 100 works in graph–LLM integration; and (3) distillation of five critical future research directions. This work establishes a theoretical foundation, technical methodology, and practical paradigm for the emerging intersection of graph learning and large language modeling.

Complex Graph DataGraph LearningLarge Language Models

Siren Federate: Bridging document, relational, and graph models for exploratory graph analysis

Apr 10, 2025
GB
Georgeta Bordea
🏛️ La Rochelle University | Siren

To address high query latency, poor scalability, and weak multi-source coordination in interactive exploration of large-scale heterogeneous knowledge graphs, this paper proposes a unified federated query system supporting document, relational, and graph data models. Our approach introduces three key contributions: (1) a semi-join decomposition technique that effectively curbs exponential blowup of intermediate results in path queries; (2) the first integration of query plan folding, semantic caching, and adaptive query planning—collectively enhancing execution efficiency and resource utilization; and (3) a novel distributed join algorithm enabling cross-modal collaborative analysis. Experimental evaluation demonstrates near-linear scalability across three dimensions—data volume, concurrent user count, and number of compute nodes—while sustaining sub-100ms end-to-end query latency.

Addressing exponential intermediate results in path-based queriesBridging document, relational, and graph models for explorationEnabling interactive analysis on large heterogeneous knowledge graphs

To address the challenge of processing multi-source, heterogeneous data in social networks, this paper proposes a unified batch-stream-graph analytics framework built upon the Hadoop-Spark ecosystem. The method systematically integrates Hive (for SQL-based batch processing), HBase (for low-latency key-value lookups), and GraphX (for scalable graph computation) under a single Spark execution layer. It supports three core analytical tasks: user influence assessment, high-frequency term statistics, and community relationship mining. Leveraging HDFS for distributed storage, YARN for resource orchestration, and multi-language APIs, the framework achieves loosely coupled integration of computation and storage. End-to-end experiments on real-world social datasets demonstrate that the hybrid architecture accelerates complex relational analysis by 1.8–3.2× compared to single-component baselines, significantly improving both processing efficiency and system flexibility.

Analyzing user influence and relationships in networksMeasuring performance of polyglot data tasks in clustersProcessing social media data with Hadoop-Spark ecosystem

How Do Large Language Models Understand Graph Patterns? A Benchmark for Graph Pattern Comprehension

Oct 04, 2024
XD
Xinnan Dai
🏛️ Michigan State University | The Hong Kong Polytechnic University | Microsoft Research | Independent Researcher

This work systematically evaluates the capabilities and limitations of large language models (LLMs) in graph pattern understanding—particularly graph pattern mining—a task largely unexplored for LLMs. Method: We introduce the first dedicated benchmark comprising three task categories—terminology comprehension, topological description, and autonomous discovery—covering synthetic and real-world graph datasets, seven mainstream LLMs, and eleven subtasks. Our zero-shot and few-shot evaluation framework integrates structured graph representations with natural language prompts and adopts a multidimensional assessment paradigm that jointly measures descriptive understanding and generative capability. Contribution/Results: Experiments reveal, for the first time, that LLMs possess preliminary graph pattern understanding ability (with O1-mini achieving best overall performance), exhibiting reasoning pathways fundamentally distinct from traditional algorithms. Performance improves substantially when input formats align with pretraining knowledge. This work establishes foundational benchmarks, methodological frameworks, and empirical evidence for interdisciplinary research at the intersection of graph AI and foundation models.

Assessing LLMs' ability to understand graph patterns via benchmarksEvaluating LLMs' performance on terminological and topological graph descriptionsExploring LLMs' potential in graph pattern mining tasks

Can Large Language Models Analyze Graphs like Professionals? A Benchmark, Datasets and Models

Sep 29, 2024
XL
Xin Li
🏛️ Beijing University of Posts and Telecommunications | Tsinghua University | China University of Petroleum

Existing LLM-based graph analysis benchmarks rely on direct structural reasoning, limiting scalability to large graphs; in contrast, human experts routinely solve such tasks programmatically using graph libraries (e.g., NetworkX, PyTorch Geometric). Method: We propose ProGraph—the first programming-centric benchmark for graph analysis—comprising three expert-level task categories, multi-scale real-world graphs, and six mainstream graph libraries. We further introduce LLMS4Graph, a high-quality dataset featuring authoritative documentation and automatically generated code. To enhance API comprehension and code generation, we employ documentation-augmented retrieval and fine-tune open-source LLMs on this data. Contribution/Results: Our approach yields 11–32% absolute accuracy gains on ProGraph, with the best model achieving 36% accuracy. All components—including the ProGraph benchmark, LLMS4Graph dataset, and enhanced models—are publicly released to advance LLMs’ programmatic understanding of structured graph data.

Enhance LLMs with programming-based solutionsEvaluate LLMs' graph analysis capabilitiesPropose datasets for improved graph task accuracy

Latest Papers

What's happening recently
View more

This work addresses the need for unified and efficient knowledge provisioning in large language models by proposing a novel architecture that integrates relational and property graph data models. The approach leverages record addresses from log files as immutable reference values in place of traditional foreign keys, enabling efficient graph-style link traversal instead of costly join queries while natively supporting triple-based knowledge representation. The resulting unified knowledge service framework combines the structural rigor of relational models with the flexible associative capabilities of graph models, significantly enhancing knowledge retrieval efficiency and effectively supporting knowledge integration and invocation in generative AI systems.

data integrationknowledge representationlarge language models

This work addresses the high computational cost of existing pretraining data selection methods, which often rely on auxiliary models or labeled data. The authors propose WebGraphMix, a novel framework that leverages the host-level web graph topology derived from Common Crawl to guide data mixing without requiring additional training or annotations. By employing unsupervised centrality scores, WebGraphMix adaptively balances the proportion of core and peripheral documents, revealing their complementary learning value in the web graph—a signal orthogonal to existing content-quality metrics. Evaluated on the DataComp-LM benchmark, a simple 1:1 mixing strategy achieves an average score of 41.4% on 400M and 1B language models, outperforming uniform sampling (39.8%). Further gains are realized by combining this approach with content-quality scoring, yielding a performance of 43.8%.

Common Crawldata curationlanguage models

This work addresses the challenges of high construction costs and difficulties in language grounding for natural language query interfaces over enterprise private knowledge graphs, particularly underperforming in short-query and schema-paraphrasing scenarios. The authors propose KG2Cypher, a novel data-centric Text-to-Cypher framework that automatically generates high-quality text–Cypher pairs from the knowledge graph itself. The approach integrates candidate-aware supervised fine-tuning (SFT) data generation, category-conditional schema prompting, entity retrieval, and efficient LoRA-based fine-tuning, complemented by an LLM-driven automatic evaluation pipeline and human validation. Evaluated on Korean enterprise use cases, the method achieves execution F1 scores of 0.950 and 0.920 for broadcast program and company queries, respectively, and attains 95.2% exact match accuracy, 99.9% execution rate, and 0.964 F1 on an 11-way classification task.

enterprise knowledge graphslanguage groundingnatural language interfaces

This work addresses data contamination caused by irreversible entity merging and ontology misclassification based on name fragments in knowledge graph construction. The authors propose a “review-before-linking” mechanism featuring an identity-ladder strategy—leveraging identifiers, names, and type scopes—to enable controlled deduplication, alongside anchor-evidence constraints that govern multi-class ontology label assignment. This approach corrects the evidential asymmetry arising when names are treated as instance labels rather than type assertions. Integrated into a system combining automated merging, evidence validation, and a human review queue, the method was evaluated on a knowledge graph comprising 537,157 entities and 2,198,567 relations. It reduced role assignment errors from 36 to zero, requiring only 775 manual decisions to resolve 48,403 merge proposals, thereby significantly mitigating risks of over-merging and misclassification.

data curationentity identityentity resolution

Hot Scholars

CZ

Chengzhi Zhang

Nanjing University of Science and Technology
Text MiningNatural Language ProcessingScience of Science
KS

Kijung Shin

Associate Professor, KAIST
Data MiningGraph MiningNetwork Science
FK

Foutse Khomh

NSERC Arthur B. McDonald Fellow, CRC Tier 1, Canada CIFAR AI Chair, FRQ-IVADO Chair, Full Professor
Software engineeringMachine learning systems engineeringMining software repositoriesReverse
AS

Akrati Saxena

LIACS, Leiden University, Netherlands
Social Network AnalysisComplex NetworksMachine LearningSocial Computing