pmi co-occurrence graph

Designs, builds, and analyzes graph representations whose nodes are discrete items (e.g., words, entities, or events) and whose edges are weighted by pointwise mutual information (PMI) computed from co-occurrence counts in a corpus or dataset. Uses PMI-weighted co-occurrence graphs to quantify pairwise association strength, detect unexpected or deviant co-occurrence patterns relative to normative statistics, and evaluate local or global coherence or consistency in sequences or texts.

pmico-occurrencegraph

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.24
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Clustering coefficient reflecting pairwise relationships within hyperedges

Oct 31, 2024
RM
Rikuya Miyashita
🏛️ Tokyo Institute of Technology | Kyoto University

Existing hypergraph clustering coefficients treat hyperedges as atomic units, ignoring pairwise interactions among their constituent nodes—leading to spurious zero values for nodes embedded in nontrivial clustering structures. Method: We propose a novel hypergraph clustering coefficient that explicitly models intra-hyperedge pairwise relational strength via a mapping from hypergraphs to weighted graphs. Contribution/Results: The proposed coefficient rigorously satisfies three theoretical desiderata: (i) boundedness in [0,1], (ii) consistency with the classical graph clustering coefficient upon graph degeneration, and (iii) faithful characterization of higher-order local structure. Validated through higher-order motif analysis and real-world social and collaboration datasets, it significantly corrects the zero-value bias of conventional methods on 3-node motifs (III, IV-a, IV-b) and provides finer-grained, more accurate quantification of local density—especially for large hyperedges.

Current methods fail to capture meaningful clustering patterns in nodesExisting hypergraph clustering coefficients ignore pairwise relationships within hyperedgesLack of accurate local density measurement in complex group interactions

A Survey on Hypergraph Mining: Patterns, Tools, and Generators

Jan 16, 2024
GL
Geon Lee
🏛️ KAIST | Northeastern University

This work systematically investigates recurrent higher-order structural patterns across domains in hypergraphs and establishes an analytical framework capable of generating realistic synthetic hypergraphs. Method: We propose the first unified tripartite taxonomy for hypergraph mining—comprising pattern discovery, analytical tools, and generative models—integrating graph theory, random hypergraph models, statistical significance testing, and higher-order metrics (e.g., hypergraph transitivity). Our toolkit includes null models, substructure identification algorithms, and structural measures; we further design a feature-driven synthetic generator grounded in empirical hypergraph characteristics. Contribution/Results: We introduce the first multidimensional, fine-grained classification scheme and comprehensive research survey of hypergraph mining, explicitly identifying open challenges and interdisciplinary application pathways. This work lays a theoretical foundation and provides practical guidelines for higher-order network analysis, advancing both methodological rigor and real-world applicability in hypergraph science.

Develop tools and generators for hypergraph mining.Explore group interactions using hypergraph modeling.Identify recurring structural patterns in hypergraphs.

Existing graph mining methods primarily focus on topological subgraph discovery and lack a unified mechanism for jointly modeling syntactic and semantic aspects of association rules over attributed graphs. This paper proposes the MINE GRAPH RULE operator, the first to enable integrated syntactic–semantic expression of graph association rules in an attributed graph database—implemented as an extension to Neo4j. Syntactically, conditions are specified via Cypher-like queries; semantically, rule quality is evaluated using support and confidence metrics. The operator tightly couples graph structure with attribute semantics, leverages Neo4j’s native query optimization, and incorporates relational association rule pruning strategies to ensure efficiency and portability. Experiments demonstrate strong scalability across multidimensional parameters. An open-source plugin implementing the operator significantly enhances both the expressiveness and practical utility of graph association rule mining.

Defining MINE GRAPH RULE operator for property graph miningExpressing syntax and semantics for graph association rulesImplementing and evaluating operator on Neo4j with real-world data

LLM-Enhanced User-Item Interactions: Leveraging Edge Information for Optimized Recommendations

Feb 14, 2024
XW
Xinyuan Wang
🏛️ Arizona State University | Linkedin | HKUST (Guangzhou)

To address the limitations of large language models (LLMs) in modeling user–item edge relationships and integrating graph-structural information, this paper proposes the first graph-relational natural language prompting framework. It explicitly endows LLMs with edge-awareness and graph connectivity reasoning capabilities, enabling collaborative optimization between LLMs and graph neural networks (GNNs) for edge-level relational mining. The method synergistically integrates LLMs (e.g., LLaMA, ChatGLM) with GNNs (e.g., GCN, GAT), augmented by customized graph-relational prompt engineering and targeted fine-tuning strategies. Evaluated on real-world datasets including Amazon and Yelp, the approach achieves 3.2–5.8% AUC improvement, substantially enhancing recommendation accuracy—particularly for long-tail items—and interpretability. This work marks the first instance of semantic-topological joint modeling of LLMs and GNNs in recommender systems.

Bridging graph-based and LLM-based recommendation methodsEnhancing recommendation relevance via graph-aware attention mechanismsIntegrating graph edge information into LLM prompts

What Do LLMs Need to Understand Graphs: A Survey of Parametric Representation of Graphs

Oct 16, 2024
DF
Dongqi Fu
🏛️ University of Illinois Urbana-Champaign

Large language models (LLMs) struggle to effectively comprehend graph-structured data due to their inherent sequence-based architecture and lack of native graph-aware representations. Method: This paper introduces *graph laws*—statistically derived, topologically parameterized features that are interpretable as natural language descriptions—establishing a novel paradigm for representing graphs as LLM-compatible inputs. We systematically construct a multi-dimensional graph law framework spanning macro/micro scales, low/high orders, and static/dynamic properties, integrating graph-theoretic analysis, multi-scale observational modeling, and natural language alignment techniques, while establishing semantic mappings to downstream graph tasks and retrieval-augmented generation (RAG) scenarios. Results: Experiments demonstrate that graph laws substantially mitigate LLM hallucination, overcome context-length limitations, and enable end-to-end graph reasoning. The approach achieves strong generalization across diverse domains, including molecular design, recommender systems, and protein structure modeling.

LLMs require effective graph understanding for reasoning.Parametric graph representation aids LLMs in data input.Survey explores graph laws for LLM-compatible representations.

Latest Papers

What's happening recently
View more

This study addresses the challenge of applying conventional statistical methods to collections of heterogeneous networks that vary in size and type and lack node correspondence. To overcome this, the authors propose a functional Topological Data Analysis (funTDA) framework that uniquely integrates functional data analysis with persistent homology to extract topological features from networks. This approach enables standard statistical operations—including mean and variance estimation, principal component analysis, and hypothesis testing—despite the non-Euclidean nature of network structures, thereby establishing a unified inferential framework. Empirical evaluations demonstrate that funTDA effectively discriminates networks with distinct connectivity patterns and successfully uncovers significant topological differences in real-world applications, such as literary co-occurrence networks and influenza gene regulatory networks.

functional data analysisnetwork collectionsnetwork comparison

This study investigates systematic disparities in coverage between mainstream and marginal Indian news outlets during the 2020–21 and 2024 farmer protests, particularly regarding the underrepresentation of key actors such as farmer leaders. By constructing entity co-occurrence–based media networks and integrating GraphSAGE with complex network analysis, the research characterizes media behavior through three structural dimensions: centrality, community structure, and a novel metric proposed herein—link predictability. This new indicator enables scalable, unsupervised media analysis solely from relational structures, without reliance on textual labels. Findings reveal a consistent and significant underestimation of farmer leaders’ visibility across diverse media outlets, exposing deep-seated structural biases in protest coverage.

entity co-occurrencefarmers protestsmedia bias

This study addresses the lack of coherence in longitudinal analyses of scientific knowledge graphs, which often stems from inconsistent approaches to topic identification and cross-temporal linkage. To overcome this limitation, the authors propose a unified relational framework that integrates cross-sectional topic detection and longitudinal lineage reconstruction within a single weighted network structure. For the first time, topic evolution is modeled as structural reconfiguration rather than lexical continuity within a relational paradigm. By incorporating relational clustering, directional coverage, centrality-weighted measures of lineage strength, and document membership modeling, the method substantially enhances methodological consistency and interpretive robustness in longitudinal science mapping. This approach enables a more accurate and nuanced understanding of the dynamic mechanisms underlying scientific topic evolution.

longitudinal analysisrelational clusteringscience mapping

Existing graph analysis systems struggle to effectively integrate topological structure with node attributes, limiting the discovery of patterns driven by their interaction. This work proposes ZipLine, a novel system that, for the first time, unifies predicate logic to express topology, node attributes, and neighborhood relationships within a single formalism. ZipLine introduces an interaction-driven predicate learning algorithm that enables cross-space collaborative reasoning and iterative analysis. By integrating coordinated views, subgraph selection, and attribute brushing techniques, the system facilitates expressive and efficient exploration of complex patterns in multivariate graphs. Empirical evaluation across three real-world domains—energy infrastructure, cybersecurity, and drug discovery—demonstrates ZipLine’s effectiveness in significantly enhancing the expressiveness and discoverability of intricate graph patterns.

integrated analysismultivariate graphsnode attributes

Traditional hypergraph incidence matrices employ binary representations, which struggle to capture the complex higher-order relationships between nodes and hyperedges arising from their size disparities. This work proposes the first continuous node–hyperedge proximity matrix grounded in a resource allocation mechanism, overcoming the limitations of binary encoding and enabling a more refined modeling of higher-order interactions. The resulting proximity matrix can be seamlessly integrated into lightweight algorithms to effectively support downstream tasks such as link prediction, key node identification, and community detection. Extensive experiments on multiple real-world hypergraph datasets demonstrate that methods leveraging this matrix consistently outperform existing baselines across all three core tasks, confirming both its expressive power and practical utility.

higher-order interactionshypergraphincidence matrix

Hot Scholars

DZ

Doudou Zhou

National University of Singapore
High-dimensional StatisticsEHR Data AnalysisChange-point DetectionTransfer Learning
TC

Tianxi Cai

Harvard University
statisticsbiostatisticsmodelingprediction
ZG

Ziming Gan

PhD in statistics, University of Chicago
EHR datasingle cell
AW

Alan Wee-Chung Liew

Professor, School of ICT, Griffith University
Machine learningmedical imagingcomputer visionensemble learning