From Enumeration to Covering: Near-Optimal Densest P-Partite Subgraph Search over Large Heterogeneous Information Networks

📅 2026-08-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the problem of dense subgraph discovery in heterogeneous information networks based on meta-paths. Given a meta-path $P$ of length $i$, the goal is to find a dense $P$-partite subgraph spanning the $i$ node-type layers defined by $P$ that maximizes the ratio between the number of meta-path instances and the geometric mean of the node counts across layers. To this end, the authors propose a coverage-based weight selection strategy that requires only a polylogarithmic number of representative weight sets, casting each fixed-weight subproblem as a weighted hypermodular dense subgraph optimization. They innovatively design an adaptive peeling algorithm that avoids explicit enumeration of meta-path instances and, for the first time, achieves a tunable $(1-\delta)/(1+\eta)$ approximation to near-optimal density when $i>2$. Experiments on five real-world networks demonstrate significant superiority over enumeration-based baselines, offering both substantial efficiency gains and verifiable near-optimality of solutions.
📝 Abstract
Given a heterogeneous information network (HIN) and a query meta-path P of length i, the densest P-partite subgraph problem finds the subgraph, spanning the i typed layers of P, that maximizes a parameter-free density: the number of meta-path instances over the geometric mean of the layer sizes. It has applications across bibliographic, e-commerce, and biomedical networks. The state-of-the-art approximation linearizes the geometric-mean objective by fixing per-layer weights, but solves one subproblem for every feasible weight set, of which there are $O((n/i)^i)$, and on each achieves only a $1/i$ approximation. We show that neither the exhaustive enumeration nor the loose guarantee is necessary. First, we replace enumeration by covering: polylogarithmically many representative weight sets, localized further by a data-dependent bound, cover all feasible ones while losing only a tunable factor $1+η$ in density. Second, we cast each fixed-weight subproblem as a weighted supermodular densest-subgraph instance and solve it near-optimally, lifting the overall guarantee to $(1-δ)/(1+η)$. To our knowledge, this is the first near-optimal density approximation beyond the bipartite ($i=2$) case, and it yields a PTAS for every fixed i. Algorithmically, our solver is an adaptive peeling scheme that never materializes the meta-path instances, whose number can exceed the graph size by orders of magnitude. An incumbent-driven reduction further discards representative weight sets before their subproblems are solved. Experiments on five real HINs show that our algorithms achieve substantial speedups over enumeration-based baselines and can further certify the near-optimality of the returned subgraph.
Problem

Research questions and friction points this paper is trying to address.

densest P-partite subgraph
heterogeneous information network
meta-path
density maximization
subgraph search
Innovation

Methods, ideas, or system contributions that make the work stand out.

heterogeneous information networks
densest subgraph
meta-path
supermodular optimization
approximation algorithm
🔎 Similar Papers
No similar papers found.