🤖 AI Summary
This work studies the learning-augmented densest subgraph problem: given a partial solution from a lightweight classifier—covering at least a $(1-varepsilon)$ fraction of the nodes in the optimal subgraph—we propose the first linear-time algorithm with theoretical guarantee of outputting a $(1-varepsilon)$-approximate solution. Our method abandons traditional LP or maximum-flow solvers, instead integrating greedy refinement with density-driven pruning to tightly couple prediction signals with combinatorial optimization. The key contribution is the first rigorous integration of supervised node classification with a minimal combinatorial algorithm, achieving both strong approximation guarantees and low computational overhead. The framework naturally extends to directed graphs and NP-hard variants—including constrained and weighted densest subgraph problems. On the Twitch Ego Nets dataset, our algorithm significantly outperforms Charikar’s algorithm and pure prediction baselines, demonstrating high accuracy, efficiency, and generalization across problem variants.
📝 Abstract
We study the densest subgraph problem and its variants through the lens of learning-augmented algorithms. For this problem, the greedy algorithm by Charikar (APPROX 2000) provides a linear-time $ 1/2 $-approximation, while computing the exact solution typically requires solving a linear program or performing maximum flow computations.We show that given a partial solution, i.e., one produced by a machine learning classifier that captures at least a $ (1 - epsilon) $-fraction of nodes in the optimal subgraph, it is possible to design an extremely simple linear-time algorithm that achieves a provable $ (1 - epsilon) $-approximation. Our approach also naturally extends to the directed densest subgraph problem and several NP-hard variants.An experiment on the Twitch Ego Nets dataset shows that our learning-augmented algorithm outperforms Charikar's greedy algorithm and a baseline that directly returns the predicted densest subgraph without additional algorithmic processing.