spanning forest sampling

Design, implement, and analyze algorithms that generate random spanning forests of a graph — including converging/arborescent forests in directed graphs — and manage the sampling process and probabilities. Use those sampled forests to compute or estimate node-level statistics and related matrix quantities (for example per-node sampling probabilities or diagonal entries of graph-related matrices) and to validate sampling bias and correctness.

spanningforestsampling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.17
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Sampling tree-weighted partitions without sampling trees

Aug 14, 2025
SC
Sarah Cannon
🏛️ Claremont McKenna College | Carnegie Mellon University | Yale University

Efficient exact sampling of balanced tree-weighted 2-partitions on planar graphs remains computationally challenging; existing methods rely on spanning-tree sampling followed by rejection steps, yielding near-optimal time complexity of approximately $O(n log^2 n)$ (approximate) and $O(n log n)$ (exact). Method: We propose the first direct sampling algorithm that bypasses precomputing spanning trees. Leveraging combinatorial probability analysis and structural properties of planar graphs, we construct a carefully designed Markov chain to sample from the target conditional distribution. Contribution/Results: We prove that our algorithm achieves $O(n)$ expected running time on a broad class of planar graphs—marking the first linear expected-time exact sampler for this problem. This improves upon the prior $O(n log n)$ barrier and provides a scalable theoretical tool for applications such as political districting and graph partitioning.

Avoiding computational bottleneck in spanning tree samplingImproving speed for exact and approximate tree samplingSampling balanced tree-weighted partitions efficiently

Splittable Spanning Trees and Balanced Forests in Dense Random Graphs

Jul 16, 2025
DG
David Gillman
🏛️ New College of Florida | Georgia Institute of Technology

This paper studies the weighted fair partition problem on dense random graphs. Addressing fundamental limitations of the existing ReCom algorithm—namely, lack of irreducibility and exponentially slow rejection sampling—we propose a novel spanning-tree-based sampling framework. First, we prove that in dense random graphs, a uniformly random spanning tree can be split into *k* balanced subtrees (i.e., a balanced forest) with inverse-polynomial probability. Building on this, we design an efficient algorithm that combines uniform spanning tree sampling, removal of *k−1* edges, acceptance-probability tuning, and upper-lower walk-based approximate uniform sampling over forests. Our method yields the first provably fast, approximately uniform sampler for fair *k*-partitions on dense random graphs. It advances the theoretical understanding of sampling complexity for graph-balanced partitions and introduces a new toolset for algorithmic fairness research.

Analyze equitable partitioning via spanning-tree metricsIdentify limitations in ReCom algorithm for partitioningStudy splittable spanning trees in dense random graphs

Spanning tree methods for sampling graph partitions

Oct 04, 2022
SC
Sarah Cannon
🏛️ Claremont McKenna College | University of Chicago | Georgia Institute of Technology | Tufts University

This work addresses the lack of neutral, interpretable benchmarks for detecting partisan gerrymandering in political redistricting. We propose RevReCom—the first reversible Markov chain sampling method for districting with a closed-form stationary distribution. Its stationary probability is proportional to the product of spanning tree counts across districts, inherently encoding contiguity, population balance, and community coherence. Unlike standard ReCom, RevReCom yields a theoretically tractable target distribution, guarantees strict convergence verification, and enables efficient diagnostics (e.g., effective sample size, trace plots). On real-world redistricting instances, it generates high-quality ensembles within one hour. This establishes the first mathematically rigorous yet computationally feasible benchmark framework for statistical fairness testing in redistricting. Moreover, RevReCom serves as a verifiable ground-truth reference for evaluating other graph partitioning samplers.

Develops new sampling methods for political districting validationEstablishes spanning tree distribution as principled comparison baselineProvides powerful null model to detect gerrymandering scientifically

On the Number of Steps of CyclePopping in Weakly Inconsistent U(1)-Connection Graphs

Apr 23, 2024
MF
M. Fanuel
🏛️ Université de Lille | CNRS | Centrale Lille | UMR 9189 - CRIStAL

This work addresses the probabilistic sampling of cycle-rooted spanning forests (CRSFs) on U(1)-connection graphs, focusing on the termination-time complexity of the CyclePopping algorithm under weak inconsistency. Employing loop measure theory, determinantal point processes, and the Viennot-type loop pyramid structure, we provide the first elementary and complete correctness proof of the algorithm. We derive an explicit distribution for its termination time—characterized as a Poisson point process on the graph’s directed loops. Furthermore, we establish a unified framework applicable to both deterministic and random CRSF distributions, explicitly identifying a sufficient condition for efficient convergence: low loop weights. This yields foundational theoretical guarantees and precise complexity bounds for sampling random combinatorial structures on connection graphs.

Analyzing time complexity of the sampling algorithmGeneralizing Wilson's algorithm for cycle-rooted spanning forestsProviding detailed proof of correctness for sampling CRSFs

Algorithms for spanning trees of unweighted networks

May 13, 2022
LŠ
Lovro Šubelj
🏛️ University of Ljubljana

This work addresses structural distortion in spanning tree generation for unweighted networks. Traditional Prim’s and Kruskal’s algorithms lack theoretical justification on unweighted graphs, while DFS often yields highly imbalanced, deep trees. We systematically evaluate the applicability of classical algorithms and propose a novel BFS-based spanning tree construction framework that inherently preserves shortest-path distances between nodes and network diameter, while yielding an approximately power-law degree distribution—balancing topological fidelity and tree balance. Extensive experiments across over one thousand real-world and synthetic networks demonstrate that BFS-generated spanning trees significantly outperform baseline methods in both distance preservation and compactness. Our approach provides a theoretically grounded, computationally efficient tool for network backbone extraction, simplified sampling, and structural analysis.

Comparing spanning tree algorithms for unweighted networksEvaluating network structure preservation in spanning treesRecommending BFS for compact, balanced unweighted network trees

Latest Papers

What's happening recently
View more

This work addresses the challenge that existing fast Laplacian solvers struggle to efficiently compute the diagonal of the convergent forest matrix for directed graphs. We propose three novel sampling-based algorithms—SCF, SCFV, and SCFV+—which extend Wilson’s algorithm for forest sampling and incorporate two variance reduction techniques to enhance estimation accuracy and efficiency. The key innovation lies in the first integration of opinion dynamics–inspired matrix-vector iterations and a new iterative formulation into forest sampling. Notably, SCFV+ achieves linear time complexity and reduced variance by eliminating cross terms. Experimental results demonstrate that our approach delivers high accuracy and scalability across diverse real-world networks, supporting graphs with over twenty million nodes—both directed and undirected—and significantly outperforms state-of-the-art methods.

diagonal computationdigraphsforest matrix

This work addresses the challenge of efficiently generating representative ensembles of districting plans by enabling independent sampling from the space of graph partitions. The authors propose a novel method that, for the first time, explicitly constructs a probability distribution over graph partitions under exact population balance constraints, thereby achieving truly independent samples and circumventing the mixing difficulties inherent in traditional Markov chain approaches. By integrating probabilistic modeling with an efficient sampling algorithm, the method demonstrates substantially improved sampling efficiency and diversity compared to existing Markov chain baselines, as validated on both grid graphs and real-world congressional and state legislative districting maps across U.S. states. This advance breaks away from the conventional paradigm reliant on sequential chain-based sampling.

district plansensemble generationgraph partitions

This work addresses the lack of systematic approaches for quantifying and exploring local discrepancies between sampled graphs and their original counterparts at the node, edge, and structural levels—a gap that hinders effective evaluation and selection of graph sampling strategies. To bridge this gap, the authors propose three general quantitative metrics—neighborhood-, path-, and structure-based—to measure local fidelity. They further introduce DiffLens, an interactive visualization system that, for the first time, incorporates lens-based views tailored to these three types of differences, enabling users to focus on regions of interest. Case studies on real-world network datasets and user experiments demonstrate the framework’s effectiveness and practicality in supporting intuitive comparison of sampling outcomes and enhancing understanding of localized structural variations.

graph comparisongraph samplinglocal differences

This study addresses the error structure and inter-tree interaction mechanisms of CART-based random forests under feature subsampling. To this end, it introduces the CART-ROSA framework, which, for the first time, models feature subsampling as a stochastic set of admissible actions and formalizes the entire process as a sequential resource allocation problem. By integrating stochastic control theory, CART splitting rules, and an information-split counting process, the framework disentangles two key design dimensions: the "information opportunity rate" and the "splitting policy contraction strength." Theoretical analysis reveals that while the CART policy is locally stable, it may be globally suboptimal. Under linear models, the work further derives an explicit mean squared error risk expansion, effectively bridging the gap between algorithmic description and theoretical analysis.

black-box interpretabilityCART random forestsfeature subsampling

This work addresses the challenge of maintaining high prediction accuracy in random forest inference under time interruptions on resource-constrained systems, where execution may be halted before all trees complete. To enable effective anytime behavior, we reformulate random forests as fine-grained anytime algorithms by treating individual internal nodes of decision trees as scheduling units and optimizing their execution order to maximize average accuracy under interruption. We introduce an exponential-time algorithm to find the optimal node ordering and propose two efficient polynomial-time heuristics—Forward and Backward Squirrel Order. Experimental results demonstrate that Backward Squirrel Order achieves 94% of the optimal performance and outperforms existing strategies by approximately 99%, significantly enhancing the robustness and accuracy of anytime random forest inference.

anytime algorithminferenceprediction confidence

Hot Scholars

MS

Muhammad Shafique

Professor, ECE, New York University (AD-UAE, Tandon-USA), Director eBRAIN Lab
Embedded Machine LearningBrain-Inspired ComputingRobust & Energy-Efficient System DesignSmart
PF

Phablo F. S. Moura

KU Leuven
AlgorithmsCombinatorial OptimizationGraph TheoryOperations Research
AB

Abdul Basit

Research Engineer, New York University Abu Dhabi
Artificial IntelligenceDeep LearningRobotics