topological ordering sampling

Design, implement, and analyze algorithms and heuristics that generate samples of topological orderings of a directed acyclic graph (including methods that aim for uniform sampling over all valid topological orders). Use those samplers and their outputs to construct estimators or approximations of graph-level quantities, and evaluate their correctness, sampling bias, variance, and computational efficiency.

topologicalorderingsampling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.41
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the asymptotic enumeration and efficient uniform random sampling of ordered directed acyclic graphs (DAGs)—i.e., DAGs equipped with a total order on outgoing edges—under both edge-count-constrained and unconstrained settings. Methodologically, we establish the first exact asymptotic counting formula for this novel model, devise the first labelled DAG generator that *exactly* controls the number of edges in $O(n^2)$ time, and unify theoretical analysis with algorithm design via a synthesis of combinatorial analysis, exponential generating functions, the Boltzmann sampling framework, and recursive dynamic programming. Our contributions advance the combinatorial understanding of DAG structures and provide provably optimal generators for simulating diverse data structures—including priority queues, scheduling graphs, and dependency networks—where edge-ordering captures essential operational semantics.

Develop efficient random sampling for ordered DAGsEnable labeled DAG generation with edge controlProvide asymptotic analysis of ordered DAG counts

This study addresses the problem of efficiently generating uniformly random directed acyclic graphs (DAGs) of a fixed size. The authors propose two novel algorithms, one of which constitutes the first asymptotically optimal exact-size sampler for DAGs, achieving an expected time complexity of $\frac{n^2}{2} + o(n^2)$. This method extends the Boltzmann sampling framework by integrating structural decompositions of DAGs with generating function techniques and an optimized strategy for random number usage. Compared to the current state-of-the-art, the proposed approach yields significant improvements both in theoretical time complexity and practical runtime performance, thereby enabling, for the first time, uniform and efficient sampling of DAGs at a prescribed scale.

Boltzmann samplingdirected acyclic graphsexact-size sampling

Efficient Estimation of Shortest-Path Distance Distributions to Samples in Graphs

Feb 21, 2025
AZ
Alan Zhu
🏛️ University of California, Berkeley | University of Illinois | University of Michigan

Estimating the impact of graph sampling on shortest-path distance distributions is challenging without performing actual sampling or computing all-pairs shortest paths. Method: This paper proposes an analytical evaluation framework that avoids both full-graph shortest-path computation and empirical sampling. Its core innovation is the first closed-form estimation of the shortest-path distance distribution from sampled to unsampled nodes—derived solely from the node-degree distribution—and extended to community-structured graphs via random-graph modeling and community-aware approximations. Contribution/Results: Compared to simulation-based empirical methods, our approach achieves over 10× speedup while maintaining an average error below 8% across diverse real-world and synthetic graphs. It also demonstrates high consistency in downstream bias-comparison tasks. The implementation is publicly available.

Estimates shortest-path distance distributions in graphs.Evaluates sampling methods impact on graph representativeness.Handles graphs with community structures efficiently.

Geometric Sampling

Aug 15, 2023
BP
Bardia Panahbehagh
🏛️ Kharazmi University

Designing optimal sampling schemes for finite populations is challenging due to complex mathematical constraints and intractable optimization. Method: This paper proposes a novel geometric sampling paradigm based on two-dimensional representation: first-order inclusion probabilities are modeled as adjustable rectangular bars, enabling intuitive parameterization and diverse design generation. We introduce the first geometric visualization framework for sampling design—bypassing traditional reliance on intricate analytical derivations—and incorporate greedy best-first search to jointly optimize entropy maximization and design optimality, without prescribing algorithmic structure. Contribution/Results: The approach significantly enhances design flexibility and computational efficiency. Experiments demonstrate superior performance over classical designs across key metrics—including entropy, balance, and variance control—establishing it as an interpretable, user-friendly, and efficient tool for finite-population sampling.

Developing a graphical framework for finite population sampling designsIntegrating intelligent algorithms to optimize complex sampling challengesRepresenting inclusion probabilities as manipulable bars on graphs

Graph sub-sampling for divide-and-conquer algorithms in large networks

Sep 11, 2024
EY
Eric Yanchenko
🏛️ Akita International University

Scalability bottlenecks hinder community detection and core-periphery (CP) structure inference in large-scale networks. Method: We systematically evaluate seven graph subsampling strategies within a divide-and-conquer framework and, for the first time, derive statistically grounded upper bounds on estimation error for CP structure identification under subsampling. Contribution/Results: Theory and experiments reveal task-dependent optimality: uniform node sampling achieves the best community detection performance, whereas core-biased sampling significantly improves CP identification accuracy (up to +37%) and computational efficiency. Compared to full-network algorithms, core-biased subsampling delivers both high efficiency and robustness on real-world and synthetic networks. This work bridges a critical gap in statistical theory by establishing the first formal analysis of task-specific subsampling adaptivity for network structural inference, yielding interpretable, principled guidelines for selecting optimal sampling strategies in large-scale network analysis.

Compares graph sub-sampling algorithms for large networks.Evaluates divide-and-conquer methods for community and CP structures.Identifies optimal sub-sampling strategies for specific network tasks.

Latest Papers

What's happening recently
View more

This study addresses a critical limitation in the evaluation of causal discovery algorithms, which often rely on randomly generated directed acyclic graphs (DAGs) whose implicit topological properties may distort performance assessments. The authors observe that in common random DAG models—such as Erdős–Rényi and scale-free graphs—the number of “relatives” (nodes reachable via open paths) for each node strictly increases along the true causal order. They prove that this monotonicity property renders the Markov equivalence class degenerate, collapsing it to a single unique graph. Leveraging this insight, they propose a causal ordering recovery criterion based on estimating relative counts and introduce a corresponding time-ordered DAG sampling scheme. Experiments demonstrate that the method efficiently approximates the true causal order across diverse simulation settings, while also exposing inherent limitations in current synthetic data evaluation paradigms.

causal discoverydirected acyclic graphsMarkov equivalence class

This work addresses the problem of efficiently sampling Eulerian circuits approximately uniformly from directed Eulerian multigraphs. For sparse graphs with $m$ arcs, the authors propose a randomized algorithm based on a novel local Markov chain called the “flip–repair walk,” which integrates a dynamic chord data structure with a degree-reduction framework. This approach overcomes the $O(mn)$ time bottleneck inherent in classical arborescence-based sampling methods. The algorithm achieves a worst-case time complexity of $\widetilde{O}(m^{3/2})$, significantly improving upon existing techniques. Approximate uniformity of the sampling distribution is rigorously guaranteed through a hybrid analysis combining switch-network reductions and linear-algebraic arguments.

directed graphsEulerian toursgraph algorithms

This work addresses the lack of systematic approaches for quantifying and exploring local discrepancies between sampled graphs and their original counterparts at the node, edge, and structural levels—a gap that hinders effective evaluation and selection of graph sampling strategies. To bridge this gap, the authors propose three general quantitative metrics—neighborhood-, path-, and structure-based—to measure local fidelity. They further introduce DiffLens, an interactive visualization system that, for the first time, incorporates lens-based views tailored to these three types of differences, enabling users to focus on regions of interest. Case studies on real-world network datasets and user experiments demonstrate the framework’s effectiveness and practicality in supporting intuitive comparison of sampling outcomes and enhancing understanding of localized structural variations.

graph comparisongraph samplinglocal differences

This work addresses the computational challenge of uniformly sampling proper $k$-colorings of graphs when the maximum degree $\Delta$ is large. The authors propose a novel approach based on partial rejection sampling (PRS), which introduces tunable soft coloring constraints that are progressively tightened to achieve exact uniform sampling. By integrating a recursive divide-and-conquer strategy, the original problem is decomposed into $O(\log n)$ independent subproblems of reduced size, each solvable in parallel by any exact sampler. This method is the first to combine soft coloring with PRS, enabling parallelization and achieving a runtime of $O(L^{\log^* n} \cdot n\Delta)$ when the number of relaxation levels $L$ is independent of $n$, improving upon the best-known algorithms that require $k > 3\Delta$. Empirical evidence suggests $L$ is likely constant, indicating potential for linear-time performance.

combinatorial samplingexact samplinggraph coloring

Hot Scholars

EJ

Eui-Jin Kim

Ajou University
Artificial IntelligenceTravel BehaviorSmart Mobility
NL

Nati Linial

Professor of Computer Science, The Hebrew University of Jerusalem
CombinatoricsTheoretical Computer ScienceBioinformatics
HX

Hui Xiong

Senior Scientist, Candela Corporation
Ultrafast dynamicsatomic molecular physicsfree electron laser