Score
Design, implement, and analyze algorithms and heuristics that generate samples of topological orderings of a directed acyclic graph (including methods that aim for uniform sampling over all valid topological orders). Use those samplers and their outputs to construct estimators or approximations of graph-level quantities, and evaluate their correctness, sampling bias, variance, and computational efficiency.
This work addresses the asymptotic enumeration and efficient uniform random sampling of ordered directed acyclic graphs (DAGs)—i.e., DAGs equipped with a total order on outgoing edges—under both edge-count-constrained and unconstrained settings. Methodologically, we establish the first exact asymptotic counting formula for this novel model, devise the first labelled DAG generator that *exactly* controls the number of edges in $O(n^2)$ time, and unify theoretical analysis with algorithm design via a synthesis of combinatorial analysis, exponential generating functions, the Boltzmann sampling framework, and recursive dynamic programming. Our contributions advance the combinatorial understanding of DAG structures and provide provably optimal generators for simulating diverse data structures—including priority queues, scheduling graphs, and dependency networks—where edge-ordering captures essential operational semantics.
This study addresses the problem of efficiently generating uniformly random directed acyclic graphs (DAGs) of a fixed size. The authors propose two novel algorithms, one of which constitutes the first asymptotically optimal exact-size sampler for DAGs, achieving an expected time complexity of $\frac{n^2}{2} + o(n^2)$. This method extends the Boltzmann sampling framework by integrating structural decompositions of DAGs with generating function techniques and an optimized strategy for random number usage. Compared to the current state-of-the-art, the proposed approach yields significant improvements both in theoretical time complexity and practical runtime performance, thereby enabling, for the first time, uniform and efficient sampling of DAGs at a prescribed scale.
Estimating the impact of graph sampling on shortest-path distance distributions is challenging without performing actual sampling or computing all-pairs shortest paths. Method: This paper proposes an analytical evaluation framework that avoids both full-graph shortest-path computation and empirical sampling. Its core innovation is the first closed-form estimation of the shortest-path distance distribution from sampled to unsampled nodes—derived solely from the node-degree distribution—and extended to community-structured graphs via random-graph modeling and community-aware approximations. Contribution/Results: Compared to simulation-based empirical methods, our approach achieves over 10× speedup while maintaining an average error below 8% across diverse real-world and synthetic graphs. It also demonstrates high consistency in downstream bias-comparison tasks. The implementation is publicly available.
Designing optimal sampling schemes for finite populations is challenging due to complex mathematical constraints and intractable optimization. Method: This paper proposes a novel geometric sampling paradigm based on two-dimensional representation: first-order inclusion probabilities are modeled as adjustable rectangular bars, enabling intuitive parameterization and diverse design generation. We introduce the first geometric visualization framework for sampling design—bypassing traditional reliance on intricate analytical derivations—and incorporate greedy best-first search to jointly optimize entropy maximization and design optimality, without prescribing algorithmic structure. Contribution/Results: The approach significantly enhances design flexibility and computational efficiency. Experiments demonstrate superior performance over classical designs across key metrics—including entropy, balance, and variance control—establishing it as an interpretable, user-friendly, and efficient tool for finite-population sampling.
Scalability bottlenecks hinder community detection and core-periphery (CP) structure inference in large-scale networks. Method: We systematically evaluate seven graph subsampling strategies within a divide-and-conquer framework and, for the first time, derive statistically grounded upper bounds on estimation error for CP structure identification under subsampling. Contribution/Results: Theory and experiments reveal task-dependent optimality: uniform node sampling achieves the best community detection performance, whereas core-biased sampling significantly improves CP identification accuracy (up to +37%) and computational efficiency. Compared to full-network algorithms, core-biased subsampling delivers both high efficiency and robustness on real-world and synthetic networks. This work bridges a critical gap in statistical theory by establishing the first formal analysis of task-specific subsampling adaptivity for network structural inference, yielding interpretable, principled guidelines for selecting optimal sampling strategies in large-scale network analysis.
This study addresses a critical limitation in the evaluation of causal discovery algorithms, which often rely on randomly generated directed acyclic graphs (DAGs) whose implicit topological properties may distort performance assessments. The authors observe that in common random DAG models—such as Erdős–Rényi and scale-free graphs—the number of “relatives” (nodes reachable via open paths) for each node strictly increases along the true causal order. They prove that this monotonicity property renders the Markov equivalence class degenerate, collapsing it to a single unique graph. Leveraging this insight, they propose a causal ordering recovery criterion based on estimating relative counts and introduce a corresponding time-ordered DAG sampling scheme. Experiments demonstrate that the method efficiently approximates the true causal order across diverse simulation settings, while also exposing inherent limitations in current synthetic data evaluation paradigms.
This work addresses the problem of efficiently sampling Eulerian circuits approximately uniformly from directed Eulerian multigraphs. For sparse graphs with $m$ arcs, the authors propose a randomized algorithm based on a novel local Markov chain called the “flip–repair walk,” which integrates a dynamic chord data structure with a degree-reduction framework. This approach overcomes the $O(mn)$ time bottleneck inherent in classical arborescence-based sampling methods. The algorithm achieves a worst-case time complexity of $\widetilde{O}(m^{3/2})$, significantly improving upon existing techniques. Approximate uniformity of the sampling distribution is rigorously guaranteed through a hybrid analysis combining switch-network reductions and linear-algebraic arguments.
This work addresses the lack of systematic approaches for quantifying and exploring local discrepancies between sampled graphs and their original counterparts at the node, edge, and structural levels—a gap that hinders effective evaluation and selection of graph sampling strategies. To bridge this gap, the authors propose three general quantitative metrics—neighborhood-, path-, and structure-based—to measure local fidelity. They further introduce DiffLens, an interactive visualization system that, for the first time, incorporates lens-based views tailored to these three types of differences, enabling users to focus on regions of interest. Case studies on real-world network datasets and user experiments demonstrate the framework’s effectiveness and practicality in supporting intuitive comparison of sampling outcomes and enhancing understanding of localized structural variations.
This work addresses the computational challenge of uniformly sampling proper $k$-colorings of graphs when the maximum degree $\Delta$ is large. The authors propose a novel approach based on partial rejection sampling (PRS), which introduces tunable soft coloring constraints that are progressively tightened to achieve exact uniform sampling. By integrating a recursive divide-and-conquer strategy, the original problem is decomposed into $O(\log n)$ independent subproblems of reduced size, each solvable in parallel by any exact sampler. This method is the first to combine soft coloring with PRS, enabling parallelization and achieving a runtime of $O(L^{\log^* n} \cdot n\Delta)$ when the number of relaxation levels $L$ is independent of $n$, improving upon the best-known algorithms that require $k > 3\Delta$. Empirical evidence suggests $L$ is likely constant, indicating potential for linear-time performance.
研究二值矩阵固定行和列和的均匀采样问题,提出无拒绝蛇算法,通过生长交替路径并翻转环路,实现高效采样。