Score
Designs and implements sampling methods that generate multi-hop paths through structured graphs or relational structures, using topology and schema information to produce semantically coherent and structure-consistent path instances. These methods control path difficulty and granularity, support cross-structure multi-hop queries, and select or generate paths that reduce lexical overlap with associated evidence.
Estimating the impact of graph sampling on shortest-path distance distributions is challenging without performing actual sampling or computing all-pairs shortest paths. Method: This paper proposes an analytical evaluation framework that avoids both full-graph shortest-path computation and empirical sampling. Its core innovation is the first closed-form estimation of the shortest-path distance distribution from sampled to unsampled nodes—derived solely from the node-degree distribution—and extended to community-structured graphs via random-graph modeling and community-aware approximations. Contribution/Results: Compared to simulation-based empirical methods, our approach achieves over 10× speedup while maintaining an average error below 8% across diverse real-world and synthetic graphs. It also demonstrates high consistency in downstream bias-comparison tasks. The implementation is publicly available.
This paper addresses hypothesis testing at the node, edge, and path levels in large-scale attributed graphs, proposing the first formal multi-granularity graph hypothesis testing framework. To overcome the limitations of conventional tabular methods on graph-structured data, we design PHASE—a path-hypothesis-aware random walk sampler—and its optimized variant PHASE-opt, jointly optimizing hypothesis-driven sampling and computational efficiency. Our approach integrates graph sampling theory, *m*-dimensional random walk modeling, and rigorous time-complexity analysis to ensure statistical validity while enhancing scalability. Experiments on three real-world attributed graph datasets demonstrate that, compared to generic sampling baselines, our framework improves testing accuracy by 12.7% and accelerates runtime by 3.8×, significantly strengthening statistical inference capabilities for large-scale attributed graphs.
This work addresses the lack of systematic approaches for quantifying and exploring local discrepancies between sampled graphs and their original counterparts at the node, edge, and structural levels—a gap that hinders effective evaluation and selection of graph sampling strategies. To bridge this gap, the authors propose three general quantitative metrics—neighborhood-, path-, and structure-based—to measure local fidelity. They further introduce DiffLens, an interactive visualization system that, for the first time, incorporates lens-based views tailored to these three types of differences, enabling users to focus on regions of interest. Case studies on real-world network datasets and user experiments demonstrate the framework’s effectiveness and practicality in supporting intuitive comparison of sampling outcomes and enhancing understanding of localized structural variations.
Efficiently supporting edge and adjacency queries on large-scale Barabási–Albert (BA) graphs (with out-degree one) and random recursive trees without storing or pre-generating the entire graph. Method: We propose the first on-demand, sublinear random-access framework, built upon probabilistic preprocessing, hierarchical indexing, and inverse sampling. It constructs a deterministic auxiliary structure and employs a lightweight online sampling algorithm. Results: Our framework guarantees polylogarithmic time, space, and random-bit complexity per query—i.e., $ ext{polylog}(n) $—while ensuring that query outputs are *exactly* distributed according to the standard BA model or random recursive tree. With probability $ 1 - 1/ ext{poly}(n) $, it enables faithful simulation of sublinear-time graph algorithms and accurate estimation of graph properties, significantly reducing computational, storage, and randomness overheads compared to full-graph approaches.
Path matching in graph query languages (e.g., Cypher, SQL/PGQ, GQL) lacks a unified and efficient processing mechanism—particularly when supporting complex path semantics (e.g., shortest paths, simple paths) and regular-expression constraints on edge labels—posing dual challenges in expressive power and performance. This paper introduces the first cross-language, general-purpose path-solving framework. It features a compact symbolic path representation and integrates dynamic-programming-based enumeration, incremental pipelined execution, and regex compilation optimizations to enable unified modeling and efficient evaluation of diverse path semantics and edge-label constraints. Experimental evaluation on real-world datasets and complex queries demonstrates an order-of-magnitude speedup over state-of-the-art graph engines, while maintaining high expressiveness, strong scalability, and behavioral stability.
This work addresses the underexplored impact of the implicit two-stage sampling mechanism—used in partitioning training, validation, and test sets—on link prediction performance. The authors propose a β-sampling strategy, wherein the probability of a link being sampled is proportional to the β-th power of the product of its endpoint node degrees, enabling systematic investigation of how second-stage sampling affects predictive accuracy. Large-scale experiments across 45 real-world networks demonstrate that prediction performance improves significantly when missing links are more likely to connect high-degree nodes. The findings reveal that the optimal sampling strategy is neither uniform random nor degree-preserving, underscoring the critical role of structural characteristics inherent in missing links for effective link prediction.
This work addresses the challenge of efficiently generating representative ensembles of districting plans by enabling independent sampling from the space of graph partitions. The authors propose a novel method that, for the first time, explicitly constructs a probability distribution over graph partitions under exact population balance constraints, thereby achieving truly independent samples and circumventing the mixing difficulties inherent in traditional Markov chain approaches. By integrating probabilistic modeling with an efficient sampling algorithm, the method demonstrates substantially improved sampling efficiency and diversity compared to existing Markov chain baselines, as validated on both grid graphs and real-world congressional and state legislative districting maps across U.S. states. This advance breaks away from the conventional paradigm reliant on sequential chain-based sampling.
This work addresses the challenge of efficiently constructing high-quality multi-hop reasoning training data from unstructured, expert-level documents lacking annotated structure. The authors propose a graph-constrained path selection mechanism that first constructs an offline keyword-context centroid graph and then applies five geometric admissibility constraints to identify plausible reasoning paths. To mitigate embedding drift, Gram matrix analysis is integrated into the pipeline. Validated paths are subsequently transformed into question-answer pairs by a teacher model, decoupling logical reasoning from language generation. Evaluated on the CUAD legal contract corpus, the method synthesizes 80,000 training samples, boosting Qwen3-32B’s closed-book Token F1 score from 21.66% to 38.58% and expanding usable training data by a factor of 4.4, particularly enhancing performance on templated and cross-referential documents.
In multi-hop graph retrieval, the original query often fails to fully capture the complete information need distributed across multiple reasoning steps, resulting in insufficient retrieval signals. This work proposes an evidence-guided query reformulation mechanism that decouples query refinement from evidence aggregation on the graph: a residual query is generated from already retrieved passages to characterize unmet information needs, and the retrieval signals from the original and residual queries are separately normalized and then fused, propagating through shared entities across propositions. By moving beyond the conventional reliance solely on the initial query, the method achieves substantial gains, improving Recall@5 by up to 5.59 points and F1 by up to 4.50 points on 2WikiMultiHopQA, HotpotQA, and MuSiQue.
This work addresses the limitations of traditional knowledge graph construction approaches, wherein structural decisions are hard-coded into rigid pipelines, resulting in tight coupling between schema and construction process and hindering support for ontology-level tasks. To overcome this, the authors propose an ontology-oriented construction framework featuring a novel intrinsic-relational routing mechanism. This mechanism dynamically assigns attributes to corresponding schema modules through iterative attribute classification, enabling a declarative and reusable decoupled design. The pipeline integrates rule-based cleaning, tool-augmented large language model–assisted annotation, and human review. Evaluated on Wikidata (January 2026), the resulting graph comprises 34 million nodes and 61.2 million edges, achieving 93.3% schema coverage and 98.0% module assignment accuracy, effectively supporting five ontology-level applications.