Score
Designs and implements indexing structures for interval-valued records that organize entries by interval duration and endpoint (e.g., a duration-endpoint tree or "tide" index). Such work builds two-layer, append-only B+-tree variants that support very fast append-heavy insertions and enable efficient interval query pruning via corner-space (duration×endpoint) classification.
This work addresses the lack of disk-based index structures for large-scale interval data in time-series databases that simultaneously support efficient querying, fast insertion, and compact storage. The authors propose CEB and TIDE, two novel two-level append-only B⁺-tree indexes that, for the first time, map intervals into a two-dimensional corner space by leveraging interval centroids or durations together with monotonically increasing end timestamps to establish ordering and indexing mechanisms. This approach significantly improves insertion throughput and reduces storage overhead while preserving high query efficiency. Experimental results demonstrate that both CEB and TIDE outperform existing solutions in terms of index size and insertion speed, with TIDE consistently achieving superior query performance—accelerating queries by up to several orders of magnitude.
Spatial conjunctive queries—such as range emptiness, counting, and nearest-neighbor search—suffer from inefficient query response times and excessive index space overhead when evaluated over join results. Method: We propose the first general-purpose indexing framework that simultaneously achieves time and space optimality. Theoretically, we establish the first tight space–time trade-off lower bounds for *k*-star and *k*-path queries. Methodologically, we design a compact index structure based on generalized hypertree decompositions, which provably meets these lower bounds and extends to arbitrary joins and hierarchical queries. Results: Experiments demonstrate significant index size reduction alongside sublinear query response times, substantially accelerating spatially aware relational query processing. Our core contribution is the establishment of theoretical optimality for spatial conjunctive queries over joins, coupled with index construction and query algorithms that are provably efficient.
This paper addresses the problem of efficiently accessing suffix arrays (SAs) when they cannot be stored explicitly. It establishes, for the first time, a *bidirectional equivalence*—in space, query time, and construction efficiency—between SA access and prefix selection, unifying the complexity analysis of fundamental string indexing operations. Through systematic reductions, the authors identify six pairs of intrinsically equivalent problems and prove that nearly all optimal SA representations can be realized via prefix selection structures. Leveraging this equivalence, they design a data structure supporting sublinear construction: for binary text, it achieves *O(n)* bits of space, *O(n/√log n)* preprocessing time, and *O(log^ε n)* query time—*matching and closing a long-standing complexity gap* in the field.
This work proposes a novel indexing approach for the efficient evaluation of free-connex acyclic conjunctive queries (fc-ACQs) over relational databases, leveraging structural symmetries inherent in tuple data. By introducing an auxiliary database $D_{col}$ and employing a relation coloring refinement technique, the method constructs a compact structural index that enables linear-time preprocessing and constant-delay enumeration or counting. This is the first approach to exploit internal structural symmetries in relational data, departing from conventional value- or order-based indexing paradigms. The resulting index achieves significant compression on canonical structures such as binary trees and regular graphs—while maintaining worst-case linear size—and supports efficient evaluation of all fc-ACQs in time complexity strictly better than the size of the underlying database.
This paper studies efficient enumeration of results for projected star- and path-shaped conjunctive queries (CQs), aiming for provably low enumeration delay after preprocessing. We propose an instance-adaptive combinatorial enumeration framework. For star queries, we achieve—first time—linear preprocessing time and sublinear (optimal) enumeration delay. For path queries, we establish new preprocessing–delay trade-off bounds that significantly improve upon prior approaches. Furthermore, by modeling Boolean matrix multiplication as a projected CQ, we derive the first sparse, output-sensitive matrix multiplication algorithm with guaranteed enumeration delay. Our core techniques include incremental join computation, delay-aware enumeration design, and instance-specific analysis—yielding substantial advances in both theoretical guarantees and practical efficiency.
This work addresses the inefficiency and inaccuracy of traditional interval join algorithms, which disregard overlap duration and consequently generate excessive spurious results. To overcome this limitation, the paper investigates interval joins under explicit overlap duration constraints and proposes the first efficient algorithm that natively supports such constraints, thereby eliminating the need for costly post-filtering. By constructing a dedicated interval index, devising effective pruning strategies, and incorporating a mechanism to pre-estimate overlap durations, the approach substantially reduces intermediate result sizes. Experimental evaluation on three real-world datasets demonstrates that the proposed method significantly outperforms existing techniques, achieving notable reductions in both computation time and output size.
Existing methods struggle to efficiently support diverse interval-aware approximate nearest neighbor queries using a single index, often resulting in redundant indices and high memory overhead. This work proposes a unified interval-aware Relative Neighborhood Graph framework (URNG) and its practical graph index, UG, which for the first time enables a single graph structure to simultaneously accommodate a variety of interval-constrained query semantics. By integrating unified pruning, iterative repair, and query-specific strategies, the approach ensures both monotonic searchability and hereditary subgraph structure during querying. Experimental results demonstrate that UG consistently achieves an excellent trade-off between accuracy and efficiency across five datasets, while maintaining competitive index construction costs and memory footprint.
This work addresses the problem of efficiently supporting substring frequency queries—returning occurrence counts rather than explicit positions. It introduces the first black-box framework that automatically transforms any reporting-based text index into a counting-based one. The approach leverages combinatorial lemmas to characterize the relationship between substring frequencies and lengths, precomputing frequencies for at most $n$ high-frequency substrings while handling low-frequency ones by converting reports from an existing index into counts. Requiring only linear space, the method achieves optimal $O(|P|)$ query time for a pattern $P$, and seamlessly extends to diverse settings—including consecutive occurrences, weighted sequences, strings with utilities, and non-overlapping matches—while maintaining optimal time and space efficiency in all cases.
This work addresses the challenge of efficiently supporting mixed queries involving interval predicates—such as containment and overlap—in approximate nearest neighbor (ANN) search, a task poorly handled by existing methods due to their reliance on coupled endpoint conditions or lack of a general indexing abstraction. The authors propose the Unified Dominance Graph (UDG), which maps both data objects and query intervals into a normalized two-dimensional dominance space and constructs a graph index annotated with dominance labels. This framework is the first to enable unified ANN search across diverse interval predicates. By leveraging semantic mappings, UDG facilitates index reuse, and it further enhances traversal efficiency under strong filtering through validity-preserving patch edges. Experiments demonstrate that UDG consistently outperforms state-of-the-art approaches across various interval relationships and workloads, achieving low indexing overhead and stable query performance.
本文提出了一种新的全文索引ZigZag Trie,用于高效处理字符串P在长文本T中的上下文查询问题,通过重组文本以群组对称增长的LR子串,支持四种类型的上下文查询。