Text Indexing: From Reporting to Counting

📅 2026-07-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the problem of efficiently supporting substring frequency queries—returning occurrence counts rather than explicit positions. It introduces the first black-box framework that automatically transforms any reporting-based text index into a counting-based one. The approach leverages combinatorial lemmas to characterize the relationship between substring frequencies and lengths, precomputing frequencies for at most $n$ high-frequency substrings while handling low-frequency ones by converting reports from an existing index into counts. Requiring only linear space, the method achieves optimal $O(|P|)$ query time for a pattern $P$, and seamlessly extends to diverse settings—including consecutive occurrences, weighted sequences, strings with utilities, and non-overlapping matches—while maintaining optimal time and space efficiency in all cases.
📝 Abstract
We prove an elementary yet powerful combinatorial lemma: in any rooted tree with $L$ leaves, the number of nodes whose depth is smaller than the number of their leaf descendants is at most $L$. For any string $T$ of length $n$, a direct application of this lemma to the suffix trie of $T$ yields that the number of substrings of $T$ whose length is smaller than their number of occurrences in $T$ is at most $n$. This combinatorial insight leads to space-efficient data structures with optimal query times for string counting problems via the following algorithmic framework: store the counts for the at most $n$ ``frequent'' substrings of $T$ in a preprocessing step, and use a reporting query to count for the ``infrequent'' substrings. Our framework acts as a convenient black box, lifting indexes with reporting time $\mathcal{O}(|P|+|\textsf{Occ}_T(P)|)$ to support counting queries in time $\mathcal{O}(|P|)$, where $P$ is the queried pattern and $\textsf{Occ}_T(P)$ is the set of occurrences of $P$ in $T$. As applications, we show efficient indexes for consecutive occurrences, weighted sequences, strings with utilities, and non-overlapping occurrences.
Problem

Research questions and friction points this paper is trying to address.

string counting
text indexing
substring occurrences
data structures
query time
Innovation

Methods, ideas, or system contributions that make the work stand out.

combinatorial lemma
string indexing
counting queries
suffix trie
space-efficient data structures
🔎 Similar Papers
No similar papers found.