🤖 AI Summary
This work addresses the problem of efficiently supporting substring frequency queries—returning occurrence counts rather than explicit positions. It introduces the first black-box framework that automatically transforms any reporting-based text index into a counting-based one. The approach leverages combinatorial lemmas to characterize the relationship between substring frequencies and lengths, precomputing frequencies for at most $n$ high-frequency substrings while handling low-frequency ones by converting reports from an existing index into counts. Requiring only linear space, the method achieves optimal $O(|P|)$ query time for a pattern $P$, and seamlessly extends to diverse settings—including consecutive occurrences, weighted sequences, strings with utilities, and non-overlapping matches—while maintaining optimal time and space efficiency in all cases.
📝 Abstract
We prove an elementary yet powerful combinatorial lemma: in any rooted tree with $L$ leaves, the number of nodes whose depth is smaller than the number of their leaf descendants is at most $L$. For any string $T$ of length $n$, a direct application of this lemma to the suffix trie of $T$ yields that the number of substrings of $T$ whose length is smaller than their number of occurrences in $T$ is at most $n$. This combinatorial insight leads to space-efficient data structures with optimal query times for string counting problems via the following algorithmic framework: store the counts for the at most $n$ ``frequent'' substrings of $T$ in a preprocessing step, and use a reporting query to count for the ``infrequent'' substrings. Our framework acts as a convenient black box, lifting indexes with reporting time $\mathcal{O}(|P|+|\textsf{Occ}_T(P)|)$ to support counting queries in time $\mathcal{O}(|P|)$, where $P$ is the queried pattern and $\textsf{Occ}_T(P)$ is the set of occurrences of $P$ in $T$. As applications, we show efficient indexes for consecutive occurrences, weighted sequences, strings with utilities, and non-overlapping occurrences.