Score
Designs and analyzes online algorithms and data structures that incrementally detect, compute, and report maximal closed substrings (MCSs) in a text as characters are appended; implementations must identify newly formed MCS occurrences and test closedness under extensions immediately, with worst-case-optimal per-append update time.
This study addresses the problem of efficiently computing all maximal closed substrings (MCS) during the online, character-by-character input of a string. To this end, the authors propose a novel data structure—the Link-Cut Suffix Tree (LCST)—which integrates an online suffix tree with a link-cut tree to dynamically maintain the rightmost occurrence information of every substring. This enables real-time detection of newly formed MCS after each character insertion. Based on the LCST, they design the first worst-case time-optimal online algorithm for MCS enumeration, achieving a total time complexity of $O(n \log n)$ and space complexity of $O(n)$. The approach further extends to related applications such as rightmost LZ77 factorization and recent match queries.
This work addresses the efficient enumeration of closed substrings and their maximal variants (maximal closed substrings, MCSs) in a string. We present the first algorithm achieving $O(n log n)$ time and space complexity for enumerating all $Theta(n^2)$ closed substrings. For MCS extraction, we propose a lightweight method based on the suffix array and LCP array, ensuring both theoretical optimality and practical efficiency. To mitigate output explosion, we introduce a compact representation that significantly reduces output size. Furthermore, we fully characterize the asymptotic behavior of MCS counts in Fibonacci words, deriving the exact asymptotic formula $sim 1.382,F_n$. This constitutes the first near-linear-time algorithm for complete closed substring enumeration.
This study addresses the problem of efficiently enumerating all maximal closed substrings (MCS) in run-length encoded (RLE) strings. The authors propose a unified approach that operates directly on the RLE representation, leveraging a compact family-based encoding to handle both periodic and aperiodic cases efficiently. Their key contributions include the first algorithm for directly enumerating MCS from RLE without decompression, a proof that all MCS can be fully represented using only O(m²) families—with this bound shown to be tight in certain instances—and an output-sensitive enumeration algorithm. By integrating techniques such as sparse suffix trees, height-driven three-sided range reporting, and McCreight’s balanced priority search trees, the method outputs a compact representation of all MCS in O(m log²m + |F| log m) time and O(m) space, where m is the length of the RLE string and |F| denotes the number of output families.
This work addresses the problem of efficiently supporting substring frequency queries—returning occurrence counts rather than explicit positions. It introduces the first black-box framework that automatically transforms any reporting-based text index into a counting-based one. The approach leverages combinatorial lemmas to characterize the relationship between substring frequencies and lengths, precomputing frequencies for at most $n$ high-frequency substrings while handling low-frequency ones by converting reports from an existing index into counts. Requiring only linear space, the method achieves optimal $O(|P|)$ query time for a pattern $P$, and seamlessly extends to diverse settings—including consecutive occurrences, weighted sequences, strings with utilities, and non-overlapping matches—while maintaining optimal time and space efficiency in all cases.
This paper studies the Longest Common Extension (LCE) problem for strings with wildcards, where each wildcard matches any character. Parameterizing by the number $G$ of wildcard groups—a natural structural parameter—we present the first smooth time–space trade-off LCE data structure: it supports $O(t)$-time queries using $O(nG/t)$ space and $O(n(G/t)log n)$ preprocessing time. Technically, our approach integrates group-parameterized design, kangaroo jumping, and sparse Boolean matrix multiplication, and establishes a tight reduction to Boolean matrix multiplication. Under the 3SUM and Set-Disjointness conjectures, we prove the conditional optimality of this trade-off. Our result yields the first combinatorial, parameter-sensitive, and asymptotically optimal primitive for approximate pattern matching and structural analysis of wildcard strings.
We revisit two well-known algorithmic problems on strings: computing a shortest unique substring (SUS) and a shortest absent substring (SAS) of a string $S$ of length $n$. Both problems admit folklore $\mathcal{O}(n)$-time solutions using the suffix tree of $S$. However, for small alphabets, this complexity is not necessarily optimal in the word RAM model, where a string of length $n$ over alphabet $[0,σ)$ can be stored in $\mathcal{O}(n \log σ/\log n)$ space and read in $\mathcal{O}(n \log σ/\log n)$ time. We present an $\mathcal{O}(n \log σ/\sqrt{\log n})$-time algorithm for computing a SUS of $S$. This algorithm decomposes the problem according to the length and the period of the sought substring and uses several tools and techniques, such as synchronizing sets, the analysis of runs, and wavelet trees, to reduce the computation of a SUS to a simple geometric problem. Further, we adapt this algorithm and combine it with an efficient construction of de Bruijn sequences in order to obtain an $\mathcal{O}(n \log σ/\sqrt{\log n})$-time algorithm for computing a SAS of $S$.
This work addresses the problem of dynamically maintaining a minimal suffix-rich set for a string under near-real-time constraints to efficiently quantify its repetitiveness. Focusing on online scenarios where characters arrive one by one—either left-to-right or right-to-left—it presents the first algorithm achieving polyloglog worst-case time per character update. The approach leverages Weiner’s suffix tree and its fundamental algorithmic primitives to establish a core maintenance mechanism, thereby enabling, for the first time, efficient dynamic maintenance of minimal suffix-rich sets under bidirectional streaming input. This breakthrough substantially extends the applicability of string repetitiveness measures to dynamic environments.
This study addresses the online computation of the Longest Repeated Suffix (LRS) and the maintenance of a minimal suffix-dense set. The authors propose a novel approach based on an incremental run-length compressed Burrows–Wheeler Transform (BWT) index, enabling efficient updates of both structures within compressed space. Two time–space trade-offs are presented, achieving per-character processing times of \(O((\log n / \log \log n)^2)\) while using only \(O(n)\) or \(O(n \log \log n)\) bits of working space on highly repetitive texts—outperforming existing methods. Notably, this work achieves the first online construction of a minimal suffix-dense set in compressed space, provides worst-case time guarantees for LRS computation, and establishes a space lower bound for deterministic online LRS algorithms.
This study addresses the vulnerability of regular expressions with backreferences (REWB) to ReDoS attacks and the lack of efficient matching algorithms. Combining fine-grained complexity theory and string algorithms, it establishes new theoretical lower bounds by proving, under the Strong Exponential Time Hypothesis (SETH), that k-REWB matching cannot be solved in O(n^{2k−ε}) time for any ε > 0. This result elevates the parameterized complexity of the problem from W[1] to W[2] and reveals a novel connection to the triangle detection problem. On the algorithmic front, the paper presents the first near-linear-time matching algorithm for 1-use REWB, achieving O(n log²n) time complexity—substantially improving upon the previous O(n²) bound—by leveraging suffix trees, transition monoids, factorization forests, and periodicity analysis.
This study addresses the problem of efficiently identifying the longest palindromic substring in strings containing wildcards, where each wildcard can match any character and thereby disrupts the symmetry exploited by classical palindrome-finding algorithms. To overcome this challenge, we present the first non-trivial algorithm that operates within linear space, leveraging a novel wildcard-based longest common extension (wildcard-LCE) technique. Our approach establishes a continuous trade-off between time and memory usage, achieving substantial improvements in the time–memory product across a range of parameter settings. Furthermore, the method generalizes effectively to the more challenging setting allowing up to k mismatches, yielding both theoretical advances and practical performance gains.