enumerate maximal closed substrings

Designs and implements algorithms that, given a string (optionally in a run-length encoded form), enumerate all maximal closed substrings — substrings (also called maximal closed repeats) that cannot be extended without changing their set of distinct contexts — and report their occurrences. This work focuses on output-sensitive enumerators with provable time and working-space bounds and on correctly handling both periodic and non-periodic cases.

enumeratemaximalclosedsubstrings

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.88
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the problem of efficiently enumerating all maximal closed substrings (MCS) in run-length encoded (RLE) strings. The authors propose a unified approach that operates directly on the RLE representation, leveraging a compact family-based encoding to handle both periodic and aperiodic cases efficiently. Their key contributions include the first algorithm for directly enumerating MCS from RLE without decompression, a proof that all MCS can be fully represented using only O(m²) families—with this bound shown to be tight in certain instances—and an output-sensitive enumeration algorithm. By integrating techniques such as sparse suffix trees, height-driven three-sided range reporting, and McCreight’s balanced priority search trees, the method outputs a compact representation of all MCS in O(m log²m + |F| log m) time and O(m) space, where m is the length of the RLE string and |F| denotes the number of output families.

closed stringcompact representationmaximal closed substring

This study addresses the problem of efficiently computing all maximal closed substrings (MCS) during the online, character-by-character input of a string. To this end, the authors propose a novel data structure—the Link-Cut Suffix Tree (LCST)—which integrates an online suffix tree with a link-cut tree to dynamically maintain the rightmost occurrence information of every substring. This enables real-time detection of newly formed MCS after each character insertion. Based on the LCST, they design the first worst-case time-optimal online algorithm for MCS enumeration, achieving a total time complexity of $O(n \log n)$ and space complexity of $O(n)$. The approach further extends to related applications such as rightmost LZ77 factorization and recent match queries.

maximal closed substringsonline computationrepetitive structures

Efficient Computation of Closed Substrings

Jun 06, 2025
SK
Samkith K Jain
🏛️ McMaster University

This work addresses the efficient enumeration of closed substrings and their maximal variants (maximal closed substrings, MCSs) in a string. We present the first algorithm achieving $O(n log n)$ time and space complexity for enumerating all $Theta(n^2)$ closed substrings. For MCS extraction, we propose a lightweight method based on the suffix array and LCP array, ensuring both theoretical optimality and practical efficiency. To mitigate output explosion, we introduce a compact representation that significantly reduces output size. Furthermore, we fully characterize the asymptotic behavior of MCS counts in Fibonacci words, deriving the exact asymptotic formula $sim 1.382,F_n$. This constitutes the first near-linear-time algorithm for complete closed substring enumeration.

Calculates exact number of maximal closed substrings in Fibonacci wordsDevelops fast algorithm to compute all closed substringsIntroduces compact representation for closed substrings

R-enum Revisited: Speedup and Extension for Context-Sensitive Repeats and Net Frequencies

Nov 14, 2025
KK
Kotaro Kimura
🏛️ Kyushu Institute of Technology

Efficient enumeration and analysis of context-sensitive repeats in strings—such as maximal repeats, super-/near-super-maximal repeats—remain computationally challenging. Method: We propose a linear-time algorithm based on the run-length compressed Burrows–Wheeler transform (RLBWT), reducing the r-enum complexity from $O(n log log_w (n/r))$ to $O(n)$. Our approach enables, for the first time, joint computation of near-super-maximal repeats along with their net occurrences and net frequencies. We prove that the total number of net occurrences is strictly less than $2r$, and construct an $O(r)$-space data structure supporting efficient net-frequency queries for arbitrary patterns. Contributions/Results: (i) Full context-sensitive repeat enumeration in $O(n)$ time and $O(r)$ space; (ii) a new upper bound of $2r$ on the number of minimal unique substrings; (iii) a unified solution to key problems—including context diversity quantification of maximal repeats, net-frequency statistics, and dynamic net-frequency querying.

Building compact data structures for net frequency queries on patternsExtending r-enum to compute context-sensitive repeats and net frequenciesImproving time complexity of r-enum algorithm for substring enumeration

Longest Common Extensions with Wildcards: Trade-off and Applications

Aug 07, 2024
GB
Gabriel Bathie
🏛️ École normale supérieure de Paris | PSL Research University | Birkbeck | University of London

This paper studies the Longest Common Extension (LCE) problem for strings with wildcards, where each wildcard matches any character. Parameterizing by the number $G$ of wildcard groups—a natural structural parameter—we present the first smooth time–space trade-off LCE data structure: it supports $O(t)$-time queries using $O(nG/t)$ space and $O(n(G/t)log n)$ preprocessing time. Technically, our approach integrates group-parameterized design, kangaroo jumping, and sparse Boolean matrix multiplication, and establishes a tight reduction to Boolean matrix multiplication. Under the 3SUM and Set-Disjointness conjectures, we prove the conditional optimality of this trade-off. Our result yields the first combinatorial, parameter-sensitive, and asymptotically optimal primitive for approximate pattern matching and structural analysis of wildcard strings.

Apply solution to pattern matching and string analysisDevelop optimal data structure for wildcard LCEStudy LCE problem in strings with wildcards

Latest Papers

What's happening recently
View more

This study resolves a long-standing open problem concerning the attainability of the string repetitiveness measure χ(w): whether there always exists a string representation of size O(χ(w)). To this end, we introduce the Substring Equation System (SES), a novel theoretical framework, and combine it with combinatorial structures such as suffix-complete sets to construct the first compression scheme capable of representing any string w within O(χ(w)) space. Our work not only establishes, for the first time, the attainability of the χ measure but also provides a new modeling paradigm for string compression that leverages structural regularities in repetitive strings.

compressed representationreachabilityrepetitiveness measure

This study addresses the problem of efficiently identifying the longest palindromic substring in strings containing wildcards, where each wildcard can match any character and thereby disrupts the symmetry exploited by classical palindrome-finding algorithms. To overcome this challenge, we present the first non-trivial algorithm that operates within linear space, leveraging a novel wildcard-based longest common extension (wildcard-LCE) technique. Our approach establishes a continuous trade-off between time and memory usage, achieving substantial improvements in the time–memory product across a range of parameter settings. Furthermore, the method generalizes effectively to the more challenging setting allowing up to k mismatches, yielding both theoretical advances and practical performance gains.

k-mismatchesmaximal palindromestime-memory tradeoffs

We revisit two well-known algorithmic problems on strings: computing a shortest unique substring (SUS) and a shortest absent substring (SAS) of a string $S$ of length $n$. Both problems admit folklore $\mathcal{O}(n)$-time solutions using the suffix tree of $S$. However, for small alphabets, this complexity is not necessarily optimal in the word RAM model, where a string of length $n$ over alphabet $[0,σ)$ can be stored in $\mathcal{O}(n \log σ/\log n)$ space and read in $\mathcal{O}(n \log σ/\log n)$ time. We present an $\mathcal{O}(n \log σ/\sqrt{\log n})$-time algorithm for computing a SUS of $S$. This algorithm decomposes the problem according to the length and the period of the sought substring and uses several tools and techniques, such as synchronizing sets, the analysis of runs, and wavelet trees, to reduce the computation of a SUS to a simple geometric problem. Further, we adapt this algorithm and combine it with an efficient construction of de Bruijn sequences in order to obtain an $\mathcal{O}(n \log σ/\sqrt{\log n})$-time algorithm for computing a SAS of $S$.

shortest absent substringshortest unique substringstring algorithms

This study investigates the instability of mainstream repetitiveness measures—such as the number of runs in the Burrows–Wheeler Transform (\(r\)), the size of the LZ77 parsing (\(z\)), and the LZ-end size (\(v\))—under string reversal. By constructing specific families of strings and employing combinatorial analysis, asymptotic methods, and information-theoretic tools, the authors establish the first tight additive sensitivity bound of \(\Theta(n)\) for \(r\) under reversal, substantially improving the previous \(\Omega(\log n)\) lower bound. They also demonstrate that the ratio \(z(w^R)/z(w)\) can approach 3 and that the multiplicative sensitivity of \(v\) is \(\Theta(\log n)\). All derived bounds are asymptotically tight, revealing the pronounced vulnerability of these measures to such a simple transformation.

Burrows-Wheeler transformLempel-Ziv parsingrepetitiveness measures

This work addresses the problem of efficiently supporting substring frequency queries—returning occurrence counts rather than explicit positions. It introduces the first black-box framework that automatically transforms any reporting-based text index into a counting-based one. The approach leverages combinatorial lemmas to characterize the relationship between substring frequencies and lengths, precomputing frequencies for at most $n$ high-frequency substrings while handling low-frequency ones by converting reports from an existing index into counts. Requiring only linear space, the method achieves optimal $O(|P|)$ query time for a pattern $P$, and seamlessly extends to diverse settings—including consecutive occurrences, weighted sequences, strings with utilities, and non-overlapping matches—while maintaining optimal time and space efficiency in all cases.

data structuresquery timestring counting

Hot Scholars

SI

Shunsuke Inenaga

Professor, Department of Informatics, Kyushu University
Algorithms and Data StructuresString AlgorithmsCompressionCombinatorics on Words
JL

Jana Lasser

University of Graz
data analysismachine learningcomputational social sciencecomplex systems
RS

Ravid Shwartz-Ziv

New York University
machine learningdeep learningrepresentation learning theoryneuroscience
TM

Takuya Mieno

The University of Electro-Communications
Stringology
GG

Gennian Ge

Capital Normal University
CombinatoricsCoding theoryInformation Security