detect patterns online

Design and implement algorithms and data structures that incrementally detect and maintain pattern-related structures and properties as data arrive (e.g., updating structures on append), including discovery of newly maximal substrings, maintenance of online substring indexes, and characterization of substring closures.

detectpatternsonline

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.36
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

String Matching with a Dynamic Pattern

Jun 12, 2025
BM
Bruno Monteiro
🏛️ Federal University of Minas Gerais

This paper addresses the problem of real-time exact pattern matching for a dynamic pattern string (P) against a static text (T), supporting single-character insertions/deletions, substring deletions, shifts, and copy operations on (P), with immediate reporting of (P)’s occurrence count in (T) after each edit. We propose the first lightweight dynamic string index based on the suffix array, integrating binary search with efficient interval maintenance. The index requires (O(|T|)) preprocessing time, achieves (O(log |T|)) amortized time per edit operation, and supports online text expansion. Unlike prior approaches, our framework is the first to achieve logarithmic update complexity for multiple edit types within a unified index structure. Experimental results demonstrate high efficiency and low latency on large-scale texts.

Extending solution to online text with amortized boundsHandling dynamic pattern updates in string matchingSupporting character edits and substring operations efficiently

Incremental Computation: What Is the Essence? (Invited Contribution)

Dec 13, 2023
YA
Yanhong A. Liu
🏛️ Stony Brook University

Repetitive computation in change-sensitive programs—such as database queries, compilers, and real-time analytics—incurs substantial overhead and undermines complexity control. Method: We propose the “incrementalization” paradigm, formalizing incremental computation as a discrete analogue of differentiation and establishing its theoretical foundation in discrete computation. Our approach introduces an “iterate–incrementalize–implement” design framework, featuring a novel meta-level abstraction-driven model for algorithmic complexity refinement, integrating higher-order abstractions over data, control flow, and modules with formal incrementalization transformations. Contribution/Results: We deliver a reusable, formally verifiable incrementalization methodology that guarantees correctness while significantly improving computational efficiency and enhancing controllability of algorithmic complexity. The framework enables systematic, principled application of incremental computation across diverse domains, bridging theory and practice in program optimization and reactive systems.

Applying incremental computation to algorithms and distributed systemsEfficiently computing output changes for input modificationsSystematic method for incrementalization in program optimization

This work proposes a formalization of algorithms within an intensional computability framework and clarifies their relationship to implementations in computational models. Treating computational models as monoid actions on configuration spaces, programs are modeled as dynamical systems constrained by such actions. Algorithms are defined as finite directed graphs of partial maps over edge-labeled abstract data structures, explicitly separating control flow from data operations. By leveraging tools from category theory, dynamical systems theory, and graph theory, the approach constructs a rigorous semantic framework that, for the first time, treats algorithms as abstract specifications of computational behavior and precisely characterizes the structure-preserving implementation relation between programs and algorithms, thereby deepening our understanding of the nature of computation.

abstract data structurealgorithmcomputability

Existing algorithm identification methods often suffer from poor usability, limited scalability, and insufficient evaluation. This work proposes a novel paradigm that integrates domain-specific languages (DSLs) with abstract syntax tree (AST) pattern matching: algorithmic characteristics are formally specified using a DSL to construct a reusable library of AST patterns, enabling automatic recognition of common algorithm implementations in source code. Evaluated on a subset of BigCloneEval, the approach achieves an average F1 score of 0.74, substantially outperforming CodeLlama (0.35) and state-of-the-art code clone detectors, which attain a recall of only 0.20 compared to our method’s 0.62. This advance represents a dual improvement in both precision and practical applicability for algorithm identification.

abstract syntax treealgorithm recognitionautomated analysis

Existing graph databases lack effective support for the tree-shaped substructures commonly found in property graphs. This work addresses this limitation by treating such tree substructures as first-class citizens and proposes a systematic management framework encompassing modeling, indexing, and query optimization. Drawing inspiration from XML structural indexing techniques, the approach enables efficient path queries within a relational graph database backend. Experimental evaluation demonstrates that the proposed method significantly improves path query performance, thereby validating the potential of structural indexing to enhance graph data management.

graph schemasproperty graphsquery languages

Latest Papers

What's happening recently
View more

This work addresses the problem of efficiently supporting substring frequency queries—returning occurrence counts rather than explicit positions. It introduces the first black-box framework that automatically transforms any reporting-based text index into a counting-based one. The approach leverages combinatorial lemmas to characterize the relationship between substring frequencies and lengths, precomputing frequencies for at most $n$ high-frequency substrings while handling low-frequency ones by converting reports from an existing index into counts. Requiring only linear space, the method achieves optimal $O(|P|)$ query time for a pattern $P$, and seamlessly extends to diverse settings—including consecutive occurrences, weighted sequences, strings with utilities, and non-overlapping matches—while maintaining optimal time and space efficiency in all cases.

data structuresquery timestring counting

This work addresses the problem of efficiently maintaining the rank, a column basis, and a maximum-rank submatrix of a matrix under dynamic updates—either to individual entries or entire columns—and applies these techniques to the dynamic maximum matching problem in graphs. The paper presents the first dynamic algorithm whose update time depends on the current rank \( r \) rather than the matrix dimension \( n \). By integrating sparse update strategies with rank-sensitive complexity analysis, it achieves an amortized update time of \( \tilde{O}(r^{1.405}) \) for single-entry modifications and \( \tilde{O}(r^{1.528} + z) \) for column updates, where \( z \) denotes the number of changed entries. This approach is the first to simultaneously support dynamic maintenance of rank, basis, and maximum-rank submatrix, yielding an edge update time of \( \tilde{O}(|M|^{1.405}) \) for dynamic graph matching and significantly improving upon prior methods.

dynamic graphsdynamic rankfull-rank submatrix

This work addresses the string constraint satisfaction problem involving relational string constraints—such as replaceAll—and length constraints, proposing an efficient solving approach. It introduces a novel extension of automata-based stabilization techniques to finite-state transducer representations of relational constraints, thereby substantially reducing the costly concatenation-elimination operations inherent in prior methods. The approach further integrates strong heuristic strategies to optimize both the handling of length constraints and the overall search process. Experimental evaluation demonstrates that the proposed method significantly outperforms existing solvers on benchmarks featuring relational constraints, solving more instances and achieving speedups of several orders of magnitude.

concatenation eliminationfinite-state transducerslength constraints

This study addresses the problem of efficiently computing all maximal closed substrings (MCS) during the online, character-by-character input of a string. To this end, the authors propose a novel data structure—the Link-Cut Suffix Tree (LCST)—which integrates an online suffix tree with a link-cut tree to dynamically maintain the rightmost occurrence information of every substring. This enables real-time detection of newly formed MCS after each character insertion. Based on the LCST, they design the first worst-case time-optimal online algorithm for MCS enumeration, achieving a total time complexity of $O(n \log n)$ and space complexity of $O(n)$. The approach further extends to related applications such as rightmost LZ77 factorization and recent match queries.

maximal closed substringsonline computationrepetitive structures

This study investigates the inclusion depth of pattern languages—the length of the longest strict inclusion chain from the universal pattern language to a given pattern-generated language—a measure that captures the cognitive complexity, in terms of mind changes, required to identify a pattern from positive examples. The authors propose the conjectural formula ID_Σ(p) = 2|p| − #var(p) − 1 and develop a linear-time algorithm based on this hypothesis, thereby establishing a profound connection between formal language theory and learning complexity. By integrating combinatorics, algorithmic learning theory, and mind-change analysis, this work not only highlights the fundamental open nature of the inclusion depth problem but also lays a new theoretical foundation for efficient pattern recognition: if the conjecture holds, inclusion depth becomes computable in linear time.

algorithmic learning theorycomputabilityinclusion depth

Hot Scholars

HN

Haruki Nishimura

Toyota Research Institute
roboticsmachine learningplanning under uncertaintystatistics
YL

Yuqi Li

The City College of New York, the City University of New York
Model CompressComputer Vision
SI

Shunsuke Inenaga

Professor, Department of Informatics, Kyushu University
Algorithms and Data StructuresString AlgorithmsCompressionCombinatorics on Words
YJ

Yuanliang Ju

CS Ph.D. Student @ University of Toronto
Robotics3D Vision