🤖 AI Summary
This work addresses the inefficiency of traditional backtracking-based regular expression engines, which suffer from poor worst-case performance and struggle to support advanced features such as lookaheads, negative lookaheads, and submatch extraction. Focusing on regular expressions with lookahead and negative lookahead (REwLA), the paper presents the first finite automaton transformation framework capable of handling submatches. By leveraging formal language theory and automata construction techniques, it converts an REwLA of size \( m \) into a deterministic finite automaton (DFA) with \( \tilde{O}(2^{2^m}) \) states, and further extends this approach to weighted REwLA by constructing a weighted NFA with comparable state complexity. This method overcomes the performance limitations of backtracking and establishes both a theoretical foundation and a practical pathway for efficiently implementing advanced regular expression functionalities.
📝 Abstract
Most of the conventional implementations of regular expressions are based on backtracking. Such implementations are slow in the worst case, and thus, we would like to develop a better matching algorithm. However, it is nontrivial to provide an efficient matching algorithm that can deal with practical extensions including submatch addressing. This paper studies regular expression with lookaheads and negative lookaheads, abbreviated to REwLA. First, we propose a transformation from a REwLA of size $m$ to a deterministic finite automaton of $\mr{O}(2^{2^m})$ states. Next, we consider weighted regular expressions, which enable us to calculate submatch addressing. We propose a transformation from a weighted REwLA of size $m$ to a weighted nondeterministic finite automaton of $\mr{O}(2^{2^m})$ states.