🤖 AI Summary
This study investigates how the memory capacity and rule generalization capabilities of neural architectures over symbolic sequences are constrained by sequence complexity. To this end, it introduces the ROTE benchmark, which leverages LZW compression to precisely modulate the algorithmic complexity of sequences, enabling a systematic comparison of RNNs, attention mechanisms, and hybrid recurrent-attention models on prediction and closed-loop unrolling tasks. The primary contribution lies in establishing a formal link between algorithmic complexity and network memory capacity, thereby elucidating the trade-offs among memory quality, stability, and computational cost. Furthermore, this work quantifies architectural differences across string distance metrics, training time, and memory footprint, and releases fully reproducible experimental code to facilitate future research.
📝 Abstract
We introduce ROTE (RollOut Testing of Exact memorization), a benchmarking protocol for evaluating symbolic memorization of neural architectures. We study memorization and the extension of symbolic rules in neural sequence models by using sequences whose complexity is regulated by Lempel--Ziv--Welch (LZW) compression. Under ROTE, each architecture is trained as the same finite-context conditional predictor and is evaluated using teacher-forced one-step prediction as well as closed-loop rollout on the withheld symbols. Following a shared prediction-and-rollout evaluation routine, the benchmark evaluates gated recurrent, minimal recurrent, attention-based, and hybrid recurrent-attention models with their native computational characteristics preserved. Beyond standard predictive metrics, the benchmark reports normalized string distances, training time, memory usage, and parameter count across an LZW-complexity sweep. The study establishes a connection between the complexity of algorithmic sequences and the memorization capacity of neural architectures, revealing the trade-offs involving memorization quality, rollout stability, and computational expense. Our software and reproducible experimental code can be obtained from https://github.com/nla-group/rote.