🤖 AI Summary
This study addresses the challenge of efficiently distributing persistent representations and inheriting recurrent states in long-sequence modeling. To this end, it proposes an attention-aligned recurrent architecture that distributes persistent representations via low-dimensional trajectory rollout. The method enhances causal flow by recursively accumulating delay increments through fixed additive phases, integrating local rightward and leftward attention reads with a stride-based prefix sum algorithm to support parallel prefilling and bounded-buffer continuation. As a primary contribution, this work introduces the first attention-aligned recurrence mechanism, achieving 99% accuracy on the Equal Repeats task and significantly outperforming non-inheriting baselines. Furthermore, it reveals that distinct tasks impose differential requirements on recurrence dimensions, offering new insights into the design of efficient long-context models.
📝 Abstract
We present TraceRelay, an attention-aligned recurrent architecture that distributes persistent representations over a rolling sequence of low-dimensional traces. Local right looking attention forms increments from lower-layer representations; delivery is delayed until all attended inputs are in the causal past. A fixed additive phase recurrence accumulates the delayed increments, and left-looking attention reads the resulting residual augmented stream. A stride-wise prefix sum supports parallel prefill and bounded-buffer continuation. We study 36 small-model runs on Equal Repeats, bounded Dyck closing-type prediction, and causal Most-Freq generation, using three seeds per setting. At trained length 256, Equal Repeats models with recurrent phase inheritance reach 98.81-99.69% accuracy versus 50.73-51.63% for separately trained variants without inheritance, despite the latter receiving more updates. Accuracy drops sharply at lengths beyond the training range. At the longest evaluated lengths, models with more dimensions in the middle layer's recurrent traces perform better on Dyck (76.34% versus 55.49% close accuracy at length 4096), whereas models with fewer trace dimensions perform better on five-symbol Most-Freq (70.74% versus 55.60% exact generation at length 1024). These contrasting cases motivate further study of how the size of recurrent representations should be chosen for different tasks, without establishing a general rule across tasks or model configurations.