🤖 AI Summary
This study addresses the limitation of state space models (SSMs) in retaining long-sequence information due to exponential forgetting. To overcome this, we propose FRAC, an architecture that introduces the power-law decay property of fractional dynamics into selective SSMs for the first time, replacing conventional exponential forgetting mechanisms to achieve long-range memory. Furthermore, FRAC approximates heavy-tailed kernels using a sum of logarithmically spaced exponentials, constructing a bounded-state recurrent module that supports parallel training while maintaining computational efficiency. Experimental results demonstrate that, in 1.3B-parameter language modeling tasks, FRAC significantly enhances long-context performance while preserving competitive accuracy on short contexts.
📝 Abstract
State Space Models (SSMs) compress sequence history into a bounded recurrent state, making the resulting memory law a central architectural choice for long-context performance. Most modern SSMs rely on ODE-based dynamics that lead to exponential forgetting, limiting their ability to retain information over broad temporal ranges. We introduce FRAC, a selective SSM architecture derived from fractional dynamics that replaces this exponential decay with power-law long memory. To make fractional dynamics practical, FRAC approximates the heavy-tailed target kernel with a finite-state, log-spaced sum of exponential modes. This construction turns fractional memory into an efficient recurrent module with parallel training and prefill, while retaining bounded-state autoregressive decoding. Extensive experiments, including 1.3B-parameter language modeling, demonstrate that FRAC consistently improves long-context performance over state-of-the-art SSM baselines while staying competitive on short-context. These results show that fractional dynamics provide a practical and effective prior for long-context SSMs.