π€ AI Summary
This study addresses the difficulty of traditional recurrent neural networks (RNNs) in modeling long-range dependencies due to vanishing gradients and limited receptive fields. To overcome these limitations, this work proposes a second-order recurrent model that replaces conventional neuronal communication mechanisms with a spatial evolution field governed by partial differential equations (PDEs). By conditioning the state on the complete historical trajectory, the architecture achieves an infinite receptive field under fixed parameters. Furthermore, marginal stability conditions are analytically derived to theoretically eliminate gradient pathologies. The primary contribution lies in a higher-order framework that integrates spatial neural computation with PDE discretization. This architecture significantly outperforms existing recurrent models on long-sequence benchmarks, achieving an effective balance between inference efficiency and long-term memory capacity while utilizing fewer parameters.
π Abstract
Recurrent neural networks (RNNs) offer linear-time scaling with sequence length while requiring only constant memory, yet they struggle to capture long-range dependencies due to vanishing gradients and limited receptive fields. To address these limitations, we introduce a second-order recurrent model in which the standard neuron-to-neuron communication is replaced by a spatially evolving field governed by (discretized) partial differential equations. Drawing inspiration from the role of cortical waves in brain computation, this mechanism allows structured spatiotemporal patterns to serve as an implicit, high-capacity memory. We show that the resulting model is equivalent to a structured infinite-order RNN in which the current state depends explicitly on its entire history of past states, yielding an effectively unbounded receptive field with a fixed number of parameters. We further derive constructive conditions to ensure marginal stability, constraining the gradient spectrum on the unit circle and thereby eliminating vanishing and exploding gradients. Empirically, the proposed architecture outperforms other recurrent models on long-horizon benchmarks while using substantially fewer parameters, demonstrating that spatial dynamics can effectively bridge the gap between efficient inference and long-term memory.