🤖 AI Summary
This study addresses the challenge that efficient sequence models struggle to simultaneously achieve data adaptivity, long-range dependency modeling, GPU parallelism, and nonlinear reasoning. To this end, we propose ADPTNet, which integrates linear attention, Riemannian optimization, and state space models to construct local topological conjugacy for nonlinear yet predictable long-term behavior. Furthermore, this work introduces dynamical systems theory for the first time to constrain timescale parameters and presents Conv DEER, an efficient Jacobian-free parallel algorithm. The framework is subsequently extended to neuromorphic spiking networks as SpikingADPTNet. Experimental results demonstrate that our model outperforms existing methods on Selective Copying and CIFAR-10 benchmarks, while SpikingADPTNet achieves a new state-of-the-art accuracy of 83.56% on the Speech Commands dataset.
📝 Abstract
A central aim of neuromorphic computing is to provide a viable alternative to highly energy-intensive Transformer-based AI. However, efficient alternatives struggle to capture the set of qualities that have secured the Transformer's status as the de facto standard in sequence modelling. Any realistic contender must be data-adaptive, able to capture long-range dependencies, and GPU-parallelisable, but also non-linearly recurrent to enable complex reasoning. Based on evidence suggesting the auditory cortex operates on fixed timescales, this work proposes the ADaptive with Prescriptive Timescales Network (ADPTNet) as a potential solution to achieving all four properties simultaneously. ADPTNet is built around local topological conjugates, obtained by a novel combination of linear attention and Riemannian optimisation, applied to static global dynamics. This enables non-linear yet predictable long-term behaviour. Dynamical systems theory proofs provide theoretical guarantees for the parametric control of ADPTNet's timescales (its Lyapunov spectrum). ADPTNet improves performance on Selective Copying over Hawk, the existing method balancing long-range memory and adaptability, while also improving state tracking over linear SSMs like Mamba. On sequential CIFAR-10, ADPTNet matches linear SSM accuracy and outperforms existing selective models (incl. the Transformer), using fewer parameters. We also introduce a neuromorphic SpikingADPTNet, which achieves a new state-of-the-art accuracy on the Spiking Speech Commands dataset ($83.56\%\pm0.15$). Finally, ADPTNet's constant timescales enable two efficient, Jacobian-free extensions to the DEER parallel simulation algorithm (Conv and Forward DEER) that retain the same average convergence. Conv DEER adds no computational overhead beyond the network's forward pass and enables non-linear RNN parallelisation via iterated convolutions for the first time.