Score
Designs and implements recurrent neural network architectures and their components—such as recurrent memory modules, update rules, and per-layer activation dynamics—and builds recurrent policy networks and state estimators. Develops and applies training procedures and evaluates streaming efficiency, stability, and stepwise behavior for models that process sequential or streaming data.
Recurrent Neural Networks (RNNs) and their variants—despite shared sequential modeling objectives—exhibit significant heterogeneity in architecture, objective functions, and learning algorithms, leading to conceptual ambiguity and fragmented understanding. Method: This survey introduces the first unified taxonomy grounded in three orthogonal dimensions: network architecture, training objective, and optimization algorithm. It systematically categorizes mainstream models—including LSTMs, convolutional recurrent networks, graph/tree-structured RNNs, higher-order RNNs, and memory-augmented architectures—while analyzing their interdependencies and evolutionary trajectories. Contribution/Results: The work establishes a generalizable modeling paradigm for complex sequence, speech, and image tasks; clarifies the fundamental distinctions and synergies between recursive and recurrent paradigms; synthesizes representative applications in NLP, automatic speech recognition, and image understanding; and identifies emerging research directions—including differentiable neural architecture search, dynamic computation graphs, and neuro-symbolic integration.
Existing RNN variants suffer from fragmented implementations across deep learning frameworks, lacking standardized interfaces and rigorous cross-framework evaluation—leading to poor reproducibility, redundant development, and inefficient experimentation. Method: We introduce the first standardized, multi-framework RNN interface, implemented in open-source libraries: `torchrecurrent` (Python, for PyTorch), `RecurrentLayers.jl`, and `LuxRecurrentLayers.jl` (Julia, for JAX/Flux). These provide unified, semantically consistent implementations of diverse recurrent units—including LSTM, GRU, mGRU, and Delta-RNN—as well as higher-level architectures. Contribution/Results: Our work establishes the first modular, extensible, and framework-agnostic RNN API, enabling flexible customization and seamless model migration across ecosystems. All libraries are MIT-licensed, actively maintained, and rigorously tested. By abstracting framework-specific complexities, they significantly lower implementation overhead and experimental barriers, thereby promoting structural standardization, reproducible research, and collaborative advancement in recurrent modeling.
Long-sequence modeling faces fundamental challenges including limited context length, difficulty in capturing long-range dependencies, and low efficiency in online learning. To address these, this work systematically reviews the resurgence of state-space models (SSMs) and recurrent computation, proposing a novel local forward-computation paradigm tailored for real-world online learning—thereby circumventing the temporal backtracking constraints inherent in standard backpropagation through time (BPTT). We introduce the first unified taxonomy encompassing both deep SSMs and large-context Transformers. Our framework integrates structured linear attention, enhanced RNN architectures, local recurrence mechanisms, and online optimization algorithms. The study rigorously clarifies the theoretical representational advantages and practical sequential reasoning benefits of recurrent modeling over alternatives. Collectively, this work delivers a scalable technical roadmap for low-latency, highly extensible long-sequence modeling.
To address gradient explosion/vanishing, training instability, and high computational cost inherent in backpropagation through time (BPTT) for RNNs, this paper proposes a gradient-free end-to-end training paradigm. It fixes random hidden-layer weights and models temporal dynamics via Koopman operator theory, leveraging extended dynamic mode decomposition (EDMD) to analytically solve for output-layer weights. This work is the first to integrate Koopman theory with randomized feature networks, yielding theoretically interpretable and numerically stable RNN constructions. Evaluated on chaotic system forecasting, meteorological modeling, and control tasks, the method achieves significantly faster training, surpasses state-of-the-art gradient-based methods in prediction accuracy, and exhibits strong convergence robustness.
This work addresses the challenge that existing parallelizable sequence models, such as Transformers and state space models, struggle to simultaneously achieve efficient training and strong long-term memory retention. To overcome this limitation, the authors propose the Memory Recurrent Unit (MRU), which introduces multistable dynamics into a parallelizable RNN architecture for the first time. By eliminating transient responses, MRU enables persistent memory storage while leveraging parallel scan algorithms to ensure training efficiency. An instantiation of this framework, termed BMRU, demonstrates superior performance on long-range dependency tasks. Furthermore, BMRU can be effectively integrated with state space models to yield a hybrid architecture that combines both transient and steady-state memory capabilities, achieving high efficiency without compromising representational power.
This work addresses the fragmentation in current linear recurrent neural network (LRNN) research, which is hindered by scattered implementations, strong framework dependencies, and a lack of standardized interfaces—impeding reproducibility, comparison, and extension. To bridge this gap, we introduce lrnnx, the first unified open-source library that integrates multiple modern LRNN architectures. Built on mainstream deep learning frameworks, lrnnx supports diverse parameterizations and discretization schemes without requiring custom CUDA kernels. The library offers a multi-level modular interface, ranging from low-level components to high-level models, substantially lowering the barrier to both usage and development. This design enables flexible control, facilitates fair benchmarking, and establishes a much-needed standardized toolkit for the LRNN community.
This work addresses the challenge of maintaining stable and coherent internal representations in long-sequence modeling, particularly under out-of-distribution (OOD) generalization settings where existing methods often underperform. To this end, the authors propose an alternating fast-slow recurrent mechanism that interleaves slow observational updates with rapid, self-organizing latent state updates. This design enables the model to internally “reason” while processing inputs, dynamically constructing clustered yet temporally coherent long-range representations. By integrating self-organizing representation learning with sequential modeling, the approach significantly outperforms established baselines—including LSTMs, state space models, and Transformers—in both reinforcement learning and algorithmic tasks, demonstrating markedly improved OOD generalization over extended time horizons.
This work addresses the question of why linear recurrent memory units are effective in partially observable reinforcement learning by proposing two classes of linear filters that offer theoretical justification. The first class exactly reproduces the pre-softmax logits of the belief state in a hidden Markov model (HMM), while the second achieves near-zero state decoding error under nearly deterministic transitions. This study establishes the first theoretical foundation for the empirical success of linear recurrent memory in this setting, proving that such representations constitute sufficient statistics for learning optimal policies. The framework is further extended to action-conditioned HMMs. Empirical evaluations confirm the efficacy of the proposed filters and demonstrate their strong feature extraction capabilities in small-scale reinforcement learning environments.
This work addresses the fundamental limitation of conventional neural networks, which rely on repeated data access and epoch-based training, thereby losing long-term consistency in irreversible data streams and degenerating into reactive filters devoid of temporal memory. To overcome this, the paper introduces Streaming Neural Networks (StNN), equipped with a stream-native execution algorithm (SNA) and streaming neurons that maintain persistent temporal states, enabling continuous, epoch-free learning. The study establishes the first minimal computational foundation for neural processing of irreversible streaming data, theoretically proving that the network’s state dynamics exhibit boundedness and contractivity. Through phase-space analysis and temporal modeling, the authors further demonstrate the stability and effectiveness of StNN in long-horizon stream processing tasks.