Score
Designs and trains LSTM-based recurrent neural network models to emulate, predict, or reconstruct sequences and time-series; maps static or time-varying input parameters to multi-step sequence outputs (for example, time-evolving summary observables) and evaluates temporal accuracy. Builds sequence emulators and recurrent predictors that produce high-fidelity multi-step forecasts or summary traces, covering LSTM modeling, sequence emulation, and time-series prediction tasks.
Recurrent Neural Networks (RNNs) and their variants—despite shared sequential modeling objectives—exhibit significant heterogeneity in architecture, objective functions, and learning algorithms, leading to conceptual ambiguity and fragmented understanding. Method: This survey introduces the first unified taxonomy grounded in three orthogonal dimensions: network architecture, training objective, and optimization algorithm. It systematically categorizes mainstream models—including LSTMs, convolutional recurrent networks, graph/tree-structured RNNs, higher-order RNNs, and memory-augmented architectures—while analyzing their interdependencies and evolutionary trajectories. Contribution/Results: The work establishes a generalizable modeling paradigm for complex sequence, speech, and image tasks; clarifies the fundamental distinctions and synergies between recursive and recurrent paradigms; synthesizes representative applications in NLP, automatic speech recognition, and image understanding; and identifies emerging research directions—including differentiable neural architecture search, dynamic computation graphs, and neuro-symbolic integration.
This work addresses the lack of systematic benchmarks for deep learning model selection in real-world scenarios. We propose the first unified experimental framework to conduct a horizontal evaluation of six model architectures—CNN, Simple RNN, LSTM, Bi-LSTM, GRU, and Bi-GRU—across three modalities: image classification (Fruit-360), time-series classification (ARAS), and text sentiment analysis (IMDB). Implemented in PyTorch, the evaluation employs multiple metrics: accuracy, F1-score, and training efficiency. Results reveal that Bi-LSTM achieves a 4.2% average accuracy gain on time-series tasks; CNN attains 98.7% accuracy on Fruit-360; and GRU matches LSTM’s performance with ~20% fewer parameters. The study uncovers principled correspondences between architectural characteristics (e.g., gating mechanisms, bidirectionality) and data modality properties, yielding reproducible empirical guidelines and theoretical insights for informed model selection.
To address the limitation of sLSTM—its poor short-term memory retention, which hinders direct application to long-horizon time series forecasting (TSF)—this paper proposes P-sLSTM, the first model to synergistically integrate sequence patching and channel independence. P-sLSTM enhances local pattern modeling via patch-based encoding, while leveraging channel-wise decoupling, exponential gating, and memory mixing to mitigate short-term memory degradation without compromising sLSTM’s inherent capacity for long-range dependency capture. Theoretically interpretable, structurally lightweight, and parameter-efficient, P-sLSTM achieves state-of-the-art performance across multiple standard long-horizon TSF benchmarks, reducing average prediction error by 12.7% compared to prior methods. Moreover, it outperforms mainstream Transformer- and RNN-based variants in inference speed, demonstrating both accuracy and efficiency advantages.
This paper addresses poor model reproducibility and insufficient open-source implementations in time-series forecasting by proposing a lightweight, fully reproducible LSTM/GRU modeling paradigm. Methodologically, it constructs univariate sequence samples via sliding windows and evaluates performance using two metrics—RMSE and directional accuracy (DA)—on both synthetic activity data (Activities) and real-world financial data (BSE BANKEX). A key finding is that effective training requires only a single time series exhibiting repetitive patterns, without complex preprocessing or large-scale datasets. Experiments show that the proposed implementation significantly outperforms the “repeat last value” baseline for 1-step and 20-step predictions on Activities, while achieving comparable performance on BSE BANKEX. All code, datasets, and complete experimental configurations are publicly released to ensure full reproducibility and out-of-the-box usability.
To address the limitations of RNNs and LSTMs in multistep time-series forecasting—including insufficient capacity to model complex nonlinear patterns, poor interpretability, and low computational efficiency—this paper proposes the TKAN architecture. Its core innovation is the Recurrent Kolmogorov–Arnold Network (RKAN) layer, which integrates learnable spline-based activations from Kolmogorov–Arnold Networks (KANs) into a gated recurrent structure for the first time. This design unifies dynamic weight evolution with long-term dependency modeling while ensuring end-to-end differentiability and strong universal approximation capability grounded in the Kolmogorov–Arnold representation theorem. Extensive experiments demonstrate that TKAN achieves significantly higher prediction accuracy and faster inference speed than state-of-the-art MLP, RNN, and LSTM baselines, without sacrificing model interpretability.
Long-sequence modeling faces fundamental challenges including limited context length, difficulty in capturing long-range dependencies, and low efficiency in online learning. To address these, this work systematically reviews the resurgence of state-space models (SSMs) and recurrent computation, proposing a novel local forward-computation paradigm tailored for real-world online learning—thereby circumventing the temporal backtracking constraints inherent in standard backpropagation through time (BPTT). We introduce the first unified taxonomy encompassing both deep SSMs and large-context Transformers. Our framework integrates structured linear attention, enhanced RNN architectures, local recurrence mechanisms, and online optimization algorithms. The study rigorously clarifies the theoretical representational advantages and practical sequential reasoning benefits of recurrent modeling over alternatives. Collectively, this work delivers a scalable technical roadmap for low-latency, highly extensible long-sequence modeling.
This paper addresses the issues of model redundancy and poor interpretability in Linear Recurrent Neural Networks (LRNNs) for time-series approximation. We propose Predictive Linear Recurrent Neural Networks (PLRNNs), which jointly optimize network weights and sparse architecture via least-squares solving of linear systems—enabling, for the first time, integrated parameter learning and structural pruning. PLRNNs exhibit elliptical dynamical trajectories, ensuring both interpretability and compact functional representation. Structural pruning is guided by principal component analysis, significantly reducing model size. On the Multi-Sinusoidal Oscillator (MSO) benchmark, PLRNN achieves state-of-the-art performance with the fewest neurons. Furthermore, it demonstrates strong generalization on real-world tasks: robotic soccer motion modeling and stock price prediction.
This study addresses the lack of systematic evaluation comparing data-driven models, such as Long Short-Term Memory (LSTM) networks, with model-based approaches in structured time series classification tasks. The authors construct a controlled evaluation framework using Monte Carlo simulations to compare LSTM against an Expectation-Maximization (EM)-based classifier within linear Gaussian state-space models, across varying task difficulties, sequence lengths, and training set sizes. The theoretical performance upper bound is established by the likelihood ratio test derived from Kalman filter outputs under known model parameters. For the first time in a structured setting, this work quantifies the performance gap between these paradigms, revealing that LSTM exhibits performance saturation when classification relies solely on differences in measurement noise, failing to surpass the theoretical optimum even with increased data or longer sequences. In contrast, the EM-based classifier consistently approaches the upper bound when model assumptions hold.
This work addresses the challenge of transfer learning for modeling physical systems across disparate temporal scales under data-scarce conditions. The authors propose a time-warping–based recurrent neural network (RNN) transfer approach that adapts a pretrained long short-term memory (LSTM) model to target systems with different dynamic time scales by rescaling the time axis. Theoretically, they demonstrate that LSTMs can accurately approximate a class of delay differential equations and retain this approximation capability under time-warping transformations. By incorporating time warping into RNN-based transfer learning—a novel contribution—the method achieves high-accuracy predictions across time spans ranging from one hour to one thousand hours in fuel moisture content forecasting, matching or surpassing state-of-the-art transfer techniques while fine-tuning only a small subset of parameters.
This work addresses the challenges of long-sequence modeling, where Transformers suffer from quadratic computational complexity and conventional RNNs are hindered by sequential computation that impedes parallelization. The authors propose PR-LSTM, a hierarchical recurrent architecture that transforms the nonlinear recurrence of hidden states into a parallelizable tree reduction process via a balanced computation tree. By uniquely integrating nonlinear gated state updates with logarithmic-depth parallelism, PR-LSTM overcomes the inherent serial bottleneck of RNNs while avoiding the high computational cost of attention mechanisms. Efficient hierarchical state composition is achieved through a fixed schedule based on parallel scan and a fused recursive gating module. Experiments demonstrate that PR-LSTM significantly outperforms standard RNNs, LSTMs, and Transformers on formal language tasks and exhibits strong length generalization capabilities.
This study addresses the challenge of long-term prediction in chaotic dynamical systems—such as the Lorenz attractor—where minute errors grow exponentially, rendering forecasts unreliable. The authors systematically evaluate seven recurrent and convolutional architectures, including LSTM, BiLSTM, TCN, and their variants, under unified preprocessing, sequence length, and rollback configurations. Their findings reveal that bidirectional contextual modeling combined with Huber robust loss significantly enhances prediction stability, whereas attention mechanisms and CNN-based frontends degrade performance. Notably, a BiLSTM trained with Huber loss achieves the best results, scoring between 45.72 and 58.81 on the AI-DEEDS 2026 Challenge and substantially outperforming competing methods on difficult test pairs.
This work addresses the challenge of maintaining stable and coherent internal representations in long-sequence modeling, particularly under out-of-distribution (OOD) generalization settings where existing methods often underperform. To this end, the authors propose an alternating fast-slow recurrent mechanism that interleaves slow observational updates with rapid, self-organizing latent state updates. This design enables the model to internally “reason” while processing inputs, dynamically constructing clustered yet temporally coherent long-range representations. By integrating self-organizing representation learning with sequential modeling, the approach significantly outperforms established baselines—including LSTMs, state space models, and Transformers—in both reinforcement learning and algorithmic tasks, demonstrating markedly improved OOD generalization over extended time horizons.