lstm sequence modeling

Designs and trains LSTM-based recurrent neural network models to emulate, predict, or reconstruct sequences and time-series; maps static or time-varying input parameters to multi-step sequence outputs (for example, time-evolving summary observables) and evaluates temporal accuracy. Builds sequence emulators and recurrent predictors that produce high-fidelity multi-step forecasts or summary traces, covering LSTM modeling, sequence emulation, and time-series prediction tasks.

lstmsequencemodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.13
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$199K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

This work addresses the lack of systematic benchmarks for deep learning model selection in real-world scenarios. We propose the first unified experimental framework to conduct a horizontal evaluation of six model architectures—CNN, Simple RNN, LSTM, Bi-LSTM, GRU, and Bi-GRU—across three modalities: image classification (Fruit-360), time-series classification (ARAS), and text sentiment analysis (IMDB). Implemented in PyTorch, the evaluation employs multiple metrics: accuracy, F1-score, and training efficiency. Results reveal that Bi-LSTM achieves a 4.2% average accuracy gain on time-series tasks; CNN attains 98.7% accuracy on Fruit-360; and GRU matches LSTM’s performance with ~20% fewer parameters. The study uncovers principled correspondences between architectural characteristics (e.g., gating mechanisms, bidirectionality) and data modality properties, yielding reproducible empirical guidelines and theoretical insights for informed model selection.

Comparative analysis of deep learning models like CNN, RNN, LSTM, GRU.Evaluation of model performance using IMDB, ARAS, and Fruit-360 datasets.Exploration of benefits and limitations of various deep learning architectures.

Must-Read Papers

Most classic and influential ideas
View more

Unlocking the Power of LSTM for Long Term Time Series Forecasting

Aug 19, 2024
YK
Yaxuan Kong
🏛️ University of Oxford | University of Pennsylvania | Princeton University | Alibaba Group | The Hong Kong University of Science and Technology | Duke Kunshan University | Squirrel AI

To address the limitation of sLSTM—its poor short-term memory retention, which hinders direct application to long-horizon time series forecasting (TSF)—this paper proposes P-sLSTM, the first model to synergistically integrate sequence patching and channel independence. P-sLSTM enhances local pattern modeling via patch-based encoding, while leveraging channel-wise decoupling, exponential gating, and memory mixing to mitigate short-term memory degradation without compromising sLSTM’s inherent capacity for long-range dependency capture. Theoretically interpretable, structurally lightweight, and parameter-efficient, P-sLSTM achieves state-of-the-art performance across multiple standard long-horizon TSF benchmarks, reducing average prediction error by 12.7% compared to prior methods. Moreover, it outperforms mainstream Transformer- and RNN-based variants in inference speed, demonstrating both accuracy and efficiency advantages.

Address short memory in sLSTMEnhance sLSTM for long-term TSFPropose P-sLSTM with patching method

This paper addresses poor model reproducibility and insufficient open-source implementations in time-series forecasting by proposing a lightweight, fully reproducible LSTM/GRU modeling paradigm. Methodologically, it constructs univariate sequence samples via sliding windows and evaluates performance using two metrics—RMSE and directional accuracy (DA)—on both synthetic activity data (Activities) and real-world financial data (BSE BANKEX). A key finding is that effective training requires only a single time series exhibiting repetitive patterns, without complex preprocessing or large-scale datasets. Experiments show that the proposed implementation significantly outperforms the “repeat last value” baseline for 1-step and 20-step predictions on Activities, while achieving comparable performance on BSE BANKEX. All code, datasets, and complete experimental configurations are publicly released to ensure full reproducibility and out-of-the-box usability.

Comparing forecasting accuracy against a simple baseline modelEvaluating performance on financial and synthetic activity datasetsImplementing open-source LSTM and GRU for time series forecasting

To address the limitations of RNNs and LSTMs in multistep time-series forecasting—including insufficient capacity to model complex nonlinear patterns, poor interpretability, and low computational efficiency—this paper proposes the TKAN architecture. Its core innovation is the Recurrent Kolmogorov–Arnold Network (RKAN) layer, which integrates learnable spline-based activations from Kolmogorov–Arnold Networks (KANs) into a gated recurrent structure for the first time. This design unifies dynamic weight evolution with long-term dependency modeling while ensuring end-to-end differentiability and strong universal approximation capability grounded in the Kolmogorov–Arnold representation theorem. Extensive experiments demonstrate that TKAN achieves significantly higher prediction accuracy and faster inference speed than state-of-the-art MLP, RNN, and LSTM baselines, without sacrificing model interpretability.

Address limitations in complex sequential pattern handlingCombine strengths of KAN and LSTM networksEnhance multi-step time series forecasting accuracy

State-Space Modeling in Long Sequence Processing: A Survey on Recurrence in the Transformer Era

Jun 13, 2024
MT
Matteo Tiezzi
🏛️ IIT | University of Siena | IMT

Long-sequence modeling faces fundamental challenges including limited context length, difficulty in capturing long-range dependencies, and low efficiency in online learning. To address these, this work systematically reviews the resurgence of state-space models (SSMs) and recurrent computation, proposing a novel local forward-computation paradigm tailored for real-world online learning—thereby circumventing the temporal backtracking constraints inherent in standard backpropagation through time (BPTT). We introduce the first unified taxonomy encompassing both deep SSMs and large-context Transformers. Our framework integrates structured linear attention, enhanced RNN architectures, local recurrence mechanisms, and online optimization algorithms. The study rigorously clarifies the theoretical representational advantages and practical sequential reasoning benefits of recurrent modeling over alternatives. Collectively, this work delivers a scalable technical roadmap for low-latency, highly extensible long-sequence modeling.

Addressing limitations of Transformers with state-space modelsExploring efficient online learning beyond backpropagation through timeSurveying recurrent models for long sequence processing

The Power of Linear Recurrent Neural Networks.

Feb 09, 2018
FS
Frieder Stolzenburg
🏛️ Harz University of Applied Sciences | Western Sydney University

This paper addresses the issues of model redundancy and poor interpretability in Linear Recurrent Neural Networks (LRNNs) for time-series approximation. We propose Predictive Linear Recurrent Neural Networks (PLRNNs), which jointly optimize network weights and sparse architecture via least-squares solving of linear systems—enabling, for the first time, integrated parameter learning and structural pruning. PLRNNs exhibit elliptical dynamical trajectories, ensuring both interpretability and compact functional representation. Structural pruning is guided by principal component analysis, significantly reducing model size. On the Multi-Sinusoidal Oscillator (MSO) benchmark, PLRNN achieves state-of-the-art performance with the fewest neurons. Furthermore, it demonstrates strong generalization on real-world tasks: robotic soccer motion modeling and stock price prediction.

Approximating time-dependent functions with linear recurrent networksLearning network architecture through eigenvalue spectrum analysisPredicting time-series values with compact linear representations

Latest Papers

What's happening recently
View more

This study addresses the lack of systematic evaluation comparing data-driven models, such as Long Short-Term Memory (LSTM) networks, with model-based approaches in structured time series classification tasks. The authors construct a controlled evaluation framework using Monte Carlo simulations to compare LSTM against an Expectation-Maximization (EM)-based classifier within linear Gaussian state-space models, across varying task difficulties, sequence lengths, and training set sizes. The theoretical performance upper bound is established by the likelihood ratio test derived from Kalman filter outputs under known model parameters. For the first time in a structured setting, this work quantifies the performance gap between these paradigms, revealing that LSTM exhibits performance saturation when classification relies solely on differences in measurement noise, failing to surpass the theoretical optimum even with increased data or longer sequences. In contrast, the EM-based classifier consistently approaches the upper bound when model assumptions hold.

deep learningLSTMmodel-based methods

This work addresses the challenge of transfer learning for modeling physical systems across disparate temporal scales under data-scarce conditions. The authors propose a time-warping–based recurrent neural network (RNN) transfer approach that adapts a pretrained long short-term memory (LSTM) model to target systems with different dynamic time scales by rescaling the time axis. Theoretically, they demonstrate that LSTMs can accurately approximate a class of delay differential equations and retain this approximation capability under time-warping transformations. By incorporating time warping into RNN-based transfer learning—a novel contribution—the method achieves high-accuracy predictions across time spans ranging from one hour to one thousand hours in fuel moisture content forecasting, matching or surpassing state-of-the-art transfer techniques while fine-tuning only a small subset of parameters.

dynamical systemsfuel moisture contentrecurrent neural networks

This work addresses the challenges of long-sequence modeling, where Transformers suffer from quadratic computational complexity and conventional RNNs are hindered by sequential computation that impedes parallelization. The authors propose PR-LSTM, a hierarchical recurrent architecture that transforms the nonlinear recurrence of hidden states into a parallelizable tree reduction process via a balanced computation tree. By uniquely integrating nonlinear gated state updates with logarithmic-depth parallelism, PR-LSTM overcomes the inherent serial bottleneck of RNNs while avoiding the high computational cost of attention mechanisms. Efficient hierarchical state composition is achieved through a fixed schedule based on parallel scan and a fused recursive gating module. Experiments demonstrate that PR-LSTM significantly outperforms standard RNNs, LSTMs, and Transformers on formal language tasks and exhibits strong length generalization capabilities.

long-context efficiencyparallelismquadratic complexity

This study addresses the challenge of long-term prediction in chaotic dynamical systems—such as the Lorenz attractor—where minute errors grow exponentially, rendering forecasts unreliable. The authors systematically evaluate seven recurrent and convolutional architectures, including LSTM, BiLSTM, TCN, and their variants, under unified preprocessing, sequence length, and rollback configurations. Their findings reveal that bidirectional contextual modeling combined with Huber robust loss significantly enhances prediction stability, whereas attention mechanisms and CNN-based frontends degrade performance. Notably, a BiLSTM trained with Huber loss achieves the best results, scoring between 45.72 and 58.81 on the AI-DEEDS 2026 Challenge and substantially outperforming competing methods on difficult test pairs.

chaotic dynamical systemserror amplificationlong-term forecasting

This work addresses the challenge of maintaining stable and coherent internal representations in long-sequence modeling, particularly under out-of-distribution (OOD) generalization settings where existing methods often underperform. To this end, the authors propose an alternating fast-slow recurrent mechanism that interleaves slow observational updates with rapid, self-organizing latent state updates. This design enables the model to internally “reason” while processing inputs, dynamically constructing clustered yet temporally coherent long-range representations. By integrating self-organizing representation learning with sequential modeling, the approach significantly outperforms established baselines—including LSTMs, state space models, and Transformers—in both reinforcement learning and algorithmic tasks, demonstrating markedly improved OOD generalization over extended time horizons.

internal representation stabilitylatent recurrencelong-horizon sequential modeling

Hot Scholars

MS

Mohammad Shojafar

Associate Professor, University of Surrey, EU Marie Curie Alumni, ACM Distinguished Speaker
Network SecurityFog Computing5G/6GFuture Internet
JT

Jason T. L. Wang

Professor of Computer Science, New Jersey Institute of Technology
Data MiningMachine LearningDeep LearningComputational Biology
CM

Catherine M. Elias

German University in Cairo
System ArchitectureCooperative SystemsIntelligent Transportation SystemsConnected and Automated Vehicles (CAVs)
MB

Marc Bernacki

Professor, MINES ParisTech, PSL
Materials science - Multiscale modeling - Numerical Metallurgy - Computational Mechanics - FEM - HPC
TR

Taufiq Rahman

National Research Council Canada
MechatronicsRoboticsConnected & Autonomous Vehicles