LSTM Variants for Chaotic Dynamical Systems: An Empirical Study on the Lorenz Attractor

📅 2026-06-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of long-term prediction in chaotic dynamical systems—such as the Lorenz attractor—where minute errors grow exponentially, rendering forecasts unreliable. The authors systematically evaluate seven recurrent and convolutional architectures, including LSTM, BiLSTM, TCN, and their variants, under unified preprocessing, sequence length, and rollback configurations. Their findings reveal that bidirectional contextual modeling combined with Huber robust loss significantly enhances prediction stability, whereas attention mechanisms and CNN-based frontends degrade performance. Notably, a BiLSTM trained with Huber loss achieves the best results, scoring between 45.72 and 58.81 on the AI-DEEDS 2026 Challenge and substantially outperforming competing methods on difficult test pairs.
📝 Abstract
Forecasting chaotic dynamical systems such as the Lorenz attractor is notoriously difficult: small numerical errors are amplified exponentially over long autoregressive rollouts. We study seven recurrent and convolutional architectures for the AI-DEEDS 2026 Chaotic Systems Challenge: a vanilla LSTM, an LSTM with additive attention, a Bidirectional LSTM (BiLSTM), a BiLSTM trained with the Huber loss, a Temporal Convolutional Network (TCN), a CNN front-end followed by an LSTM, and a CNN front-end followed by a BiLSTM. All models share the same pre-processing, sequence length, and rollout procedure, isolating the contribution of each design choice. The challenge scores predictions on a 0-100 scale where higher is better. We obtain leaderboard scores between 45.72 and 58.81, with the BiLSTM trained with Huber loss being the strongest configuration. Two findings stand out: (i) adding additive attention to the unidirectional baseline degraded performance by over ten points, and (ii) prepending a CNN front-end to either an LSTM or a BiLSTM did not help and slightly hurt the score. Per-pair RMSE measurements confirm that the BiLSTM family generalizes better in the harder pairs (6-7), while the LSTM + Attention model collapses there (RMSE up to 8.94 on pair 6). We discuss why bidirectional context and a robust loss help in chaotic regimes while attention and CNN front-ends fail in this setting.
Problem

Research questions and friction points this paper is trying to address.

chaotic dynamical systems
Lorenz attractor
long-term forecasting
error amplification
recurrent neural networks
Innovation

Methods, ideas, or system contributions that make the work stand out.

chaotic dynamical systems
Bidirectional LSTM
Huber loss
attention mechanism
empirical evaluation
R
Ruslan Gokhman
Yeshiva University