Score
Designs, implements, and trains bidirectional LSTM (BiLSTM) recurrent neural network models and their training pipelines to process sequential data by combining forward and backward LSTM passes. This includes architecture selection, state handling, loss and optimizer configuration, and evaluation practices to capture past and future context, improve long-rollout stability, and reduce prediction error.
Recurrent Neural Networks (RNNs) and their variants—despite shared sequential modeling objectives—exhibit significant heterogeneity in architecture, objective functions, and learning algorithms, leading to conceptual ambiguity and fragmented understanding. Method: This survey introduces the first unified taxonomy grounded in three orthogonal dimensions: network architecture, training objective, and optimization algorithm. It systematically categorizes mainstream models—including LSTMs, convolutional recurrent networks, graph/tree-structured RNNs, higher-order RNNs, and memory-augmented architectures—while analyzing their interdependencies and evolutionary trajectories. Contribution/Results: The work establishes a generalizable modeling paradigm for complex sequence, speech, and image tasks; clarifies the fundamental distinctions and synergies between recursive and recurrent paradigms; synthesizes representative applications in NLP, automatic speech recognition, and image understanding; and identifies emerging research directions—including differentiable neural architecture search, dynamic computation graphs, and neuro-symbolic integration.
This study addresses the challenge of long-term prediction in chaotic dynamical systems—such as the Lorenz attractor—where minute errors grow exponentially, rendering forecasts unreliable. The authors systematically evaluate seven recurrent and convolutional architectures, including LSTM, BiLSTM, TCN, and their variants, under unified preprocessing, sequence length, and rollback configurations. Their findings reveal that bidirectional contextual modeling combined with Huber robust loss significantly enhances prediction stability, whereas attention mechanisms and CNN-based frontends degrade performance. Notably, a BiLSTM trained with Huber loss achieves the best results, scoring between 45.72 and 58.81 on the AI-DEEDS 2026 Challenge and substantially outperforming competing methods on difficult test pairs.
To address the limitation of sLSTM—its poor short-term memory retention, which hinders direct application to long-horizon time series forecasting (TSF)—this paper proposes P-sLSTM, the first model to synergistically integrate sequence patching and channel independence. P-sLSTM enhances local pattern modeling via patch-based encoding, while leveraging channel-wise decoupling, exponential gating, and memory mixing to mitigate short-term memory degradation without compromising sLSTM’s inherent capacity for long-range dependency capture. Theoretically interpretable, structurally lightweight, and parameter-efficient, P-sLSTM achieves state-of-the-art performance across multiple standard long-horizon TSF benchmarks, reducing average prediction error by 12.7% compared to prior methods. Moreover, it outperforms mainstream Transformer- and RNN-based variants in inference speed, demonstrating both accuracy and efficiency advantages.
To address BiLSTM’s limitation in capturing global sentence-level semantics for sequence labeling, this paper observes that its initial and final hidden states naturally encode holistic sentence representations. Leveraging this insight, we propose a lightweight, plug-and-play global context gating mechanism that dynamically injects sentence-level information into token-level representations at each time step. The method requires no architectural modification to the underlying RNN and is fully compatible with pretrained embeddings (e.g., BERT). Compared to various RNN variants, it achieves faster inference and training, and greater integration flexibility. Extensive experiments across nine benchmark datasets—including NER, POS tagging, and end-to-end aspect-based sentiment analysis—demonstrate consistent improvements: average F1-score and accuracy gains of 1.2–2.8 percentage points. These results validate the effectiveness and generalizability of synergistic global–local modeling for sequence labeling.
Identifying flight phases from unstructured aviation accident narratives remains challenging due to implicit temporal and operational cues. Method: This study proposes a deep learning–based automated classification framework using LSTM, GRU, and BiLSTM architectures—along with their pairwise hybrid variants (GRU-LSTM, LSTM-BiLSTM, GRU-BiLSTM)—integrated with text preprocessing and sequence modeling techniques. Contribution/Results: It presents the first systematic comparison of standalone versus hybrid RNN models for this task: BiLSTM achieves 64% accuracy, while the optimal hybrid model (LSTM-BiLSTM) reaches 67%, significantly outperforming baseline methods. Results demonstrate that deep learning can effectively infer latent flight phases from accident reports; moreover, synergistic integration of bidirectional context and gating mechanisms enhances representational capacity. This work establishes a reproducible NLP methodology and provides empirical evidence supporting its application in aviation safety analysis.
Standard backpropagation through time (BPTT) and its truncated or higher-order approximations suffer from significant gradient bias and unstable convergence in RNNs due to long-range dependencies. To address this, we propose an exact backward propagation method grounded in discrete forward sensitivity equations (DFSE). This is the first work to integrate DFSE into RNN training, enabling unbiased, full-sequence gradient computation while natively supporting time-varying parameters and multi-cycle coupled architectures. By performing precise Jacobian chain propagation, our method eliminates truncation errors and avoids cumulative bias from higher-order approximations. Experiments on long-sequence tasks demonstrate substantial improvements in gradient accuracy and training stability. Our approach establishes a new paradigm for modeling strong long-term dependencies in recurrent systems.
This work systematically investigates the joint sensitivity of BiLSTM-based time series forecasting models to input sequence length and additive noise, elucidating their impact on model robustness and generalization. Through controlled experiments on multi-frequency real-world datasets, we establish a standardized, modular, and reproducible evaluation pipeline. Our key findings are: (1) excessive sequence length induces overfitting and data leakage; (2) additive noise consistently degrades prediction accuracy; and (3) the combined effect causes the most severe performance degradation—even high-frequency data, though comparatively more robust, remains vulnerable. The main contributions are: (1) empirical identification of the synergistic degradation effect between sequence length and noise; (2) proposal of a “data-aware modeling” paradigm, advocating pre-modeling analysis of data characteristics—especially under low-sample and high-noise regimes—to guide architecture design; and (3) open-sourcing an extensible evaluation framework to facilitate future research on robust time series modeling.
This study addresses the challenge of insufficient accuracy and robustness in remaining useful life (RUL) prediction for aircraft engines under varying operating conditions and noisy measurements. To this end, the authors propose a bidirectional residual-corrected LSTM (Bi-cLSTM) model that integrates bidirectional temporal modeling with an adaptive residual correction mechanism. The approach further incorporates a condition-aware preprocessing pipeline—comprising operating-condition segmentation-based normalization, feature selection, and exponential smoothing—to significantly enhance generalization and noise resilience across complex, multi-regime scenarios. Experimental results on all four NASA C-MAPSS datasets demonstrate that the proposed method outperforms existing LSTM-based baselines, achieving state-of-the-art performance, with particularly pronounced gains in multi-condition settings.
This work addresses the challenges of long-sequence modeling, where Transformers suffer from quadratic computational complexity and conventional RNNs are hindered by sequential computation that impedes parallelization. The authors propose PR-LSTM, a hierarchical recurrent architecture that transforms the nonlinear recurrence of hidden states into a parallelizable tree reduction process via a balanced computation tree. By uniquely integrating nonlinear gated state updates with logarithmic-depth parallelism, PR-LSTM overcomes the inherent serial bottleneck of RNNs while avoiding the high computational cost of attention mechanisms. Efficient hierarchical state composition is achieved through a fixed schedule based on parallel scan and a fused recursive gating module. Experiments demonstrate that PR-LSTM significantly outperforms standard RNNs, LSTMs, and Transformers on formal language tasks and exhibits strong length generalization capabilities.
Standard backpropagation through time (BPTT) incurs substantial computational costs in long-sequence tasks, while truncated BPTT improves efficiency at the expense of performance degradation. This work establishes, for the first time, a theoretical error bound characterizing the performance loss induced by truncated BPTT, revealing that the burn-in phase is a critical hyperparameter governing model performance. Building on this insight, the authors propose a targeted tuning strategy for the burn-in duration. Experiments on system identification and time series prediction benchmarks demonstrate that appropriately configuring the burn-in phase can reduce both training and test prediction errors by over 60%, offering a novel perspective for achieving both efficiency and high performance in recurrent neural network training.
Addressing the challenge of simultaneously achieving structural interpretability and data adaptability in S&P 500 volatility forecasting, this paper proposes the SV-LSTM hybrid model—the first end-to-end jointly trained framework integrating stochastic volatility (SV) latent-variable modeling with long short-term memory (LSTM) networks. The model synergistically combines the statistical rigor of SV models with LSTM’s capacity to capture nonlinear temporal dynamics, enhanced by rolling-window training and Monte Carlo likelihood estimation for robustness. Empirical evaluation over 1998–2024 demonstrates that SV-LSTM reduces mean squared error (MSE) by 23.7% relative to standalone SV or LSTM baselines and achieves a 96.4% pass rate in Value-at-Risk (VaR) backtesting. These results indicate substantially improved stability in extreme-event volatility forecasting, offering a theoretically grounded and empirically effective tool for risk management and dynamic asset allocation.