Score
Designs and implements gradient computation through time for models with temporal recurrence by unrolling computations across time steps and backpropagating errors to obtain time-accumulated parameter gradients. This includes choosing and implementing full or truncated backpropagation through time, managing memory and compute trade-offs, and applying techniques to stabilize recurrent gradients (for example gradient clipping or other regularization).
Recurrent Neural Networks (RNNs) and their variants—despite shared sequential modeling objectives—exhibit significant heterogeneity in architecture, objective functions, and learning algorithms, leading to conceptual ambiguity and fragmented understanding. Method: This survey introduces the first unified taxonomy grounded in three orthogonal dimensions: network architecture, training objective, and optimization algorithm. It systematically categorizes mainstream models—including LSTMs, convolutional recurrent networks, graph/tree-structured RNNs, higher-order RNNs, and memory-augmented architectures—while analyzing their interdependencies and evolutionary trajectories. Contribution/Results: The work establishes a generalizable modeling paradigm for complex sequence, speech, and image tasks; clarifies the fundamental distinctions and synergies between recursive and recurrent paradigms; synthesizes representative applications in NLP, automatic speech recognition, and image understanding; and identifies emerging research directions—including differentiable neural architecture search, dynamic computation graphs, and neuro-symbolic integration.
Standard backpropagation through time (BPTT) and its truncated or higher-order approximations suffer from significant gradient bias and unstable convergence in RNNs due to long-range dependencies. To address this, we propose an exact backward propagation method grounded in discrete forward sensitivity equations (DFSE). This is the first work to integrate DFSE into RNN training, enabling unbiased, full-sequence gradient computation while natively supporting time-varying parameters and multi-cycle coupled architectures. By performing precise Jacobian chain propagation, our method eliminates truncation errors and avoids cumulative bias from higher-order approximations. Experiments on long-sequence tasks demonstrate substantial improvements in gradient accuracy and training stability. Our approach establishes a new paradigm for modeling strong long-term dependencies in recurrent systems.
How to achieve efficient credit assignment under spatiotemporal locality constraints in physical neural networks remains a fundamental challenge in neuromorphic computing. This paper introduces the Generalized Latent Equilibrium (GLE) framework, which for the first time couples energy minimization with neuron-level local mismatch dynamics to derive biologically plausible forward and backward continuous-time dynamics. By incorporating dendritic morphology modeling and membrane potential phase modulation, GLE implements spatiotemporal convolution and temporal reversal of feedback signals. Crucially, GLE relies exclusively on local synaptic plasticity—requiring no global timing coordination or external error broadcasting. Experiments demonstrate that, under strict locality constraints, GLE approximates the performance of backpropagation through time (BPTT), enables real-time online learning, incurs minimal memory overhead, and provides an interpretable, biologically realistic credit assignment mechanism for deep cortical networks.
Neural ODE training suffers from high computational cost, excessive memory consumption, and numerical instability during backpropagation. To address these challenges, this paper introduces the Algebraically Invertible ODE Solver family, grounded in algebraically invertible numerical integration. The method integrates high-order implicit/explicit reversible schemes, adjoint-state techniques, and memory–computation co-optimization to achieve, for the first time, high-order accuracy, strict numerical stability, and exact gradient computation in backpropagation. Unlike recursive checkpointing, our approach achieves strictly superior time and memory complexity bounds. Extensive evaluation on multiple benchmark ODE tasks demonstrates a 2.1× reduction in training latency and a 68% decrease in GPU memory usage, while preserving gradient precision and numerical robustness.
This work addresses the challenge of exactly implementing backpropagation within a physically realizable continuous-time dynamical system in finite time, circumventing the conventional reliance of energy-based models on symmetric weights or asymptotic convergence. By modeling feedforward inference as a continuous process, the authors introduce a non-conservative Lagrangian framework and construct a two-state energy functional encompassing both activations and sensitivities. The saddle-point dynamics of this functional enable simultaneous inference and credit assignment. Crucially, the study provides the first rigorous proof that standard backpropagation can be precisely replicated by a physical relaxation process in at most 2L steps for an L-layer network—without requiring weight symmetry, infinitesimal perturbations, or asymptotic assumptions—thereby enabling exact, finite-time gradient computation and offering a theoretical foundation for brain-inspired and analog hardware implementations.
Data correlations induce non-orthogonality in the parameter space after linear transformations across neural network layers, severely degrading the efficiency and stability of gradient descent optimization. This work identifies the covariance structure of intra-layer neuronal responses as a fundamental bottleneck to training performance and establishes— for the first time—a rigorous equivalence between layer-wise dynamic decorrelation across the entire network and the natural gradient optimization objective. Building on this insight, we propose a lightweight intra-layer response decorrelation algorithm based on online whitening and adaptive covariance correction, fully compatible with distributed training and neuromorphic hardware. Experiments demonstrate that our method significantly accelerates convergence of standard backpropagation; more critically, it restores high accuracy and robust convergence to multiple approximate backpropagation algorithms previously rendered infeasible due to accuracy collapse. This enables efficient, low-power deep learning training on brain-inspired hardware, establishing a novel paradigm for neuromorphic AI.
This study investigates the learning dynamics of local gradient approximation algorithms—such as Random Feedback Local Online (RFLO) and truncated Backpropagation Through Time (tBPTT)—in brain-inspired recurrent neural networks constrained by spatiotemporal locality, and elucidates their fundamental differences from standard BPTT. Leveraging dynamical systems theory, orthogonal mode decomposition, and linear RNN modeling, the work provides the first dynamical-systems-based characterization of the steady-state solutions, stability, and convergence behavior of such local learning rules. It reveals that RFLO solutions are confined to low-rank perturbations of the initial parameters, exhibiting an intrinsic low-rank update structure, thereby exposing a fundamental limitation imposed by locality constraints on the network’s representational capacity. These findings establish a new theoretical foundation for neuromorphic computing and biologically plausible plasticity models.
This work investigates the implicit acceleration phenomenon of gradient descent (GD) in training two-layer neural networks. While GD on linear models suffers from an Ω(d) iteration complexity lower bound, we establish, for the first time, a rigorous equivalence between GD and the generalized perceptron algorithm under logistic loss—thereby mapping the nonlinear optimization dynamics to geometrically tractable perceptron updates. Leveraging classical linear algebra and theoretical analysis, complemented by numerical experiments, we prove this equivalence yields a √d-speedup: under minimal realistic assumptions, the iteration complexity improves from Ω(d) to Õ(√d). Our result provides the first analytical explanation for rapid convergence in neural network training and reveals that nonlinearity itself inherently encodes optimization acceleration—bypassing fundamental limitations of linear model theory.