Score
Designs and analyzes learning algorithms that couple a forward model with an adjoint (backward) system to produce purely local update rules that implement gradient descent on a specified loss; these algorithms are constructed and analyzed in both discrete- and continuous-time formulations. The skill also covers proving convergence properties of the coupled adjoint dynamics under coercivity-type conditions.
This work addresses the lack of theoretical guarantees for local convergence in existing physical learning methods—such as Equilibrium Propagation (EP) and Coupled Learning (CL)—when applied to linear circuit training in the small-perturbation limit. The authors propose a unified convergence analysis framework, introducing for the first time a coercivity condition based on the network’s interconnection structure, formulated as a matrix rank condition. Under this condition, they prove exponential decay of the loss function and convergence of parameters to the solution manifold. Using Sard’s theorem, they further show that this coercivity condition holds for almost all target outputs, demonstrating that degeneracies caused by symmetry are non-generic. The analysis also clarifies that both EP and the newly proposed Adjoint Coupled Learning implement exact gradient descent on the natural loss, whereas standard CL incurs third-order correction terms. Theoretical findings are validated through a kite-circuit example.
本文提出了一种从数据中学习控制系统的线性算子的结构化方法,利用(半)群框架和反问题框架分析算法,以获得收敛估计器。
This work investigates the convergence of continuous-time stochastic gradient descent (SGD) for minimizing population expected loss, extending analysis to overparameterized linear deep neural networks. Methodologically, it establishes a continuous-time modeling framework based on stochastic differential equations and integrates Lyapunov stability theory with nonconvex optimization principles. The key contribution is the first general convergence criterion applicable to stochastic dynamical systems—overcoming the limitation of Chatterjee (2022), which applies only to deterministic gradient descent. Theoretically, under mild regularity conditions, SGD trajectories are proven to converge almost surely to the global optimum. For linear deep networks, the paper derives verifiable sufficient conditions for convergence and elucidates the intrinsic mechanism by which noise perturbations preserve stability—thereby providing novel theoretical foundations for optimization in overparameterized models.
This paper addresses the problem of improving dynamical system identification accuracy for a target linear system using auxiliary data from a similar—but non-homologous—linear system, particularly under limited target-data regimes. To tackle the challenge of fusing heterogeneous main and auxiliary datasets with scarce target samples, we propose a weighted least-squares identification method. Our key theoretical contribution is the first derivation of a computable, data-dependent finite-sample error upper bound. We rigorously characterize a fundamental trade-off between noise suppression and model discrepancy, yielding an explicit error bound that provably demonstrates substantial reduction in noise-induced estimation error through auxiliary data. Furthermore, we establish an adaptive weighting criterion that optimizes this trade-off. Extensive simulations confirm that the proposed method reduces identification error by over 30% across representative scenarios, providing a robust, analytically tractable framework for cross-system knowledge transfer in dynamical modeling.
This work addresses the challenge of reward fine-tuning in diffusion models and Boltzmann distribution sampling by formulating generative model optimization as a stochastic optimal control problem governed by stochastic differential equations. Leveraging the Stochastic Maximum Principle (SMP), the paper rigorously derives, for the first time, a general Hamiltonian adjoint matching objective applicable to settings where both drift and diffusion coefficients depend on the control, and establishes its intrinsic connection to the Hamilton–Jacobi–Bellman (HJB) equation. By integrating the adjoint system with a continuous-time successive approximation algorithm, the method recovers the lightweight adjoint loss when the diffusion coefficient is state-independent, validates the necessity of higher-order terms in state-dependent cases, and provides a tractable iterative scheme based on SMP that circumvents intractable martingale terms.
This work addresses the finite-time convergence of stochastic iterative algorithms for fixed-point equations accessible only through a noisy oracle. The authors propose a norm-independent, unified Lyapunov function framework constructed via a generalized Moreau envelope, which integrates Lyapunov stability theory with stochastic approximation analysis. This framework accommodates complex settings such as Markovian noise, seminorm contractive operators, and dissipative operators, yielding sharp non-asymptotic convergence bounds in both high-probability and mean-square senses. As a result, it provides a unified and refined finite-time convergence guarantee for a broad class of algorithms, including stochastic gradient descent, linear stochastic approximation, Q-learning, and temporal difference learning.
This study investigates the training dynamics of coupled learning (CL) and equilibrium propagation (EP) in the continuous-time, small-perturbation limit, revealing a parameter conservation law in physically realizable systems that parallels mass conservation. By leveraging continuous-time dynamical systems analysis and perturbation theory, the work establishes— for the first time—that this conservation law holds universally across a broad range of physical settings. Furthermore, it elucidates how this constraint governs the convergence behavior of learning in linear circuits. The findings not only enhance the reliability of CL and EP training but also provide a rigorous theoretical foundation and practical guidance for achieving efficient and stable learning in neuromorphic hardware implementations.
本文通过算法稳定性建立样本外边界,利用耗散性论点为学习动力系统提供了一种系统理论解释,并通过优化算法依赖的动力增益来认证和比较学习动态的泛化能力。
This work addresses the lack of a clear last-iterate convergence rate for the Follow-the-Leader with Backward Regret and Multiplicative Weight Updates (FLBR-MWU) dynamics by proposing a novel variant inspired by the extragradient method. The proposed algorithm employs distinct learning rates in intermediate and actual update steps, enabling a refined analysis of its last-iterate convergence behavior in zero-sum games. For the first time, this study establishes a geometric convergence rate of $O(c^t)$—with $c < 1$ independent of time—for the duality gap under FLBR-MWU. Leveraging tools from dynamical systems theory, matrix spectral analysis, and numerical experiments, the paper theoretically proves this geometric convergence and demonstrates that the method matches or even surpasses the performance of state-of-the-art algorithms such as Optimistic Gradient Descent Ascent (OGDA).
This work addresses the limitations of traditional iterative methods for solving large-scale systems of equations, including low efficiency, poor robustness, and difficulties in parameter selection. To overcome these challenges, it proposes a "hybrid iteration" paradigm that integrates classical numerical algorithms with machine learning. Building upon conventional iterative schemes such as Newton's method, this approach incorporates deep learning-based optimization strategies to enable adaptive hyperparameter tuning, thereby combining the flexibility of data-driven methods with the theoretical reliability and interpretability of traditional algorithms. Furthermore, this project systematically reviews state-of-the-art methodologies in this domain and identifies key open challenges. Ultimately, it establishes a clear research trajectory for developing efficient and robust solvers for scientific computing.