Score
Design and analyze time-evolving learning processes and their limiting behavior by producing mathematical descriptions in which the model (primal) variables are represented as time-dependent loss minimizers and hypothesis sets are characterized via subdifferentials. Build and analyze dynamics expressed as mirror-flow or differential-inclusion style evolutions that capture staged or incremental data updates, relate domain support functions to progression, and prove convergence or qualitative properties of the incremental trajectory.
This work investigates the mechanism by which mirror flows achieve incremental learning in dynamically evolving hypothesis spaces. Focusing on mirror flows driven by convex quadratic losses and general convex, lower semicontinuous mirror potentials, the study shows that when initialized near the boundary of the potential’s domain, the rescaled trajectory converges to a limiting mirror flow whose potential is the indicator function of the domain. In this regime, the primal variable continuously minimizes the loss over a time-varying hypothesis set induced by the dual variable. By leveraging tools from convex analysis, subdifferential calculus, and mirror descent dynamics, the authors establish a theoretical framework linking mirror flows to incremental learning dynamics. This reveals a geometric optimization mechanism—activated under specific initialization—that naturally enables continual learning, offering a novel theoretical perspective on incremental learning systems.
该文通过统一的数学视角整合了连续时间机器学习的主要分支,探讨了其数学关系、设计权衡及训练算法,并指出了未来研究方向。
This paper addresses the problem of next-state prediction for model-free dynamical systems—i.e., predicting future states without knowledge of the underlying evolution function and without parametric assumptions. Methodologically, it introduces novel combinatorial metrics and dimensions to characterize the fundamental statistical complexity of model-free time-series prediction for the first time. Leveraging tools from combinatorial learning theory, dynamical systems analysis, and online learning regret analysis, the paper rigorously derives tight optimal mistake bounds (in the realizable setting) and regret bounds (in the agnostic setting). The results establish the first systematic theoretical foundation for model-free time-series prediction, precisely quantifying its intrinsic statistical difficulty and delineating fundamental algorithmic limits.
This work proposes a sparse and unified theoretical framework to systematically uncover the core mechanisms underlying learning, optimization, and modeling. It conceptualizes learning as a multi-level process arising from the coupling of problem formulation, method selection, and optimization dynamics. By precisely defining “solvable problems” and “parameterized methods,” the framework reduces complex learning theory to a few fundamental concepts rooted in dynamical systems, differential geometry, and foundational physics. The approach yields a general convergence theorem and establishes a universal theoretical foundation for cross-domain modeling and algorithm design, substantially enhancing both the parsimony and explanatory power of learning theory.
To address the coupled generalization-forgetting error problem arising from catastrophic forgetting in continual learning, this paper pioneers modeling the temporal evolution of the loss function as a parabolic partial differential equation (PDE), with the memory buffer serving as a dynamic boundary condition—explicitly capturing long-range dependencies and error propagation. Leveraging the intrinsic physical regularity of PDEs, we formulate a spatiotemporal constrained optimization framework driven by boundary conditions, enabling analyzable and interpretable dynamic regularization. Theoretically, we derive a tight coupled bound on forgetting and generalization errors. Empirically, our method significantly reduces forgetting across multiple standard benchmarks, and the theoretical error bound closely aligns with observed performance—demonstrating the effectiveness, stability, and analytical tractability of PDE-based regularization.
This work addresses the lack of a unified theoretical framework for reinforcement learning, which has hindered systematic analysis of its convergence, sample complexity, and generalization. Building upon Markov decision processes and Bellman operators, the paper introduces a cohesive analytical framework that integrates tools from operator theory, stochastic approximation, convex duality, and function approximation. This framework encompasses a broad range of algorithms, including value iteration, policy iteration, temporal difference methods, off-policy learning, and constrained MDPs. By leveraging contraction mappings, monotone operators, martingale techniques, mirror/proximal optimization, concentration inequalities, and mixing process theory, the study establishes finite-sample performance bounds and asymptotic convergence guarantees for diverse reinforcement learning algorithms, thereby forging a rigorous theoretical bridge between probability theory, optimization, and statistics.
This study investigates how the geometric structure of decision boundaries in deep classifiers evolves with layer depth and the dynamical mechanisms governing this process. By modeling feedforward networks as non-autonomous discrete dynamical systems, we employ finite-time maximum Lyapunov exponents (FTMLE) to analyze data trajectories, revealing the dynamical characteristics of decision boundaries across probability, logit, and hidden layers. This work establishes the first mathematical connection between the Lyapunov spectrum and decision boundaries, proving that probability-level FTMLE encodes normal geometric information of the boundary. Building on this insight, we propose a geometry-aware fine-tuning strategy that reconstructs sensitivity distributions within hidden layers, thereby providing a theoretical foundation for layer-aware regularization.
This work investigates the phenomenon of memory in generative model training—where models persistently output similar samples—and elucidates its underlying mechanism through the lens of dynamical systems theory. By integrating the two-timescale dynamics of stochastic gradient descent (SGD) with structural properties of the loss landscape, the study offers the first unified explanation linking memory effects, double descent, and mode collapse, emphasizing the pivotal role of training dynamics themselves. Building upon Austin’s (2016) loss modeling, Borkar’s (2025a, 2026) theories of collapse and double descent, and recent advances by Azizian et al. (2024) on constant-stepsize SGD, the authors construct a dynamical framework that reveals the fundamental causes of output stagnation during training.
本文提出了一种从数据中学习控制系统的线性算子的结构化方法,利用(半)群框架和反问题框架分析算法,以获得收敛估计器。
This work addresses the lack of a unified mathematical characterization of fast-slow coupling dynamics in population-based neural network training. We model the neural network population as a two-timescale interacting agent system, where parameters evolve via fast stochastic gradient updates and hyperparameters evolve through a slow selection-mutation mechanism. For the first time, we rigorously derive the evolution equation governing the joint distribution in the large-population limit and, under strong timescale separation, obtain a selection-mutation equation for the hyperparameter density. This reveals its intrinsic connection to Boltzmann–Gibbs measures and an effective fitness function. Our theoretical analysis bridges population learning, bilevel optimization, and replicator-mutator models, elucidating the roles of noise and diversity in the exploration–exploitation trade-off. Experiments confirm the validity of the reduced dynamics and demonstrate that leveraging the effective fitness significantly enhances hyperparameter optimization performance.