Score
Deriving and applying probabilistic, non-asymptotic bounds (e.g., martingale or matrix concentration) that control deviations of random quantities to prove uniform guarantees, phase transitions, or finite-sample error bounds in heterogeneous or sparse settings.
This work addresses the absence of finite-sample error bounds and concentration inequalities for nonlinear stochastic approximation algorithms under the Wasserstein-p distance. By coupling the discrete-time iterative process with its Ornstein–Uhlenbeck diffusion limit, the paper establishes the first non-asymptotic distributional convergence rates in Wasserstein distance under general noise conditions—such as martingale differences and ergodic Markov chains. The main contributions include proving that the last iterate converges to a Gaussian distribution at a rate of γₙ^{1/6}, while the Polyak–Ruppert averaged iterate achieves a rate of n^{-1/6}. Moreover, the analysis yields high-probability concentration inequalities that improve upon those derived via classical moment-based methods. The proposed framework applies broadly to canonical algorithms, including linear stochastic approximation and stochastic gradient descent.
This work investigates the non-asymptotic pathwise approximation accuracy of stochastic iterative algorithms—such as SGD and SGLD—to the Ornstein–Uhlenbeck process in the univariate setting. Addressing the lack of quantifiable, path-level error bounds in existing theory, we introduce a novel analytical framework for path space by integrating infinite-dimensional Stein’s method with exchangeable pair techniques. This yields explicit convergence rates under both the Lévy–Prokhorov metric and the bounded Wasserstein distance, delivering tight, non-asymptotic upper bounds on pathwise approximation error. We rigorously establish weak convergence and provide quantitative error control for both the iterates’ mean and variance. The framework thus furnishes a foundational toolset for extending the analysis to multivariate settings and more complex stochastic optimization algorithms.
This work addresses the lack of high-probability convergence guarantees for stochastic gradient descent (SGD) under Markovian and martingale difference noise within a unified timescale, even when the objective satisfies the Polyak–Łojasiewicz (PL) condition. The paper establishes, for the first time, high-probability convergence bounds for SGD with Markov noise under the PL condition, allowing the noise magnitude to scale with the function value—a setting relevant to decentralized optimization, privacy-amplified sampling, and online system identification. By characterizing the Markov noise via the Poisson equation and employing a probabilistic induction argument to circumvent the absence of almost-sure bounds on the objective, the authors prove that the expected suboptimality decays at the optimal rate of $1/k$. The theoretical findings are validated through experiments on token-based decentralized linear regression, privacy-amplified supervised learning, and online system identification.
Statistical Model Checking (SMC) often yields inflated error rates in probabilistic and expected reward estimation due to insufficient statistical rigor. To address this, we propose a robust estimation framework with rigorous theoretical guarantees: (i) we extend the Dvoretzky–Kiefer–Wolfowitz (DKW) inequality to expected reward estimation for the first time; (ii) we introduce a limit-PAC (Probably Approximately Correct) procedure ensuring controllable estimation error; and (iii) we derive a computable upper bound on reachability rewards and enhance practicality via path truncation and distribution bounding. Our method is implemented in the *modes* tool. Experimental evaluation demonstrates a substantial reduction in erroneous conclusions while maintaining high precision, thereby ensuring both statistical correctness and engineering applicability.
High-probability analysis of learning algorithms involving light-tailed (e.g., sub-exponential, sub-Gaussian) but possibly unbounded random variables poses significant technical challenges due to the lack of uniform concentration tools across distribution families. Method: We propose a generic black-box reduction that systematically transforms high-probability analysis of any algorithm relying on light-tailed randomness into the corresponding analysis under bounded-variable assumptions, incurring only controllable logarithmic-factor overheads. Contribution/Results: This is the first unified framework handling diverse light-tailed distributions without ad hoc concentration inequalities—greatly simplifying theoretical analysis. As applications, we reconstruct a generalized Azuma’s inequality and derive tight high-probability convergence bounds for stochastic optimization algorithms under light-tailed noise, demonstrating both the method’s effectiveness and broad applicability.
This work investigates the concentration of iteration errors in stochastic approximation algorithms driven by heavy-tailed Markov noise, covering both expansive and non-expansive operator settings. Under a framework involving a finite-state Markov component and martingale difference noise, the authors construct a novel Lyapunov function via the moment-generating function of the solution to the Poisson equation, complemented by auxiliary projection and black-box truncation techniques to reduce unbounded noise to a bounded setting. The study provides the first systematic characterization of the fine structure of error tails: under bounded noise, tails can be sub-Gaussian, sub-Weibull, or intermediate between Pareto and Weibull; under unbounded noise, if the operator is almost surely non-expansive, the error tail is at most three times heavier than that of the noise, whereas if the operator is expansive with positive probability, significantly heavier tails may arise, with sharp worst-case examples demonstrating the tightness of these bounds.
This work addresses the limitations of classical algorithmic stability theory, which typically relies on strong assumptions such as bounded loss functions or sub-Gaussian/sub-Weibull tail behavior—conditions often violated in heavy-tailed or unbounded loss settings. The paper introduces a novel $L_p$ stability framework that requires only finite $L_p$ moments of the loss function, thereby relaxing the conventional bounded differences condition. By extending McDiarmid’s inequality to accommodate $L_p$ constraints, the authors derive sharp high-probability generalization bounds under this significantly weaker assumption. This approach is shown to be broadly applicable across multiple learning paradigms, including empirical risk minimization, transductive regression, and meta-learning, demonstrating that robust generalization guarantees can still be achieved even when losses are unbounded, provided $L_p$ stability holds.
This work addresses stochastic approximation problems involving multiplicative noise and unbounded iterates under compression, proposing a unified and elementary analytical framework. By directly constructing a first-order Lyapunov drift inequality for the error norm, the approach avoids sophisticated tools such as smoothing or Moreau envelopes. Combining averaged noise sequences with auxiliary iterates, the method employs inductive expectation arguments to derive mean-square error bounds and, for the first time under multiplicative noise, establishes a full-trajectory maximal concentration bound with sub-Gaussian tails via probabilistic induction and the Azuma–Hoeffding inequality. The stepsize schedule depends only logarithmically on the confidence parameter, eliminating the need for multi-stage warm-up procedures or smooth Lyapunov functions, thereby significantly simplifying the analysis. The framework is successfully applied to ℓ∞-contractive operators in reinforcement learning and exhibits strong generalizability.
This work addresses kernel regression under non-Gaussian noise—including sub-Gaussian, bounded, sub-exponential, and moment-bounded types—by establishing a unified probabilistic uniform error bound within a non-asymptotic framework. It overcomes the prevailing limitation of existing approaches that assume conditionally independent sub-Gaussian noise, and for the first time simultaneously accommodates multiple non-Gaussian noise distributions and dependent noise structures. By integrating concentration inequalities with reproducing kernel Hilbert space theory, the derived non-conservative error bounds substantially enhance the reliability of uncertainty quantification. This leads to significantly tighter confidence regions in safety-critical control tasks, outperforming current state-of-the-art methods in both theoretical rigor and practical performance.
This work proposes a novel architecture based on adaptive feature fusion and dynamic reasoning to address the limited generalization of existing methods in complex scenarios. By integrating a multi-scale context-aware module and a learnable strategy for selecting inference paths, the approach significantly enhances model adaptability and robustness to heterogeneous data. Extensive experiments demonstrate that the proposed method consistently outperforms state-of-the-art techniques across multiple benchmark datasets, achieving particularly strong performance under low-resource and cross-domain settings. Beyond advancing the frontier of general-purpose intelligent reasoning, this study also introduces and open-sources the first training framework capable of dynamic structural optimization, thereby fostering the development of efficient and scalable AI systems.