Score
Designs and analyzes probabilistic bounds, limit theorems, and decompositions for stochastic processes and their functionals by representing quantities as martingales and applying martingale concentration inequalities and central-limit results; uses martingale decompositions and related techniques to derive concentration bounds, variance-dependent error or regret bounds, and conditional asymptotic validity.
This work addresses the asymptotic convergence and non-asymptotic maximal inequalities for semidefinite matrix-valued martingales and reverse submartingales. Motivated by the lack of systematic convergence analysis and adaptive stopping-time control under the Loewner order in existing theory, we establish, for the first time, a Loewner-order-based martingale convergence theorem and adaptive maximal inequalities for stopping times—unifying light-tailed, heavy-tailed, and self-normalized settings. Our approach integrates matrix probability theory, random matrix theory, martingale analysis, and dependence modeling to derive novel matrix concentration inequalities. Crucially, our results overcome the limitations of classical Chernoff-type bounds—which require fixed sample sizes and independence—by accommodating arbitrary (possibly data-dependent) sample sizes and stopping times. This significantly broadens theoretical applicability and practical utility in high-dimensional statistical inference and online machine learning.
High-probability analysis of learning algorithms involving light-tailed (e.g., sub-exponential, sub-Gaussian) but possibly unbounded random variables poses significant technical challenges due to the lack of uniform concentration tools across distribution families. Method: We propose a generic black-box reduction that systematically transforms high-probability analysis of any algorithm relying on light-tailed randomness into the corresponding analysis under bounded-variable assumptions, incurring only controllable logarithmic-factor overheads. Contribution/Results: This is the first unified framework handling diverse light-tailed distributions without ad hoc concentration inequalities—greatly simplifying theoretical analysis. As applications, we reconstruct a generalized Azuma’s inequality and derive tight high-probability convergence bounds for stochastic optimization algorithms under light-tailed noise, demonstrating both the method’s effectiveness and broad applicability.
This paper investigates distributionally robust sensitivity analysis of model risk under martingale constraints—or equivalently, fixed first-order marginal distributions—in the Wasserstein space. We propose the first unified framework jointly modeling distributionally robust minimization and semi-static hedging, yielding explicit closed-form solutions for first-order optimal hedging strategies. Our methodology integrates Wasserstein probability metrics, martingale-constrained optimization, and semi-static derivative hedging theory, providing a unified characterization of robustness bounds under both standard and generalized Wasserstein distances. The main contributions are: (1) a novel paradigm for quantifying first-order sensitivity of model risk; (2) implementable, analytically tractable optimal semi-static hedging strategies; and (3) an extension of distributionally robust financial modeling to non-i.i.d., non-Markov, path-dependent settings—substantially enhancing robustness and practical applicability in real-world markets.
This work establishes Rosenthal- and Bernstein-type concentration inequalities for additive functionals of geometrically ergodic Markov chains, explicitly characterizing the dependence of deviation bounds on mixing time. Methodologically, it pioneers the extension of the classical Rosenthal inequality to the Markov-dependent setting via a novel analytical framework based on Poisson equation decomposition, which precisely links mixing constants, martingale Rosenthal constants, and deviation bounds. Integrating martingale techniques, geometric ergodicity analysis, and quantitative mixing time estimation, the approach yields computable, explicit, and tight upper bounds—significantly improving the polynomial dependence on mixing time present in prior results. The derived inequalities provide a rigorous theoretical foundation for error control in MCMC algorithms, sequential Monte Carlo estimation, and large-sample inference for non-i.i.d. statistics.
This work investigates the concentration of iteration errors in stochastic approximation algorithms driven by heavy-tailed Markov noise, covering both expansive and non-expansive operator settings. Under a framework involving a finite-state Markov component and martingale difference noise, the authors construct a novel Lyapunov function via the moment-generating function of the solution to the Poisson equation, complemented by auxiliary projection and black-box truncation techniques to reduce unbounded noise to a bounded setting. The study provides the first systematic characterization of the fine structure of error tails: under bounded noise, tails can be sub-Gaussian, sub-Weibull, or intermediate between Pareto and Weibull; under unbounded noise, if the operator is almost surely non-expansive, the error tail is at most three times heavier than that of the noise, whereas if the operator is expansive with positive probability, significantly heavier tails may arise, with sharp worst-case examples demonstrating the tightness of these bounds.
This work addresses the absence of finite-sample error bounds and concentration inequalities for nonlinear stochastic approximation algorithms under the Wasserstein-p distance. By coupling the discrete-time iterative process with its Ornstein–Uhlenbeck diffusion limit, the paper establishes the first non-asymptotic distributional convergence rates in Wasserstein distance under general noise conditions—such as martingale differences and ergodic Markov chains. The main contributions include proving that the last iterate converges to a Gaussian distribution at a rate of γₙ^{1/6}, while the Polyak–Ruppert averaged iterate achieves a rate of n^{-1/6}. Moreover, the analysis yields high-probability concentration inequalities that improve upon those derived via classical moment-based methods. The proposed framework applies broadly to canonical algorithms, including linear stochastic approximation and stochastic gradient descent.
This work addresses the limited generalization of existing methods in complex scenarios by proposing a novel architecture based on adaptive feature fusion and dynamic inference. The approach effectively integrates local details and global semantic information through a multi-scale context-aware module and a learnable routing strategy, enabling the model to dynamically adjust its computational pathway during inference according to input content. Experimental results demonstrate that the proposed model significantly outperforms current state-of-the-art methods across multiple benchmark datasets while maintaining low computational overhead. The primary contribution lies in the introduction of a general and efficient dynamic inference framework, offering a new perspective for enhancing model robustness on out-of-distribution data.
This work addresses the inconsistency between training and inference in existing speculative decoding methods, where training optimizes only a single greedy path while inference requires verifying multiple sampled paths. To bridge this gap, we introduce variational inference into speculative decoding for the first time, reformulating draft model training as posterior inference over latent proposal paths by maximizing the marginal probability of acceptance under the target model. We propose a path-level utility function, an EM-based optimization framework, and two novel mechanisms: Adaptive Rejection Weighting (ARW) and Confidence-Aware Regularization (CAR). Experiments demonstrate that our approach achieves up to 9.6% higher speedup than EAGLE-3 and a 7.9% improvement in acceptance rate over ViSpec across various large language and multimodal models, significantly enhancing inference efficiency.
This study addresses the challenge of enhancing both computational efficiency and solution accuracy in primal optimal stopping problems by leveraging dual martingales. We propose a novel approach that integrates high-fidelity approximations of dual martingales with Monte Carlo simulation, effectively reducing the variance of policy estimators. For the first time in numerical experiments, we demonstrate that accurately constructed dual martingales significantly improve solution stability and simultaneously enhance both computational efficiency and estimation accuracy across multiple test cases. Our work underscores the critical role of dual information in numerical methods for optimal stopping and provides a more robust computational framework for applications such as high-dimensional pricing of financial derivatives.