algorithmic stability analysis

Analyze and derive quantitative characterizations of how algorithm outputs change under small perturbations or adversarial conditions, including proving matching upper and lower stability bounds; specifically develop stability bounds and proofs that account for local differential privacy constraints and Byzantine-robust aggregation, and relate those stability measures to downstream generalization error for learning algorithms.

algorithmicstabilityanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.14
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work aims to provide a more precise characterization of the generalization error of differentially private learning algorithms. By integrating information-theoretic tools with typicality analysis and leveraging the stability properties inherent to differential privacy, we rigorously refine existing mutual information-based upper bounds and establish, for the first time, a computable and tight upper bound on maximal leakage. The proposed approach applies uniformly to both expected and high-probability analyses of generalization error. The resulting bounds are not only tighter than prior results but also exhibit strong computability, thereby offering enhanced guarantees on generalization performance.

differentially private algorithmsgeneralization errorinformation-theoretic bounds

Stability and Generalization in Free Adversarial Training

Apr 13, 2024
XC
Xiwei Cheng
🏛️ The Chinese University of Hong Kong | Purdue University

This work investigates the intrinsic relationship between generalization performance and optimization stability in Free Adversarial Training (FreeAT). Addressing the large generalization gap and poor training stability inherent in standard adversarial training, we first establish—within the algorithmic stability framework—that FreeAT achieves a tighter generalization error bound by jointly optimizing perturbations and model parameters. We further reveal that its synchronous min-max optimization mechanism is critical for narrowing the train-test accuracy gap. Theoretical analysis demonstrates that FreeAT’s generalization upper bound is significantly lower than that of standard adversarial training. Empirical evaluation confirms that, under identical iteration budgets, FreeAT consistently reduces the generalization gap by 15–22% across diverse benchmarks. Our implementation is publicly available.

Adversarial TrainingGeneralizationStability

Quantitative Verification With Neural Networks For Probabilistic Programs and Stochastic Systems

Jan 15, 2023
AA
Alessandro Abate
🏛️ University of Oxford | University of Birmingham

This work addresses the quantitative verification of probabilistic programs and stochastic dynamical systems, specifically aiming to rigorously infer upper bounds on the probability that a stochastic process reaches a target condition within a finite number of steps. We propose a neuro-symbolic approach: supermartingale certificates are parameterized using differentiable neural networks; training employs stochastic optimization, while formal verification leverages SMT solvers (e.g., Z3); and an counterexample-guided inductive synthesis (CEGIS) framework enables iterative refinement. To our knowledge, this is the first method to embed neural networks directly into supermartingale construction—balancing expressive power with formal verifiability—and thereby significantly improves bound tightness and reliability. Evaluated on diverse benchmarks, our computed probability bounds match or surpass those of state-of-the-art techniques. Notably, we successfully verify high-dimensional, nonlinear stochastic models that defy analysis by conventional symbolic methods.

Computing tight probability bounds using neural networksHandling reachability, safety, and termination analysis efficientlyQuantitative verification of probabilistic programs and stochastic models

Generalization Bounds for Dependent Data using Online-to-Batch Conversion

May 22, 2024
SC
Sagnik Chatterjee
🏛️ Indraprastha Institute of Information Technology | Delhi (IIIT -D)

This work addresses the problem of deriving generalization error upper bounds for batch learning algorithms under mixing stochastic processes (i.e., dependent data), without imposing any stability assumptions on the batch learner. The method introduces a novel analytical framework based on Online-to-Batch conversion: stability requirements are shifted to an associated online learner, and a new notion of online algorithm stability—defined via the first-order Wasserstein distance—is proposed for the first time. It is shown that the Exponentially Weighted Average (EWA) algorithm satisfies this stability condition. Consequently, the framework yields both expectation- and high-probability generalization bounds applicable to *any* batch learning algorithm; under i.i.d. assumptions, the bounds reduce to classical forms up to correction terms governed by the mixing decay rate. The resulting bounds are explicitly computable, substantially broadening the applicability and practical utility of learning theory under data dependence.

Dependent data sourcesGeneralization error boundsOnline-to-Batch conversion framework

Latest Papers

What's happening recently
View more

This work addresses the limitations of classical algorithmic stability theory, which typically relies on strong assumptions such as bounded loss functions or sub-Gaussian/sub-Weibull tail behavior—conditions often violated in heavy-tailed or unbounded loss settings. The paper introduces a novel $L_p$ stability framework that requires only finite $L_p$ moments of the loss function, thereby relaxing the conventional bounded differences condition. By extending McDiarmid’s inequality to accommodate $L_p$ constraints, the authors derive sharp high-probability generalization bounds under this significantly weaker assumption. This approach is shown to be broadly applicable across multiple learning paradigms, including empirical risk minimization, transductive regression, and meta-learning, demonstrating that robust generalization guarantees can still be achieved even when losses are unbounded, provided $L_p$ stability holds.

algorithmic stabilitygeneralization boundsheavy-tailed losses

This study investigates the impact of local differential privacy (LDP) on generalization error in Byzantine-robust distributed learning. Leveraging algorithmic stability theory, it reveals—for the first time—a non-monotonic relationship between LDP noise magnitude and generalization error: stronger privacy can reduce generalization error, whereas weaker privacy may exacerbate it. This finding challenges the conventional wisdom that privacy, robustness, and generalization are inherently at odds. The theoretical analysis establishes matching upper and lower bounds and incorporates a Byzantine-robust aggregation mechanism. Empirical experiments corroborate the existence of this non-monotonic effect and elucidate its underlying mechanisms.

Byzantine robustnessdistributed learninggeneralization

This work investigates the generalization error and stability of gradient descent (GD) and stochastic gradient descent (SGD) under deterministic or random rounding in discrete parameter spaces. Leveraging frameworks of uniform stability and parameter uniform stability, and assuming convexity, Lipschitz continuity, and smoothness, the authors derive generalization bounds for both algorithms under rounding operations. Their main contributions include showing that deterministic rounding degrades GD’s generalization error to $O(T/\sqrt{n})$ and undermines its stability; demonstrating that SGD retains nontrivial stability under deterministic rounding, with bounds of $O(T/n)$ in one dimension and $O(T^2/n)$ in higher dimensions; and establishing a tight upper bound on parameter stability for random rounding under coordinate-wise separable losses.

discrete parameter spacesgeneralization errorgradient descent

This work addresses the limitations of classical algorithmic stability analyses in non-convex optimization, which typically require a stringent learning rate decay of \(O(1/t)\)—a condition that hampers optimization efficiency and contradicts common empirical practice. Focusing on homogeneous neural networks, such as fully connected or convolutional architectures with ReLU or LeakyReLU activations, the paper bridges algorithmic stability theory with non-convex optimization analysis to relax this constraint. Under mild assumptions, it establishes generalization bounds that permit a significantly slower learning rate decay of \(\Omega(1/\sqrt{t})\). The resulting framework not only accommodates non-Lipschitz settings but also markedly improves the alignment between theoretical guarantees and practical training dynamics.

algorithmic stabilitygeneralization boundshomogeneous neural networks

Hot Scholars

YL

Yunwen Lei

The University of Hong Kong
Statistical Learning TheoryStochastic OptimizationMachine Learning
FH

Feihu Huang

Professor of Nanjing University of Aeronautics & Astronautics
Machine LearningOptimization
XC

Xiaochun Cao

Sun Yat-sen University
Computer VisionArtificial IntelligenceMultimediaMachine Learning
SC

Songcan Chen

Nanjing University of Aeronautics & Astronautics
Machine LearningPattern recognition
AB

Aurélien Bellet

Research scientist at Inria
Machine LearningArtificial IntelligenceData SciencePrivacy