Score
Analyze and derive quantitative characterizations of how algorithm outputs change under small perturbations or adversarial conditions, including proving matching upper and lower stability bounds; specifically develop stability bounds and proofs that account for local differential privacy constraints and Byzantine-robust aggregation, and relate those stability measures to downstream generalization error for learning algorithms.
This work aims to provide a more precise characterization of the generalization error of differentially private learning algorithms. By integrating information-theoretic tools with typicality analysis and leveraging the stability properties inherent to differential privacy, we rigorously refine existing mutual information-based upper bounds and establish, for the first time, a computable and tight upper bound on maximal leakage. The proposed approach applies uniformly to both expected and high-probability analyses of generalization error. The resulting bounds are not only tighter than prior results but also exhibit strong computability, thereby offering enhanced guarantees on generalization performance.
This work investigates the intrinsic relationship between generalization performance and optimization stability in Free Adversarial Training (FreeAT). Addressing the large generalization gap and poor training stability inherent in standard adversarial training, we first establish—within the algorithmic stability framework—that FreeAT achieves a tighter generalization error bound by jointly optimizing perturbations and model parameters. We further reveal that its synchronous min-max optimization mechanism is critical for narrowing the train-test accuracy gap. Theoretical analysis demonstrates that FreeAT’s generalization upper bound is significantly lower than that of standard adversarial training. Empirical evaluation confirms that, under identical iteration budgets, FreeAT consistently reduces the generalization gap by 15–22% across diverse benchmarks. Our implementation is publicly available.
This work addresses the quantitative verification of probabilistic programs and stochastic dynamical systems, specifically aiming to rigorously infer upper bounds on the probability that a stochastic process reaches a target condition within a finite number of steps. We propose a neuro-symbolic approach: supermartingale certificates are parameterized using differentiable neural networks; training employs stochastic optimization, while formal verification leverages SMT solvers (e.g., Z3); and an counterexample-guided inductive synthesis (CEGIS) framework enables iterative refinement. To our knowledge, this is the first method to embed neural networks directly into supermartingale construction—balancing expressive power with formal verifiability—and thereby significantly improves bound tightness and reliability. Evaluated on diverse benchmarks, our computed probability bounds match or surpass those of state-of-the-art techniques. Notably, we successfully verify high-dimensional, nonlinear stochastic models that defy analysis by conventional symbolic methods.
This work addresses the problem of deriving generalization error upper bounds for batch learning algorithms under mixing stochastic processes (i.e., dependent data), without imposing any stability assumptions on the batch learner. The method introduces a novel analytical framework based on Online-to-Batch conversion: stability requirements are shifted to an associated online learner, and a new notion of online algorithm stability—defined via the first-order Wasserstein distance—is proposed for the first time. It is shown that the Exponentially Weighted Average (EWA) algorithm satisfies this stability condition. Consequently, the framework yields both expectation- and high-probability generalization bounds applicable to *any* batch learning algorithm; under i.i.d. assumptions, the bounds reduce to classical forms up to correction terms governed by the mixing decay rate. The resulting bounds are explicitly computable, substantially broadening the applicability and practical utility of learning theory under data dependence.
This work addresses the limitations of classical algorithmic stability theory, which typically relies on strong assumptions such as bounded loss functions or sub-Gaussian/sub-Weibull tail behavior—conditions often violated in heavy-tailed or unbounded loss settings. The paper introduces a novel $L_p$ stability framework that requires only finite $L_p$ moments of the loss function, thereby relaxing the conventional bounded differences condition. By extending McDiarmid’s inequality to accommodate $L_p$ constraints, the authors derive sharp high-probability generalization bounds under this significantly weaker assumption. This approach is shown to be broadly applicable across multiple learning paradigms, including empirical risk minimization, transductive regression, and meta-learning, demonstrating that robust generalization guarantees can still be achieved even when losses are unbounded, provided $L_p$ stability holds.
This study investigates the impact of local differential privacy (LDP) on generalization error in Byzantine-robust distributed learning. Leveraging algorithmic stability theory, it reveals—for the first time—a non-monotonic relationship between LDP noise magnitude and generalization error: stronger privacy can reduce generalization error, whereas weaker privacy may exacerbate it. This finding challenges the conventional wisdom that privacy, robustness, and generalization are inherently at odds. The theoretical analysis establishes matching upper and lower bounds and incorporates a Byzantine-robust aggregation mechanism. Empirical experiments corroborate the existence of this non-monotonic effect and elucidate its underlying mechanisms.
This work investigates the generalization error and stability of gradient descent (GD) and stochastic gradient descent (SGD) under deterministic or random rounding in discrete parameter spaces. Leveraging frameworks of uniform stability and parameter uniform stability, and assuming convexity, Lipschitz continuity, and smoothness, the authors derive generalization bounds for both algorithms under rounding operations. Their main contributions include showing that deterministic rounding degrades GD’s generalization error to $O(T/\sqrt{n})$ and undermines its stability; demonstrating that SGD retains nontrivial stability under deterministic rounding, with bounds of $O(T/n)$ in one dimension and $O(T^2/n)$ in higher dimensions; and establishing a tight upper bound on parameter stability for random rounding under coordinate-wise separable losses.
This work addresses the limitations of classical algorithmic stability analyses in non-convex optimization, which typically require a stringent learning rate decay of \(O(1/t)\)—a condition that hampers optimization efficiency and contradicts common empirical practice. Focusing on homogeneous neural networks, such as fully connected or convolutional architectures with ReLU or LeakyReLU activations, the paper bridges algorithmic stability theory with non-convex optimization analysis to relax this constraint. Under mild assumptions, it establishes generalization bounds that permit a significantly slower learning rate decay of \(\Omega(1/\sqrt{t})\). The resulting framework not only accommodates non-Lipschitz settings but also markedly improves the alignment between theoretical guarantees and practical training dynamics.