Score
Designs, implements, and analyses optimization regularizers that use the Jensen–Shannon (JS) divergence to penalize or constrain differences between two probability distributions; this includes formulating JS-based loss terms, integrating them into training objectives (JS-regularized optimization), choosing weighting schedules, and measuring their effect on training dynamics. These techniques are applied to discourage large deviations from a reference distribution, mitigate bias toward uniform references, and help preserve diversity in model outputs while maintaining optimization performance.
This study investigates the sample complexity required to distinguish two probability distributions based on their Jensen–Shannon divergence (JSD). Focusing on independent and identically distributed samples, the authors analyze the logarithmic likelihood ratio classifier and the majority vote classifier under a fixed JSD. By leveraging tools from information theory and statistical learning theory, they establish distinct scaling laws for the sample complexity of the two classifiers: it grows as $1/\text{JSD}$ for the likelihood ratio classifier and as $1/\text{JSD}^2$ for the majority vote classifier. These findings provide an operational statistical interpretation of JSD and, for the first time, quantify its precise relationship with the distinguishability of probability distributions.
This work investigates the incorporation of $f$-divergence regularization into empirical risk minimization to enhance generalization in expected risk. By establishing equivalence conditions between $f$-divergence-regularized empirical risk minimization and expected risk minimization under an $f$-divergence constraint, the study introduces the notion of a “normalizing function,” which is characterized as a nonlinear ordinary differential equation (ODE). This characterization reveals structural equivalences across different $f$-divergence regularizations. Leveraging duality theory, ODE analysis, and numerical approximation techniques, the authors develop a unified computational framework applicable to a broad class of $f$-divergences. Numerical experiments demonstrate the practical impact of various $f$-functions on training and test risks, thereby extending the range of tractable divergences and strengthening the theoretical and algorithmic coherence of the approach.
This work uncovers a critical asymmetry in relative entropy regularization for empirical risk minimization (ERM). We systematically analyze two distinct ERM-relative entropy regularization (ERM-RER) formulations: Type-I, where the optimized measure is regularized relative to a fixed reference measure, and Type-II, where the reference measure is regularized relative to the optimized one. We provide the first structural characterization of Type-II solutions, revealing that they necessarily concentrate—i.e., their support collapses—onto the support of the reference measure, inducing a strong inductive bias. Theoretically, we prove that both types enforce support contraction of the solution, and further establish that Type-II regularization is strictly equivalent to a Type-I formulation under a specific transformation of the loss function. This equivalence refutes the conventional view of Type-I and Type-II as symmetric variants, and advances an information-geometric understanding of how entropy regularization governs generalization.
In data-driven optimization, decision samples often exhibit optimistic bias relative to true performance due to the “optimizer’s curse.” To address this, we propose a first-order bias correction method that avoids re-optimization. We introduce the Optimizer’s Information Criterion (OIC), the first information-theoretic criterion tailored for decision selection in data-driven optimization—generalizing the Akaike Information Criterion (AIC) to encompass empirical models, parametric models, regularization, and contextual optimization. Leveraging asymptotic statistical analysis, we derive an analytical bias expression that explicitly captures the coupling between optimization and learning, eliminating the need for cross-validation. Evaluated on both synthetic and real-world datasets, our method achieves more accurate bias estimation and significantly lower computational overhead, while providing rigorous theoretical guarantees.
This work addresses the limited generalization of genetic programming in symbolic regression due to overfitting. The authors propose a novel evolutionary feature construction method based on neighborhood risk decomposition, which, for the first time, incorporates the neighborhood Jensen gap as a regularization term to jointly optimize empirical risk and the Jensen gap. To enhance robustness, the approach integrates dynamic regularization strength adjustment, manifold intrusion detection, and noise perturbation mechanisms, effectively mitigating the generation of unrealistic samples caused by data augmentation. Extensive experiments on 58 benchmark datasets demonstrate that the proposed method outperforms existing complexity-controlling metrics and significantly improves symbolic regression performance compared to 15 state-of-the-art machine learning algorithms.
This study addresses the challenges of bounding generalization error and elucidating regularization mechanisms in neural network quantization. We systematically analyze the generalization behavior of OPTQ and its variants under test distributions, deriving theoretical bounds on the expected squared error. Through rigorous theoretical proofs and stochastic optimization analysis, we uncover how regularization terms influence dependence on the calibration set, and accordingly propose a novel regularization parameter selection strategy. Experimental results demonstrate that this strategy significantly outperforms existing methods, effectively reducing model quantization error.
This study addresses the lack of comparability in empirical Jensen-Shannon divergence (JSD) estimates used for evaluating synthetic tabular data fidelity, stemming from inconsistent estimation protocols. The authors systematically analyze the behavior of two prevalent JSD estimators under finite-sample conditions: marginal-based and classifier-based estimators. They reveal that marginal estimators neglect variable dependencies and introduce prior-shift bias, while classifier-based estimators suffer from class imbalance and sensitivity in high dimensions. To mitigate these issues, the work proposes a closed-form posterior correction for classifier-based estimators and advocates for explicit declaration of estimation protocols to ensure reproducibility and comparability. Through controlled experiments, benchmark divergences, and real-world synthetic datasets, the study delivers practical guidelines and open-source tools to enable estimator-aware, reliable fidelity evaluation.
该研究通过局部曲率分析,探讨了基于f-散度的正则化与锐度感知最小化(SAM)之间的关系,并提出了一种新的输入空间扰动框架。
This study addresses the lack of systematic investigation into the statistical properties and testing performance of Jensen–Shannon divergence (JSD) and Kullback–Leibler (KL) divergence in credit risk model monitoring. It derives, for the first time, chi-squared asymptotic reference distributions for both divergences under distributional shift using asymptotic theory, and conducts a comprehensive Monte Carlo simulation to evaluate their Type I error control and statistical power relative to the Population Stability Index (PSI). The results demonstrate that JSD exhibits superior Type I error control, closely attaining the nominal 5% level, yet shows limited power (27%) in small samples (n=200); in contrast, KL divergence and PSI achieve higher power (32%). These findings provide both theoretical grounding and empirical guidance for selecting appropriate divergence metrics in practical model monitoring.
This work addresses the reliance on empirical design and poor generalizability of hand-crafted optimizers in gradient-based learning. Methodologically, it formulates the optimizer as a learnable functional mapping from gradients to parameter updates, and—novelty—systematically recasts this as a sequence of analytically solvable convex optimization problems. This unified framework yields closed-form derivations of mainstream optimizers (e.g., SGD, Adam) along with their theoretically optimal hyperparameters. Furthermore, it incorporates a runtime gradient statistics mechanism enabling dynamic, adaptive tuning during training. Experiments demonstrate substantial improvements in convergence speed and training stability, while preserving theoretical rigor and practical deployability.