js divergence regularization

Designs, implements, and analyses optimization regularizers that use the Jensen–Shannon (JS) divergence to penalize or constrain differences between two probability distributions; this includes formulating JS-based loss terms, integrating them into training objectives (JS-regularized optimization), choosing weighting schedules, and measuring their effect on training dynamics. These techniques are applied to discourage large deviations from a reference distribution, mitigate bias toward uniform references, and help preserve diversity in model outputs while maintaining optimization performance.

jsdivergenceregularization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.16
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study investigates the sample complexity required to distinguish two probability distributions based on their Jensen–Shannon divergence (JSD). Focusing on independent and identically distributed samples, the authors analyze the logarithmic likelihood ratio classifier and the majority vote classifier under a fixed JSD. By leveraging tools from information theory and statistical learning theory, they establish distinct scaling laws for the sample complexity of the two classifiers: it grows as $1/\text{JSD}$ for the likelihood ratio classifier and as $1/\text{JSD}^2$ for the majority vote classifier. These findings provide an operational statistical interpretation of JSD and, for the first time, quantify its precise relationship with the distinguishability of probability distributions.

classification errordistribution distinguishabilityJensen-Shannon divergence

This work investigates the incorporation of $f$-divergence regularization into empirical risk minimization to enhance generalization in expected risk. By establishing equivalence conditions between $f$-divergence-regularized empirical risk minimization and expected risk minimization under an $f$-divergence constraint, the study introduces the notion of a “normalizing function,” which is characterized as a nonlinear ordinary differential equation (ODE). This characterization reveals structural equivalences across different $f$-divergence regularizations. Leveraging duality theory, ODE analysis, and numerical approximation techniques, the authors develop a unified computational framework applicable to a broad class of $f$-divergences. Numerical experiments demonstrate the practical impact of various $f$-functions on training and test risks, thereby extending the range of tractable divergences and strengthening the theoretical and algorithmic coherence of the approach.

Constraint OptimizationEmpirical Risk MinimizationExpected Risk

Asymmetry of the Relative Entropy in the Regularization of Empirical Risk Minimization

Oct 02, 2024
FD
Francisco Daunas
🏛️ The University of Sheffield | INRIA | Princeton University | University of French Polynesia

This work uncovers a critical asymmetry in relative entropy regularization for empirical risk minimization (ERM). We systematically analyze two distinct ERM-relative entropy regularization (ERM-RER) formulations: Type-I, where the optimized measure is regularized relative to a fixed reference measure, and Type-II, where the reference measure is regularized relative to the optimized one. We provide the first structural characterization of Type-II solutions, revealing that they necessarily concentrate—i.e., their support collapses—onto the support of the reference measure, inducing a strong inductive bias. Theoretically, we prove that both types enforce support contraction of the solution, and further establish that Type-II regularization is strictly equivalent to a Type-I formulation under a specific transformation of the loss function. This equivalence refutes the conventional view of Type-I and Type-II as symmetric variants, and advances an information-geometric understanding of how entropy regularization governs generalization.

Analyzes relative entropy asymmetry in empirical risk minimizationCompares Type-I and Type-II ERM-RER regularization effectsShows regularization collapses solution support to reference measure

In data-driven optimization, decision samples often exhibit optimistic bias relative to true performance due to the “optimizer’s curse.” To address this, we propose a first-order bias correction method that avoids re-optimization. We introduce the Optimizer’s Information Criterion (OIC), the first information-theoretic criterion tailored for decision selection in data-driven optimization—generalizing the Akaike Information Criterion (AIC) to encompass empirical models, parametric models, regularization, and contextual optimization. Leveraging asymptotic statistical analysis, we derive an analytical bias expression that explicitly captures the coupling between optimization and learning, eliminating the need for cross-validation. Evaluated on both synthetic and real-world datasets, our method achieves more accurate bias estimation and significantly lower computational overhead, while providing rigorous theoretical guarantees.

Correcting optimistic bias in data-driven optimization decisionsGeneralizing Akaike Information Criterion for optimization performanceReducing computational cost of bias correction methods

This work addresses the limited generalization of genetic programming in symbolic regression due to overfitting. The authors propose a novel evolutionary feature construction method based on neighborhood risk decomposition, which, for the first time, incorporates the neighborhood Jensen gap as a regularization term to jointly optimize empirical risk and the Jensen gap. To enhance robustness, the approach integrates dynamic regularization strength adjustment, manifold intrusion detection, and noise perturbation mechanisms, effectively mitigating the generation of unrealistic samples caused by data augmentation. Extensive experiments on 58 benchmark datasets demonstrate that the proposed method outperforms existing complexity-controlling metrics and significantly improves symbolic regression performance compared to 15 state-of-the-art machine learning algorithms.

feature constructiongeneralizationgenetic programming

Latest Papers

What's happening recently
View more

This study addresses the challenges of bounding generalization error and elucidating regularization mechanisms in neural network quantization. We systematically analyze the generalization behavior of OPTQ and its variants under test distributions, deriving theoretical bounds on the expected squared error. Through rigorous theoretical proofs and stochastic optimization analysis, we uncover how regularization terms influence dependence on the calibration set, and accordingly propose a novel regularization parameter selection strategy. Experimental results demonstrate that this strategy significantly outperforms existing methods, effectively reducing model quantization error.

generalization errorneural network compressionOPTQ

This study addresses the lack of comparability in empirical Jensen-Shannon divergence (JSD) estimates used for evaluating synthetic tabular data fidelity, stemming from inconsistent estimation protocols. The authors systematically analyze the behavior of two prevalent JSD estimators under finite-sample conditions: marginal-based and classifier-based estimators. They reveal that marginal estimators neglect variable dependencies and introduce prior-shift bias, while classifier-based estimators suffer from class imbalance and sensitivity in high dimensions. To mitigate these issues, the work proposes a closed-form posterior correction for classifier-based estimators and advocates for explicit declaration of estimation protocols to ensure reproducibility and comparability. Through controlled experiments, benchmark divergences, and real-world synthetic datasets, the study delivers practical guidelines and open-source tools to enable estimator-aware, reliable fidelity evaluation.

divergence estimationestimator dependenceJensen-Shannon divergence

This study addresses the lack of systematic investigation into the statistical properties and testing performance of Jensen–Shannon divergence (JSD) and Kullback–Leibler (KL) divergence in credit risk model monitoring. It derives, for the first time, chi-squared asymptotic reference distributions for both divergences under distributional shift using asymptotic theory, and conducts a comprehensive Monte Carlo simulation to evaluate their Type I error control and statistical power relative to the Population Stability Index (PSI). The results demonstrate that JSD exhibits superior Type I error control, closely attaining the nominal 5% level, yet shows limited power (27%) in small samples (n=200); in contrast, KL divergence and PSI achieve higher power (32%). These findings provide both theoretical grounding and empirical guidance for selecting appropriate divergence metrics in practical model monitoring.

credit riskdistributional shiftdivergence measures

Optimizing Optimizers for Fast Gradient-Based Learning

Dec 06, 2025
JL
Jaerin Lee
🏛️ Seoul National University

This work addresses the reliance on empirical design and poor generalizability of hand-crafted optimizers in gradient-based learning. Methodologically, it formulates the optimizer as a learnable functional mapping from gradients to parameter updates, and—novelty—systematically recasts this as a sequence of analytically solvable convex optimization problems. This unified framework yields closed-form derivations of mainstream optimizers (e.g., SGD, Adam) along with their theoretically optimal hyperparameters. Furthermore, it incorporates a runtime gradient statistics mechanism enabling dynamic, adaptive tuning during training. Experiments demonstrate substantial improvements in convergence speed and training stability, while preserving theoretical rigor and practical deployability.

Automating optimizer design in gradient-based learningDynamically tuning hyperparameters based on gradient statisticsMaximizing instantaneous loss decrease via optimizer formulation

Hot Scholars

MD

Marc Dymetman

Independent Researcher (Prev. Principal Scientist, NAVER Labs Europe)
Natural Language ProcessingMachine Learning
SN

Shuaicheng Niu

Nanyang Technological University
Machine LearningDomain AdaptationRobustnessAutoML
GC

Guohao Chen

South China University of Technology
Transfer LearningDomain AdaptationTest-Time Adaptation
MT

Mingkui Tan

South China University of Technology
Machine LearningLarge-scale Optimization
TZ

Tunyu Zhang

Rutgers University. PhD Student
Machine LearningLLMDiffusion ModelGenerative AI