Score
Designs, implements, and analyzes methods for quantifying differences between probability distributions, including creating and evaluating divergence metrics and estimators (f‑divergences, KL, Jensen–Shannon), deriving their mathematical properties and bounds, and constructing statistical tests for divergence. Builds algorithms and optimization procedures to calculate, estimate, or minimize divergence (KL divergence estimation/manipulation/optimization), and produces principled scalar measures to report distributional change, compare impacts, or detect outlyingness.
This study addresses the lack of systematic investigation into the statistical properties and testing performance of Jensen–Shannon divergence (JSD) and Kullback–Leibler (KL) divergence in credit risk model monitoring. It derives, for the first time, chi-squared asymptotic reference distributions for both divergences under distributional shift using asymptotic theory, and conducts a comprehensive Monte Carlo simulation to evaluate their Type I error control and statistical power relative to the Population Stability Index (PSI). The results demonstrate that JSD exhibits superior Type I error control, closely attaining the nominal 5% level, yet shows limited power (27%) in small samples (n=200); in contrast, KL divergence and PSI achieve higher power (32%). These findings provide both theoretical grounding and empirical guidance for selecting appropriate divergence metrics in practical model monitoring.
This study clarifies the informational nature of the Kullback–Leibler (KL) divergence difference (Δ_KL) between discrete empirical distributions and corrects the common misinterpretation of its sign as indicating support set inclusion or coverage. By analytically decomposing the mathematical structure of Δ_KL, the work reveals that it fundamentally serves as an asymmetric contrastive measure of weighted log-probability ratios across categories, rather than a metric of distributional breadth or set containment. The theoretical insights are substantiated through a bibliometric case study examining topic distributions in COVID-19 preprints, offering both intuitive interpretation and empirical validation. This research advances the information-theoretic understanding of asymmetric distributional discrepancies and provides a rigorous foundation—supported by concrete examples—for the proper interpretation of Δ_KL in practical applications.
This paper addresses the diagnostic ambiguity in multivariate KL divergence arising from the entanglement of marginal mismatch and statistical dependencies. We propose an algebraically exact, fully additive hierarchical decomposition method based on Möbius inversion over the subset lattice—yielding the first closed-form, complete analytic decomposition of KL divergence. The total divergence between a joint distribution and its product reference is rigorously disentangled into a sum of independent marginal mismatch terms and all-order (r-wise) statistical dependency terms, expressed solely in terms of standard Shannon information measures, without approximations or modeling assumptions. The framework unifies higher-order mutual information, total correlation, and information-geometric principles. Numerical experiments confirm machine-precision accuracy across diverse systems, substantially enhancing attribution-based diagnostic capability for divergence sources in machine learning, econometrics, and complex systems analysis.
The original kernelized Kullback–Leibler (KL) divergence is ill-defined when the supports of compared distributions are disjoint—a fundamental limitation. To address this, the paper proposes a Tikhonov-regularized kernel KL divergence, constructed via covariance operator embeddings in a reproducing kernel Hilbert space (RKHS). This metric is well-defined for arbitrary probability distributions—including discrete, continuous, and mutually singular ones—and provides theoretical guarantees: a bias bound relative to the true KL divergence, finite-sample convergence rates, and a closed-form solution for discrete distributions. Furthermore, the authors formulate a Wasserstein gradient flow optimization framework for the proposed divergence, ensuring theoretical convergence, and design an efficient algorithm applicable to discrete structures such as point clouds. Experiments on point cloud transport tasks demonstrate that the method outperforms existing kernelized and Wasserstein-based approaches, achieving superior stability and robustness.
Wasserstein and Cramér distances lack directional interpretability—i.e., they do not distinguish between location shifts and scale changes in distributional discrepancies. To address this, we propose an interpretable framework based on geometric decomposition of quantile functions, uniquely disentangling each distance into directed shift (location) and dispersion (scale) components. This decomposition is the first to satisfy naturalness, uniqueness, and additivity—key statistical desiderata—within the location-scale family. We further derive explicit sensitivity expressions of the distances with respect to location and dispersion parameters, and establish a weak stochastic order theory that jointly characterizes both location and dispersion orderings. Empirically validated on extreme temperature forecasting evaluation and probabilistic survey design in economics, our method substantially enhances semantic clarity in interpreting distributional differences and strengthens decision-support capabilities.
This study investigates the sample complexity required to distinguish two probability distributions based on their Jensen–Shannon divergence (JSD). Focusing on independent and identically distributed samples, the authors analyze the logarithmic likelihood ratio classifier and the majority vote classifier under a fixed JSD. By leveraging tools from information theory and statistical learning theory, they establish distinct scaling laws for the sample complexity of the two classifiers: it grows as $1/\text{JSD}$ for the likelihood ratio classifier and as $1/\text{JSD}^2$ for the majority vote classifier. These findings provide an operational statistical interpretation of JSD and, for the first time, quantify its precise relationship with the distinguishability of probability distributions.
This study addresses the critical role of utility function selection in Bayesian experimental design under model misspecification. Through theoretical analysis and sequential source inversion experiments, the authors systematically compare the performance of Kullback–Leibler (KL) divergence and Wasserstein distance as utility functions. They find that KL divergence yields faster convergence in the absence of model bias, whereas Wasserstein distance exhibits greater robustness when significant model misspecification is present. The work further reveals that the Wasserstein distance may introduce spurious rewards under non-informative priors, potentially misleading the design process. By elucidating the trade-offs between these two criteria across varying degrees of model bias, the study provides both theoretical justification and practical guidance for selecting appropriate utility functions in Bayesian experimental design.
This work addresses why the Kullback–Leibler (KL) divergence is uniquely suited for inference by formalizing inference as the selection of a minimal element within a preorder of positive measures, where divergences serve merely as numerical representations. Building on the axiom of reconstruction invariance—which requires that inference outcomes remain unchanged under equivalent problem formulations—the authors show that KL divergence emerges uniquely without invoking additional assumptions. This framework unifies maximum entropy, Bayesian updating, and exponential family estimation, extending classical axiomatic characterizations from finite alphabets to general measurable spaces. By integrating category theory, f-divergence theory, preorder structures, and Čencov’s category of statistical models, the paper establishes a rigorous mathematical foundation wherein inference operators arise naturally as covariant functors, applicable uniformly across both discrete and continuous settings.
This study addresses the lack of comparability in empirical Jensen-Shannon divergence (JSD) estimates used for evaluating synthetic tabular data fidelity, stemming from inconsistent estimation protocols. The authors systematically analyze the behavior of two prevalent JSD estimators under finite-sample conditions: marginal-based and classifier-based estimators. They reveal that marginal estimators neglect variable dependencies and introduce prior-shift bias, while classifier-based estimators suffer from class imbalance and sensitivity in high dimensions. To mitigate these issues, the work proposes a closed-form posterior correction for classifier-based estimators and advocates for explicit declaration of estimation protocols to ensure reproducibility and comparability. Through controlled experiments, benchmark divergences, and real-world synthetic datasets, the study delivers practical guidelines and open-source tools to enable estimator-aware, reliable fidelity evaluation.
This work proposes a class of structure-aware divergences that explicitly incorporate geometric relationships among elements in the support set of probability distributions—addressing a key limitation of classical information-theoretic measures such as Shannon entropy and f-divergences, which disregard structural similarities. By integrating the underlying geometry of the support set, the authors define a structure-aware entropy and derive corresponding Bregman divergences that retain desirable properties of the Kullback–Leibler divergence and Shannon entropy while embedding pairwise similarities directly into the divergence formulation. The approach successfully uncovers structural patterns missed by conventional methods in synthetic clustering tasks, achieves computational efficiency several orders of magnitude higher than optimal transport, and reproduces and extends established findings in applications to economic geography and ecology.