Score
Computing, estimating, and manipulating Kullback–Leibler divergence between distributions and using KL-based tilting or importance-weighting analyses to design sampling schemes, justify correctors, and interpret inverse-probability weights in generalized Bayesian frameworks.
This paper addresses the diagnostic ambiguity in multivariate KL divergence arising from the entanglement of marginal mismatch and statistical dependencies. We propose an algebraically exact, fully additive hierarchical decomposition method based on Möbius inversion over the subset lattice—yielding the first closed-form, complete analytic decomposition of KL divergence. The total divergence between a joint distribution and its product reference is rigorously disentangled into a sum of independent marginal mismatch terms and all-order (r-wise) statistical dependency terms, expressed solely in terms of standard Shannon information measures, without approximations or modeling assumptions. The framework unifies higher-order mutual information, total correlation, and information-geometric principles. Numerical experiments confirm machine-precision accuracy across diverse systems, substantially enhancing attribution-based diagnostic capability for divergence sources in machine learning, econometrics, and complex systems analysis.
This study addresses the critical role of utility function selection in Bayesian experimental design under model misspecification. Through theoretical analysis and sequential source inversion experiments, the authors systematically compare the performance of Kullback–Leibler (KL) divergence and Wasserstein distance as utility functions. They find that KL divergence yields faster convergence in the absence of model bias, whereas Wasserstein distance exhibits greater robustness when significant model misspecification is present. The work further reveals that the Wasserstein distance may introduce spurious rewards under non-informative priors, potentially misleading the design process. By elucidating the trade-offs between these two criteria across varying degrees of model bias, the study provides both theoretical justification and practical guidance for selecting appropriate utility functions in Bayesian experimental design.
This study clarifies the informational nature of the Kullback–Leibler (KL) divergence difference (Δ_KL) between discrete empirical distributions and corrects the common misinterpretation of its sign as indicating support set inclusion or coverage. By analytically decomposing the mathematical structure of Δ_KL, the work reveals that it fundamentally serves as an asymmetric contrastive measure of weighted log-probability ratios across categories, rather than a metric of distributional breadth or set containment. The theoretical insights are substantiated through a bibliometric case study examining topic distributions in COVID-19 preprints, offering both intuitive interpretation and empirical validation. This research advances the information-theoretic understanding of asymmetric distributional discrepancies and provides a rigorous foundation—supported by concrete examples—for the proper interpretation of Δ_KL in practical applications.
Gaussian variational approximations suffer from limited expressiveness in high-dimensional Bayesian inference, while KL divergence-based objectives exhibit substantial bias for non-Gaussian posteriors. Method: We propose replacing KL divergence with weighted Fisher divergence as the variational objective to improve local structural fidelity via gradient matching. We systematically introduce weighted Fisher divergence into sparse-precision Gaussian variational inference—breaking the mean-field assumption—and integrate conditional independence structure modeling with stochastic gradient-based sparsity optimization. The framework combines reparameterization, mini-batch objective approximation, and score-based optimization. Contribution/Results: Experiments on logistic regression, generalized linear mixed models, and stochastic volatility models demonstrate significant improvements in posterior gradient estimation accuracy, while retaining computational efficiency and scalability.
The original kernelized Kullback–Leibler (KL) divergence is ill-defined when the supports of compared distributions are disjoint—a fundamental limitation. To address this, the paper proposes a Tikhonov-regularized kernel KL divergence, constructed via covariance operator embeddings in a reproducing kernel Hilbert space (RKHS). This metric is well-defined for arbitrary probability distributions—including discrete, continuous, and mutually singular ones—and provides theoretical guarantees: a bias bound relative to the true KL divergence, finite-sample convergence rates, and a closed-form solution for discrete distributions. Furthermore, the authors formulate a Wasserstein gradient flow optimization framework for the proposed divergence, ensuring theoretical convergence, and design an efficient algorithm applicable to discrete structures such as point clouds. Experiments on point cloud transport tasks demonstrate that the method outperforms existing kernelized and Wasserstein-based approaches, achieving superior stability and robustness.
This work addresses the limitation of traditional Bayesian inference, which relies on global divergences such as the Kullback–Leibler (KL) divergence and fails to capture the local quality behavior of posterior distributions. The authors propose a quality exponent and a regularized extended KL divergence (RE-KL), enabling the first systematic analysis of polynomial and logarithmic decay rates of local small-ball mass under Bayesian updating. Their framework explicitly models parameter-dependent support sets and singular components. On the theoretical front, they establish absolute, relative, and directional inequalities characterizing local mass variation. Empirical results demonstrate that the proposed approach effectively controls local posterior behavior, offering a novel localized analytical framework for Bayesian inference in high-dimensional or singular settings.
This study addresses the challenge of accurate parameter estimation in Bayesian inference when selection bias and systematic bias are present. The authors propose a generalized Bayesian approach that, for the first time, integrates inverse probability weighting (IPW) into the Bayesian framework. By interpreting IPW as a reweighting of the Kullback–Leibler divergence between the model and the true data-generating mechanism, they construct a posterior distribution that simultaneously inherits the bias-correction properties of frequentist IPW and maintains Bayesian coherence. Theoretical analysis establishes favorable asymptotic convergence properties of the proposed posterior. Empirical validation on both simulated data and a large-scale prostate cancer registry dataset—used to predict mortality based on PSA levels—demonstrates the method’s effectiveness, substantially extending the applicability of Bayesian inference to settings with biased observations.
This study addresses the lack of systematic investigation into the statistical properties and testing performance of Jensen–Shannon divergence (JSD) and Kullback–Leibler (KL) divergence in credit risk model monitoring. It derives, for the first time, chi-squared asymptotic reference distributions for both divergences under distributional shift using asymptotic theory, and conducts a comprehensive Monte Carlo simulation to evaluate their Type I error control and statistical power relative to the Population Stability Index (PSI). The results demonstrate that JSD exhibits superior Type I error control, closely attaining the nominal 5% level, yet shows limited power (27%) in small samples (n=200); in contrast, KL divergence and PSI achieve higher power (32%). These findings provide both theoretical grounding and empirical guidance for selecting appropriate divergence metrics in practical model monitoring.
Under model misspecification, conventional Bayesian inference yields posterior credible sets with inadequate uncertainty quantification due to the failure of the information identity. This work addresses the learning rate calibration problem in generalized Bayesian inference by proposing an optimization criterion based on the weighted Fisher divergence. The proposed approach derives a closed-form solution for the learning rate by minimizing the discrepancy between the asymptotic distribution of the generalized posterior and a normal distribution with sandwich covariance. This solution encompasses the Fisher information-matching learning rate as a special case and is shown to be no greater than this benchmark in important scenarios. Theoretical analysis, supported by numerical experiments and real-data applications, demonstrates the superior performance of the proposed method in posterior calibration and uncertainty quantification.
This work investigates the stability of Kullback–Leibler (KL) divergence under Gaussian perturbations, extending beyond prior studies restricted to Gaussian distributions. For any distribution \( P \) with finite second moments and Gaussian distributions \( N_1 \) and \( N_2 \), it establishes that if \( \mathrm{KL}(P \| N_1) \) is large and \( \mathrm{KL}(N_1 \| N_2) \leq \varepsilon \), then \( \mathrm{KL}(P \| N_2) \geq \mathrm{KL}(P \| N_1) - O(\sqrt{\varepsilon}) \). The key contribution lies in generalizing the relaxed triangle inequality for KL divergence from the Gaussian family to arbitrary distributions with finite variance, while proving the optimality of the \( \sqrt{\varepsilon} \) rate. This result provides rigorous theoretical support for out-of-distribution detection in flow-based models under non-Gaussian settings.