🤖 AI Summary
This study clarifies the informational nature of the Kullback–Leibler (KL) divergence difference (Δ_KL) between discrete empirical distributions and corrects the common misinterpretation of its sign as indicating support set inclusion or coverage. By analytically decomposing the mathematical structure of Δ_KL, the work reveals that it fundamentally serves as an asymmetric contrastive measure of weighted log-probability ratios across categories, rather than a metric of distributional breadth or set containment. The theoretical insights are substantiated through a bibliometric case study examining topic distributions in COVID-19 preprints, offering both intuitive interpretation and empirical validation. This research advances the information-theoretic understanding of asymmetric distributional discrepancies and provides a rigorous foundation—supported by concrete examples—for the proper interpretation of Δ_KL in practical applications.
📝 Abstract
When empirical objects are represented as discrete probability distributions, within-distribution summaries such as Shannon entropy and Hill-type diversity indices describe how probability mass is spread inside each object, while Kullback-Leibler (KL) divergence provides pairwise asymmetric information. This note focuses on the KL difference $Δ_{\mathrm{KL}}(p,q)=D_{\mathrm{KL}}(p|q)-D_{\mathrm{KL}}(q|p)$. Although $Δ_{\mathrm{KL}}$ can add information beyond within-distribution summaries and symmetric overlap, its sign does not, by itself, establish support inclusion, coverage, or breadth. It is better understood as a weighted category-wise log-ratio contrast reflecting asymmetric probability-mass placement. The point becomes clear once the definition is written out. The aim of this note is therefore to present it in a compact, example-based form, together with a descriptive bibliometric illustration based on COVID-19-related preprint-server topic distributions.