🤖 AI Summary
This study addresses the lack of systematic investigation into the statistical properties and testing performance of Jensen–Shannon divergence (JSD) and Kullback–Leibler (KL) divergence in credit risk model monitoring. It derives, for the first time, chi-squared asymptotic reference distributions for both divergences under distributional shift using asymptotic theory, and conducts a comprehensive Monte Carlo simulation to evaluate their Type I error control and statistical power relative to the Population Stability Index (PSI). The results demonstrate that JSD exhibits superior Type I error control, closely attaining the nominal 5% level, yet shows limited power (27%) in small samples (n=200); in contrast, KL divergence and PSI achieve higher power (32%). These findings provide both theoretical grounding and empirical guidance for selecting appropriate divergence metrics in practical model monitoring.
📝 Abstract
Divergence measures are essential tools for detecting distributional shifts in model monitoring, particularly crucial given the volatility of financial data. While the Population Stability Index is the most widely used measure, Jensen-Shannon Divergence and Kullback-Leibler Divergence offer distinct advantages. Jensen-Shannon Divergence handles mixture models, addresses zero-binning problems, and is symmetric, while Kullback-Leibler Divergence excels in Bayesian model comparison.
This study extends the work of Yurdakul and Naranjo (2020) with two primary contributions. First, we derive the statistical properties and chi-square benchmark values for Jensen-Shannon Divergence and Kullback-Leibler Divergence. Second, we demonstrate their applicability by detecting distributional changes in credit default probabilities from Merton, Merton with jump, and stochastic volatility with jump models.
Our results establish that Jensen-Shannon Divergence and Kullback-Leibler Divergence follow chi-square distributions and reveal important practical trade-offs. Jensen-Shannon Divergence exhibits superior Type I error control, maintaining rejection rates closest to 5%, thereby minimizing false positives. However, this conservatism reduces statistical power at small samples (27% versus 32% for Population Stability Index and Kullback-Leibler Divergence at n = m = 200), requiring larger samples for reliable detection. This trade-off enables practitioners to select measures based on whether minimizing false alarms or maximizing detection sensitivity is the priority.