🤖 AI Summary
This paper addresses the challenge of quantifying uncertainty in Apparent Shannon Mutual Information (ASI)—a key metric for predictive quality assessment—whose estimation is hindered by the long-range dependencies, asymmetry, and heavy-tailed nature of the joint distribution $j(x,y)$, rendering classical mutual information inapplicable. We propose the first Bayesian ASI assessment framework: (1) We formally prove, for the first time, that ASI is the unique predictive quality measure satisfying both predictive and discriminative independence axioms; (2) We introduce a Dirichlet process mixture of skewed Student’s $t$-distributions, enabling the first Bayesian modeling of heavy-tailed, asymmetric $j(x,y)$ and principled uncertainty quantification of ASI; (3) We validate the framework on prostate cancer recurrence time prediction data, demonstrating its generality, robustness, and computational feasibility—particularly in real-world settings where $j(x,y)$ lacks closed-form expression.
📝 Abstract
Shannon defined the mutual information between two variables. We illustrate why the true mutual information between a variable and the predictions made by a prediction algorithm is not a suitable measure of prediction quality, but the apparent Shannon mutual information (ASI) is; indeed it is the unique prediction quality measure with either of two very different lists of desirable properties, as previously shown by de Finetti and other authors. However, estimating the uncertainty of the ASI is a difficult problem, because of long and non-symmetric heavy tails to the distribution of the individual values of $j(x,y)=logfrac{Q_y(x)}{P(x)}$ We propose a Bayesian modelling method for the distribution of $j(x,y)$, from the posterior distribution of which the uncertainty in the ASI can be inferred. This method is based on Dirichlet-based mixtures of skew-Student distributions. We illustrate its use on data from a Bayesian model for prediction of the recurrence time of prostate cancer. We believe that this approach is generally appropriate for most problems, where it is infeasible to derive the explicit distribution of the samples of $j(x,y)$, though the precise modelling parameters may need adjustment to suit particular cases.