🤖 AI Summary
This work addresses the challenge of validating calibration in regression models—particularly critical in applications like insurance pricing where cross-subsidization across groups must be avoided—under limited and noisy data. It introduces, for the first time, gradient boosting trees into the calibration testing framework and proposes a nonparametric regression-based statistical inference method that leverages expected consistency checks to assess whether a model satisfies necessary conditions for both calibration and self-calibration. The approach provides a practical diagnostic tool for high-dimensional, nonlinear models and demonstrates high statistical power on large-scale insurance datasets, effectively identifying miscalibrated models.
📝 Abstract
The main goal in regression modelling consists in approximating the conditional mean of a response given a set of features. A regression function is said to be calibrated if the resulting mean estimates match the true conditional means for almost every set of features. Aiming for calibration seems not achievable in practice as one typically deals with finite samples of noisy observations. A weaker notion of calibration is auto-calibration, and it means that the expectation of responses being given the same mean estimate matches this estimate. This notion is important, e.g., in insurance pricing as it ensures no cross-subsidization between different price cohorts. In this paper, we show that boosting trees can be used to test necessary conditions for calibration and auto-calibration, respectively. The practical relevance of our approach is supported by a numerical example, in which the proposed tests prove to be very powerful on a large insurance dataset.