Benchmarking Goodness-of-Fit and Calibration Algorithms for Logistic Regression Classifiers: A Large-Scale Simulation Study under Sparse Data

📅 2026-07-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the well-documented limitations of traditional goodness-of-fit and calibration tests—such as Pearson’s chi-square and Hosmer–Lemeshow—in settings with sparse data and continuous predictors, where they often exhibit poor control of Type I error or low power. Through large-scale, reproducible simulations (10,000 replicates per configuration), the authors systematically evaluate over twenty tests across five covariate distributions and four model misspecification scenarios, complemented by empirical validation on a low birth weight dataset. They propose a unified classification framework and provide the first comprehensive benchmark under sparsity, implemented in the open-source R package ebrahim.gof. Results identify McCullagh, Osius–Rojek, le Cessie–van Houwelingen, Stute–Zhu, and GiViTI tests as offering superior Type I error control and high power, substantially outperforming Hosmer–Lemeshow; their integration with calibration plots effectively detects omitted interaction effects.
📝 Abstract
Binary logistic regression is among the most widely used classification algorithms, yet a classifier is only trustworthy if its predicted probabilities are well calibrated. The classical checks -- the Pearson chi-square and deviance statistics -- break down precisely in the modern setting where predictors are continuous and the data are sparse (one covariate pattern per observation). Four decades of research have produced dozens of alternative goodness-of-fit and calibration algorithms, yet practitioners still default to the Hosmer-Lemeshow test because it ships with their software. This paper provides a unified taxonomy and a large-scale, reproducible simulation benchmark; more than twenty tests are implemented in the open-source R package ebrahim.gof. We evaluate them across five covariate distributions and four misspecification scenarios, with 10,000 replications each, measuring both Type I error and power. Several classical tests prove liberal, rejecting correct models far too often, while others have little power. A compact core -- McCullagh, Osius-Rojek, le Cessie-van Houwelingen, Stute-Zhu, and the GiViTI calibration test -- delivers the best balance of correct size and high power, and is consistently more powerful than the ubiquitous Hosmer-Lemeshow test. A low-birth-weight application reinforces the point: a model with omitted interactions slips past nearly every test, exposed only by pairing sensitive tests with a calibration (reliability) curve. We translate these findings into practical, evidence-based guidance for assessing logistic regression fit.
Problem

Research questions and friction points this paper is trying to address.

goodness-of-fit
calibration
logistic regression
sparse data
Hosmer-Lemeshow test
Innovation

Methods, ideas, or system contributions that make the work stand out.

goodness-of-fit
calibration
logistic regression
sparse data
simulation benchmark
💼 Related Jobs
No related jobs found.
E
Ebrahim Khaled Ebrahim
Department of Applied Statistics, Faculty of Business, Alexandria University, Alexandria, Egypt
A
Ahmed El-Kotory
Department of Applied Statistics, Faculty of Business, Alexandria University, Alexandria, Egypt