Score
Designs, fits, and evaluates generalized linear classification models using a logit or softmax link to predict binary or multiclass outcomes, producing calibrated class-probability estimates and interpretable linear decision boundaries; incorporates observation weights and interaction terms where needed and assesses coefficient significance and predictive performance (including cross-validation and accuracy/ROC-style measures).
Existing methods struggle to effectively and interpretably evaluate and recalibrate the probabilistic outputs of black-box multiclass models without internal access. This work proposes the Multiclass Linear Log-Odds (MCLLO) recalibration framework, which, for the first time, enables calibration assessment and adjustment using only predicted probabilities from a single model. By modeling calibration through a linear transformation in log-odds space and employing a likelihood ratio test for direct calibration evaluation, MCLLO achieves both interpretability and broad applicability. Experiments across three real-world domains—image classification, obesity analysis, and ecological modeling—demonstrate that MCLLO matches or outperforms four state-of-the-art recalibration methods in terms of calibration performance.
Generalized linear models (GLMs) are widely used in actuarial modeling, yet their nonlinear link functions render existing fairness diagnostics—typically grounded in linear intuition—inapplicable. This work proposes the first fairness decomposition framework tailored to GLMs by extending the Wasserstein barycenter criterion to the distributional level and integrating moment decomposition with curvature analysis of the inverse link function. The resulting interpretable “four-channel plus dual-curvature” effect decomposition yields explicit formulas for logistic, Poisson, and Tweedie models. Empirical application to healthcare expenditure data demonstrates that the method effectively disentangles sources of prediction disparities, including the direct effect of sensitive attributes, mediation through proxy variables, differences in covariance structure, and amplification effects induced by nonlinear link coupling, thereby offering a practical tool for fairness auditing in actuarial practice.
This study addresses the well-documented limitations of traditional goodness-of-fit and calibration tests—such as Pearson’s chi-square and Hosmer–Lemeshow—in settings with sparse data and continuous predictors, where they often exhibit poor control of Type I error or low power. Through large-scale, reproducible simulations (10,000 replicates per configuration), the authors systematically evaluate over twenty tests across five covariate distributions and four model misspecification scenarios, complemented by empirical validation on a low birth weight dataset. They propose a unified classification framework and provide the first comprehensive benchmark under sparsity, implemented in the open-source R package ebrahim.gof. Results identify McCullagh, Osius–Rojek, le Cessie–van Houwelingen, Stute–Zhu, and GiViTI tests as offering superior Type I error control and high power, substantially outperforming Hosmer–Lemeshow; their integration with calibration plots effectively detects omitted interaction effects.
This study addresses parameter estimation in binary choice models under severe class imbalance by proposing a support vector machine (SVM)-based approach that adjusts class weights to achieve consistent estimation of slope parameters under the linear conditional mean assumption, subsequently recovering the intercept. Theoretical analysis demonstrates that the proposed method is asymptotically equivalent to logistic regression and possesses desirable consistency properties. Finite-sample simulations indicate that its performance is comparable to that of quasi-maximum likelihood estimation (QMLE), with each method exhibiting relative strengths depending on the context. This work thus offers a novel theoretical perspective and a practical tool for modeling binary choice outcomes in imbalanced data settings.
This work investigates the intrinsic mechanisms by which logit regularization—such as label smoothing—enhances model calibration and generalization. Through theoretical analysis of convex penalties in logit space within linear classification, we uncover an implicit bias that induces logits to cluster around sample-specific targets. We establish, for the first time, that this clustering behavior aligns the weight vectors precisely with the Fisher linear discriminant direction. Under a signal-plus-noise model, this alignment substantially reduces sample complexity, triggers grokking phenomena, and improves noise-robust generalization. In the small-noise regime, the method achieves a halving of sample complexity while ensuring stable generalization, thereby offering a deeper understanding of the foundational principles underlying logit regularization.
This study addresses the problem of assessing whether observed data are “sufficiently close” to a binary generalized linear model—such as logistic regression—with fully categorical covariates, rather than requiring exact model fit. To this end, the authors propose a formal equivalence testing framework based on minimum distance methodology. The approach leverages both asymptotic theory and bootstrap procedures to compute critical values, thereby filling a critical gap left by conventional goodness-of-fit tests, which are ill-suited for evaluating practical equivalence. Through extensive simulation studies and analyses of two real-world datasets, the proposed method demonstrates strong finite-sample performance and practical utility, offering a robust tool for model adequacy assessment in applied settings.
This work addresses the potential nonexistence of a global optimum in linear ensembles of multiple binary classifiers by proposing a theoretical framework grounded in truth-table logical structuring and equivalence class partitioning, which establishes sufficient conditions for the existence of a convexified empirical risk minimizer. By introducing a multidimensional generalization of classification-calibrated loss functions and the notion of φ-frontiers, the study analyzes solution stability in relation to data quality. Under exponential (Boost) and logistic (Logit) losses, the authors derive, for the first time, explicit closed-form expressions for the optimal ensemble weights and fully characterize all solution regimes in the three-classifier setting. This approach circumvents iterative optimization, thereby substantially enhancing both the interpretability and computational efficiency of ensemble models.
This study addresses the limitations of covariate logistic models in latent class analysis, specifically their difficulty in capturing complex interactions and providing sufficient interpretability. To overcome these challenges, this work proposes a novel framework that directly integrates decision trees into latent class analysis. By leveraging interpretable tree structures combined with pruning and binary splitting techniques, the method effectively models covariate effects without requiring additional assumptions. Empirical validation demonstrates the approach's efficacy, yielding highly interpretable classification paths. Consequently, this research significantly expands the methodological toolkit for latent class analysis, establishing a new paradigm for handling complex covariate relationships while enhancing model transparency and practical utility.
This study addresses variable selection for binary-response generalized linear models in high-dimensional settings by proposing a novel method, termed Boosting with Multiple Testing correction (BMT), which integrates multiple testing adjustment within a nonlinear boosting framework. At each iteration, BMT incorporates only the covariate exhibiting the strongest conditional significance while progressively constructing a sparse model through rigorous control of the multiplicity-induced error rate. The approach uniquely embeds a formal multiple hypothesis testing procedure into the boosting paradigm, offering theoretical guarantees of selection consistency and oracle properties for parameter estimation. Empirical evaluations demonstrate that BMT outperforms existing methods in both variable selection accuracy and estimation precision, and it achieves superior out-of-sample predictive performance in forecasting U.S. inflation.
This study addresses the challenge of estimating training sample requirements and performance ceilings in machine learning by proposing Seer, a system that leverages learning curve modeling to predict both the sample size needed for target accuracy and the maximum achievable performance. Methodologically, Seer constructs an optimally constrained statistical model and employs an efficient maximum likelihood estimation algorithm, integrating nonlinear regression with iterative optimization for parameter inference. Extensive experiments across nearly one hundred real-world datasets spanning three domains demonstrate that the system accurately characterizes learning dynamics and forecasts model performance. Notably, its predictive accuracy approaches theoretical limits, thereby providing a reliable quantitative foundation for computational resource planning.