A subsampling approach for large data sets when the Generalised Linear Model is potentially misspecified

📅 2025-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
For generalized linear models (GLMs) under potential model misspecification in massive-data settings, existing subsampling methods suffer from low inferential efficiency and poor robustness. To address this, we propose a robust subsampling framework guided by prediction mean squared error (PMSE). Unlike conventional approaches assuming correct model specification, our method explicitly accommodates GLM misspecification by dynamically allocating sampling probabilities based on local PMSE estimates—thereby jointly accounting for data informativeness and model uncertainty. Theoretical analysis establishes consistency and asymptotic normality of the resulting estimator. Extensive simulations and real-world large-scale experiments demonstrate that our approach significantly outperforms state-of-the-art subsampling methods under model deviation, achieving a superior trade-off between computational efficiency and statistical accuracy. Consequently, it enhances both the reliability and practicality of statistical inference in large-scale, imperfectly specified modeling scenarios.

Technology Category

Machine Learning: Ensemble MethodsReasoning under Uncertainty: Relational Probabilistic ModelsIntelligent Robots: State Estimation

Application Category

Web Mining and Content Analysis: Robustness and generalizability of Web mining methodsUser Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalizationGraph Algorithms and Modeling for the Web: Representation, reconstruction, and subgraph or motif discovery in Web-related graphs
📝 Abstract
Subsampling is a computationally efficient and scalable method to draw inference in large data settings based on a subset of the data rather than needing to consider the whole dataset. When employing subsampling techniques, a crucial consideration is how to select an informative subset based on the queries posed by the data analyst. A recently proposed method for this purpose involves randomly selecting samples from the large dataset based on subsampling probabilities. However, a major drawback of this approach is that the derived subsampling probabilities are typically based on an assumed statistical model which may be difficult to correctly specify in practice. To address this limitation, we propose to determine subsampling probabilities based on a statistical model that we acknowledge may be misspecified. To do so, we propose to evaluate the subsampling probabilities based on the Mean Squared Error (MSE) of the predictions from a model that is not assumed to completely describe the large dataset. We apply our subsampling approach in a simulation study and for the analysis of two real-world large datasets, where its performance is benchmarked against existing subsampling techniques. The findings suggest that there is value in adopting our approach over current practice.
Problem

Research questions and friction points this paper is trying to address.

Addresses model misspecification in subsampling for large datasets
Develops subsampling probabilities using prediction Mean Squared Error
Improves inference accuracy when statistical models are potentially wrong
Innovation

Methods, ideas, or system contributions that make the work stand out.

Subsampling with misspecified model probabilities
MSE-based evaluation for prediction accuracy
Benchmarked performance against existing techniques
🔎 Similar Papers
A
Amalan Mahendran
School of Mathematical Sciences, Queensland University of Technology, Brisbane, 4000, Queensland, Australia. Centre for Data Science, Queensland University of Technology, Brisbane, 4000, Queensland, Australia.
H
Helen Thompson
School of Mathematical Sciences, Queensland University of Technology, Brisbane, 4000, Queensland, Australia. Centre for Data Science, Queensland University of Technology, Brisbane, 4000, Queensland, Australia.
J
James M. McGree
School of Mathematical Sciences, Queensland University of Technology, Brisbane, 4000, Queensland, Australia. Centre for Data Science, Queensland University of Technology, Brisbane, 4000, Queensland, Australia.