🤖 AI Summary
This study addresses the limitations of leave-one-covariate-out (LOCO) feature importance inference for black-box models, which typically relies on data splitting or refitting and suffers degraded predictive performance in high-dimensional sparse settings. We propose an adaptive mini-patch ensemble framework based on LOCO importance. This method introduces a novel adaptive feature sampling mechanism that leverages leave-two-out perturbation bounds to handle complex feature dependencies, while integrating leave-one-covariate-out extrapolation with subsampling stability analysis. Experimental results demonstrate that the proposed framework achieves asymptotically valid feature inference without requiring data splitting. Furthermore, it significantly outperforms existing methods in predictive accuracy, inferential validity, and stability across both synthetic and real-world datasets.
📝 Abstract
As black-box machine learning models become increasingly common, extracting interpretations with uncertainty quantification has become a critical challenge. One popular type of interpretation is leave-one-covariate-out (LOCO) feature importance, while prior LOCO inference methods often require data-splitting or model-refitting. A recent ensemble framework, LOCO-MP, addresses these challenges using minipatches that subsample both observations and features, but massive feature subsampling can hurt prediction in high-dimensional sparse settings. Motivated by this limitation, we consider minipatch ensembles with adaptive feature sampling guided by LOCO importance, and propose LOCO-AdaMP, which enables free LOCO inference for the resulting adaptive minipatch ensemble. We show that LOCO-AdaMP yields substantially improved predictive models while retaining asymptotically valid feature importance inference without data-splitting, despite the complex dependence between the adaptive sampling distribution and the LOCO importance statistics. Our analysis relies on a careful leave-two-out perturbation bound for the iteratively updated sampling probabilities together with the stability of LOCO scores induced by observation subsampling. Empirical results on synthetic and real datasets demonstrate advantages of LOCO-AdaMP over existing methods in predictive performance, inferential power, and stability. Overall, LOCO-AdaMP provides a flexible ensemble framework (agnostic to base models) that delivers both strong predictive performance and asymptotically valid, powerful feature importance inference for regression.