๐ค AI Summary
This paper addresses uncertainty quantification for risk-minimizing estimators in machine learning, overcoming limitations of classical approaches that rely on restrictive distributional assumptions and asymptotic theory. We propose the first general-purpose, finite-sample, distribution-free, and frequentist-valid inference framework applicable to *any* risk minimizer. Our method is grounded in the generalized likelihood ratio test, integrated with empirical process analysis and data-driven tuning, and inherently supports anytime-valid inference. Theoretically, it guarantees exact coverage of confidence sets for *all* finite sample sizesโwithout asymptotic approximations. Empirically, it consistently outperforms classical asymptotic methods across diverse tasks, demonstrating both high accuracy and strong robustness. This work establishes a new paradigm for model-agnostic statistical inference.
๐ Abstract
A common goal in statistics and machine learning is estimation of unknowns. Point estimates alone are of little value without an accompanying measure of uncertainty, but traditional uncertainty quantification methods, such as confidence sets and p-values, often require strong distributional or structural assumptions that may not be justified in modern problems. The present paper considers a very common case in machine learning, where the quantity of interest is the minimizer of a given risk (expected loss) function. For such cases, we propose a generalization of the recently developed universal inference procedure that is designed for inference on risk minimizers. Notably, our generalized universal inference attains finite-sample frequentist validity guarantees under a condition common in the statistical learning literature. One version of our procedure is also anytime-valid in the sense that it maintains the finite-sample validity properties regardless of the stopping rule used for the data collection process, thereby providing a link between safe inference and fast convergence rates in statistical learning. Practical use of our proposal requires tuning, and we offer a data-driven procedure with strong empirical performance across a broad range of challenging statistical and machine learning examples.