🤖 AI Summary
Traditional statistical inference requires problem-specific estimators, resulting in cumbersome workflows and substantial computational overhead. To address this limitation, this work proposes TabCon, a system that introduces a novel paradigm for amortized confidence interval construction based on tabular foundation models. The proposed method employs a sparse mixture-of-experts architecture combined with reinforcement learning-based post-training for calibration, enabling the generation of confidence intervals through a single forward pass. Experimental results demonstrate that TabCon achieves near-nominal coverage rates and shorter interval lengths across benchmark datasets. Furthermore, it delivers a 50-fold improvement in inference speed compared to classical bootstrap methods, offering an efficient new solution for scalable statistical inference.
📝 Abstract
For decades, statistical inference has largely been developed one problem at a time. Given a scientific target, such as a treatment effect or a regression function, statisticians design a problem-specific estimator together with a procedure for quantifying its uncertainty. This paper proposes a different paradigm. We focus on a classical problem in statistical inference, confidence interval construction, and develop TabCon, an amortized inference system built on a tabular foundation model that produces confidence intervals for new datasets through a simple forward pass. The key methodological ingredients of TabCon are a sparse mixture-of-experts architecture and reinforcement-learning-based post-training that calibrate the resulting confidence intervals to a desired coverage level. Across a wide range of benchmark datasets, TabCon attains near-nominal coverage while producing short confidence intervals. At inference time, it also offers considerably greater computational efficiency, running 50 times faster than the classical bootstrap procedure, even when the latter uses only 50 bootstrap samples.