🤖 AI Summary
This study addresses the challenge of balancing privacy protection with predictive efficiency in tabular data, where conventional methods remain inefficient and foundation models lack formal privacy guarantees. We propose PrivTab, a tabular foundation model with differential privacy (DP) intrinsically embedded into its architecture. This approach pioneers the integration of DP within the model structure, enabling “learning to learn under privacy” through pre-training on synthetic datasets. During inference, PrivTab generates provably private summaries via in-context learning combined with DP-enabled forward propagation. Experimental results demonstrate that PrivTab outperforms existing baselines under moderate-to-strong privacy constraints while nearly eliminating member inference leakage risks. Furthermore, the model exhibits well-calibrated predictions and reduces fitting time by four orders of magnitude compared to prior approaches.
📝 Abstract
Tabular data underpin prediction and decision-making in medicine, finance, government and science, but often contain sensitive individual-level information, creating a need for accurate prediction while preserving privacy. Traditional private learning provides formal privacy guarantees, but requires slow dataset-specific optimisation, suffers substantial utility loss under strong privacy, and is often difficult to apply correctly. Tabular foundation models adapt rapidly to new datasets, but existing models lack formal privacy guarantees, and are highly vulnerable to membership-inference attacks, limiting their use on sensitive data. Here we introduce PrivTab, an easy to use tabular foundation model for differentially private classification that embeds a privacy mechanism within its architecture. Pretrained on simulated datasets, PrivTab uses in-context learning to transform sensitive rows into compact, provably private summaries---effectively learning how to learn under privacy. PrivTab outperforms private linear and neural-network baselines under moderate-to-strong privacy, shows negligible membership leakage, maintains well-calibrated predictions under strong privacy, and reduces dataset fitting time by 10,000 times, requiring only a single forward pass. By combining formal privacy, speed, and easy of use, PrivTab brings recent advances in AI to applications where sensitive individual-level data have limited their adoption.