Efficient Provably Private Classification with a Tabular Foundation Model

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of balancing privacy protection with predictive efficiency in tabular data, where conventional methods remain inefficient and foundation models lack formal privacy guarantees. We propose PrivTab, a tabular foundation model with differential privacy (DP) intrinsically embedded into its architecture. This approach pioneers the integration of DP within the model structure, enabling “learning to learn under privacy” through pre-training on synthetic datasets. During inference, PrivTab generates provably private summaries via in-context learning combined with DP-enabled forward propagation. Experimental results demonstrate that PrivTab outperforms existing baselines under moderate-to-strong privacy constraints while nearly eliminating member inference leakage risks. Furthermore, the model exhibits well-calibrated predictions and reduces fitting time by four orders of magnitude compared to prior approaches.
📝 Abstract
Tabular data underpin prediction and decision-making in medicine, finance, government and science, but often contain sensitive individual-level information, creating a need for accurate prediction while preserving privacy. Traditional private learning provides formal privacy guarantees, but requires slow dataset-specific optimisation, suffers substantial utility loss under strong privacy, and is often difficult to apply correctly. Tabular foundation models adapt rapidly to new datasets, but existing models lack formal privacy guarantees, and are highly vulnerable to membership-inference attacks, limiting their use on sensitive data. Here we introduce PrivTab, an easy to use tabular foundation model for differentially private classification that embeds a privacy mechanism within its architecture. Pretrained on simulated datasets, PrivTab uses in-context learning to transform sensitive rows into compact, provably private summaries---effectively learning how to learn under privacy. PrivTab outperforms private linear and neural-network baselines under moderate-to-strong privacy, shows negligible membership leakage, maintains well-calibrated predictions under strong privacy, and reduces dataset fitting time by 10,000 times, requiring only a single forward pass. By combining formal privacy, speed, and easy of use, PrivTab brings recent advances in AI to applications where sensitive individual-level data have limited their adoption.
Problem

Research questions and friction points this paper is trying to address.

tabular data
differential privacy
foundation model
membership inference attack
private classification
Innovation

Methods, ideas, or system contributions that make the work stand out.

Tabular Foundation Model
Differential Privacy
In-Context Learning
Membership Inference Attack
PrivTab
🔎 Similar Papers
No similar papers found.
T
Talal Alrawajfeh
Department of Computer Science, University of Helsinki, Helsinki, Finland
Cristiana Diaconu
Cristiana Diaconu
PhD Student, University of Cambridge
probabilistic machine learningdiffusion modelsneural processes
O
Ossi Räisä
CISPA Helmholtz Center for Information Security, Saarbrücken, Germany
S
Sebastian Rodriguez Beltran
Current address: Vienna, Austria
Y
Yuan He
Department of Computer Science, University of Helsinki, Helsinki, Finland
John Bronskill
John Bronskill
University of Cambridge
R
Richard E. Turner
Department of Engineering, University of Cambridge, Cambridge, United Kingdom
Antti Honkela
Antti Honkela
Professor, University of Helsinki
Machine LearningDifferential PrivacyBayesian InferenceBioinformatics#UnivHelsinkiCS