Differentially Private Mixing of Public Datasets Improves Private Learning

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the degradation of model utility in differentially private training and the challenge of determining optimal mixing ratios for public data pretraining. To this end, it proposes the first dynamic public data mixing pipeline under privacy constraints. By privately learning a low-dimensional linear regression model, the approach automatically derives the optimal mixing strategy across multiple public data sources to enhance downstream pretraining. Empirical evaluations demonstrate that this method improves macro-AUC by 0.037 (a 22.8% relative gain) on the NIH chest X-ray classification task and reduces perplexity by 16% on ENRON language modeling. These results indicate substantial improvements in model utility within differentially private settings.
📝 Abstract
Many machine learning applications involve sensitive data and therefore require training under differential privacy (DP). However, DP training often degrades model utility. In some cases, first pre-training the model on"public"data before finetuning with DP on the sensitive data can reduce the drop in utility. However, the success of this depends on how relevant the selected public dataset is to the sensitive data. We introduce the first pipeline that privately learns the mixture of several public datasets to pretrain on for a given sensitive downstream task. Our key insight is that we can privately find the best mixture of multiple public datasets by privately learning a low-dimensional linear model. We tested our method on the NIH dataset for X-ray classification and the ENRON email dataset for language modeling. Applying our method to find tailored mixtures of X-ray datasets to pretrain on for diseases in the NIH ChestX-ray14 dataset, we improved macro AUC by up to 0.037 across privacy budgets compared to the baselines, with gains as large as +22.8% relative AUC on Cardiomegaly at $\epsilon=1$. For DP training on the ENRON dataset, pre-training on our mixture of The Common Pile (a collection of public-domain text datasets) decreased test perplexity by 16% relative to the baseline mixtures.
Problem

Research questions and friction points this paper is trying to address.

Differential Privacy
Private Learning
Public Dataset Mixing
Pre-training
Model Utility
Innovation

Methods, ideas, or system contributions that make the work stand out.

Differential Privacy
Public Dataset Mixing
Private Pre-training
Low-dimensional Linear Model
Utility Improvement
🔎 Similar Papers
No similar papers found.