High-Dimensional Asymptotics and Dataset Selection for Private Transfer Learning

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses key challenges in privacy-preserving transfer learning, where evaluating the utility of external data solely through summary statistics, mitigating negative transfer, and balancing privacy noise against predictive performance remain difficult. To this end, the authors develop an error theory grounded in high-dimensional asymptotic analysis and weighted ridge regression that captures the interplay among sample size, covariance structure, model shift, and privacy noise. By integrating zero-concentrated differential privacy (zCDP) with certainty equivalence techniques, they construct a test error expression from aggregated statistics to guide hyperparameter optimization. This framework enables utility assessment and decision support without requiring access to raw data. Extensive experiments on both synthetic and real-world datasets validate the theoretical findings, establishing a computationally tractable foundation for data selection under strict privacy constraints.
📝 Abstract
To commit to buying external data or participate in collaborative learning, one must decide whether the additional data will improve prediction enough to justify the cost. This comes with several challenges: (i) the decision often relies only on aggregated statistics available publicly, rather than individual-level data; (ii) covariate and model shifts can induce negative transfer, so the additional data deteriorates rather than improves performance; (iii) if the data is sensitive, its privatization requires the injection of noise, which can also offset the benefit of a larger sample size. In this paper, we model the problem of dataset selection through high-dimensional regression with multiple heterogeneous sources and a weighted ridge estimator. Our approach uses only summary statistics and it gives privacy guarantees either on labels only or jointly on features and labels, in terms of $ρ$-zero-concentrated differential privacy. The main technical contribution is a deterministic equivalent of the test error, which captures the interactions between sample size, covariance structure, model shift, regularization and privacy noise. Our theory allows to optimize hyperparameters (weights and ridge regularizers) and, more broadly, to decide when private external datasets are useful without accessing the data itself but only relying on population-level quantities. This provides a theoretically tractable foundation for private transfer learning, which we support via experiments on both synthetic and real-world datasets.
Problem

Research questions and friction points this paper is trying to address.

Private Transfer Learning
Dataset Selection
Negative Transfer
High-Dimensional Asymptotics
Differential Privacy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Private Transfer Learning
High-Dimensional Asymptotics
Dataset Selection
Differential Privacy
Deterministic Equivalent
🔎 Similar Papers
F
Filip Kovačević
Institute of Science and Technology Austria
Edwige Cyffers
Edwige Cyffers
Postdoctoral Fellow, ISTA
Machine LearningTrustworthy Machine LearningFederated LearningPrivacy
S
Stefano Sarao Mannelli
Computer Science and Engineering Department, Chalmers University of Technology and University of Gothenburg; School of Computer Science and Applied Mathematics, University of the Witwatersrand
Marco Mondelli
Marco Mondelli
Professor, IST Austria
Machine LearningData ScienceCoding TheoryInformation theory