🤖 AI Summary
Membership inference attacks (MIAs) threaten the privacy of model training data, yet existing approaches rely heavily on shadow models or extensive query budgets, limiting practical applicability. This paper proposes an efficient and interpretable MIA framework based on Gaussian process meta-modeling: it requires only a single set of posterior metrics from the target model—such as prediction accuracy, entropy, gradient norm, and dataset-level statistical features—to construct a calibrated, uncertainty-aware binary classifier, eliminating the need for shadow training or repeated queries. Our core innovation lies in formulating membership inference as an uncertainty-quantified meta-learning task, enabling high attack accuracy, strong generalization across diverse domains, and intrinsic interpretability. We validate the method on synthetic data, fraud detection benchmarks, CIFAR-10, and WikiText-2, demonstrating consistent performance gains and significantly improved practicality and deployment feasibility.
📝 Abstract
Membership inference attacks (MIAs) test whether a data point was part of a model's training set, posing serious privacy risks. Existing methods often depend on shadow models or heavy query access, which limits their practicality. We propose GP-MIA, an efficient and interpretable approach based on Gaussian process (GP) meta-modeling. Using post-hoc metrics such as accuracy, entropy, dataset statistics, and optional sensitivity features (e.g. gradients, NTK measures) from a single trained model, GP-MIA trains a GP classifier to distinguish members from non-members while providing calibrated uncertainty estimates. Experiments on synthetic data, real-world fraud detection data, CIFAR-10, and WikiText-2 show that GP-MIA achieves high accuracy and generalizability, offering a practical alternative to existing MIAs.