Score
Designs and evaluates membership-inference attacks that leverage a model's attention distributions or concentration patterns to score whether specific inputs or context examples were part of the training data, including variants that operate without shadow models. Work comprises building attention-based scoring functions or classifiers and analyzing their sensitivity and effectiveness relative to confidence-based and other membership-inference methods.
This work addresses the privacy and intellectual property risks posed by large language models (LLMs) memorizing training data, a vulnerability inadequately exploited by existing membership inference attacks. The authors propose a novel membership inference method that leverages the Transformer’s self-attention mechanism: by analyzing information flow patterns across multiple attention heads and integrating perturbation-induced divergence metrics, they construct a highly discriminative classifier. This approach is the first to utilize attention signals for membership inference, revealing that while attention mechanisms enhance model interpretability, they simultaneously exacerbate privacy leakage. Evaluated on benchmarks such as WikiMIA-32, the method significantly outperforms prior techniques—achieving a 0.996 ROC AUC and 87.9% TPR@1%FPR on Llama2-13b—and demonstrates strong generalization across datasets and model architectures, thereby substantially strengthening training data extraction attacks.
This paper addresses a more realistic distribution shift scenario in membership inference attacks—where the auditor/adversary has access to only a subset of classes (e.g., 90% of classes missing)—causing conventional shadow-model-based attacks, which rely on complete background class distributions, to suffer severe performance degradation. We propose the first systematic theoretical framework modeling distribution shift under class dropout. Innovatively, we design a quantile regression–based attack that does not require alignment with the target model’s training distribution, and we theoretically characterize its feasibility and performance bounds. On unseen classes of CIFAR-100, our method achieves a true positive rate (TPR) 11× higher than standard shadow-model attacks; on ImageNet with 90% of classes removed, it retains a non-zero TPR—demonstrating significantly improved robustness and generalization under extreme distribution shift.
This paper reveals that existing language model membership inference attacks remain highly vulnerable to data poisoning—even under semantic neighborhood definitions—where “members” are extended to semantically similar samples. Attackers can significantly degrade inference accuracy using carefully crafted poisoned inputs. Method: The authors theoretically establish a fundamental trade-off between membership inference accuracy and poisoning robustness, and propose the first provably effective semantic-aware poisoning method. It jointly optimizes gradient manipulation and semantic similarity constraints, preserving model utility while degrading inference performance. Contribution/Results: Empirical evaluation demonstrates that state-of-the-art membership inference methods fall below the 50% random baseline under the proposed attack. This exposes a structural vulnerability in real-world deployments of membership inference, providing critical theoretical insights for trustworthy AI assessment and establishing a practical, semantically grounded benchmark for adversarial evaluation.
This work identifies a fundamental flaw in current evaluation methodologies for membership inference (MI) attacks against foundation models: member and non-member samples are typically drawn from disparate distributions, causing standard metrics—such as AUC—to reflect data distribution shift rather than genuine model memorization or privacy leakage. To address this, the authors propose the first model-agnostic “blind baseline” for MI—namely, zero-knowledge classifiers leveraging text statistics or embedding distances—requiring no access to the target model. They systematically evaluate it across eight public MI benchmark datasets. Results show that this blind baseline consistently achieves significantly higher AUC than state-of-the-art MI attacks on all datasets, with remarkable cross-dataset stability. The study demonstrates that prevailing MI evaluation paradigms primarily capture distributional discrepancies—not true membership information leakage—thereby challenging their validity as privacy assessment tools and providing both theoretical grounding and an empirical benchmark for developing more robust privacy evaluation frameworks.
Current membership inference attack (MIA) research against large language models (LLMs) suffers from severe methodological flaws: mainstream studies rely on posteriorly constructed datasets, inducing substantial distributional shifts between member and non-member samples—thereby distorting privacy leakage assessments. Method: The authors first quantify such biases across six widely used MIA benchmarks; then propose a reproducible evaluation paradigm comprising four key components—randomized test-splitting, unique sequence injection, random fine-tuning, and posterior control—and establish both sequence-level and document-level MIA benchmarking frameworks. Results: Experiments demonstrate that most reported MIA success stems from dataset construction artifacts rather than genuine model memorization. This work establishes rigorous, trustworthy evaluation principles for LLM MIAs and provides a robust methodological foundation for studying LLM privacy and memorization.
This work addresses the privacy risks in tabular foundation models, which rely on in-context examples containing sensitive data during inference, potentially leaking membership information through their attention mechanisms. The study is the first to identify and quantify membership signals embedded in attention patterns and introduces Attention-based Membership Inference Attack (AMIA), a novel attack that operates without requiring shadow models. To mitigate this threat, the authors propose a defense mechanism that applies k-anonymity–inspired de-uniquification to context keys at inference time, selectively protecting only high-risk queries. Experimental results demonstrate that AMIA outperforms conventional confidence-based attacks by an average of 7.7% in attack performance. The proposed defense reduces membership leakage by 50% under AMIA (and by 25% against confidence-based attacks) while incurring only a modest 3.9% drop in model utility.
This study challenges the prevailing assumption that a model’s generalization capability is unrelated to its vulnerability to membership inference attacks. Through large-scale controlled experiments, we systematically investigate the true relationship between these two properties by training over a thousand models within a unified framework, incorporating standard generalization techniques such as data augmentation and early stopping, and evaluating multiple membership inference attack strategies. Our work provides the first empirical evidence that stronger generalization directly suppresses the success of membership inference attacks. Moreover, we demonstrate that the training randomness introduced by combining generalization-enhancing strategies can drastically reduce attack effectiveness—by up to two orders of magnitude. These findings establish improved generalization as both an effective and practical defense against membership inference attacks.
This work addresses the privacy risks of safety classifiers trained on sensitive data involving self-harm and mental health, which are vulnerable to membership inference attacks that can leak user information. The authors propose a boundary-targeting strategy that, for the first time, integrates low-confidence samples with membership inference attacks to expose how models rely on memorization rather than generalization when handling ambiguous inputs, thereby amplifying membership signals. Experimental results demonstrate that, at a 5% false positive rate, the method successfully recovers 19% of conversations labeled as indicating emotional distress—achieving 3.5 times the performance of the current state-of-the-art approach. The study further validates that adding noise effectively mitigates this privacy vulnerability.
This work addresses the lack of a systematic evaluation framework in existing membership inference attack (MIA) research, which hinders accurate characterization of privacy risks in real-world scenarios. The paper proposes the first end-to-end MIA evaluation framework encompassing data, model architectures, training algorithms, and post-training modules. Under a unified formal threat model, it introduces multidimensional metrics—such as balanced accuracy and true positive rate at low false positive rates—to accommodate both symmetric and asymmetric misclassification costs. Through large-scale empirical analysis across diverse configurations, the study reveals the strong dependence of MIA performance on the choice of threat model and evaluation metrics, leading to practical guidelines for privacy assessment. An open-source, ready-to-use auditing toolkit is released to significantly enhance the reliability and reproducibility of privacy risk evaluations in real-world deployments.
Current evaluations of membership inference attacks (MIAs) against language models suffer from statistical invalidity due to distributional shifts between member and non-member data, hindering fair comparisons. This work proposes an unbiased evaluation benchmark that leverages the temporal in-distribution property observed during model training: by utilizing intermediate checkpoints of open-source large language models (e.g., Pythia, OLMo), it constructs temporally adjacent member and non-member datasets that share the same underlying distribution. We introduce Pandora_LLM, a modular and open-source MIA attack library, and conduct systematic evaluations across models ranging from 70M to 7B parameters. Our experiments reveal the true performance of various MIA methods under this unbiased setting, thereby advancing standardized and reliable assessment of membership privacy risks in language models.