Score
Designs, implements, and evaluates processes, algorithms, and toolchains that remove, mask, or otherwise obfuscate direct and indirect personal identifiers in datasets while preserving their statistical or analytical utility. Work includes selecting and applying de-identification techniques (e.g., redaction, pseudonymization, generalization, differential-privacy or other privacy-preserving transforms), implementing local or on-device pipelines to avoid external transfers, measuring re-identification risk versus data utility, and producing reproducible documentation of processing and consent steps.
This paper identifies a pervasive “privacy illusion” in text anonymization: existing methods—such as PII removal or synthetic data generation—only mitigate explicit identifier leakage, yet remain vulnerable to semantic re-identification attacks. To address this, we propose the first risk assessment framework for semantic-level privacy leakage, integrating empirical re-identification attacks, benchmarking against mainstream commercial PII detection tools (e.g., Azure), differential privacy baselines, and evaluation on the MedQA medical QA dataset. Results reveal that Azure fails to protect 74% of sensitive information in MedQA; while differential privacy reduces re-identification risk, it severely degrades textual utility. Our work challenges prevailing assumptions about anonymization efficacy and establishes a verifiable, empirically grounded evaluation paradigm for semantic privacy protection.
The widespread use of personal data has intensified privacy leakage risks—particularly under strong re-identification attacks and mounting regulatory compliance pressures. Differential privacy (DP), a mathematically rigorous privacy-preserving paradigm, has emerged as a foundational mitigation strategy. This paper systematically surveys DP’s theoretical foundations, mainstream mechanisms—including Laplace/Gaussian noise injection, privacy budget allocation, and sensitive query perturbation—as well as its cutting-edge applications in privacy-preserving machine learning and synthetic data generation. It critically examines practical challenges: utility–privacy trade-offs, cross-domain adaptability, and user comprehension barriers. Building on this analysis, the paper proposes a practice-oriented framework centered on enhancing system transparency, interpretability, and usability. Designed for both researchers and practitioners, the framework bridges theoretical rigor with engineering feasibility, facilitating trustworthy DP deployment in high-stakes domains such as healthcare and cybersecurity.
This paper identifies and defines “anonymity-washing”—the erroneous claim that data is anonymized despite failing to meet legal or technical anonymity standards—exposing systemic flaws in data privacy governance, including fragmented legal interpretations, technical misconceptions, and regulatory lag. Methodologically, it synthesizes EU GDPR case law, global regulatory guidance, and mainstream anonymization technical documentation to construct the first legal-technical attribution framework for anonymity assessment. Through multi-source systematic review and empirical analysis, the study uncovers root causes of pseudonymization misuse and widespread reliance on obsolete anonymization techniques. The contribution comprises three novel, actionable policy pathways: (1) targeted professional education to bridge technical-legal knowledge gaps; (2) mechanism-driven, dynamic updating of anonymization guidelines; and (3) cross-sectoral governance coordination. These pathways collectively enhance the credibility, accountability, and practical enforceability of anonymization practices in compliance frameworks. (149 words)
To address the lack of standardized re-identification risk assessment methodologies for anonymized datasets—hindering compliance verification with privacy regulations such as the GDPR—this study proposes a practical, deployable risk assessment framework. Methodologically, it pioneers the adaptation of the EBIOS (Expression des Besoins et Identification des Objectifs de Sécurité) risk analysis paradigm from cybersecurity to privacy contexts, integrating real-world attack-pattern-driven threat modeling with an attribute-level exposure quantification model to jointly evaluate attack feasibility and individual impact. Key contributions include: (1) the first EBIOS-based privacy risk assessment workflow; (2) a computable, tiered attribute exposure model; and (3) ready-to-use operational guidelines and tooling support. Empirical validation demonstrates that the framework enables organizations to conduct compliant, reproducible assessments of anonymization effectiveness.
Existing face de-identification methods suffer from significant limitations in preserving attribute details and maintaining robustness against occlusions, often introducing noticeable editing artifacts that compromise visual authenticity and fidelity. To address these issues, we propose a “decoupling-first” two-stage framework. First, a Contrastive Identity-Decoupling (CID) module explicitly disentangles identity-specific features from attribute-related features via contrastive learning. Second, a Key-driven Reversible Anonymous Representation (KRIA) mechanism—integrated with a Multi-scale Attention-based Attribute Refinement (MAAR) module—enables high-fidelity, fine-grained attribute preservation under occluded conditions. The entire framework supports end-to-end joint optimization. Extensive experiments demonstrate that our method achieves a 12.6% improvement in attribute fidelity and a 23.4% gain in editing quality under occlusion, while substantially suppressing artificial artifacts. It outperforms state-of-the-art approaches across multiple quantitative and qualitative metrics.
This study addresses pervasive data leakage and contamination issues in current evaluations of personally identifiable information (PII) anonymization techniques, which lead to a significant overestimation of privacy protection efficacy. We systematically analyze methodological flaws in existing assessment practices and, for the first time, expose data contamination introduced by the use of non-real private data in experimental designs. Our critical evaluation, combined with data leakage pathway analysis and adversarial scenario modeling, demonstrates that most reported attack successes rely on unrealistic data assumptions. We argue that only evaluations grounded in authentic private data can reliably validate the security of anonymization methods. This work thus establishes a theoretical foundation and practical direction for developing trustworthy, reproducible frameworks for privacy-preserving technology assessment.
This study addresses the persistent challenge of effectively ensuring privacy in scientific data anonymization, which is often hindered by a lack of actionable guidance. It introduces, for the first time, a systematic red team–blue team adversarial framework to this domain: the red team simulates realistic re-identification attacks, while the blue team iteratively refines anonymization strategies. The approach is empirically validated on real-world datasets using mixed-methods research. The work demonstrates that red teaming efficiently uncovers vulnerabilities in anonymization protocols and further delivers a reusable, publicly released framework and toolset for researchers. This contribution substantially enhances both the practicality and security of data anonymization practices in scientific research.
This study addresses the challenge that existing identity document verification systems struggle to detect high-fidelity localized forgeries generated by modern AI techniques, further compounded by privacy regulations that restrict access to real ID documents containing authentic security features for training. To bridge this gap, the authors introduce FakeIDet3-DB—the first digital tampering benchmark built from genuine government-issued IDs—encompassing both classical and generative AI–based attacks. They propose the PACE algorithm, which extracts semantically rich pseudo-anonymous image patches while complying with GDPR and other privacy mandates, leveraging geometric constraints, integral image mapping, and distance-driven non-maximum suppression. Evaluated on 5.2 million image patches derived from over 6,400 IDs, the approach effectively narrows the domain gap between synthetic and real data, revealing that state-of-the-art models achieve a detection equal error rate of 32.45% and a localization AUC-ROC of 83.48% on this benchmark.
This study addresses the trade-off between privacy preservation and model utility in Retrieval-Augmented Generation (RAG) systems when handling personally identifiable information (PII). It presents the first systematic evaluation of how applying anonymization at different stages of the RAG pipeline—specifically at the input data versus the generated output—affects both privacy protection and task performance. Through quantitative analysis, the research demonstrates that the placement of anonymization significantly influences the privacy-utility balance: anonymizing at the input stage offers stronger privacy guarantees, whereas anonymization at the output stage better preserves the quality of generated text. These findings provide empirical evidence and practical design guidance for mitigating privacy risks in RAG systems without unduly compromising their functional effectiveness.
Current evaluations of synthetic data privacy lack quantifiable, comparable metrics due to ambiguous privacy definitions and existing measures’ inability to reflect real-world disclosure risks. Method: We propose the first benchmark framework based on deliberate risk insertion—integrating legal theory with a black-box threat model—to enable reproducible, cross-method assessment of privacy-utility trade-offs. Our approach systematically controls perturbations, models diverse black-box attacks, maps outputs to regulatory compliance criteria, and validates findings on public datasets. Contribution/Results: Empirical evaluation reveals substantial discrepancies between mainstream privacy metrics (e.g., k-anonymity, differential privacy estimates) and actual re-identification risks under realistic attack scenarios. This work establishes the first evaluation paradigm for privacy-enhancing technologies (PETs) that is simultaneously interpretable, empirically grounded, and aligned with regulatory requirements—thereby bridging theoretical guarantees, practical security, and legal accountability.