Score
Designs, implements, and evaluates processes, tools, and data pipelines that transform datasets to reduce re‑identification risk by removing or masking direct identifiers and replacing them with consistent pseudonyms while preserving temporal order and entity consistency. This includes selecting and applying pseudonymization techniques (e.g., deterministic HMAC hashing), managing pseudonym mappings and keys, and enforcing declared privacy boundaries and policy controls to limit identifier exposure.
To address the lack of standardized re-identification risk assessment methodologies for anonymized datasets—hindering compliance verification with privacy regulations such as the GDPR—this study proposes a practical, deployable risk assessment framework. Methodologically, it pioneers the adaptation of the EBIOS (Expression des Besoins et Identification des Objectifs de Sécurité) risk analysis paradigm from cybersecurity to privacy contexts, integrating real-world attack-pattern-driven threat modeling with an attribute-level exposure quantification model to jointly evaluate attack feasibility and individual impact. Key contributions include: (1) the first EBIOS-based privacy risk assessment workflow; (2) a computable, tiered attribute exposure model; and (3) ready-to-use operational guidelines and tooling support. Empirical validation demonstrates that the framework enables organizations to conduct compliant, reproducible assessments of anonymization effectiveness.
This work addresses the challenge of anonymizing sensitive personal information in textual data while preserving utility for downstream tasks. Methodologically, it leverages large language models (LLMs) in dual roles—both for anonymization and re-identification—to inform a multi-layered anonymization paradigm grounded in named entity recognition and author identity obfuscation. The approach integrates differential privacy, risk-aware frameworks, and domain-specific customization. Key contributions include: (1) the first comprehensive taxonomy of text anonymization techniques spanning cross-domain challenges; (2) a reproducible benchmark dataset, an open-source toolkit, and practical deployment guidelines; and (3) a novel evaluation framework that jointly employs formal privacy guarantees and empirical risk assessment. Collectively, these advances provide both theoretical foundations and actionable standards for academic research and industrial deployment of privacy-preserving text processing.
This paper identifies and defines “anonymity-washing”—the erroneous claim that data is anonymized despite failing to meet legal or technical anonymity standards—exposing systemic flaws in data privacy governance, including fragmented legal interpretations, technical misconceptions, and regulatory lag. Methodologically, it synthesizes EU GDPR case law, global regulatory guidance, and mainstream anonymization technical documentation to construct the first legal-technical attribution framework for anonymity assessment. Through multi-source systematic review and empirical analysis, the study uncovers root causes of pseudonymization misuse and widespread reliance on obsolete anonymization techniques. The contribution comprises three novel, actionable policy pathways: (1) targeted professional education to bridge technical-legal knowledge gaps; (2) mechanism-driven, dynamic updating of anonymization guidelines; and (3) cross-sectoral governance coordination. These pathways collectively enhance the credibility, accountability, and practical enforceability of anonymization practices in compliance frameworks. (149 words)
This study addresses the trade-off between privacy preservation and model utility in Retrieval-Augmented Generation (RAG) systems when handling personally identifiable information (PII). It presents the first systematic evaluation of how applying anonymization at different stages of the RAG pipeline—specifically at the input data versus the generated output—affects both privacy protection and task performance. Through quantitative analysis, the research demonstrates that the placement of anonymization significantly influences the privacy-utility balance: anonymizing at the input stage offers stronger privacy guarantees, whereas anonymization at the output stage better preserves the quality of generated text. These findings provide empirical evidence and practical design guidance for mitigating privacy risks in RAG systems without unduly compromising their functional effectiveness.
This paper identifies a pervasive “privacy illusion” in text anonymization: existing methods—such as PII removal or synthetic data generation—only mitigate explicit identifier leakage, yet remain vulnerable to semantic re-identification attacks. To address this, we propose the first risk assessment framework for semantic-level privacy leakage, integrating empirical re-identification attacks, benchmarking against mainstream commercial PII detection tools (e.g., Azure), differential privacy baselines, and evaluation on the MedQA medical QA dataset. Results reveal that Azure fails to protect 74% of sensitive information in MedQA; while differential privacy reduces re-identification risk, it severely degrades textual utility. Our work challenges prevailing assumptions about anonymization efficacy and establishes a verifiable, empirically grounded evaluation paradigm for semantic privacy protection.
This work addresses the challenge of conducting efficient and scalable post-hoc privacy audits on deployed large language models, which existing methods struggle to achieve without either injecting synthetic data during training or requiring private holdout datasets drawn from the same distribution as the training data. To overcome these limitations, the paper introduces Natural Identifiers (NIDs)—structured random strings inherently present in training data, such as hash values or shortened URLs—as intrinsic signals for auditing. Leveraging NIDs, the authors develop the first general-purpose, post-training privacy auditing framework that requires no modification to the training process and no access to private reserved data. The approach supports both differential privacy verification and dataset inference, and its practicality is further enhanced through distribution-consistent synthetic data, enabling effective and scalable privacy evaluation across diverse scenarios.
Current data preparation pipelines typically assess privacy risks only at the final release stage, overlooking the cumulative privacy implications of intermediate steps. This work proposes an inference-aware privacy modeling approach that conceptualizes data preparation as an interactive, guided process. Leveraging a compatibility-set semantic framework, it characterizes how deterministic curation operations affect an observer’s inferential capabilities. The framework distinguishes between two operation types—“evidence removal” and “ambiguity elimination”—revealing the non-monotonic nature of their privacy effects and enabling prefix-level privacy feedback under a given disclosure budget. By formalizing these dynamics, this study lays the theoretical groundwork for end-to-end interactive privacy-preserving systems and identifies key challenges in their realization.
This study addresses the risk of network topology data leakage when CSIRTs fine-tune small language models (1B–3B parameters). It empirically investigates the combined effect of differentially private SGD (DP-SGD, ε=2–8) and HMAC-based pseudonymization. Evaluating 96 configurations across four model families using LoRA/QLoRA adapters, and employing four extraction attacks alongside a two-stage auditing framework, the work quantifies for the first time that the number of optimizer updates predominantly drives memorization reduction. While DP-SGD offers formal privacy guarantees, it does not enhance memorization suppression; in contrast, HMAC effectively reduces the exposure surface by 40%–61% without introducing new memorization targets. Across all settings, model utility remains limited, with F1 scores ranging only from 0.19 to 0.28 under the adopted privacy budgets.