Score
Designs and implements metrics, experiments, and benchmarks to quantify how privacy-preserving transformations affect the usefulness of data, models, or artifacts. Builds comparative analyses and evaluation procedures that measure privacy–utility tradeoffs, derive measurable privacy–utility boundaries, and compare performance (including human versus model) across anonymization or protection methods.
Text anonymization lacks reliable, regulation-compliant, and user-aligned privacy assessment methods. This work bridges this gap through a systematic literature review and interdisciplinary manual analysis—integrating NLP, legal text analysis, and HCI empirical findings—to jointly examine GDPR, HIPAA, and other regulatory frameworks alongside user mental models. We identify and critically analyze six privacy conceptualizations in terms of their effectiveness for risk characterization. Our analysis reveals substantial misalignments among existing metrics, legal requirements, and user expectations. We propose the first privacy metric framework explicitly designed for both regulatory compliance and human-centered design, delineating method-specific applicability boundaries and concrete improvement pathways. The framework delivers actionable guidance for practitioners and advances text privacy evaluation toward greater rigor, cross-study comparability, and interpretability. (149 words)
Recent critiques have challenged the differential privacy guarantees of PATE-GAN and PrivBayes, questioning the validity of their privacy-utility trade-offs. However, these critiques rely on restrictive assumptions—such as synthetic or simplistic data distributions—and limited experimental settings, potentially biasing their conclusions. Method: We propose a more general privacy-utility evaluation framework that integrates privacy game analysis and theoretical verification, and conduct k-anonymity benchmarking experiments on real-world datasets without distributional assumptions. Contribution/Results: Under identical privacy budgets, both PATE-GAN and PrivBayes significantly outperform k-anonymity in statistical utility while maintaining strong differential privacy guarantees. We demonstrate that prior claims of “privacy failure” stem from flawed evaluation premises—specifically, the absence of rigorous privacy accounting and realistic data assumptions. Our empirical analysis refutes these criticisms and establishes synthetic data generation as a robust and effective privacy-enhancing technology.
Current privacy evaluation of synthetic data lacks standardized benchmarks, and existing metrics inadequately capture real-world adversarial risks. Method: This paper presents the first systematic empirical evaluation of mainstream privacy metrics—including adversarial attack simulation and membership inference success rates—in generative models. It comparatively analyzes practical efficacy of privacy-enhancing techniques such as differential privacy integration and privacy-aware generation, and quantifies the privacy–utility trade-off. Contribution/Results: We propose a deployment-oriented synthetic data privacy assessment framework featuring a reproducible evaluation pipeline, a standardized metric suite, and implementation guidelines. The study establishes a rigorous benchmarking methodology for academia and delivers an actionable, practice-driven privacy assurance evaluation paradigm—with concrete best practices—for industry adoption.
This work addresses the challenge of securely sharing electronic health records (EHRs) across institutions, which is hindered by privacy concerns, governance constraints, and interoperability limitations that impede multicenter research and medical AI development. The authors propose a novel EHR transformation framework based on irreversible geometric operators, designed under a rigorous threat model through collaborative strategy formulation between human experts and an AI agent (SciencePal). The approach integrates hybrid mechanisms tailored for high-risk scenarios and is theoretically grounded, with robustness validated against diverse privacy attacks—including reconstruction, linkage, and membership inference. Experimental results demonstrate that the transformed data effectively resist such attacks while preserving clinical interpretability and utility for machine learning tasks, thereby establishing a secure and efficient foundation for large-scale medical AI training.
Current evaluations of synthetic data privacy lack quantifiable, comparable metrics due to ambiguous privacy definitions and existing measures’ inability to reflect real-world disclosure risks. Method: We propose the first benchmark framework based on deliberate risk insertion—integrating legal theory with a black-box threat model—to enable reproducible, cross-method assessment of privacy-utility trade-offs. Our approach systematically controls perturbations, models diverse black-box attacks, maps outputs to regulatory compliance criteria, and validates findings on public datasets. Contribution/Results: Empirical evaluation reveals substantial discrepancies between mainstream privacy metrics (e.g., k-anonymity, differential privacy estimates) and actual re-identification risks under realistic attack scenarios. This work establishes the first evaluation paradigm for privacy-enhancing technologies (PETs) that is simultaneously interpretable, empirically grounded, and aligned with regulatory requirements—thereby bridging theoretical guarantees, practical security, and legal accountability.
Existing tool-calling benchmarks lack the capability to audit privacy violations arising from purpose-constrained information flows within multi-tool execution trajectories. This work introduces the concept of “need-to-know disclosure boundaries” and constructs a benchmark comprising 2,150 cases to enable trajectory-level auditing of excessive privacy disclosures by LLM agents during multi-tool interactions. By comparing agent behaviors against a policy knowledge base and simulated backend business logs, the framework supports systematic evaluation of nine mainstream agents. Empirical results reveal that successful task completion does not guarantee privacy compliance—several agents disclose unnecessary private information in intermediate steps. These findings demonstrate the effectiveness and necessity of the proposed benchmark in addressing the critical gap in purpose-bound privacy assessment.
Real-world Security Operations Center (SOC) data is rarely accessible for research due to privacy constraints, leading existing studies to rely on synthetic or outdated datasets. This work proposes a high-fidelity anonymization method that extracts and structures SIEM logs from a financial-sector SOC, preserving temporal ordering and entity consistency while enforcing strict privacy guarantees—thereby establishing the first quantifiable privacy-utility trade-off boundary. Leveraging this approach, we construct 37 HIKARI evaluation challenges and develop a deterministic validator alongside a large language model (LLM) behavioral compliance detection mechanism. In experiments involving 200 SOCpilot incidents, our framework uncovered LLM non-compliant actions undetected by human baselines, enabling reproducible and verifiable evaluation of autonomous defense systems.
This study addresses the trade-off between privacy preservation and model utility in Retrieval-Augmented Generation (RAG) systems when handling personally identifiable information (PII). It presents the first systematic evaluation of how applying anonymization at different stages of the RAG pipeline—specifically at the input data versus the generated output—affects both privacy protection and task performance. Through quantitative analysis, the research demonstrates that the placement of anonymization significantly influences the privacy-utility balance: anonymizing at the input stage offers stronger privacy guarantees, whereas anonymization at the output stage better preserves the quality of generated text. These findings provide empirical evidence and practical design guidance for mitigating privacy risks in RAG systems without unduly compromising their functional effectiveness.
This study addresses the challenge of privacy communication in human–robot collaboration systems within Industry 5.0, where sensitive data monitoring raises significant privacy concerns that are often obscured by technical complexity, leading to mistrust and resistance among non-technical stakeholders. To bridge this gap, the authors propose a novel conceptual framework that integrates Privacy by Design principles with large language models (LLMs), leveraging LLMs for the first time in the requirements engineering process to automatically generate natural-language privacy reports tailored for non-technical audiences from representative human–robot monitoring scenarios. Evaluation across two industrial use cases demonstrates that the approach substantially enhances the comprehensibility of privacy information and supports informed decision-making, thereby addressing a critical accessibility gap in existing privacy communication mechanisms.