🤖 AI Summary
This study addresses the conflation of personally identifiable information recognition with privacy semantics in existing redaction benchmarks, which often neglect the critical role of context in privacy judgments. Drawing on contextual integrity theory, the authors construct RedactionBench—a human-annotated benchmark comprising 200 real-world documents across 11 domains—and introduce R-Score, a character-level evaluation metric that distinguishes semantically equivalent redactions while disregarding superficial formatting differences. The work presents the first context-aware redaction evaluation framework and systematically assesses 35 models, including named entity recognition systems, small language models, and large language models with agent-based reasoning. Human evaluations reveal low inter-annotator agreement (47.7%) on context-sensitive redaction decisions, and all evaluated models struggle significantly with this task, highlighting fundamental limitations in current approaches and underscoring the need for standardized evaluation of privacy-preserving systems.
📝 Abstract
Large Language Models are increasingly applied to sensitive domains that require redaction of personally identifiable information (PII). While redacting PII is a data cleaning prerequisite, existing benchmarks conflate extraction mechanics with privacy semantics. A public phone number is not equivalent to a phone number in a medical record. Whether information constitutes a violation depends heavily on who holds it, why, and in what context, fundamentally differentiating redaction from simple entity recognition. Grounded in contextual integrity, we introduce RedactionBench, a manually annotated benchmark comprising 200 diverse documents across 11 domains, mostly seeded from real-world sources. We also introduce R-Score, a novel character-level metric that treats semantically similar redactions equally and nullifies shallow formatting choices, such as varying masking styles for phone numbers. Evaluations across Named Entity Recognition models, entity extraction Small Language Models, and frontier models equipped with agentic tools demonstrate that contextual redaction remains an unsolved problem. A human evaluation with over 80 users on RedactionBench reveals a stark dichotomy in privacy perceptions. Annotators show consensus with target labels for mandatory redactions (89.4 percent) and safe text preservations (94.1 percent), but fail to agree on contextual redactions (47.7 percent). This variance demonstrates the subjective nature of contextual privacy and motivates R-Score, which decouples contextual ambiguity from strict precision. We compare 35 models across families and report their performance in redacting PII. Finally, we release RedactionBench to establish a baseline for future privacy-preserving systems, hoping to inspire efficient model design and standardized evaluations.