REDACT: A Systematically Controlled Multilingual Benchmark for Personal Information Detection

📅 2026-06-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing PII detection benchmarks, which suffer from narrow entity coverage and uncontrolled generation conditions that obscure failure mechanisms of detectors. We present the first systematically controlled multilingual PII detection benchmark, spanning 25 languages, 51 entity types, and 4,127 surface form patterns. Generation is governed by a strength-2 covering array sampler that modulates nine dimensions, complemented by a GDPR-aligned sensitivity stratification mechanism. Our approach innovatively integrates multilingual synthetic data generation, entity-level metadata annotation, LLM-as-judge evaluation, and hierarchical test set construction. Evaluation reveals that rule-based systems exhibit alarmingly low recall—dropping to 0.07—for high-sensitivity categories, whereas large language models demonstrate greater robustness; sensitivity stratification emerges as the most challenging dimension. The full benchmark and associated tools are publicly released.
📝 Abstract
Benchmark infrastructure for personally identifiable information (PII) detection remains limited: existing corpora cover few entity types, use ad hoc generation conditions, and do not show which surface conditions cause detector failures. We present REDACT, a systematically controlled multilingual PII benchmark with 13,427 records, 324,078 entity annotations, 51 entity types, 4,127 surface-form patterns, and 25 languages across 9 scripts. A strength-2 covering-array sampler controls nine generation axes: domain, format, difficulty, length, density, code-switching, language, adjacency, and co-occurrence. Three entity-level metadata fields (disclosure status, disclosure form, and a GDPR-aligned sensitivity tier) enable stratified evaluation beyond aggregate or per-type F1. From the full benchmark, we evaluate five detectors (Presidio, GLiNER, the OpenAI Privacy Filter, GPT-4.1, and Claude Sonnet 4.6) on a locked, language-stratified sample of 1,000 records. Aggregate F1 masks an architecture-dependent failure structure: the rule-based detector performs poorly on the highest-stakes data, including HIGH-sensitivity categories (recall 0.07) and non-verbatim disclosure forms, while the LLM detectors remain more robust, with the HIGH tier as their strongest sensitivity slice. A three-model reference-free LLM-as-judge assessment corroborates that sensitivity-tier assignment is the task's hardest axis. We release the benchmark, schema, prompts, and stratified evaluation harness.
Problem

Research questions and friction points this paper is trying to address.

personally identifiable information
PII detection
multilingual benchmark
entity types
surface-form patterns
Innovation

Methods, ideas, or system contributions that make the work stand out.

systematically controlled benchmark
multilingual PII detection
sensitivity-tier evaluation
covering-array sampling
LLM-as-judge
G
Guneesh Vats
ServiceNow
A
Anubha Agrawal
ServiceNow
S
Shikha Singhal
ServiceNow
A
Ajita Dash
ServiceNow
P
Praison Selvaraj
ServiceNow
V
Vidhan Jhawar
ServiceNow
R
Ranga Prasad Chenna
ServiceNow
B
Bharadwaj Y M G
ServiceNow