🤖 AI Summary
To address performance bottlenecks in PII detection for low-resource languages—caused by scarce annotated data and linguistic diversity—this paper proposes RECAP, a hybrid framework integrating deterministic regular expressions with context-aware large language models (LLMs) within a modular, three-stage refinement pipeline. RECAP enables zero-shot generalization across 13 languages and 300+ entity types without retraining for new categories. It first applies regex-based coarse filtering, followed by LLM-driven fine-grained recognition, and finally rule-guided disambiguation and boundary refinement. This design significantly enhances boundary detection and ambiguity resolution. On the nervaluate benchmark, RECAP achieves a weighted F1-score of 89.7%, outperforming fine-tuned NER models by 82% and zero-shot LLM baselines by 17%. The framework delivers an efficient, scalable solution for multilingual privacy compliance.
📝 Abstract
The detection of Personally Identifiable Information (PII) is critical for privacy compliance but remains challenging in low-resource languages due to linguistic diversity and limited annotated data. We present RECAP, a hybrid framework that combines deterministic regular expressions with context-aware large language models (LLMs) for scalable PII detection across 13 low-resource locales. RECAP's modular design supports over 300 entity types without retraining, using a three-phase refinement pipeline for disambiguation and filtering. Benchmarked with nervaluate, our system outperforms fine-tuned NER models by 82% and zero-shot LLMs by 17% in weighted F1-score. This work offers a scalable and adaptable solution for efficient PII detection in compliance-focused applications.