An Evaluation Study of Hybrid Methods for Multilingual PII Detection

📅 2025-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address performance bottlenecks in PII detection for low-resource languages—caused by scarce annotated data and linguistic diversity—this paper proposes RECAP, a hybrid framework integrating deterministic regular expressions with context-aware large language models (LLMs) within a modular, three-stage refinement pipeline. RECAP enables zero-shot generalization across 13 languages and 300+ entity types without retraining for new categories. It first applies regex-based coarse filtering, followed by LLM-driven fine-grained recognition, and finally rule-guided disambiguation and boundary refinement. This design significantly enhances boundary detection and ambiguity resolution. On the nervaluate benchmark, RECAP achieves a weighted F1-score of 89.7%, outperforming fine-tuned NER models by 82% and zero-shot LLM baselines by 17%. The framework delivers an efficient, scalable solution for multilingual privacy compliance.

Technology Category

Natural Language Processing: Language Grounding & Multi-modal NLPMachine Learning: Large Multimodal Models (LMMs)Computer Vision: Language and Vision

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationEconomics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labeling
📝 Abstract
The detection of Personally Identifiable Information (PII) is critical for privacy compliance but remains challenging in low-resource languages due to linguistic diversity and limited annotated data. We present RECAP, a hybrid framework that combines deterministic regular expressions with context-aware large language models (LLMs) for scalable PII detection across 13 low-resource locales. RECAP's modular design supports over 300 entity types without retraining, using a three-phase refinement pipeline for disambiguation and filtering. Benchmarked with nervaluate, our system outperforms fine-tuned NER models by 82% and zero-shot LLMs by 17% in weighted F1-score. This work offers a scalable and adaptable solution for efficient PII detection in compliance-focused applications.
Problem

Research questions and friction points this paper is trying to address.

Detecting PII in low-resource languages with limited data
Combining regex and LLMs for scalable multilingual PII detection
Improving accuracy over existing NER models and zero-shot LLMs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hybrid framework combines regex with LLMs
Modular design supports 300+ entity types
Three-phase refinement pipeline for disambiguation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Centific Global Solutions Inc.
H
Harshit Rajgarhia
Centific Global Solutions Inc.
S
Suryam Gupta
Centific Global Solutions Inc.
A
Asif Shaik
Centific Global Solutions Inc.
G
Gulipalli Praveen Kumar
Centific Global Solutions Inc.
Y
Y Santhoshraj
Centific Global Solutions Inc.
S
Sanka Nithya Tanvy Nishitha
Centific Global Solutions Inc.
A
Abhishek Mukherji
Centific Global Solutions Inc.