🤖 AI Summary
This study addresses the limited generalization of existing personally identifiable information (PII) detection systems, which often suffer from narrow training data domains and struggle to achieve broad coverage across heterogeneous text. Leveraging a refined multi-source PIIBench dataset, the authors systematically evaluate three DeBERTa fine-tuning strategies—direct fine-tuning, source-conditional hierarchical modeling (SC+H), and a three-stage curriculum learning approach (SC+H+Curr)—across 82 fine-grained PII entity types. Results demonstrate that direct fine-tuning, using only diverse task data and a simple weighted cross-entropy loss, achieves F1 scores of 0.6476 and 0.6455 on the test_5k and full 100k test sets, respectively, significantly outperforming current methods. This strategy leads in 54 fine-grained and all 10 coarse-grained categories, underscoring the critical role of data diversity and a concise objective function in enhancing model generalization.
📝 Abstract
Personally identifiable information (PII) detection systems are frequently trained within narrow source or domain boundaries, limiting coverage when deployed on heterogeneous text. We study model fine-tuning on a corrected multi-source PIIBench preparation spanning 82 retained entity types across ten source datasets. We evaluate three DeBERTa-based approaches: direct token classification fine-tuning, a source-conditioned hierarchical model (SC+H), and a three-phase curriculum extension (SC+H+Curr). Against eight published comparator systems on a reproducible 5,000-record held-out subset (test_5k), direct fine-tuned DeBERTa achieves F1 0.6476, while SC+H and the curriculum variant achieve 0.5899 and 0.2772 respectively; the strongest published comparator reaches only 0.1723. Because validation initially favoured SC+H, we perform a final streamed evaluation on the complete 100,002-record held-out split. Direct fine-tuning remains superior, achieving F1 0.6455 versus 0.5894 for SC+H. Entity-level analysis shows that direct fine tuning wins 54 of 82 fine entity types and all ten coarse groups by support-weighted entity F1, while SC+H retains localised advantages on 28 types. The results indicate that diverse task-specific training data and a simple weighted cross-entropy objective contribute more to broad-coverage PII detection than the tested architectural and curriculum complexity.