Addressing Labelled Data Scarcity: Taxonomy-Agnostic Annotation of PII Values in HTTP Traffic using LLMs

📅 2026-05-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limited generalization and poor adaptability of privacy auditing models caused by scarce annotated data and rigid, predefined personally identifiable information (PII) taxonomies. To overcome these challenges, the authors propose a multi-stage large language model (LLM) pipeline that enables explicit PII value annotation under arbitrary PII classification schemes at runtime. The approach integrates deterministic preprocessing, label-level classification, instance-level annotation, and output validation, achieving taxonomy-agnostic dynamic PII labeling for the first time. Additionally, they introduce an LLM-based synthetic HTTP traffic generation technique that facilitates controlled evaluation and prompt engineering without relying on real sensitive data. Experiments across three diverse PII taxonomies—varying in domain and granularity—demonstrate the method’s effectiveness in accurately identifying PII types and extracting corresponding values, highlighting the potential of LLMs for flexible privacy annotation and synthetic data generation.
📝 Abstract
Automated privacy audits of web and mobile applications often analyse outbound HTTP traffic to detect Personally Identifiable Information (PII) leakage. However, existing learning-based detectors typically depend on scarce, manually labelled traffic and are tightly coupled to fixed label taxonomies, limiting transferability across domains and evolving definitions of PII. This paper investigates whether Large Language Models (LLMs) can support taxonomy-agnostic annotation of explicitly transmitted PII values in HTTP message bodies when the taxonomy is provided at runtime. We introduce a multi-stage LLM-based pipeline that combines deterministic pre-processing with label-level classification, targeted instance-level value annotation, and output validation. To enable controlled evaluation and exemplar-based prompting without relying on sensitive real-user captures, we further propose an LLM-based generator for synthetic HTTP traffic with manually validated, taxonomy-derived PII annotations. We evaluate the approach across three taxonomies spanning different PII domains and granularity levels. Results show that the pipeline accurately detects PII types and extracts corresponding values for concrete PII taxonomies. Overall, our findings position LLMs as a promising foundation for flexible, taxonomy-agnostic traffic annotation and for creating labelled data under evolving privacy taxonomies.
Problem

Research questions and friction points this paper is trying to address.

PII detection
labelled data scarcity
taxonomy-agnostic annotation
HTTP traffic analysis
privacy audits
Innovation

Methods, ideas, or system contributions that make the work stand out.

taxonomy-agnostic
LLM-based annotation
PII detection
synthetic HTTP traffic
privacy auditing
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
T
Thomas Cory
Technische Universität Berlin
A
Axel Küpper
Technische Universität Berlin