🤖 AI Summary
This study addresses the re-identification privacy risks in anonymized text exacerbated by web-augmented large language models (LLMs), which existing methods struggle to quantify. We construct a synthetic interview dataset containing ground-truth identities and employ LLM agents equipped with web search tools to systematically evaluate both retrieval and parametric memory capabilities across diverse model configurations. By disentangling the contributions of retrieval from parametric memory, this work uncovers the failure mechanisms of privacy instructions and proposes an attack-aware desensitization pipeline. Experimental results demonstrate that once retrieval succeeds, the re-identification rate exceeds 88%. Notably, even when entity masking is applied, 85.3% of samples remain vulnerable to re-identification, highlighting the fragility of current defensive strategies against web-enabled inference attacks.
📝 Abstract
As Large Language Models (LLMs) gain tool use capabilities such as web search, they can retrieve and cross-reference public information, creating privacy risks beyond memorization. One manifestation is re-identification: linking an anonymized interview transcript to a named individual. Yet without ground-truth identities, the coverage of such attacks and the protection offered by a defense cannot be reliably measured. We introduce AgentDOXX, an evaluation suite of 822 synthetic interview transcripts grounded in public information about real individuals with known identities. We evaluate fifteen configurations of open-weight and proprietary models, isolating the effect of web search, and analyze their search trajectories to distinguish retrieval-driven from parametric identifications. Ground-truth identities reveal that re-identification risk is distributed across an agent's execution: retrieval and parametric recall both contribute, with open-weight models identifying 15-28% of transcripts without search; identification succeeds in over 88% of cases once the target appears in a retrieved result; entity masking leaves at least one attacker successful on 85.3% of a stratified sample; and privacy instructions suppress naming but not retrieval, with configurations scoring 0% accuracy yet retrieving the subject in up to 62% of transcripts. We further show that observed attack trajectories can provide supervision for localizing identifying spans, offering a path toward attack-informed anonymization.