π€ AI Summary
This study addresses the risk of personally identifiable information (PII) leakage in large language models under cross-lingual scenarios, where non-English prompts can bypass existing safeguards to extract memorized data. We construct a multi-domain PII dataset translated into Italian, Spanish, French, and German, and employ multilingual training data extraction attacks combined with intermediate-layer activation representation alignment analysis to reveal the latent cross-lingual bridging mechanisms formed during native multilingual pretraining. Our findings demonstrate that translated prompts can successfully recover PII even when the original text was never published online, and that stronger multilingual capabilities correlate with more severe information leakage. These results underscore the urgent need for language-agnostic, robust data sanitization strategies to mitigate cross-lingual privacy vulnerabilities in large language models.
π Abstract
The robustness of Personally Identifiable Information (PII) protection in Large Language Models (LLMs) is a critical concern, yet the risks associated with cross-lingual data extraction remain under-explored. This study evaluates the vulnerability of English-centric and multilingual models to Training Data Extraction (TDE) attacks when prompted in non-English languages. We construct a multi-domain PII dataset comprising social media handles, email addresses, and phone numbers and translate the attack contexts into Italian, Spanish, French, and German. Our results show that TDE attacks against both English-centric and multilingual models transfer to different languages: the attacks are successful on translated prompts, even though only the original English prompt might have been included in the pre-training data. A web-presence check on a sample of the translations confirms that they are not available online. The share of English leaks recovered in other languages grows with the multilingual capability of the model, and it drops sharply when the original wording is lost, even without a change of language. This suggests that native multilingual pre-training facilitates the emergence of latent cross-linguistic bridges that simplify the retrieval of personally identifiable information (PII). We analyze the activations of multilingual large language models (LLMs) and find that different translations of the same prompt are bridged in similar representations, with the strongest alignment in the middle layers. Our results highlight a fundamental security gap in modern LLMs, necessitating more robust, language-agnostic sanitization strategies for future model alignment.