Linguistic Loopholes in LLM Unlearning: From a 174-Language Benchmark to Coverage-Aware Unlearning

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the cross-lingual forgetting vulnerability in large language models, wherein knowledge deleted in one language remains accessible via others. To tackle this, it introduces a language-budgeted multilingual unlearning task and proposes the COVER method. By establishing a pioneering 174-language benchmark, the authors reveal the failure of strong source-language combinations. COVER employs training-free inference using frozen models and calibration data, coupled with a coverage optimization algorithm to select critical source-language subsets that maximize cross-lingual coverage under limited supervision. Compared to uniform selection, the approach reduces residual access rates by 7.8%–27.3%, with effectiveness further validated on the low-resource LORELEI dataset. This work ultimately enables efficient and secure cross-lingual machine unlearning.
📝 Abstract
Unlearning a fact in one language does not guarantee its removal in others as changing the query or even the requested answer language can reopen seemingly forgotten knowledge -- a cross-lingual loophole. The most straightforward solution to this challenge -- unlearning in all languages -- is neither scalable nor desirable as it amplifies damage to unrelated model capabilities. We introduce the task of language budgeted multilingual unlearning where the goal is to select a subset of languages that maximizes cross-lingual erasure. To study this task we introduce the Cross-Lingual Unlearning Tensor, an unlearning benchmark that spans 174 language--script pairs and 25 atomic paraphrase types to examine when forgetting generalizes across linguistic expressions of the same knowledge. We further propose COVER, which selects source languages to maximize predicted COVERage of languages receiving no forget supervision, enabling unlearning on a language budget. Surprisingly, we find naively selecting strong individual sources does not reliably compose into strong source sets motivating our development of COVER. At deployment COVER only requires benign calibration data and access to the frozen model. Across three model families and two disjoint forget sets, COVER reduces mean held-out residual access by 7.8--27.3% relative to uniform source selection. We find these gains extend beyond synthetic benchmarks to real news documents in low-resource language settings using human translated data from the Low Resource Languages for Emergent Incidents (LORELEI) corpus.
Problem

Research questions and friction points this paper is trying to address.

LLM unlearning
cross-lingual loophole
multilingual unlearning
language budget
knowledge erasure
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multilingual Unlearning
Cross-Lingual Loopholes
Language Budget
Coverage-Aware Selection
Unlearning Benchmark
🔎 Similar Papers
2024-06-22International Conference on Computational LinguisticsCitations: 4