TULIP: Targeted LLM Unlearning at Layers Identified Per-Input

๐Ÿ“… 2026-09-28
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the limitation of existing large language model unlearning methods that intervene at fixed layers, overlooking the input-dependent nature of knowledge formation. To this end, we propose TULIP, a framework that introduces the first Logit Lens-based, input-adaptive dynamic layer selection mechanism to precisely localize the boundaries of knowledge formation and readout. TULIP achieves targeted unlearning by eliminating representational alignment between intermediate hidden states and the forgetting targets. Experimental results demonstrate that TULIP significantly outperforms existing baselines across multiple benchmarks, exhibits strong robustness against paraphrasing and quantization attacks, and can serve as a plug-and-play module to enhance the unlearning performance of other methods.
๐Ÿ“ Abstract
Representation-level unlearning intervenes on the intermediate hidden states of LLMs. Although knowledge is distributed across layers, existing methods operate at a single fixed layer for the entire forget set. We ask whether such a fixed layer is sufficient. To answer this, we design a hijacking experiment that grafts hidden states of the target model into an oracle trained only on the retain set. The oracle cannot produce the forget answer on its own, yet it produces the answer from the grafted state. Thus, the answer is formed at an intermediate layer and merely read out afterward, so unlearning should focus on formation, not readout. Moreover, the layer where formation ends varies widely across inputs. Motivated by these findings, we propose Targeted Unlearning at Layers Identified Per-input (TULIP). For each input, TULIP uses the logit lens to locate the formation-readout boundary and removes the hidden state's alignment with the forget answer's unembedding vector there. TULIP consistently outperforms output- and representation-level baselines on TOFU, PISTOL, and WMDP across Llama, Qwen, and Zephyr models. It also remains robust to paraphrase and quantization attacks. Beyond standalone use, its per-input layer selection serves as a plug-and-play component that further improves existing methods.
Problem

Research questions and friction points this paper is trying to address.

LLM unlearning
representation-level unlearning
per-input layer selection
knowledge formation
hidden states
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM Unlearning
Representation-level Intervention
Logit Lens
Per-input Layer Selection
Hidden State Alignment
๐Ÿ”Ž Similar Papers
2024-06-22International Conference on Computational LinguisticsCitations: 4