UnlearningSoup: Is Repeated Tuning Necessary for Large Language Model Unlearning?

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of balancing forgetting and retention in large language model unlearning, which renders hyperparameter tuning difficult and poorly transferable. Observing that distinct training runs share performance basins in weight space, this work proposes UnlearningSoup, a framework that replaces repetitive tuning with model soups. It pioneers both efficiency- and performance-oriented merging strategies tailored for unlearning scenarios, integrating binary search interpolation and reweighting techniques to enable rapid exploration of the parameter space and unlock performance potential. Experimental results demonstrate that the proposed method improves hyperparameter selection efficiency by 2.4× to 3.3× across diverse datasets and models while consistently enhancing overall unlearning performance.
📝 Abstract
Large language models trained on vast corpora inherently risk memorizing harmful content that may later re-emerge in their outputs. To mitigate this issue, existing unlearning methods typically rely on training-based parameter updates, such as gradient ascent and its variants, to delete targeted content while preserving other knowledge. However, balancing the competing goals of forgetting and retention makes hyperparameter choices for these methods particularly difficult, often requiring repeated tuning to obtain a strong model that still leaves substantial room for improvement and transfers poorly across models and datasets. To address this challenge, we investigate whether unlearning runs exhibit exploitable structure in weight space, and observe that models from different runs still lie in a shared evaluation-performance basin. This suggests that stronger models may be recovered through an unlearning-tailored soup strategy, reducing the need for repeated tuning for further improvement or new settings. Motivated by this, we propose UnlearningSoup, a unified framework that provides two strategies: EfficientSoup uses binary-search-based interpolation to quickly discover a well-performing model in the early stage, where repeated tuning would otherwise make strong model selection costly. PerformanceSoup uses reweighted souping to efficiently unlock the remaining performance potential in the later stage, where repeated tuning becomes increasingly inefficient. Extensive experiments across diverse datasets and models show that UnlearningSoup delivers 2.4x to 3.3x efficiency gains in hyperparameter selection, while consistently improving performance across settings.
Problem

Research questions and friction points this paper is trying to address.

Large Language Model Unlearning
Hyperparameter Tuning
Forgetting-Retention Trade-off
Model Soups
Innovation

Methods, ideas, or system contributions that make the work stand out.

Machine Unlearning
Model Souping
Large Language Models
Hyperparameter Tuning
Weight Space