Machine Unlearning for Large Language Models: Foundations, Advances, and Agentic Extensions

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of machine unlearning in large language models, wherein eliminating the influence of target data often compromises retained capabilities within complex systems such as retrieval-augmented generation, long-term memory, and multi-agent frameworks. To tackle this, the work proposes a five-layer architecture linking removal requests to supporting claims alongside a seven-stage lifecycle framework, establishes a six-dimensional evidence evaluation taxonomy, and systematically reviews existing unlearning algorithms across these complex settings. The findings reveal that target construction and recovery testing critically determine unlearning efficacy. Furthermore, the authors demonstrate that current evidence remains insufficient to substantiate influence removal across external states, underscoring the necessity of dependency tracking to verify whether the influence of targeted data re-emerges over time.
📝 Abstract
Machine unlearning aims to remove target influence while preserving other capabilities. This survey compares methods, benchmarks, and evidence across large language models and systems using retrieval, memory, tools, and interacting agents. A five-layer framework connects removal requests, system boundaries, target locations, interventions, and supported claims. A seven-stage lifecycle and six evidence dimensions guide comparison. The review shows that target construction, retained data, and recovery tests affect reported outcomes. Evidence from model evaluations remains insufficient to establish removal across external state and subsequent updates, motivating evaluation that tracks dependencies and tests whether target influence returns.
Problem

Research questions and friction points this paper is trying to address.

Machine Unlearning
Large Language Models
Target Influence Removal
Evaluation Evidence
Agentic Systems
Innovation

Methods, ideas, or system contributions that make the work stand out.

Machine Unlearning
Large Language Models
Agentic Systems
Evaluation Framework
Evidence Dimensions
🔎 Similar Papers