🤖 AI Summary
This study addresses the memory bottleneck in lifelong editing of large language models caused by the linear growth of LoRA adapter storage. We propose a cascaded residual compression method that stores edited knowledge via low-rank sketches and validates performance using probe prompts. For hard-to-edit samples failing to meet predefined thresholds, the rank is progressively increased until satisfactory performance is achieved. This adaptive mechanism precisely allocates high-rank resources to difficult samples, striking an optimal balance between sparse storage and comprehensive coverage. Experimental results demonstrate that our approach reduces memory consumption by 5.2× across multiple benchmarks while maintaining high efficacy even after fifty thousand consecutive edits.
📝 Abstract
Lifelong editing of LLMs requires storing thousands of edits after acquisition. A widely used family of approaches attaches one LoRA adapter per edit, which preserves behavior but grows linearly in storage. To address this challenge, we propose LadderEdit, a method that compresses each LoRA adapter after it is acquired. Each edit is first stored at low rank as a cheap sketch. We then check whether this sketch still satisfies the rewrite, generalization, and locality contract on probe prompts. Edits that pass keep the sketch; those that fail are promoted to a higher rank along a ladder until the contract is met. Because every edit retains some representation, coverage is maintained, and only hard edits consume more rank. Across ZsRE, CounterFact, and WikiBigEdit benchmarks on LLaMA-3-8B, Mistral-7B, and Qwen2.5-7B, LadderEdit tracks exact LoRA storage at 5.2x less memory and remains effective at 50,000 sequential edits.