ChainLoRA: Geometry-Preserving Task Vector Merging for Continual Learning in LLMs

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of balancing knowledge retention, new task adaptation, and strict parameter budgets during the continual fine-tuning of large language models. To this end, we propose a replay-free continual merging framework. Methodologically, we decouple directional overlap from coefficient coupling through a geometric perspective to mitigate catastrophic forgetting. Furthermore, we achieve low-overhead incremental learning by chaining task vector updates integrated with Procrustes adaptive alignment, adaptive singular value decomposition, and parameter-efficient fine-tuning techniques. Experimental results demonstrate that our approach attains state-of-the-art performance under replay-free settings on the Large and SuperNI benchmarks, while closely approaching the average performance of replay-based methods.
📝 Abstract
Continual parameter-efficient fine-tuning for large language models (LLMs) must balance retention of previously acquired knowledge, adaptation to new tasks, and strict parameter budgets. We present \textbf{ChainLoRA}, a replay-free continual merging framework built on chain-updated task-vector geometry. From a parameter-merging perspective, we formulate a geometric view of forgetting through a measurable interaction between task updates, separating directional overlap from coefficient coupling. Building on this view, ChainLoRA combines chain-updated training with post-stream adaptive SVD merging. During training, initialization and a one-sided orthogonality proxy use only the last carrier, keeping their historical-state footprint and regularization overhead constant as the task stream grows. At merging time, Adaptive SVD extracts a shared carrier and aligns it to the latest task through Procrustes adaptation. Our theoretical analysis shows that Procrustes adaptation facilitates geometric approximate separation of shared and task-specific components. The one-sided proxy further bounds inter-task interference. An effective-rank penalty additionally promotes efficient utilization of the task subspace during continual learning. Experiments show that ChainLoRA achieves state-of-the-art performance among the evaluated replay-free methods on the Large and SuperNI benchmarks, while remaining competitive on Standard CL and attaining almost the closest average scores to the evaluated replay-based method across all three benchmarks.
Problem

Research questions and friction points this paper is trying to address.

Continual Learning
Large Language Models
Parameter-Efficient Fine-Tuning
Catastrophic Forgetting
Task Vector Merging
Innovation

Methods, ideas, or system contributions that make the work stand out.

Continual Learning
Task Vector Merging
LoRA
Adaptive SVD
Procrustes Adaptation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hang Yin
Shanghai Jiao Tong University
H
Haozhe Wang
Shanghai Jiao Tong University
Y
Yuhua Luo
Shanghai Jiao Tong University
Z
Zhangqi Pan
Huazhong University of Science and Technology
Xiaoxing Wang
Xiaoxing Wang
SJTU
Machine LearningAutoMLNeural Architecture Search
Junchi Yan
Junchi Yan
FIAPR & ICML Board Member, SJTU (2018-), SII (2024-), AWS (2019-2022), IBM (2011-2018)
Computational IntelligenceAI4ScienceMachine LearningAutonomous Driving