When the Merge Coefficient Stops Mattering: Proximity Regularized Merging for Continual LoRA Adaptation

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of LoRA merging in replay-free continual learning, where performance is highly sensitive to merging coefficients and severe interference arises among task vectors. To mitigate these issues, this work proposes a proximity-regularized merging strategy. We theoretically demonstrate that merging performance depends fundamentally on the intrinsic mergeability of task vectors rather than solely on scaling coefficients. Accordingly, a proximity penalty is introduced during training to optimize this property, complemented by a Fisher-weighting mechanism to suppress parameter interference. The proposed approach significantly improves average accuracy across diverse scenarios, effectively broadens the stable range of merging coefficients, and achieves a favorable balance between model stability and plasticity.
📝 Abstract
Rehearsal-free continual learning with parameter-efficient adapters can be cast as a sequence of task-vector write-in operations: for each new task, a low-rank adapter is learned and merged into a running model. We propose Proximity Regularized Merging (PRM), a minimal modification to sequential LoRA merging that adds a proximal penalty during task-vector training without changing the subsequent write-in rule. PRM acts as a robust task-vector regularizer: in the reported Base->+Prox diagnostics, it improves AAA across multiple write-in rules, backbones, and class-incremental settings, while its fixed-coefficient variant remains competitive with strong coefficient-based baselines. Mechanistically, matched-prefix norm controls and proximal-strength sweeps show that proximal training shrinks the task-vector radius, lowers Fisher-weighted interference, broadens the coefficient plateau, and exposes a stability-plasticity trade-off. Together, these results suggest that the effectiveness of sequential LoRA merging depends not only on how much of a task vector is written in, but also on whether the task vector itself has been trained to be mergeable.
Problem

Research questions and friction points this paper is trying to address.

Continual Learning
LoRA Merging
Task Vector
Parameter-Efficient Adaptation
Stability-Plasticity Trade-off
Innovation

Methods, ideas, or system contributions that make the work stand out.

Continual Learning
Proximity Regularized Merging
Task Vector
LoRA
Parameter-Efficient Fine-Tuning
🔎 Similar Papers
No similar papers found.