LoRASculpt: Sculpting LoRA for Harmonizing General and Specialized Knowledge in Multimodal Large Language Models

📅 2025-03-21
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address harmful redundant parameters, catastrophic forgetting of general knowledge, and degraded downstream performance in vision-instruction fine-tuning of multimodal large language models (MLLMs) using LoRA, this paper proposes a synergistic framework combining sparse parameter updates with conflict-mitigating regularization. We introduce, for the first time within the LoRA paradigm, a theoretically grounded structured sparsity mechanism for parameter updates, alongside a knowledge-conflict-aware regularizer that explicitly suppresses interference between general and task-specific knowledge at the update-trajectory level. Our method improves both general capabilities (MMMU ↑) and downstream performance (OCRBench ↑), while adding ≤5% trainable parameters—outperforming standard LoRA and other adaptation methods. It effectively mitigates catastrophic forgetting and achieves balanced optimization of generality and specialization.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Computer Vision: Large Vision ModelsNatural Language Processing: (Large) Language Models

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
While Multimodal Large Language Models (MLLMs) excel at generalizing across modalities and tasks, effectively adapting them to specific downstream tasks while simultaneously retaining both general and specialized knowledge remains challenging. Although Low-Rank Adaptation (LoRA) is widely used to efficiently acquire specialized knowledge in MLLMs, it introduces substantial harmful redundancy during visual instruction tuning, which exacerbates the forgetting of general knowledge and degrades downstream task performance. To address this issue, we propose LoRASculpt to eliminate harmful redundant parameters, thereby harmonizing general and specialized knowledge. Specifically, under theoretical guarantees, we introduce sparse updates into LoRA to discard redundant parameters effectively. Furthermore, we propose a Conflict Mitigation Regularizer to refine the update trajectory of LoRA, mitigating knowledge conflicts with the pretrained weights. Extensive experimental results demonstrate that even at very high degree of sparsity ($le$ 5%), our method simultaneously enhances generalization and downstream task performance. This confirms that our approach effectively mitigates the catastrophic forgetting issue and further promotes knowledge harmonization in MLLMs.
Problem

Research questions and friction points this paper is trying to address.

Balancing general and specialized knowledge in MLLMs
Reducing harmful redundancy in LoRA adaptation
Mitigating catastrophic forgetting during visual instruction tuning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse updates in LoRA to discard redundancy
Conflict Mitigation Regularizer for update refinement
Harmonizes general and specialized knowledge effectively
🔎 Similar Papers
No similar papers found.