🤖 AI Summary
This study addresses the challenges of catastrophic forgetting and insufficient adaptation to novel tasks in multimodal continual instruction tuning (MCIT). To this end, we propose ReCAP, a framework that pioneers the integration of externally retrieved knowledge into MCIT to guide capability reuse. Specifically, ReCAP leverages retrieval-augmented generation and large language model assistance to construct domain-specific and reasoning knowledge bases. Furthermore, it introduces an adaptive subspace reclamation mechanism within parameter-efficient fine-tuning to preserve critical parameter directions while reclaiming residual capacity, thereby enabling stable cross-stage reuse of capability modules. Extensive experiments demonstrate that ReCAP achieves state-of-the-art performance across multiple MCIT benchmarks, effectively mitigating catastrophic forgetting and significantly enhancing model adaptability to new tasks.
📝 Abstract
Multimodal continual instruction tuning (MCIT) aims to enable multimodal large language models to acquire new capabilities from sequential tasks while preserving previously learned knowledge. Existing methods primarily mitigate catastrophic forgetting by constraining parameter updates or separating task-specific adaptations. However, continual adaptation can also benefit from external knowledge that provides domain-specific information and reusable reasoning patterns for solving diverse instructions. For example, to answer "How many red cubes are to the left of the sphere?", domain knowledge can provide relevant concepts about objects and spatial relations, while reasoning knowledge can specify ordered operations such as object recognition, spatial filtering, and counting. Despite this potential, how to leverage external knowledge for continual adaptation remains largely unexplored in existing MCIT methods. To this end, we propose ReCAP, a retrieval-guided framework that leverages external knowledge to guide capability reuse during continual adaptation. At each continual stage, ReCAP uses external search and an LLM to incrementally build a knowledge base of domain, reasoning, and format knowledge based on the current-stage training data. For each instruction, retrieved domain knowledge guides generation, while retrieved reasoning knowledge selects and orders capability modules to form an instance-specific capability path. As these capability modules are reused across stages, subsequent adaptation can overwrite previously learned parameters. To enable stable cross-stage reuse, ReCAP introduces adaptive subspace recycling, which parameterizes reusable capability modules with shared bases and stage-specific cores, protects historically important directions while recycling residual capacity. Extensive experiments on MCIT benchmarks show that ReCAP achieves SOTA performance.