🤖 AI Summary
This study addresses the challenges of insufficient knowledge coverage and difficult integration of cross-modal complementary information in CLIP-based incremental learning by proposing DuLBE, a framework for exemplar-free class-incremental learning. Methodologically, it introduces a novel gradient routing mechanism that coordinates shared and residual low-rank modes, leveraging dual-mode low-rank adaptation to balance model stability and plasticity. Furthermore, the framework constructs a hyperspherical geodesic prototype bridge integrated with model ensembling techniques to compensate for deficient textual decision boundaries, thereby optimizing vision-language alignment. Experimental results demonstrate that the proposed method achieves state-of-the-art performance across various settings while maintaining high parameter efficiency.
📝 Abstract
Benefiting from transferable visual-textual alignment, CLIP has been widely adopted for class-incremental learning (CIL). However, existing learners either repeatedly update components shared across tasks, leading to knowledge overwriting, or overly isolate new-task updates, hindering the reuse of CLIP's transferable knowledge and limiting plasticity. Moreover, the text-based or bimodal classifier designs still fail to effectively integrate complementary information from the visual and textual modalities. To address these challenges, we introduce DuLBE, which couples dual-mode low-rank learning with a bridge-prototype ensemble classifier for exemplar-free CIL. DuLBE allocates two visual low-rank update modes according to the gradient demand and uses gradient routing to coordinate them: a compact and rewritable shared mode is selected from historically occupied visual directions to reuse transferable knowledge, while residual modes provide low-interference channels for task-specific variations. Building on the resulting stable inter-modal structure, we further construct geodesic bridges between visual prototypes and text embeddings on the unit hypersphere, and ensemble reliable bridge prototypes to compensate for the modality-gap limitations of textual decision boundaries. Extensive experiments under multiple settings show that DuLBE achieves state-of-the-art CIL performance while retaining the high parameter efficiency of low-rank tuning.