Dual-Mode Low-Rank Learner with Bridge-Prototype Ensemble for Vision-Language Class-Incremental Learning

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of insufficient knowledge coverage and difficult integration of cross-modal complementary information in CLIP-based incremental learning by proposing DuLBE, a framework for exemplar-free class-incremental learning. Methodologically, it introduces a novel gradient routing mechanism that coordinates shared and residual low-rank modes, leveraging dual-mode low-rank adaptation to balance model stability and plasticity. Furthermore, the framework constructs a hyperspherical geodesic prototype bridge integrated with model ensembling techniques to compensate for deficient textual decision boundaries, thereby optimizing vision-language alignment. Experimental results demonstrate that the proposed method achieves state-of-the-art performance across various settings while maintaining high parameter efficiency.
📝 Abstract
Benefiting from transferable visual-textual alignment, CLIP has been widely adopted for class-incremental learning (CIL). However, existing learners either repeatedly update components shared across tasks, leading to knowledge overwriting, or overly isolate new-task updates, hindering the reuse of CLIP's transferable knowledge and limiting plasticity. Moreover, the text-based or bimodal classifier designs still fail to effectively integrate complementary information from the visual and textual modalities. To address these challenges, we introduce DuLBE, which couples dual-mode low-rank learning with a bridge-prototype ensemble classifier for exemplar-free CIL. DuLBE allocates two visual low-rank update modes according to the gradient demand and uses gradient routing to coordinate them: a compact and rewritable shared mode is selected from historically occupied visual directions to reuse transferable knowledge, while residual modes provide low-interference channels for task-specific variations. Building on the resulting stable inter-modal structure, we further construct geodesic bridges between visual prototypes and text embeddings on the unit hypersphere, and ensemble reliable bridge prototypes to compensate for the modality-gap limitations of textual decision boundaries. Extensive experiments under multiple settings show that DuLBE achieves state-of-the-art CIL performance while retaining the high parameter efficiency of low-rank tuning.
Problem

Research questions and friction points this paper is trying to address.

Class-Incremental Learning
Vision-Language
Knowledge Overwriting
Plasticity
Modality Gap
Innovation

Methods, ideas, or system contributions that make the work stand out.

Class-Incremental Learning
Low-Rank Adaptation
Vision-Language Models
Prototype Ensemble
Geodesic Bridge
🔎 Similar Papers
No similar papers found.