🤖 AI Summary
This study addresses the semantic gap between textual and visual modalities that severely limits classification performance when adapting CLIP to class-incremental learning. To overcome this limitation, we propose VIS, a pioneering framework that establishes a purely visual pathway for CLIP-based class-incremental learning. By entirely discarding the text branch, VIS circumvents the modality gap. It constructs a unified representation space by fusing multi-layer visual features from CLIP and introduces a Kernelized Incremental Least Squares Support Vector Machine (KILS-SVM), which leverages closed-form solutions to enable efficient classifier weight updates. Notably, VIS achieves state-of-the-art performance in class-incremental learning without any textual guidance, offering a novel and highly effective paradigm for continual learning based on pretrained vision-language models.
📝 Abstract
Class-Incremental Learning (CIL) requires models to recognize new classes over time without forgetting previously learned ones. With the rise of vision-language pre-training, CLIP has become a strong foundation for CIL. A common design in CLIP-based CIL is to construct textual classifier weights by encoding class-name templates with the CLIP text encoder, and then classify visual features by image-text cosine similarity. This design is appealing: since CLIP aligns images and text in a shared embedding space, textual weights appear to provide an off-the-shelf classifier for incremental classes. However, we show that this seemingly natural design is not always beneficial, as a modality gap can still separate the two modalities and make textual classifier weights deviate from visual class distributions. Empirically, under identical task-wise CIL training, initializing the cosine classifier with visual class centers yields lower loss and better incremental accuracy than using CLIP textual features.Motivated by these observations, we propose VIS, a visual-only method for CLIP-based CIL that removes the deployed textual branch and constructs the incremental classifier entirely in the visual space. To obtain stronger task-adaptive visual representations, VISuses only base-session data to enhance CLIP's final visual representation with informative visual-layer features. Built on the enhanced visual representation, VISemploys a simple kernelized incremental least-squares SVM, whose classifier weights are solved in closed form from additive sufficient statistics. When new classes arrive, VISaccumulates their sufficient statistics and recomputes the classifier weights for all seen classes, enabling efficient incremental updates while preserving historical class knowledge. Extensive experiments show that VISachieves state-of-the-art performance without a textual branch.