🤖 AI Summary
This study addresses the challenge in continual adaptation of vision-language models, where learning new classes disrupts historical predictions and degrades zero-shot capabilities. To mitigate this, we propose a Hierarchical Response Preservation method that reconciles the competition between old and new knowledge through hierarchical probability modeling and geometrically constrained updates. This approach preserves discriminability among historical classes while enabling flexible adaptation to novel ones, further incorporating a zero-shot vision-language alignment algorithm to optimize the update trajectory. Experiments across three class-incremental settings demonstrate that our method improves average accuracy by 1.84 to 7.95 percentage points over the strongest baseline, significantly alleviating zero-shot transfer degradation.
📝 Abstract
Pretrained graph-text models align graph representations with textual semantics, enabling recognition of unseen classes and transfer across graph domains. However, as graph data and classes continually arrive, models should learn from new supervision while retaining their zero-shot transfer capabilities and historical task knowledge. Two challenges arise: (i) new classes can overturn historical predictions despite preserved distinctions among historical classes, and (ii) overly strict response preservation can stall learning of new classes. To address these challenges, we propose Hierarchical Response Preservation (HiRP). HiRP represents this competition through a hierarchical response that keeps each historical-class probability and sums new-class probabilities, preserving historical distinctions and aggregate competition while allowing distinctions within the new class group to adapt. It further uses the geometry induced by this response to guide constrained updates, retaining useful adaptation directions while controlling response drift. Across three class-incremental settings, HiRP achieves absolute gains of 1.84-7.95 percentage points in average accuracy over the strongest compared baseline in each setting, while mitigating zero-shot transfer degradation.