🤖 AI Summary
This study investigates the dynamic evolution of semantic capabilities during Vision-Language Model (VLM) fine-tuning, revealing an "acquisition-peak-decay" trajectory where over-specialization leads to performance degradation. We construct a trajectory analysis framework based on DFN and CLIP backbones to track the evolutionary patterns of identity and attribute capabilities. Furthermore, we propose a Structured Semantic Routing (SSR) mechanism combined with random name-branch dropout to effectively decouple capability acquisition from late-stage specialization. Our findings demonstrate significant temporal discrepancies in the peak performance of different semantic capabilities. Notably, SSR substantially enhances name-free attribute retrieval, highlighting the critical influence of supervision structure on capability acquisition. This work establishes a novel paradigm for mitigating over-specialization in VLM fine-tuning.
📝 Abstract
Fine-tuning vision-language models (VLMs) is typically evaluated at a single downstream checkpoint, obscuring whether a semantic capability was never acquired or emerged earlier and later declined during specialization. We ask how semantic capabilities are acquired, when they peak, how well they transfer, and what remains at deployment. We study these dynamics as a semantic capability trajectory, tracking identity- and attribute-based capabilities over training.
We formulate a trajectory-based framework that separates capability acquisition, capability-specific optima, and later specialization, and introduce Structured Semantic Routing (SSR) to study how the representation of supervision shapes what is acquired. Across six pretrained backbones spanning DFN, MetaCLIP, and OpenAI CLIP, we show that fine-tuning can acquire semantic capability beyond the pretrained state, including gains observed on held-out evaluations. Unstructured name-and-attribute supervision produces strong name-and-attribute retrieval with comparatively weak name-free attribute-profile retrieval, whereas SSR yields substantially stronger name-free attribute-profile retrieval and is further strengthened by stochastic name-branch dropout. Different capabilities can peak at different stages, so a checkpoint selected by target class-name retrieval need not coincide with a transferable semantic optimum. Continued optimization can therefore preserve strong target class-name retrieval while reducing previously acquired transferable semantic capability. In a representative diagnostic study, this late specialization is consistent with reduced cross-modal semantic accessibility while substantial image-only class structure remains available.