🤖 AI Summary
This study addresses the unclear dynamic causes underlying the widening gap between training and validation performance during pre-trained model adaptation. It proposes a dynamic structural interpretation framework, revealing that the shift of update pressure from general to narrow support constitutes the core mechanism driving this generalization gap. By establishing a theoretical link between gradient allocation heterogeneity and performance divergence, the framework enables the observation of structural evolution without requiring a validation set. Through ResMLP simulations, multimodal experiments involving RoBERTa, DeBERTa, Qwen, and ResNet-18, alongside fixed-probe techniques, the work demonstrates that the proposed monitoring metric is significantly positively correlated with the accuracy gap. These findings validate the theory's broad applicability across both natural language processing and computer vision tasks.
📝 Abstract
Train-validation separation is the evolving difference between performance on observed training examples and a finite held-out validation set. We propose a dynamic structural account of how this gap develops during adaptation of pretrained models: continued fitting can shift update demand from broadly reusable support toward narrower support with weaker held-out transfer. A conditional local model links this shift to increasing heterogeneity in gradient allocation and train-validation separation. Fixed training probes make this structural evolution observable without validation examples entering the readouts; held-out performance is used separately to evaluate its relation to the gap. In a constructed hierarchy implemented with a residual multilayer perceptron (ResMLP), increasing the target share of example-private features from $p=.3$ to $.5$ to $.7$, while preserving the relative mixture $1{:}2{:}3{:}4$ among the four shared feature levels, increases the final mean accuracy gap from $.185$ to $.331$ to $.527$ across five runs per condition. Masked-input losses measured separately on training and validation examples expose the corresponding transfer asymmetry. The natural language processing (NLP) analysis uses 10-epoch runs of RoBERTa, DeBERTa, and Qwen on six datasets (90 runs): the training-probe-weighted within-class and overall dispersion readouts each have positive raw and smoothed level correlations with the accuracy gap in all 90 runs. Raw changes paired at approximately one-epoch intervals remain positively associated in 86/90 and 87/90 runs, respectively. A 40-epoch ResNet-18 study tests both readouts on three vision datasets. Together, controlled simulation, NLP, and vision support the dynamic structural account across settings, with real-model evidence testing its observable predictions under the specified monitors.