🤖 AI Summary
This study addresses the degradation of in-context learning in Transformers under out-of-distribution (OOD) inputs and the lack of corresponding theoretical explanations. Grounded in attention dynamics theory, this work analyzes the OOD error mechanism through the temporal evolution of α-type and β-type attention weights. Furthermore, by adopting a feature interaction perspective, it elucidates how fine-tuning influences source domain forgetting. Theoretically, we demonstrate that post-fine-tuning performance on the source domain does not necessarily deteriorate; rather, it depends critically on the nature of the feature shift. Empirically, these theoretical insights are validated across both synthetic and real-world datasets. Ultimately, this research clarifies attention behavior patterns in OOD scenarios, providing a rigorous theoretical foundation for enhancing model robustness.
📝 Abstract
Transformers have demonstrated remarkable in-context learning (ICL) capabilities, enabling them to perform new tasks without additional fine-tuning. However, their performance often deteriorates when encountering out-of-distribution (OOD) inputs that deviate from the training distribution, and the underlying theory remains poorly understood. To fill this gap, we characterize the OOD error under the input distribution shift through the interplay between the dynamics of the so-called $α$-type and $β$-type attention weights, which represent the transformer's confidence in identifying the correct and incorrect features, respectively. Our results indicate that the OOD error for each feature depends on all pairwise interactions between the training features and OOD features, and under certain cases the transformer performs no better than random guessing. To improve the OOD generalization performance, we next investigate the impact of model finetuning with the OOD data, and particularly, characterize the model forgetting performance on the source domain. Interestingly, the performance on the source domain may not always degrade after finetuning, which highly depends on the nature of the feature shift: finetuning on OOD domain keeps enhancing the confidence of identifying correct features from the original distribution, while the interference from other incorrect features may either increase or decrease. Extensive experiments on both synthetic and real data are conducted to corroborate the theoretical insights.