🤖 AI Summary
This study addresses the significant performance degradation in retrieval-based in-context learning caused by distribution shifts between source and target domains. To overcome this limitation, the work formalizes the paradigm as a domain adaptation problem and introduces a flexible class of distribution shifts to quantify theoretical benefits and potential pitfalls, thereby transcending the constraints of existing analyses. By integrating domain adaptation theory with distribution shift analysis, this research systematically validates its theoretical predictions across both synthetic datasets and natural language tasks. Ultimately, it delineates the applicability boundaries and performance advantages of retrieval-based in-context learning, providing a rigorous theoretical foundation for understanding when and how retrieval augmentation yields reliable improvements under varying distributional conditions.
📝 Abstract
In-context retrieval (ICR) is a retrieval-based form of in-context learning (ICL) in which demonstrations are retrieved from a source database based on similarity to the query, rather than sampled independently. In this work, we formulate ICR as a type of domain adaptation problem, where the source distribution $P$ of the database may differ from the target distribution $Q$ of the test query-label pair. We investigate the performance of ICR under a flexible class of distributional shifts that substantially extends prior work \citep{li2024fine,guo2025retrieval}, and establish theoretical guarantees that quantify the benefits and pitfalls of this learning paradigm. Our theory is verified by experiments on synthetic and language tasks.