Transformers as Cross-Task Learners: Shared Structure Drives Sample Efficiency in In-Context Learning

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of rigorous statistical learning theory underpinning in-context learning in Transformers. To bridge this gap, it pioneers the quantification of cross-task complexity for general families of nonlinear tasks. By leveraging covering number theory and anchor function aggregation, the authors construct an identification-evaluation procedure and explicitly design a Softmax attention-based Transformer to approximate it. The work demonstrates that, given sufficient pretraining, the required prompt length becomes independent of task dimensionality, enabling the model to effectively exploit low-dimensional structures. Furthermore, it provides a quantitative theoretical explanation for the intrinsic mechanisms through which joint pretraining enhances generalization capabilities.
📝 Abstract
Transformers achieve remarkable performance by jointly learning broad families of tasks during pretraining and adapting to unseen tasks from only a short prompt. Yet a rigorous mathematical and statistical understanding of this phenomenon remains limited. This paper aims to study how Transformers exploit shared cross-task structure and how this structure affects the sample complexity of in-context learning (ICL). Specifically, we characterize task-space complexity through covering numbers under a prescribed metric, thereby quantifying the low-dimensional cross-task structure without requiring an explicit parametric representation. The resulting cover provides a set of anchor functions, which we use to introduce a task-identification-and-evaluation procedure: context observations localize an unseen task among the anchor functions, and the response at a query is predicted by aggregating the corresponding anchor function query evaluations. For approximation, we explicitly construct a Transformer with Softmax attention to approximate this procedure. For generalization, we derive an error bound that separates the effects of the number of pretraining tasks and the prompt length. The scaling with respect to the number of pretraining tasks is governed by the intrinsic dimensions of the task space and input domain; once sufficiently many tasks are available, the dependence on the prompt context length becomes dimension-free. To the best of our knowledge, this is the first work to quantify cross-task complexity for general nonlinear task families and explicitly construct a Transformer that exploits their low-dimensional structure to perform ICL. Our theory provides a quantitative explanation of how joint pretraining across related tasks improves in-context generalization.
Problem

Research questions and friction points this paper is trying to address.

Transformers
In-Context Learning
Cross-Task Structure
Sample Complexity
Task-Space Complexity
Innovation

Methods, ideas, or system contributions that make the work stand out.

In-Context Learning
Cross-Task Structure
Covering Numbers
Sample Complexity
Transformers
🔎 Similar Papers
No similar papers found.