In-context Learning of Single-index Targets: Comparing Kernel and Feature Learners

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the theoretical gap in understanding in-context learning for nonlinear single-index models by investigating how architecture selection influences learning performance. Specifically, it compares kernel-based and feature-based single-layer attention architectures, integrating fixed nonlinear mappings, learnable readout layers, and the replica method from statistical physics to derive analytical predictions of generalization error. Phase diagrams are constructed to reveal the joint effects of pretraining scale, task diversity, and context length on this error, delineating distinct regimes of architectural advantage under varying data and compute budgets while uncovering fundamentally different context-length scaling laws between the two designs. Theoretical predictions align closely with empirical results, elucidating the interplay between data characteristics and model architectures in nonlinear in-context learning.
📝 Abstract
In-context learning (ICL) enables a pretrained model to infer a task from demonstrations without updating its parameters. While much of the existing theory focuses on linear target functions, in this paper we study nonlinear cases by comparing two one-layer attention architectures on the same family of single-index tasks. A kernel learner first maps inputs through a fixed nonlinear feature map and then applies linear attention, whereas a feature learner applies attention to the original input, followed by a learned nonlinear readout. We derive predictions for their memorization and generalization errors using the replica method, retaining the effects of pretraining size, task-pool diversity, and training and inference context lengths. The resulting predictions closely match numerical experiments across a broad range of regimes. Our analysis yields phase diagrams that characterize when each architecture is advantageous as the amount of pretraining data, task diversity, and context lengths vary. We further identify qualitatively different context-length scalings for the two learners. Together, these results clarify how architectural choices interact with the dataset and govern nonlinear in-context learning.
Problem

Research questions and friction points this paper is trying to address.

in-context learning
single-index targets
nonlinear functions
kernel learner
feature learner
Innovation

Methods, ideas, or system contributions that make the work stand out.

In-context Learning
Single-index Models
Replica Method
Kernel Learner
Feature Learner
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.