Capturing In-Context Learning Dynamics with Task Operators

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the computational inefficiency and mechanistic opacity of in-context learning (ICL), which typically relies on exhaustive demonstrations. By revealing that attention head outputs exhibit stable affine transformation properties, this work proposes a "task operator" framework. Through analytical derivation of updated projection matrices, the method adapts to complex tasks without requiring fixed activation vectors. Combined with sparse circuit extraction and multi-batch operator averaging, it achieves precise zero-shot approximation of ICL. The proposed approach attains state-of-the-art performance across lexical, algorithmic, and reasoning tasks, substantially narrowing the performance gap between zero-shot inference and ICL. Ultimately, this research establishes a novel paradigm for efficient and scalable large language model reasoning.
📝 Abstract
In-context learning (ICL) enables language models to perform new tasks from demonstrations without weight updates. However, every ICL inference requires processing the full set of examples, resulting in inefficient deployments, and how ICL works mechanistically is not fully understood. Prior work compresses ICL into fixed activation vectors extracted from specific layers or positions, but these input-independent interventions fail on complex tasks where the output depends on fine-grained interactions with the input. By analyzing the ICL forward pass, we show that each attention head's output is an affine transformation of its context-masked counterpart, and that the parameters of this transformation are empirically stable across samples for a given task. Building on this, we introduce Task Operator (TO), which replays this transformation as an analytically derived update to the attention output projection. Across lexical, algorithmic, and reasoning tasks, TO achieves the best overall performance among prior methods and substantially narrows the gap between zero-shot inference and ICL. We further show that the extracted knowledge concentrates in a task-specific sparse circuit across layers and positions, and that averaging operators from disjoint demonstration batches enables effective many-shot scaling without expanding the context window. Our code is available at https://github.com/gzxiong/task_operator.
Problem

Research questions and friction points this paper is trying to address.

In-Context Learning
Inference Efficiency
Mechanistic Interpretability
Task Compression
Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

In-Context Learning
Task Operator
Affine Transformation
Sparse Circuit
Many-Shot Scaling
🔎 Similar Papers
No similar papers found.