Score
Designs and analyzes algorithms and inference procedures that iteratively infer and refine latent context representations from input exemplars at inference time, using dynamic or depth-adaptive update rules (including label-guided and contrastive mechanisms) to support multi-task in-context learning; and develops accompanying statistical analyses (e.g., excess-risk bounds and unified frameworks) that characterize when and how these adaptive in-context inference methods succeed.
This paper addresses a fundamental question in in-context learning (ICL): how large language models achieve zero-gradient generalization solely from task demonstrations in prompts. Critiquing the prevailing conceptual conflation of *skill identification* and *skill learning*, as well as fragmented analytical frameworks, we formally define both notions and propose a unified analysis framework grounded in data generation. Our framework reveals their shared reliance on implicit contextual modeling, while rigorously distinguishing identification—as pattern matching over context—versus learning—as distributional adaptation across tasks. Through conceptual abstraction, cross-method systematic review, and generative modeling paradigm reconstruction, we resolve key interpretability ambiguities in ICL, clarify the essential pathways to generalization and novel skill acquisition, and establish a theoretical foundation and methodological toolkit for controllable, predictable ICL prompt engineering.
This work addresses the challenge of enabling predictive systems to dynamically adapt their behavior based on contextual information for personalized inference. To this end, it proposes a unified framework that maps context into adaptation parameters for prediction and, for the first time, establishes a mathematical equivalence between explicit parameter adaptation and implicit expert routing under kernel ridge regression. The framework theoretically unifies diverse methodologies—including varying-coefficient models, local regression, prompt engineering, retrieval-augmented approaches, and mixture-of-experts—under fixed features and squared loss. Key contributions include deriving a general formulation for context-adaptive inference, proposing practical design principles and evaluation metrics such as adaptation efficiency and routing stability, and highlighting critical open problems concerning identifiability and robustness under distributional shifts.
This paper addresses the challenge of unifying diverse in-context learning (ICL) phenomena in large language models (LLMs)—including instruction following, role-playing, and temporal extrapolation—under a coherent theoretical framework. Method: We recast ICL as a meta-learning process: any context that nontrivially reduces subsequent prediction loss constitutes generalized ICL. Introducing the “ICL broad-spectrum view,” we integrate sequence distribution analysis with meta-learning theory to link ICL to foundational linguistic capabilities (e.g., coreference resolution, parallel structure processing) and systematically define multidimensional generalization metrics. Contribution/Results: We establish ICL as a meta-learning–driven universal adaptation mechanism—the first such unified theoretical perspective. Our framework clarifies distinct axes of generalization (e.g., task, domain, structural) and strengthens conceptual connections between ICL and emerging paradigms such as goal-directed agentic behavior. This advances both theoretical understanding and principled evaluation of LLM adaptation.
This paper addresses two fundamental challenges in foundation model research: the opaque nature of representation mechanisms and diminishing returns from scaling. To resolve these, we propose the “contexture” theory—a unified characterization of representation learning wherein optimal representations maximize mutual information between inputs and contextual variables, with peak generalization achieved at moderate contextual strength. We establish the first unified mathematical framework proving that scaling bottlenecks stem primarily from contextual *quality*, not scale. We introduce two general-purpose context-aware learning objectives—SVME and KISE—and a multi-context fusion strategy. Leveraging information theory and statistical learning theory, we derive a generalization bound for representation learning, unifying theoretical explanations across supervised, self-supervised, and generative pretraining paradigms. Empirical validation confirms that mainstream pretraining objectives implicitly optimize contexture. Our work provides both theoretical foundations and practical guidelines for designing efficient, context-driven pretraining frameworks.
Existing amortized learning approaches—including meta-learning, in-context learning, prompt tuning, and learned optimizers—face two key bottlenecks in task adaptation: (1) heterogeneous modeling of task-specific information across paradigms, and (2) poor scalability to long contexts and large-scale datasets during inference. To address these, we propose an **iterative amortized inference framework**, the first to unify in-context learning and learned optimizer paradigms under a single principled formulation. We introduce a taxonomy of amortization modes—parameterized, implicit, and explicit—and incorporate a mini-batch-based stochastic iterative update mechanism, enabling scalable adaptation to long sequences and massive datasets. Our framework significantly enhances flexibility and scalability in multi-task generalization. It establishes a novel theoretical and practical foundation for universal task adaptation, grounded in the synergistic co-design of optimization and inference.
This work investigates the behavioral limits of in-context learning (ICL) with thousands of demonstrations in ultra-long-context language models. Through systematic experiments across multiple models (e.g., Llama, Qwen) and datasets, and employing controlled analytical techniques—including random shuffling, label-based grouping, and demonstration subsampling—we find that: (1) ICL robustness to input ordering significantly increases with context length; (2) clustering examples by label degrades performance; and (3) gains do not arise from joint encoding of multiple demonstrations. Key contributions include: the first empirical demonstration that ICL performance scales continuously with demonstration count up to several thousand in large-label-space tasks; superior effectiveness over fine-tuning under low-to-moderate data regimes; and non-negligible gains achievable without fully utilizing available context capacity—challenging prevailing assumptions about ICL mechanisms.
This work investigates how context search can effectively enhance the performance of large language models in iterative reasoning, with a focus on the sampling complexity inherent in the critique-and-revise process. The authors model this process as approximate Bayesian inference over reasoning trajectories, where the base model provides a prior and self-reflection yields posterior updates. Theoretical analysis demonstrates that, under a local correction condition, only a polynomial number of sequential attempts are required to overcome the exponential decay in zero-shot success rates; moreover, when reflection reliably identifies early errors, context search yields exponential performance gains. This mechanism remains robust under approximate updates and can be efficiently learned via cross-entropy training. Key theoretical predictions are empirically validated on real large language models.
This study addresses the computational inefficiency and mechanistic opacity of in-context learning (ICL), which typically relies on exhaustive demonstrations. By revealing that attention head outputs exhibit stable affine transformation properties, this work proposes a "task operator" framework. Through analytical derivation of updated projection matrices, the method adapts to complex tasks without requiring fixed activation vectors. Combined with sparse circuit extraction and multi-batch operator averaging, it achieves precise zero-shot approximation of ICL. The proposed approach attains state-of-the-art performance across lexical, algorithmic, and reasoning tasks, substantially narrowing the performance gap between zero-shot inference and ICL. Ultimately, this research establishes a novel paradigm for efficient and scalable large language model reasoning.
Standard fine-tuning often degrades the in-context learning (ICL) capabilities of large language models and struggles to dynamically balance between ICL and in-weight learning (IWL) based on contextual relevance. This work proposes a contrastive context sampling strategy that mixes both similar and random examples during fine-tuning, incorporating multi-level context similarity to jointly train ICL and IWL. It reveals, for the first time, the critical role of structural similarity between context and target input in achieving an effective ICL–IWL trade-off. By leveraging a contrastive mechanism, the approach prevents the model from collapsing into pure copying, pure ICL, or pure IWL modes. Experiments across four large language models and diverse tasks demonstrate that the method consistently preserves stable hybrid reasoning capabilities and avoids mode collapse.
This work addresses the limitations of existing in-context learning approaches, which treat prompts merely as semantic cues and thus fail to enable task-adaptive dynamic computation, resulting in shallow and uninterpretable reasoning. To overcome this, the authors propose PromptPath, a novel framework that directly integrates prompt information into the model’s inference architecture. PromptPath employs a prompt-conditioned routing mechanism to dynamically activate and compose lightweight low-rank expert modules, thereby constructing task-specific computational pathways. This approach achieves dynamic adaptation at the computational level, significantly outperforming current methods on both 3D point cloud and 2D visual recognition benchmarks while demonstrating strong cross-domain and cross-task generalization capabilities.
This study addresses the unclear internal reasoning dynamics of large language models under strong contextual influence. Moving beyond conventional output-layer analyses, this work establishes a theoretical framework grounded in internal representations, integrating quantitative analysis with empirical validation to systematically characterize how context constrains reasoning processes. The findings reveal that predictive representations converge toward stable states, repeated assertions do not accumulate evidential weight, and predictive shifts are confined within query-dependent stable regions. These theoretical predictions align closely with observed behaviors, elucidating the intrinsic boundaries of contextual influence. Ultimately, this research provides both principled pathways and practical guidance for delineating the fundamental limits of context-driven reasoning in large language models.