Score
Design and build models or inference procedures that condition predictions on a small labeled support set at inference time so the system adapts to a new task or domain without gradient-based fine-tuning. Implement encoders, attention/retrieval mechanisms, and context-conditioned decoders that ingest support examples and produce outputs (e.g., next-step or class predictions) reflecting the specific patterns in the support set.
This work addresses the challenge of enabling predictive systems to dynamically adapt their behavior based on contextual information for personalized inference. To this end, it proposes a unified framework that maps context into adaptation parameters for prediction and, for the first time, establishes a mathematical equivalence between explicit parameter adaptation and implicit expert routing under kernel ridge regression. The framework theoretically unifies diverse methodologies—including varying-coefficient models, local regression, prompt engineering, retrieval-augmented approaches, and mixture-of-experts—under fixed features and squared loss. Key contributions include deriving a general formulation for context-adaptive inference, proposing practical design principles and evaluation metrics such as adaptation efficiency and routing stability, and highlighting critical open problems concerning identifiability and robustness under distributional shifts.
This work addresses the limited flexibility of existing time series foundation models in performing multi-task reasoning through explicit instructions and contextual examples. To overcome this, we propose an instruction-conditioned in-context learning framework for time series modeling that jointly integrates structured instructions with contextual examples for the first time. Built upon a quantile regression T5 architecture, our model employs a hierarchical Transformer to separately handle intra-example encoding, inter-example fusion, and cross-example attention. Trained via a hybrid self-supervised and supervised multi-task objective—encompassing forecasting, imputation, reconstruction, classification, and anomaly detection—the model achieves general-purpose inference without task-specific fine-tuning. Experimental results on benchmarks such as FEV-Bench and GIFT-Eval demonstrate superior performance in both point and probabilistic forecasting compared to strong baselines, while maintaining competitive accuracy in classification and anomaly detection tasks.
This work addresses the high computational cost and limited flexibility of fine-tuning for behavioral steering of large language models (LLMs). We propose a parameter-efficient steering paradigm based on neologisms—novel, task-specific tokens—optimizing only their embeddings (≈d parameters) while freezing all original model weights. This enables activation of targeted response patterns without compromising pre-trained capabilities. Our key contributions are: (i) the first formulation of neologism learning as a highly efficient alternative to conventional fine-tuning; (ii) empirical evidence that LLMs possess intrinsic semantic compositionality, enabling autonomous construction and grounding of neologism meanings; (iii) superior performance over LoRA under identical settings, with >99% reduction in trainable parameters and computational overhead; and (iv) native support for concurrent multi-behavior execution and dynamic behavioral switching. Experiments demonstrate strong controllability, exceptional efficiency, and inherent interpretability.
This paper addresses the challenge of unifying diverse in-context learning (ICL) phenomena in large language models (LLMs)—including instruction following, role-playing, and temporal extrapolation—under a coherent theoretical framework. Method: We recast ICL as a meta-learning process: any context that nontrivially reduces subsequent prediction loss constitutes generalized ICL. Introducing the “ICL broad-spectrum view,” we integrate sequence distribution analysis with meta-learning theory to link ICL to foundational linguistic capabilities (e.g., coreference resolution, parallel structure processing) and systematically define multidimensional generalization metrics. Contribution/Results: We establish ICL as a meta-learning–driven universal adaptation mechanism—the first such unified theoretical perspective. Our framework clarifies distinct axes of generalization (e.g., task, domain, structural) and strengthens conceptual connections between ICL and emerging paradigms such as goal-directed agentic behavior. This advances both theoretical understanding and principled evaluation of LLM adaptation.
This work investigates how Transformers dynamically acquire inductive capabilities during in-context learning (ICL), specifically focusing on the role of “inductive heads” in transitioning from local n-gram pattern recognition to modeling long-range dependencies. Method: We combine theoretical approximation analysis, synthetic task training dynamics modeling, attention decomposition, and mixed-objective trajectory tracking across training. Contribution/Results: We formally characterize the generalized inductive head mechanism for the first time, revealing a sharp, non-gradual phase transition—from 4-gram modeling to inductive head emergence—during training. We quantify the layer- and head-specific contributions to long-range dependency capture and demonstrate that inductive heads constitute the core architectural substrate underlying ICL emergence. Our study provides the first full-training-dynamics evidence and an interpretable framework for understanding how large language models dynamically generalize, bridging mechanistic analysis with empirical learning trajectories.
研究比较了大型语言模型通过规则和示例进行上下文学习的效果,发现模型从规则中学习比仅从示例中学习更可靠,特别是在需要代数抽象的任务中。
This work addresses the degradation of in-context learning (ICL) capabilities in large language models (LLMs) during full-parameter fine-tuning, which often undermines few-shot generalization. Leveraging a linear attention framework, the authors theoretically elucidate how standard fine-tuning disrupts the mechanisms underlying ICL. To mitigate this issue, they propose a constrained fine-tuning strategy that updates only the value matrices, thereby enhancing zero-shot performance on the target task while effectively preserving few-shot learning abilities. Further analysis demonstrates that incorporating an auxiliary few-shot loss can amplify this benefit under certain conditions. Both theoretical insights and empirical results substantiate that restricting the set of tunable parameters offers a principled and effective approach to jointly optimizing zero-shot and in-context learning performance.
This study investigates how the decodability of internal model representations dynamically evolves throughout pretraining and post-training, and whether erroneous decodability alone can reliably indicate discarded output information. Utilizing the Pythia model suite, the authors employ linear probing and steering intervention techniques to conduct cross-checkpoint comparative analyses of probe accuracy, steered responses, and error-correction mechanisms from early to late training stages. The work proposes an information-theoretic counterexample demonstrating that erroneous decodability is insufficient to establish the loss of output information. Furthermore, it reveals that while steering benefits improve progressively over the course of training, final-state decoders do not exhibit significant advantages. These findings offer novel perspectives for understanding the evolution of internal mechanisms within large language models.
Existing neural algorithmic reasoning approaches are limited by insufficient representational capacity of their encoders, hindering effective simulation of classical algorithm execution. This work proposes an encoder enhancement mechanism based on an auxiliary reconstruction task, which leverages self-supervised learning to model dependencies among internal features of the input state, thereby strengthening the preservation and expressiveness of state information. Moving beyond prior methods that solely optimize the processor module, the proposed architecture achieves substantial performance gains on standard neural algorithmic reasoning benchmarks, demonstrating its ability to learn richer and more structured state representations.
This study addresses the limitation of existing test-time scaling methods, wherein models cannot autonomously determine context allocation and reuse. To overcome this, we propose Hermes, a framework that transfers context management decisions from fixed architectures to the model itself. Through a two-stage training paradigm, Hermes enables the model to acquire adaptive reasoning strategies, thereby achieving dynamically optimized allocation for multi-window computation. Experimental results demonstrate that our approach significantly enhances the performance of smaller models while exhibiting strong generalization capabilities across diverse benchmarks. Furthermore, Hermes reveals promising scaling potential that extends beyond its training compute budget, suggesting that empowering models with autonomous context management offers an effective pathway for test-time scaling.