Score
Designs and evaluates methods that modify a pretrained model’s predictions at inference time by selecting, constructing, or optimizing the input context (prompts, demonstrations, or example sequences) rather than updating model weights. This involves building algorithms to choose or arrange examples, craft or reweight prompts, and optimize context contents or ordering to stabilize outputs, improve robustness to perturbations, and constrain the model’s local behavior.
This study addresses the insufficient exploration of relationships and trade-offs among different paradigms for cross-task adaptation of large language models by constructing the first unified analytical framework. Methodologically, it systematically integrates parameter-efficient fine-tuning, in-context learning, and embedding injection techniques, establishing a comprehensive taxonomy along the dimensions of weights, prompts, and embeddings. Employing a systematic review methodology, the work thoroughly analyzes the strengths, limitations, and intrinsic connections of each paradigm. The findings reveal the evolutionary logic underlying these three paradigms, yielding a holistic taxonomic landscape that clarifies their core advantages and constraints while identifying key open problems. Ultimately, this research provides strategic guidance for future investigations into the efficient adaptation of large models.
This work investigates the theoretical foundations and optimization limits of prompt tuning and in-context learning. Method: Building upon a Bayesian meta-learning framework, we formally characterize pre-trained language models as meta-trained Bayesian predictors, thereby elucidating their statistical mechanisms for rapid adaptation; we further derive the first theoretical criterion for optimal prompt realizability. Contribution/Results: We establish that soft prefix tuning achieves efficient adaptation unattainable by hard prompts via implicit latent-space manipulation—bridging the gap between theoretical modeling and mechanistic interpretation. Extensive experiments on both LSTM and Transformer architectures demonstrate a strong empirical correlation between Bayesian predictive behavior and in-context learning capability. Moreover, soft prefix tuning consistently outperforms hard prompting and mainstream fine-tuning methods—both on trained and untrained models—validating its robustness and superiority across diverse settings.
This work addresses the challenge of enabling predictive systems to dynamically adapt their behavior based on contextual information for personalized inference. To this end, it proposes a unified framework that maps context into adaptation parameters for prediction and, for the first time, establishes a mathematical equivalence between explicit parameter adaptation and implicit expert routing under kernel ridge regression. The framework theoretically unifies diverse methodologies—including varying-coefficient models, local regression, prompt engineering, retrieval-augmented approaches, and mixture-of-experts—under fixed features and squared loss. Key contributions include deriving a general formulation for context-adaptive inference, proposing practical design principles and evaluation metrics such as adaptation efficiency and routing stability, and highlighting critical open problems concerning identifiability and robustness under distributional shifts.
Prior work assumes example selection dominates example ordering in in-context learning (ICL), treating the latter as negligible. This study challenges that assumption by systematically investigating how example order affects large language model (LLM) performance. Method: Through controlled experiments across classification and generation tasks, we evaluate open-source models (0.5B–27B parameters) and GPT-5, isolating the impact of permutation while holding example sets constant. Contribution/Results: We demonstrate that reordering examples induces performance fluctuations comparable in magnitude to replacing the entire example set—establishing ordering as equally critical as selection. Moreover, we provide the first empirical evidence that near-optimal permutations can be efficiently discovered using only development-set labels, achieving performance close to globally optimal (test-label-dependent) ordering. This work introduces a new ICL paradigm—jointly optimizing example selection and ordering—and proposes a lightweight, practical method for order optimization, advancing prompt engineering with theoretically grounded, empirically validated insights.
This work investigates whether parameter fine-tuning can replicate the few-shot generalization and logical reasoning advantages of prompting. To this end, we propose a meta-learning framework wherein a single gradient step dynamically approximates the model’s own prompt-based outputs—enabling self-supervised objective construction without ground-truth labels. The method integrates gradient-based meta-training, prompt-driven parameter-space optimization, and a one-step update mechanism. Its core innovation lies in using the model’s own prompt responses as the target function for gradient updates—the first such formulation—thereby revealing the inherent learnability of prompting behavior via fine-tuning. Experiments demonstrate substantial improvements on logical reasoning tasks (e.g., “reverse curse”), where our approach achieves or matches prompting performance with only one gradient update. This establishes a new paradigm for low-overhead, highly generalizable model adaptation.
This study systematically investigates the effectiveness and limitations of test-time adaptation methods that do not require updating model parameters in open-source large language models. Focusing on many-shot in-context learning (ICL), it integrates dynamic and reinforcement-based ICL prompting strategies to evaluate how the number, ordering, and selection mechanisms of examples influence performance across diverse tasks and model architectures. The findings reveal that many-shot prompting substantially improves performance on structured tasks with high information gain but is highly sensitive to example selection, whereas its benefits are limited in open-ended generation tasks. This work delineates the applicability boundaries and potential risks of prompt-based test-time adaptation, offering both theoretical grounding and practical guidance for real-world deployment.
This study investigates how the decodability of internal model representations dynamically evolves throughout pretraining and post-training, and whether erroneous decodability alone can reliably indicate discarded output information. Utilizing the Pythia model suite, the authors employ linear probing and steering intervention techniques to conduct cross-checkpoint comparative analyses of probe accuracy, steered responses, and error-correction mechanisms from early to late training stages. The work proposes an information-theoretic counterexample demonstrating that erroneous decodability is insufficient to establish the loss of output information. Furthermore, it reveals that while steering benefits improve progressively over the course of training, final-state decoders do not exhibit significant advantages. These findings offer novel perspectives for understanding the evolution of internal mechanisms within large language models.
Existing theoretical frameworks struggle to explain why larger-scale pre-trained models substantially reduce sample complexity on downstream tasks. This work proposes a novel theoretical framework—termed “caulking”—inspired by parameter-efficient fine-tuning methods such as adapters, low-rank adaptation, and partial fine-tuning. It establishes, for the first time, a provable relationship between the scale of pre-trained models and the sample complexity of downstream tasks. By rigorously linking stronger pre-training capabilities to reduced data requirements in transfer learning, this study not only addresses a critical gap in current theoretical understanding but also provides a solid foundation for empirically observed scaling laws, demonstrating that enhanced pre-training capacity can significantly decrease the number of samples needed for effective downstream adaptation.
This study investigates how large language models acquire and dynamically adjust their sensitivity to contextual features—such as length, query similarity, and fluency—during instruction tuning. By systematically comparing contextual usage behaviors across supervised fine-tuning (SFT), direct preference optimization (DPO), and reinforcement learning with verifiable rewards (RLVR) on four models and three datasets, the work reveals for the first time that models actively reshape their contextual preferences throughout fine-tuning: SFT tends to favor easily interpretable contexts, while subsequent stages may either amplify or mitigate this bias. The findings underscore the decisive role of training data composition in shaping a model’s ultimate capacity for context utilization and highlight the critical importance of balanced data design in enhancing robustness.
The role of design choices in reinforcement fine-tuning remains poorly understood, leading to inconsistent findings across studies. This work proposes a minimalistic baseline—employing a single rollout, no advantage function, and a batch size of 32—and formalizes the problem as a batched contextual bandit. Through controlled ablation experiments, we systematically evaluate the marginal contribution of each component, enabling the first decoupled analysis of key factors in reinforcement fine-tuning and clearly distinguishing their effects on learning versus generalization. Extensive ablations across three base models and two datasets identify the truly decisive design elements, thereby clarifying prevailing misconceptions in current methodologies.
Existing theories struggle to explain, under weak assumptions, how factors such as demonstration selection, chain-of-thought (CoT) reasoning, the number of demonstrations, and prompt templates influence generalization in in-context learning (ICL). This work proposes a unified theoretical framework that, under mild assumptions, links demonstration quality, the model’s intrinsic ICL capability, and distribution shift to an upper bound on test loss, while modeling CoT as a task decomposition mechanism. By integrating Lipschitz-based generalization bounds, task decomposition analysis, and distribution shift metrics, the framework quantifies—for the first time—the impact of demonstration quality on generalization, identifies conditions under which CoT improves performance, and characterizes how prompt template sensitivity varies with the number of demonstrations. Both theoretical and empirical results elucidate how pretraining, CoT, and prompting jointly enable generalization to unseen tasks.