in-context tuning

Designs and evaluates methods that modify a pretrained model’s predictions at inference time by selecting, constructing, or optimizing the input context (prompts, demonstrations, or example sequences) rather than updating model weights. This involves building algorithms to choose or arrange examples, craft or reweight prompts, and optimize context contents or ordering to stabilize outputs, improve robustness to perturbations, and constrain the model’s local behavior.

in-contexttuning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.19
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Understanding Prompt Tuning and In-Context Learning via Meta-Learning

May 22, 2025
TG
Tim Genewein
🏛️ Google DeepMind

This work investigates the theoretical foundations and optimization limits of prompt tuning and in-context learning. Method: Building upon a Bayesian meta-learning framework, we formally characterize pre-trained language models as meta-trained Bayesian predictors, thereby elucidating their statistical mechanisms for rapid adaptation; we further derive the first theoretical criterion for optimal prompt realizability. Contribution/Results: We establish that soft prefix tuning achieves efficient adaptation unattainable by hard prompts via implicit latent-space manipulation—bridging the gap between theoretical modeling and mechanistic interpretation. Extensive experiments on both LSTM and Transformer architectures demonstrate a strong empirical correlation between Bayesian predictive behavior and in-context learning capability. Moreover, soft prefix tuning consistently outperforms hard prompting and mainstream fine-tuning methods—both on trained and untrained models—validating its robustness and superiority across diverse settings.

Comparing soft vs hard prompt performance in neural networksExploring limitations and effectiveness of optimal prompting strategiesUnderstanding prompt tuning via Bayesian meta-learning analysis

This work addresses the challenge of enabling predictive systems to dynamically adapt their behavior based on contextual information for personalized inference. To this end, it proposes a unified framework that maps context into adaptation parameters for prediction and, for the first time, establishes a mathematical equivalence between explicit parameter adaptation and implicit expert routing under kernel ridge regression. The framework theoretically unifies diverse methodologies—including varying-coefficient models, local regression, prompt engineering, retrieval-augmented approaches, and mixture-of-experts—under fixed features and squared loss. Key contributions include deriving a general formulation for context-adaptive inference, proposing practical design principles and evaluation metrics such as adaptation efficiency and routing stability, and highlighting critical open problems concerning identifiability and robustness under distributional shifts.

context-adaptive inferencedistribution shiftfoundation models

Order Matters: Rethinking Prompt Construction in In-Context Learning

Nov 12, 2025
WL
Warren Li
🏛️ UC San Diego | Cushing Academy

Prior work assumes example selection dominates example ordering in in-context learning (ICL), treating the latter as negligible. This study challenges that assumption by systematically investigating how example order affects large language model (LLM) performance. Method: Through controlled experiments across classification and generation tasks, we evaluate open-source models (0.5B–27B parameters) and GPT-5, isolating the impact of permutation while holding example sets constant. Contribution/Results: We demonstrate that reordering examples induces performance fluctuations comparable in magnitude to replacing the entire example set—establishing ordering as equally critical as selection. Moreover, we provide the first empirical evidence that near-optimal permutations can be efficiently discovered using only development-set labels, achieving performance close to globally optimal (test-label-dependent) ordering. This work introduces a new ICL paradigm—jointly optimizing example selection and ordering—and proposes a lightweight, practical method for order optimization, advancing prompt engineering with theoretically grounded, empirically validated insights.

Compares variance from example selection versus example ordering strategiesDemonstrates ordering effects comparable to using different example setsInvestigates how example ordering impacts in-context learning performance

Can Gradient Descent Simulate Prompting?

Jun 26, 2025
EZ
Eric Zhang
🏛️ MIT CSAIL

This work investigates whether parameter fine-tuning can replicate the few-shot generalization and logical reasoning advantages of prompting. To this end, we propose a meta-learning framework wherein a single gradient step dynamically approximates the model’s own prompt-based outputs—enabling self-supervised objective construction without ground-truth labels. The method integrates gradient-based meta-training, prompt-driven parameter-space optimization, and a one-step update mechanism. Its core innovation lies in using the model’s own prompt responses as the target function for gradient updates—the first such formulation—thereby revealing the inherent learnability of prompting behavior via fine-tuning. Experiments demonstrate substantial improvements on logical reasoning tasks (e.g., “reverse curse”), where our approach achieves or matches prompting performance with only one gradient update. This establishes a new paradigm for low-overhead, highly generalizable model adaptation.

Can gradient descent simulate prompting in language models?Improving fine-tuning to match prompting effectivenessMeta-training LMs for gradient updates mimicking prompting

This study systematically investigates the effectiveness and limitations of test-time adaptation methods that do not require updating model parameters in open-source large language models. Focusing on many-shot in-context learning (ICL), it integrates dynamic and reinforcement-based ICL prompting strategies to evaluate how the number, ordering, and selection mechanisms of examples influence performance across diverse tasks and model architectures. The findings reveal that many-shot prompting substantially improves performance on structured tasks with high information gain but is highly sensitive to example selection, whereas its benefits are limited in open-ended generation tasks. This work delineates the applicability boundaries and potential risks of prompt-based test-time adaptation, offering both theoretical grounding and practical guidance for real-world deployment.

in-context learninglarge language modelsmany-shot prompting

Latest Papers

What's happening recently
View more

This study investigates how the decodability of internal model representations dynamically evolves throughout pretraining and post-training, and whether erroneous decodability alone can reliably indicate discarded output information. Utilizing the Pythia model suite, the authors employ linear probing and steering intervention techniques to conduct cross-checkpoint comparative analyses of probe accuracy, steered responses, and error-correction mechanisms from early to late training stages. The work proposes an information-theoretic counterexample demonstrating that erroneous decodability is insufficient to establish the loss of output information. Furthermore, it reveals that while steering benefits improve progressively over the course of training, final-state decoders do not exhibit significant advantages. These findings offer novel perspectives for understanding the evolution of internal mechanisms within large language models.

in-context decodinginformation-theoretic decodabilitymodel errors

Existing theoretical frameworks struggle to explain why larger-scale pre-trained models substantially reduce sample complexity on downstream tasks. This work proposes a novel theoretical framework—termed “caulking”—inspired by parameter-efficient fine-tuning methods such as adapters, low-rank adaptation, and partial fine-tuning. It establishes, for the first time, a provable relationship between the scale of pre-trained models and the sample complexity of downstream tasks. By rigorously linking stronger pre-training capabilities to reduced data requirements in transfer learning, this study not only addresses a critical gap in current theoretical understanding but also provides a solid foundation for empirically observed scaling laws, demonstrating that enhanced pre-training capacity can significantly decrease the number of samples needed for effective downstream adaptation.

downstream taskspre-trained modelssample complexity

This study investigates how large language models acquire and dynamically adjust their sensitivity to contextual features—such as length, query similarity, and fluency—during instruction tuning. By systematically comparing contextual usage behaviors across supervised fine-tuning (SFT), direct preference optimization (DPO), and reinforcement learning with verifiable rewards (RLVR) on four models and three datasets, the work reveals for the first time that models actively reshape their contextual preferences throughout fine-tuning: SFT tends to favor easily interpretable contexts, while subsequent stages may either amplify or mitigate this bias. The findings underscore the decisive role of training data composition in shaping a model’s ultimate capacity for context utilization and highlight the critical importance of balanced data design in enhancing robustness.

context characteristicscontext usageinstruction fine-tuning

The role of design choices in reinforcement fine-tuning remains poorly understood, leading to inconsistent findings across studies. This work proposes a minimalistic baseline—employing a single rollout, no advantage function, and a batch size of 32—and formalizes the problem as a batched contextual bandit. Through controlled ablation experiments, we systematically evaluate the marginal contribution of each component, enabling the first decoupled analysis of key factors in reinforcement fine-tuning and clearly distinguishing their effects on learning versus generalization. Extensive ablations across three base models and two datasets identify the truly decisive design elements, thereby clarifying prevailing misconceptions in current methodologies.

contextual banditdesign choicesgeneralization

Existing theories struggle to explain, under weak assumptions, how factors such as demonstration selection, chain-of-thought (CoT) reasoning, the number of demonstrations, and prompt templates influence generalization in in-context learning (ICL). This work proposes a unified theoretical framework that, under mild assumptions, links demonstration quality, the model’s intrinsic ICL capability, and distribution shift to an upper bound on test loss, while modeling CoT as a task decomposition mechanism. By integrating Lipschitz-based generalization bounds, task decomposition analysis, and distribution shift metrics, the framework quantifies—for the first time—the impact of demonstration quality on generalization, identifies conditions under which CoT improves performance, and characterizes how prompt template sensitivity varies with the number of demonstrations. Both theoretical and empirical results elucidate how pretraining, CoT, and prompting jointly enable generalization to unseen tasks.

Chain-of-Thought promptingdemonstration selectiongeneralization

Hot Scholars

YG

Yi Gui

Huazhong University of Science and Technology
SX

Stella X. Yu

Professor of EECS, University of Michigan
Computer VisionRoboticsEmbodied AIDevelopmental AI
CS

Chuan Shi

Beijing University of Posts and Telecommunications
data miningmachine learningsocial network analysis
ZW

Zijun Wu

University of Alberta
Natural Language Processing (NLP)