Score
Designs and implements fine-tuning and adaptation pipelines that specialize general-purpose models to particular contexts by conditioning on contextual metadata or partitioning training data into context-specific subsets (context-aware fine-tuning / contextual adaptation). Analyzes and ranks contextual attributes (for example: speaker identity, noise characteristics, gender, language, SNR) by their impact on performance and measures the resulting changes in task-specific quality and intelligibility metrics.
Large language models (LLMs) excel at complex contextual understanding but exhibit pronounced capability asymmetry—struggling to stably generate long, equally sophisticated texts. Method: We systematically establish a unified “context engineering” framework, proposing a four-dimensional taxonomy encompassing retrieval, generation, processing, and management. Based on a systematic review and architectural analysis of 1,300+ papers, we construct the first comprehensive context engineering technology map; identify the intrinsic mechanisms underlying the understanding–generation capability mismatch; and delineate architectural integration pathways for four key application paradigms: retrieval-augmented generation, memory modeling, tool integration, and multi-agent coordination. Contribution/Results: The work delivers a standardized conceptual framework, a strategic technology roadmap, and identified critical breakthrough directions—providing both theoretical foundations and practical guidance for developing advanced context-aware AI systems.
To address the performance saturation of in-context learning (ICL) with increasing data and the poor generalization of fine-tuning under low-resource settings, this paper proposes an ICL-aware fine-tuning paradigm. It explicitly incorporates k-shot ICL structure into end-to-end fine-tuning by dynamically injecting task-relevant in-context examples into training instances. This work establishes, for the first time, a unified modeling framework bridging fine-tuning and ICL. We further design a validation-free prequential evaluation mechanism to enable efficient hyperparameter selection in low-resource scenarios. Experiments across multiple low-resource tasks demonstrate that our method consistently outperforms both standard fine-tuning and pure ICL baselines: it yields substantial gains in few-shot regimes and exhibits more stable performance improvements as training data scale increases.
To address the limited few-shot adaptation capability of large language models (LLMs) without parameter fine-tuning, this paper proposes Context Tuning—a novel prompt-based method that optimizes only task-specific prefix tokens while keeping model weights frozen. Its key innovation lies in initializing trainable prefixes with high-quality demonstration examples, thereby tightly integrating prompt learning with in-context learning (ICL) to more effectively elicit the model’s implicit knowledge. Evaluated across multiple benchmarks—including CrossFit, MMLU, and UnifiedQA—Context Tuning significantly outperforms standard ICL and hand-crafted prompts, achieving accuracy comparable to test-time fine-tuning. Crucially, it reduces training overhead by an order of magnitude, offering both computational efficiency and strong generalization across diverse tasks and domains.
This study investigates how large language models acquire and dynamically adjust their sensitivity to contextual features—such as length, query similarity, and fluency—during instruction tuning. By systematically comparing contextual usage behaviors across supervised fine-tuning (SFT), direct preference optimization (DPO), and reinforcement learning with verifiable rewards (RLVR) on four models and three datasets, the work reveals for the first time that models actively reshape their contextual preferences throughout fine-tuning: SFT tends to favor easily interpretable contexts, while subsequent stages may either amplify or mitigate this bias. The findings underscore the decisive role of training data composition in shaping a model’s ultimate capacity for context utilization and highlight the critical importance of balanced data design in enhancing robustness.
This work addresses the degradation of in-context learning (ICL) capabilities in large language models (LLMs) during full-parameter fine-tuning, which often undermines few-shot generalization. Leveraging a linear attention framework, the authors theoretically elucidate how standard fine-tuning disrupts the mechanisms underlying ICL. To mitigate this issue, they propose a constrained fine-tuning strategy that updates only the value matrices, thereby enhancing zero-shot performance on the target task while effectively preserving few-shot learning abilities. Further analysis demonstrates that incorporating an auxiliary few-shot loss can amplify this benefit under certain conditions. Both theoretical insights and empirical results substantiate that restricting the set of tunable parameters offers a principled and effective approach to jointly optimizing zero-shot and in-context learning performance.
This study investigates the effective utilization of contextual information to enhance neural speech enhancement performance under resource-constrained conditions. Through systematic evaluation of factors such as speaker identity, noise type, and language, the authors fine-tune and cross-lingually test models ranging from 10K to 5M parameters. The results demonstrate that speaker-specific adaptation yields the most substantial performance gains, significantly improving both speech intelligibility and quality. Notably, lightweight dual-specialized models—combining speaker and noise adaptation—achieve performance on par with or even surpassing that of general-purpose models ten times larger in specific scenarios. These findings underscore the efficiency and practical potential of small-scale adaptive architectures for real-world deployment.
To address severe catastrophic forgetting and low data efficiency in task adaptation of large language models (LLMs), this paper proposes Continuous Low-Rank Adaptation (CLoRA)—the first framework integrating Low-Rank Adaptation (LoRA) with continual learning. CLoRA introduces a knowledge retention module to mitigate forgetting and an adaptive parameter update strategy to enhance multi-task stability. Further augmented with knowledge distillation, it achieves both model compression and strong generalization in privacy-sensitive settings. Extensive experiments across 15 heterogeneous datasets demonstrate that CLoRA outperforms state-of-the-art baselines by +2.7%–5.3% in average task accuracy, reduces GPU memory consumption by 38%, and cuts training data requirements by 42%. The framework thus delivers superior efficiency, robustness, and practicality for continual LLM adaptation.
This work addresses the limited generalization of auditory large language models (LLMs) in low-resource or out-of-distribution settings, where conventional fine-tuning suffers from scarce annotations and distributional shifts. The authors propose SICL-AT, a post-training approach that leverages only high-resource speech data to enhance the model’s ability to adapt at inference time through in-context examples. They demonstrate, for the first time, the effectiveness of vanilla in-context learning (ICL) in multimodal auditory tasks and introduce a novel post-training strategy that requires no labeled data from the target domain. Experimental results show that SICL-AT significantly outperforms fine-tuning across diverse low-resource speech and audio understanding tasks, exhibiting superior robustness and cross-task generalization capabilities.
This work addresses the challenge of enabling predictive systems to dynamically adapt their behavior based on contextual information for personalized inference. To this end, it proposes a unified framework that maps context into adaptation parameters for prediction and, for the first time, establishes a mathematical equivalence between explicit parameter adaptation and implicit expert routing under kernel ridge regression. The framework theoretically unifies diverse methodologies—including varying-coefficient models, local regression, prompt engineering, retrieval-augmented approaches, and mixture-of-experts—under fixed features and squared loss. Key contributions include deriving a general formulation for context-adaptive inference, proposing practical design principles and evaluation metrics such as adaptation efficiency and routing stability, and highlighting critical open problems concerning identifiability and robustness under distributional shifts.
Existing PEFT methods neglect the dynamic expert routing mechanism inherent in Mixture-of-Experts (MoE) models, leading to architectural misalignment between adaptation modules and the underlying MoE structure. To address this, we propose *Routed-PEFT*, the first PEFT framework that explicitly incorporates expert routing into adapter design—dynamically assigning dedicated adapters to individual experts, thereby enabling joint optimization of expert specialization and task-specific adaptation. We systematically evaluate Routed-PEFT on OLMoE and Mixtral architectures, integrating it with LoRA, Adapter, and other PEFT variants under diverse routing strategies across commonsense and mathematical reasoning benchmarks. With only 0.1%–0.5% additional trainable parameters, Routed-PEFT achieves average accuracy gains of 2.3–5.7 percentage points over standard PEFT baselines. Moreover, our analysis uncovers task-dependent optimal routing configurations, establishing a novel paradigm for efficient fine-tuning of MoE models.
This study challenges the prevailing assumption that performance gains from supervised fine-tuning (SFT) in downstream tasks of speech foundation models stem primarily from methodological improvements. Instead, it systematically evaluates eight SFT variants across nine pretrained checkpoints of wav2vec 2.0, HuBERT, and WavLM on three SUPERB classification tasks, incorporating multiple random seeds to assess stability and transferability. The findings reveal that SFT’s apparent advantages are highly contingent on specific pretrained instances and random seeds, with optimal configurations showing little consistency or generalizability across checkpoints. These results suggest that most reported gains arise from favorable instance–seed matching rather than genuine improvements in model capacity or upper-bound performance, thereby questioning the universality of SFT enhancements.