Score
Designs, implements, and analyzes methods, prompts, and training or adaptation procedures that enable models to perform tasks from very small numbers of labeled examples or none; this includes building few-shot and zero-shot prompting workflows, exemplar selection and conditioning strategies, few-shot training/adaptation pipelines, and evaluation protocols that measure few-shot generalization and performance.
To address the deployment bottleneck of conventional machine learning models in data-scarce scenarios—such as emerging domains or urgent situations (e.g., early pandemic stages)—this paper presents a systematic survey of few-shot learning (FSL) across audio, image, text, and multimodal domains. We conduct the first cross-modal comparative analysis, identifying shared challenges (e.g., high annotation cost, modality heterogeneity) and domain-specific requirements (e.g., temporal modeling for audio, semantic alignment for text). We propose a domain-sensitive method selection framework that unifies meta-learning, metric learning, data augmentation, prompt-based fine-tuning, and multimodal alignment techniques, explicitly characterizing their applicability boundaries and failure modes. Our taxonomy encompasses 120+ studies and prescribes optimal practice pathways per domain. The survey significantly enhances the practical feasibility of FSL in low-data applications, including medical diagnosis and edge computing.
Instruction tuning often causes pretrained language models to forget foundational knowledge and over-specialize in conversational patterns, thereby degrading in-context learning (ICL) performance. This work identifies an intrinsic trade-off between instruction-following capability and ICL ability. To address it, we propose *partial adaptation*, a lightweight, parameter-efficient tuning paradigm: leveraging LoRA or Adapter modules, we freeze subsets of model parameters and progressively scale adaptation strength—without additional training or extra parameters. Evaluated across 12 canonical NLP few-shot tasks, our method improves average accuracy by 4.2%, while incurring only a marginal drop (−1.8%) in AlpacaEval instruction-following scores. Notably, this is the first systematic study to characterize performance trajectories across multiple model families and scales. Our approach offers a scalable, low-overhead pathway to balance instruction alignment with generalization—preserving ICL competence without compromising task-specific fidelity.
This study investigates the capacity of small language models to effectively use tools without relying on complex adaptation mechanisms. Focusing on Llama-3.2-3B-Instruct, the authors systematically evaluate four adaptation strategies—hypernetwork-generated LoRA weights, few-shot prompting, document-based prompting, and value-guided beam search—across four tool-use benchmarks. Experimental results demonstrate that few-shot prompting yields a 21.5% performance gain, document prompting contributes an additional 5.0%, while hypernetwork-generated LoRA weights show no significant improvement. Notably, the 3B-parameter model achieves 79.7% of GPT-5’s average performance at only one-tenth of the inference latency. These findings underscore the pivotal role of prompt engineering in enabling efficient tool use with lightweight models and offer a promising direction for resource-constrained settings.
This work systematically evaluates the generalization capability and representation robustness of small language models (SLMs) under two adaptation paradigms—few-shot prompting and supervised fine-tuning—focusing on low-resource settings, out-of-distribution (OOD) generalization, and multi-task scenarios. Methodologically, we integrate centered kernel alignment (CKA) and representational similarity analysis (RSA) for representation similarity quantification, complemented by OOD generalization benchmarks and multi-scale model comparisons. Our study is the first to characterize the knowledge internalization mechanisms of these paradigms through the lenses of representation stability and abstraction level. Results show that prompt-based learning yields more flexible representations but exhibits fragile OOD generalization; in contrast, fine-tuning achieves greater robustness yet suffers from overfitting and reduced abstraction depth. These findings provide interpretable theoretical foundations and empirically grounded guidelines for selecting adaptation strategies for SLMs in resource-constrained environments.
This study addresses the unreliability of selecting and evaluating few-shot adaptation strategies under clinical distribution shifts by proposing the Adapter and Automator architectures. Methodologically, it defines an expanded adaptation space combined with reliability rules to automatically search for optimal strategy combinations. Furthermore, evidence-based weighted fusion and reliability screening mechanisms are introduced to achieve efficient few-shot adaptation. The approach integrates techniques from large model pre-training, few-shot learning, and AutoML. Experimental results demonstrate that the proposed method attains state-of-the-art performance on critical care datasets using only minimal patient data, providing a robust and reliable solution for transfer learning in clinical scenarios.
Instruction fine-tuning (IFT) suffers from heavy reliance on large-scale annotated examples and poor few-shot cross-task generalization. To address this, we propose an instruction-driven zero-shot adapter generation framework. Our method introduces three key innovations: (1) the first end-to-end paradigm mapping natural-language instructions directly to adapter parameters; (2) a two-stage hypernetwork training scheme that decouples instruction understanding from parameter generation; and (3) the first integration of knowledge distillation into instruction learning to align instruction-level and instance-level training signals. Evaluated on Super-Natural Instructions and P3 benchmarks, our approach matches or surpasses state-of-the-art meta-trained and hypernetwork-based models in task performance, while significantly reducing inference computational overhead. This work establishes a new paradigm for efficient, low-resource generalization of large language models.
This study addresses the challenge of extracting machine learning pipeline stages, which is constrained by domain diversity and where existing methods rely on manual annotation or limited classifiers. This work systematically investigates, for the first time, the potential of small language models (SLMs) to parse ML pipeline structures leveraging their inherent code comprehension capabilities without fine-tuning, employing Cochran’s Q test, McNemar’s test, and goodness-of-fit evaluations for rigorous assessment. The findings indicate that while SLMs demonstrate robust performance, they do not surpass existing classifiers; however, the core contribution lies in revealing that different classification approaches significantly influence practical insights. Despite the limitation of high inference costs, this research establishes a novel paradigm for automated ML structure parsing.
This study addresses the difficulty language model agents face in efficiently adapting execution frameworks to diverse tasks at test time. To this end, this work proposes "framework learning," which formulates framework revision as meta-learning over executable programs. Specifically, a proposer model is trained via reinforcement learning to iteratively refine a solver's code framework using execution feedback, thereby enabling test-time adaptation without parameter updates. By integrating large language model agents with program synthesis and automated repair techniques, this approach endows agents with the capacity to continuously generalize and improve from experience. Experimental results demonstrate significant performance gains on reasoning and multi-hop question answering tasks, validating that such test-time adaptation capabilities transfer effectively to unseen tasks.
This work addresses the fragmented landscape of post-training adaptation techniques, which suffer from inconsistent terminology and a lack of unified comparative or governance frameworks. To resolve this, the paper introduces the first six-dimensional taxonomy—spanning mechanism, objective, data requirements, persistence, structural scope, and model type—that systematically integrates mainstream approaches such as fine-tuning, retrieval augmentation, prompt engineering, model editing, and machine unlearning. This framework clarifies conceptual boundaries and reveals evolutionary and compositional relationships among methods. Beyond standardizing terminology, it enables standardized technical documentation, model change tracking, and AI governance analysis. The study further identifies critical challenges, including evaluation rigor, reproducibility, continual adaptation, multimodal alignment, and governance-aware workflows.
研究探讨了客户服务LLM的多任务处理策略,通过多任务微调、顺序更新或模型合并的方法,发现多任务全微调在所有测试模型尺寸中表现最佳。
研究通过构建StrategyBench评估大型语言模型在少量示例下归纳任务规则的能力,采用明确策略归纳方法以提高适应性和鲁棒性。