design few-shot experiments

Designs and implements experimental setups that evaluate and optimize model behavior when given very small numbers of labeled examples; this includes creating and ordering few-shot prompts and demonstrations, configuring few-shot fine-tuning runs, and building heterogeneous benchmark protocols and metrics to analyze annotation-efficiency, generalization, and robustness of small-shot workflows.

designfew-shotexperiments

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.19
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the long-standing isolation among research domains such as alignment training, model organisms, and toy models, which has hindered empirical cross-pollination and led to redundant exploration and inefficiency. For the first time, it systematically transfers supervised fine-tuning (SFT) practices across these domains by integrating cross-model output training, mixed-strategy data, and benign fine-tuning to rigorously evaluate the portability of key findings. The study demonstrates three successful transfer effects: enhanced behavioral generalization, mitigation of capability degradation, and the critical insight that preserving capabilities alone is insufficient to ensure robustness in subsequent training phases. These results underscore both the efficacy and limitations of reusing methodologies across domains, thereby fostering more synergistic development across disparate research areas.

alignment traininglesson transfermodel organisms

Reasonable Experiments in Model-Based Systems Engineering

Sep 12, 2025
JC
Johan Cederbladh
🏛️ Mälardalen University | Eindhoven University of Technology | Stellenbosch University | IT University of Copenhagen | University of Oslo | Universidade Federal Rural de Pernambuco | University of Antwerp

In model-based systems engineering, low experimental data reuse efficiency and excessive redundant experiments hinder digital engineering agility. To address this, this paper proposes a case-based reasoning (CBR)-driven experimental management framework that explicitly integrates domain knowledge. The framework features structured experimental metadata modeling, digital twin–enabled scenario semantic alignment, and an interpretable similarity assessment mechanism to intelligently determine whether historical experiments can be transferred to address new verification queries. Its key innovation lies in embedding domain knowledge explicitly into both the CBR retrieval and adaptation stages, thereby enabling trustworthy cross-operating-condition and cross-configuration experimental data reuse. Evaluated on an industrial-scale vehicle energy system design case, the framework reduces redundant experiments by 37% and shortens early verification cycles by 42% on average, significantly enhancing iterative efficiency in digital engineering and advancing intelligent experimental management.

Deciding if existing experiments can answer new engineering questionsIntelligently reusing experiment-related data to avoid redundant experimentsManaging experimental configuration metadata and results efficiently

This work addresses the lack of a unified, rigorous, and realistic evaluation protocol in few-shot transfer learning, which has led to unreliable method comparisons. To this end, we introduce the FEWTRANS benchmark—comprising ten diverse datasets—and the Hyperparameter Ensemble (HPE) evaluation protocol, which effectively mitigates validation set hallucination under data scarcity. Using this framework, we systematically demonstrate for the first time that the choice of pretrained model is more critical than the complexity of the transfer algorithm. We also quantify the performance collapse of multimodal models in specialized domains due to linguistic rarity. Our analysis reveals that simple full-parameter fine-tuning consistently outperforms most sophisticated methods, owing to its ability to flexibly reshape distributed representations and high-level semantic features. The FEWTRANS benchmark is publicly released to provide the community with a reproducible evaluation standard.

benchmarkevaluation protocolfew-shot transfer

To address the challenges of excessive experimental scale, high resource consumption, and the trade-off between accuracy and efficiency in system-level LLM inference performance evaluation (e.g., throughput, latency), this paper proposes FMwork—a framework for efficient and reliable benchmarking. FMwork establishes a controlled test environment, introduces meta-metrics to quantify the cost–accuracy trade-off, designs a parameter selection strategy grounded in hardware–software interaction characteristics, and formulates a joint cost–performance optimization model. It achieves 96.6% accuracy relative to full-scale testing with only minimal samples—e.g., just 128 output tokens for Llama 3.1 8B—while improving experimental efficiency by up to 24× and delivering an additional 2.7× inference acceleration. Its core contribution is the first introduction of a meta-metric-driven sparse evaluation paradigm for LLM inference benchmarking, significantly enhancing scalability and reliability in large-scale performance analysis.

Balancing cost and accuracy in performance analysisBenchmarking LLM inference performance efficientlyReducing impractical test configurations in evaluations

A Statistical Analysis for Per-Instance Evaluation of Stochastic Optimizers: How Many Repeats Are Enough?

Mar 20, 2025
MN
Moslem Noori
🏛️ 1QB Information Technologies | Hewlett Packard Labs | Hewlett Packard Enterprise

This paper addresses the low reliability of stochastic optimizer performance evaluation due to run-to-run variability. We propose a statistically grounded, adaptive experimental design method. First, we theoretically derive a lower bound on the minimum number of independent runs required to guarantee prescribed accuracy for key performance metrics—such as best objective value and convergence iteration count. Building upon this, we design an adaptive sampling algorithm that dynamically determines the requisite number of repetitions, ensuring termination only when both a user-specified confidence level (e.g., 95%) and absolute error tolerance are simultaneously satisfied—thereby avoiding premature stopping or unnecessary resource expenditure. The method integrates confidence interval estimation, hypothesis testing, and sequential sample-size determination, substantially enhancing reproducibility and statistical rigor in optimizer benchmarking and hyperparameter tuning. Empirical evaluation demonstrates that the approach consistently confines estimation error within the prescribed threshold while reducing redundant runs by over 30% on average.

Determine required repeats for accurate optimizer evaluationDevelop statistical guidelines for performance metric confidencePropose adaptive algorithm to ensure evaluation accuracy

Latest Papers

What's happening recently
View more

This study addresses the lack of effective evaluation of large language models (LLMs) in systematic experimental design, particularly along two critical dimensions: high-level planning and low-level configuration. To bridge this gap, the authors introduce SCOPE, the first comprehensive benchmark for autonomous experimental design, encompassing 300 top-tier conference papers across 19 domains, which systematically assesses LLM performance in terms of both experimental planning completeness and configuration accuracy. Furthermore, they propose OptED, an agent-based workflow that incorporates stage isolation, tool augmentation, and rule-based constraints to substantially alleviate LLMs’ performance bottlenecks in low-level configuration. Experimental results demonstrate that prevailing LLMs struggle to generate high-quality experimental designs directly, whereas OptED significantly enhances the reasonableness and accuracy of such designs.

AI for Researchautonomous experimental designexperimental planning

This work addresses the common misconception that scaling laws apply only to large models, which often arises because small-scale models are evaluated with suboptimal hyperparameters. Through systematic analysis, the study demonstrates that scaling laws remain valid even for small models when evaluated along a properly tuned hyperparameter frontier, and further reveals that hyperparameter sensitivity diminishes as model scale increases. Building on these insights, the authors propose a new paradigm that combines small-scale experiments with efficient hyperparameter optimization. Using ablation studies, loss landscape analysis, and scaling modeling, they successfully reproduce findings typically observed only at large scales—such as the superiority of pre-normalization—thereby establishing that, under appropriate methodology, small-scale experiments can reliably predict large-model behavior.

hyperparameter sensitivitymodel scalingscaling laws

This work addresses the challenge of few-shot optimization of expensive black-box functions in real-world scenarios, particularly when high-dimensional auxiliary information and data from multiple historical tasks are available. The authors propose a context-aware neural prediction architecture that jointly models high-dimensional auxiliary observations $h(x)$ and cross-task historical data to efficiently predict the performance $f(x)$ of new design candidates. By integrating few-shot learning, context-conditioned prediction, and multi-task optimization within a neural framework, the method overcomes the information utilization bottleneck of conventional Bayesian optimization. Empirical results on robotic hardware design and neural network hyperparameter tuning demonstrate significant improvements over existing approaches, achieving more accurate performance prediction and faster convergence. The study also introduces and open-sources a new benchmark for hardware design optimization.

Auxiliary InformationBayesian OptimizationBlack-Box Optimization

This work addresses the challenge of systematically evaluating concept bottleneck models, whose applicability and failure mechanisms remain poorly understood due to the scarcity of real-world datasets with annotated concept labels. To bridge this gap, we introduce the first controllable synthetic benchmark that leverages parametric generation techniques to precisely modulate data modality, concept selection, annotation quality, and label completeness, thereby simulating diverse real-world relationships between concepts and predictions. This benchmark enables comprehensive evaluation of various concept bottleneck models across both decision-support and fully automated tasks, effectively identifying key performance determinants and characteristic failure modes. Our framework fills a critical void in the current evaluation landscape for concept-based interpretability methods.

concept bottleneck modelsconcept labelsmodel interpretability

Hot Scholars

RL

Ruixuan Li

Professor of Computer Science, Huazhong University of Science and Technology
Distributed systemssecurity and privacydata management
YZ

Yixiong Zou

Huazhong University of Science and Technology
Computer visionDomain generalizationFew-shot learningVision-language model
YL

Yuhua Li

Huazhong University of Science and Technology
data mining machine learning
FH

Faegheh Hasibi

Assistant Professor, Radboud University
Information retrievalNatural language processingConversational AI
RF

Rogerio Feris

Research Manager, MIT-IBM Watson AI Lab
Computer VisionMachine LearningArtificial Intelligence