Score
Designs and implements experimental setups that evaluate and optimize model behavior when given very small numbers of labeled examples; this includes creating and ordering few-shot prompts and demonstrations, configuring few-shot fine-tuning runs, and building heterogeneous benchmark protocols and metrics to analyze annotation-efficiency, generalization, and robustness of small-shot workflows.
This work addresses the long-standing isolation among research domains such as alignment training, model organisms, and toy models, which has hindered empirical cross-pollination and led to redundant exploration and inefficiency. For the first time, it systematically transfers supervised fine-tuning (SFT) practices across these domains by integrating cross-model output training, mixed-strategy data, and benign fine-tuning to rigorously evaluate the portability of key findings. The study demonstrates three successful transfer effects: enhanced behavioral generalization, mitigation of capability degradation, and the critical insight that preserving capabilities alone is insufficient to ensure robustness in subsequent training phases. These results underscore both the efficacy and limitations of reusing methodologies across domains, thereby fostering more synergistic development across disparate research areas.
In model-based systems engineering, low experimental data reuse efficiency and excessive redundant experiments hinder digital engineering agility. To address this, this paper proposes a case-based reasoning (CBR)-driven experimental management framework that explicitly integrates domain knowledge. The framework features structured experimental metadata modeling, digital twin–enabled scenario semantic alignment, and an interpretable similarity assessment mechanism to intelligently determine whether historical experiments can be transferred to address new verification queries. Its key innovation lies in embedding domain knowledge explicitly into both the CBR retrieval and adaptation stages, thereby enabling trustworthy cross-operating-condition and cross-configuration experimental data reuse. Evaluated on an industrial-scale vehicle energy system design case, the framework reduces redundant experiments by 37% and shortens early verification cycles by 42% on average, significantly enhancing iterative efficiency in digital engineering and advancing intelligent experimental management.
This work addresses the lack of a unified, rigorous, and realistic evaluation protocol in few-shot transfer learning, which has led to unreliable method comparisons. To this end, we introduce the FEWTRANS benchmark—comprising ten diverse datasets—and the Hyperparameter Ensemble (HPE) evaluation protocol, which effectively mitigates validation set hallucination under data scarcity. Using this framework, we systematically demonstrate for the first time that the choice of pretrained model is more critical than the complexity of the transfer algorithm. We also quantify the performance collapse of multimodal models in specialized domains due to linguistic rarity. Our analysis reveals that simple full-parameter fine-tuning consistently outperforms most sophisticated methods, owing to its ability to flexibly reshape distributed representations and high-level semantic features. The FEWTRANS benchmark is publicly released to provide the community with a reproducible evaluation standard.
To address the challenges of excessive experimental scale, high resource consumption, and the trade-off between accuracy and efficiency in system-level LLM inference performance evaluation (e.g., throughput, latency), this paper proposes FMwork—a framework for efficient and reliable benchmarking. FMwork establishes a controlled test environment, introduces meta-metrics to quantify the cost–accuracy trade-off, designs a parameter selection strategy grounded in hardware–software interaction characteristics, and formulates a joint cost–performance optimization model. It achieves 96.6% accuracy relative to full-scale testing with only minimal samples—e.g., just 128 output tokens for Llama 3.1 8B—while improving experimental efficiency by up to 24× and delivering an additional 2.7× inference acceleration. Its core contribution is the first introduction of a meta-metric-driven sparse evaluation paradigm for LLM inference benchmarking, significantly enhancing scalability and reliability in large-scale performance analysis.
This paper addresses the low reliability of stochastic optimizer performance evaluation due to run-to-run variability. We propose a statistically grounded, adaptive experimental design method. First, we theoretically derive a lower bound on the minimum number of independent runs required to guarantee prescribed accuracy for key performance metrics—such as best objective value and convergence iteration count. Building upon this, we design an adaptive sampling algorithm that dynamically determines the requisite number of repetitions, ensuring termination only when both a user-specified confidence level (e.g., 95%) and absolute error tolerance are simultaneously satisfied—thereby avoiding premature stopping or unnecessary resource expenditure. The method integrates confidence interval estimation, hypothesis testing, and sequential sample-size determination, substantially enhancing reproducibility and statistical rigor in optimizer benchmarking and hyperparameter tuning. Empirical evaluation demonstrates that the approach consistently confines estimation error within the prescribed threshold while reducing redundant runs by over 30% on average.
This study addresses the lack of effective evaluation of large language models (LLMs) in systematic experimental design, particularly along two critical dimensions: high-level planning and low-level configuration. To bridge this gap, the authors introduce SCOPE, the first comprehensive benchmark for autonomous experimental design, encompassing 300 top-tier conference papers across 19 domains, which systematically assesses LLM performance in terms of both experimental planning completeness and configuration accuracy. Furthermore, they propose OptED, an agent-based workflow that incorporates stage isolation, tool augmentation, and rule-based constraints to substantially alleviate LLMs’ performance bottlenecks in low-level configuration. Experimental results demonstrate that prevailing LLMs struggle to generate high-quality experimental designs directly, whereas OptED significantly enhances the reasonableness and accuracy of such designs.
This work addresses the common misconception that scaling laws apply only to large models, which often arises because small-scale models are evaluated with suboptimal hyperparameters. Through systematic analysis, the study demonstrates that scaling laws remain valid even for small models when evaluated along a properly tuned hyperparameter frontier, and further reveals that hyperparameter sensitivity diminishes as model scale increases. Building on these insights, the authors propose a new paradigm that combines small-scale experiments with efficient hyperparameter optimization. Using ablation studies, loss landscape analysis, and scaling modeling, they successfully reproduce findings typically observed only at large scales—such as the superiority of pre-normalization—thereby establishing that, under appropriate methodology, small-scale experiments can reliably predict large-model behavior.
This work addresses the challenge of few-shot optimization of expensive black-box functions in real-world scenarios, particularly when high-dimensional auxiliary information and data from multiple historical tasks are available. The authors propose a context-aware neural prediction architecture that jointly models high-dimensional auxiliary observations $h(x)$ and cross-task historical data to efficiently predict the performance $f(x)$ of new design candidates. By integrating few-shot learning, context-conditioned prediction, and multi-task optimization within a neural framework, the method overcomes the information utilization bottleneck of conventional Bayesian optimization. Empirical results on robotic hardware design and neural network hyperparameter tuning demonstrate significant improvements over existing approaches, achieving more accurate performance prediction and faster convergence. The study also introduces and open-sources a new benchmark for hardware design optimization.
This work addresses the challenge of systematically evaluating concept bottleneck models, whose applicability and failure mechanisms remain poorly understood due to the scarcity of real-world datasets with annotated concept labels. To bridge this gap, we introduce the first controllable synthetic benchmark that leverages parametric generation techniques to precisely modulate data modality, concept selection, annotation quality, and label completeness, thereby simulating diverse real-world relationships between concepts and predictions. This benchmark enables comprehensive evaluation of various concept bottleneck models across both decision-support and fully automated tasks, effectively identifying key performance determinants and characteristic failure modes. Our framework fills a critical void in the current evaluation landscape for concept-based interpretability methods.