Score
Designs and implements supervised fine-tuning pipelines that use rejection sampling to generate many candidate model outputs, apply automated validators or acceptance criteria to filter for valid examples, and assemble the accepted samples into an SFT training set for model refinement. Also builds analyses and tooling to set acceptance thresholds, measure sample efficiency and selection bias, and evaluate how the rejection-based dataset affects downstream generation quality and error modes.
This study addresses the fundamental alignment gap between large language models’ (LLMs) pretraining objective—next-token prediction—and human-centric instruction-following requirements. It systematically surveys instruction tuning techniques, analyzing methodological evolution, strategies for constructing high-quality instruction-output pairs, multi-stage training paradigms, and cross-modal/domain adaptation pathways. Key determinants of generalization and controllability—such as data diversity, format consistency, and task coverage—are identified. Innovatively, the work introduces the first structured, knowledge-graph-style survey integrating theoretical foundations, practical frameworks, and critical reflection. It explicitly delineates current limitations—including instruction bias and the absence of standardized evaluation metrics—and proposes future research directions: scalable alignment, dynamic instruction synthesis, and causally grounded controllable generation. The resulting synthesis has become a benchmark reference in the LLM alignment community.
Data selection for fine-tuning under scarce target-distribution samples remains challenging. Method: This paper proposes a “validation-set-driven data selection” paradigm: it swaps the conventional roles of validation set and training pool—performing lightweight fine-tuning on the validation set and selecting the most discriminative samples from the training pool based on the magnitude of prediction shifts induced by fine-tuning. The method requires no additional annotations or gradient computations, ensuring both efficiency and theoretical interpretability. Results: Evaluated on instruction tuning and named entity recognition, it significantly reduces test log-loss on the target distribution, consistently outperforming existing SOTA methods on average while improving data utilization efficiency and fine-tuning performance. Its core innovation lies in the first use of the validation set as a proxy for fine-tuning and leveraging prediction shift as the selection criterion—enabling precise identification of high-information samples under few-shot settings.
This work addresses the inefficiency of conventional rejection fine-tuning (RFT), which discards entire unsuccessful trajectories during training of large language model agents, thereby wasting potentially useful information—especially on challenging tasks. To mitigate this, the authors propose Step-level Rejection Fine-Tuning (SRFT), which employs a critic model to perform fine-grained evaluation of each step within a trajectory. SRFT introduces a step-level loss masking mechanism that suppresses the loss from erroneous steps while preserving their contextual information, enabling the model to learn error correction and recovery from partially correct behaviors. This approach innovatively leverages valid segments within unsolved trajectories rather than discarding them entirely. Experimental results demonstrate that SRFT achieves a 3.7% absolute improvement in accuracy on SWE-bench Verified (reaching 32.2%), substantially outperforming the 2.4% gain obtained by traditional RFT.
This work explores an efficient, lightweight paradigm for addressing software engineering tasks using only supervised fine-tuning (SFT), without relying on reinforcement learning or complex alignment techniques. To this end, we construct a high-quality hybrid dataset combining real-world and synthetically generated samples, and introduce several novel components: an error-masking mechanism, a software engineering–oriented curriculum learning strategy based on task difficulty, and a test-time scaling (TTS) approach integrated with trajectory validation. Our method achieves state-of-the-art performance among open-source models on SWE-bench Verified: SWE-Lego-Qwen3-8B and SWE-Lego-Qwen3-32B attain pass rates of 42.2% and 52.6%, respectively, which further improve to 49.6% and 58.8% under TTS@16.
This study systematically evaluates the applicability of the foundation model (FM) paradigm across three scientific domains—genomics, satellite imagery, and time-series analysis—to assess whether FMs can supplant traditional supervised learning. Method: We construct a cross-modal benchmark framework employing lightweight architectures (e.g., Wide ResNet, U-Net), automated hyperparameter optimization, and standardized training protocols to rigorously compare domain-specific FMs against strong supervised baselines. Contribution/Results: Across all tasks, carefully tuned supervised models match or exceed state-of-the-art domain-specific FMs; large-scale pretraining yields no consistent empirical gains. This work provides the first multi-modal scientific validation that the FM paradigm remains immature for these domains. We open-source two automated evaluation workflows and underscore the necessity—and benchmarking value—of strong supervised baselines in scientific AI assessment.
This work addresses the limitations of conventional supervised fine-tuning (SFT) data selection, which typically relies on simple top-k ranking while overlooking the critical impact of structured data recipes—comprising operations such as filtering, mixing, and deduplication—on model performance. The authors reformulate SFT data selection as a data recipe search problem within a fixed pool and propose AutoSelection, a two-level solver that, for the first time, treats recipe structure as the primary optimization target. By decoupling subset instantiation from full-scale evaluation, AutoSelection enables efficient search through cached signals from tasks, data, and models, enhanced by strategies including warm-up probing, local editing, Gaussian process-based ranking, and stagnation-aware resampling. Evaluated on a 90K instruction pool, AutoSelection consistently achieves state-of-the-art in-distribution performance across three base models, significantly outperforming full-data training, random recipes, top-k selection, and single-operator baselines, while also demonstrating strong cross-model transferability and search stability.
Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.
This study challenges the prevailing assumption that performance gains from supervised fine-tuning (SFT) in downstream tasks of speech foundation models stem primarily from methodological improvements. Instead, it systematically evaluates eight SFT variants across nine pretrained checkpoints of wav2vec 2.0, HuBERT, and WavLM on three SUPERB classification tasks, incorporating multiple random seeds to assess stability and transferability. The findings reveal that SFT’s apparent advantages are highly contingent on specific pretrained instances and random seeds, with optimal configurations showing little consistency or generalizability across checkpoints. These results suggest that most reported gains arise from favorable instance–seed matching rather than genuine improvements in model capacity or upper-bound performance, thereby questioning the universality of SFT enhancements.
This work investigates whether model ensembles within the 1–3B parameter range can enhance code generation performance through execution feedback and pipeline architectures. We construct a generate-and-refine pipeline based on small language models, incorporate an execution feedback mechanism, and employ a NEAT-inspired evolutionary algorithm to search for optimal topologies. Our experiments reveal that execution feedback is pivotal—yielding performance gains exceeding four standard deviations on HumanEval and MBPP, primarily by correcting runtime errors—whereas increased topological complexity offers no significant benefit. The refinement component’s capability outweighs the identity of the generator, and single-run evaluations tend to overestimate evolutionary improvements; early stopping proves essential to prevent performance degradation. Moreover, specialized code models consistently outperform all combinations of general-purpose models.