Score
Designs and implements training pipelines that iteratively improve a model's instruction-following by generating, filtering, and incorporating model-produced instruction–response data through supervised fine-tuning and reinforcement learning. Builds adaptive reward and evaluation mechanisms (including hybrid rewards, automated signals, and selective human validators) to minimize manual verification, stabilize gains under distribution shift, and sustain self-driven model evolution.
This work addresses the high cost, poor scalability, and diminishing effectiveness of human-supervised approaches for improving large language models, especially as model capabilities approach human-level performance. To overcome these limitations, the paper proposes a closed-loop self-improvement framework that structures the self-enhancement process into four tightly coupled stages: data acquisition, selection, model optimization, and inference refinement. A key innovation is the introduction of an autonomous evaluation layer that coordinates and guides transitions across these stages. This framework offers the first systematic, lifecycle-oriented modeling of self-improvement, unifying critical components such as self-generated data, automated evaluation, iterative training, and inference-time optimization. By comprehensively mapping existing technical pathways and their limitations, the study lays the groundwork for realizing fully autonomous, self-evolving language models.
This study addresses the fundamental alignment gap between large language models’ (LLMs) pretraining objective—next-token prediction—and human-centric instruction-following requirements. It systematically surveys instruction tuning techniques, analyzing methodological evolution, strategies for constructing high-quality instruction-output pairs, multi-stage training paradigms, and cross-modal/domain adaptation pathways. Key determinants of generalization and controllability—such as data diversity, format consistency, and task coverage—are identified. Innovatively, the work introduces the first structured, knowledge-graph-style survey integrating theoretical foundations, practical frameworks, and critical reflection. It explicitly delineates current limitations—including instruction bias and the absence of standardized evaluation metrics—and proposes future research directions: scalable alignment, dynamic instruction synthesis, and causally grounded controllable generation. The resulting synthesis has become a benchmark reference in the LLM alignment community.
This study addresses the limited instruction-following and mathematical reasoning capabilities of lightweight language models (e.g., Qwen2.5-0.5B). We systematically investigate the efficacy of reinforcement learning (RL)-based fine-tuning for alignment. To this end, we conduct the first comparative evaluation—on small-scale models—of RLOO, DPO, and supervised fine-tuning (SFT) for instruction alignment. We further propose a novel inference-time strategy: “synthetic data augmentation + external verifier-guided Best-of-N reasoning”, enabling tool-augmented, verification-aware reasoning. Experimental results show that RLOO with DeBERTa-based reward modeling achieves optimal instruction alignment, while DPO demonstrates superior robustness. Crucially, mathematical reasoning accuracy improves significantly, validating the synergistic benefit of combining RL-based fine-tuning with external verification at inference time. Our work establishes a reproducible, computationally efficient technical pathway for aligning small language models and enhancing their reliability in complex reasoning tasks.
This work challenges the implicit assumption in instruction tuning that “larger or stronger models necessarily make better teachers,” identifying and naming this phenomenon the *Large Model Paradox*. Through systematic evaluation of 20 response generators (teachers) across 5 base models, we find no monotonic positive correlation between teacher capability and student performance. To address this, we propose **Compatibility-Aware Reward (CAR)**—the first quantitative metric explicitly modeling teacher–base model compatibility, moving beyond conventional unidirectional quality-based evaluation (e.g., relying solely on teacher output quality). Via multi-model ablation, reward modeling, and empirical analysis, we demonstrate that CAR significantly outperforms existing metrics (e.g., ROUGE, BERTScore) in predicting teacher effectiveness and improving downstream instruction-following performance.
Existing reinforcement learning (RL) approaches for code generation based on unit-test feedback rely solely on sparse terminal rewards, yielding no learning signal upon complete test failure and hindering incremental optimization for complex, long-horizon tasks. To address this, we propose a line-level Process Reward Model (PRM), the first to leverage dense, per-line correctness predictions both for reward shaping and value function initialization. PRM enables fine-grained supervision during code generation and real-time policy correction. By jointly optimizing the policy and value function under this dense reward signal, PRM achieves significant improvements over state-of-the-art methods on standard benchmarks including HumanEval and MBPP. Crucially, it maintains stable convergence even in scenarios where all unit tests fail—overcoming the fundamental limitation of terminal-only feedback. This marks a key advance in enabling RL-based code generation to handle realistic, challenging programming tasks requiring iterative refinement.
This work addresses the limited capability of large language models (LLMs) to adhere to fine-grained instruction constraints—such as formatting, length, and keyword requirements—and their poor generalization across zero-shot or cross-model settings. To this end, we propose activation steering: a lightweight, inference-time intervention that computes layer-wise neural activation differences between instruction-present and instruction-absent conditions, yielding interpretable, transferable, and composable instruction vectors. Crucially, no model fine-tuning is required. Our key contribution is the first formulation of instructions as cross-model-transferable activation-difference vectors, enabling vector composition (e.g.,叠加 multiple constraints) and foundation-model enhancement. Extensive evaluation across four mainstream LLMs demonstrates substantial improvements in instruction-following accuracy. The method supports constraint-aware generation without explicit instructions, concurrent multi-constraint control, and knowledge transfer from instruction-tuned models to base models.
Existing instruction-following meta-evaluation benchmarks suffer from insufficient data coverage and oversimplified evaluation paradigms, limiting their ability to accurately reflect the performance of discriminative models in real-world alignment scenarios. To address this, this work proposes IF-RewardBench, a comprehensive benchmark encompassing diverse instruction types and constraints, which introduces—for the first time—a listwise ranking evaluation paradigm based on multi-response preference graphs. This approach better aligns with practical alignment requirements and significantly enhances the correlation between evaluation outcomes and downstream task performance. Experimental results reveal substantial deficiencies in current discriminative models’ instruction-following capabilities, while demonstrating that IF-RewardBench achieves stronger positive correlation and greater evaluative validity compared to existing benchmarks.
This work addresses the limitations of large language models in generating high-quality BPMN process models, which are constrained by supervised fine-tuning data and the absence of well-defined multidimensional reward functions. The authors propose a reinforcement learning–based optimization approach that systematically explores a reward function encompassing 38 syntactic, pragmatic, and semantic metrics. They train Llama-3.1-8B and Qwen2.5-14B models across 48 configurations and find that uniformly weighted rewards outperform targeted weighting schemes, with significant interaction effects observed between reward composition and model architecture. Leveraging Group Relative Policy Optimization and an automated evaluation framework, the method substantially improves pragmatic and syntactic quality while preserving semantic fidelity and reducing output variability by over sixfold. All code is publicly released.
This work addresses the lack of a general, auditable dynamic control mechanism in existing training systems, which typically rely on framework-specific code. The authors propose the first cross-framework, open-source control plane that exposes training interfaces through a unified protocol, integrating declarative configuration, request validation, and secure control-point scheduling within the Aim workspace to enable metric monitoring, real-time intervention, and operational traceability. The system supports safe human and automated controller interventions during training while fully logging all operational trajectories. Experiments across five NLP and reinforcement learning tasks demonstrate its effectiveness, and the open-source implementation provides a foundation for reproducible human-in-the-loop training.
This work addresses the limitation of conventional large language model training pipelines, which are unidirectional and lack feedback from post-training to pre-training, thereby hindering continuous model evolution. The authors propose an iterative self-augmentation training framework that integrates reinforcement learning during the pre-training annealing phase to dynamically reweight tokens relevant to reasoning. This approach establishes a teacher- and reference-model-free bidirectional training loop, enabling post-training signals to inform and refine pre-training. Evaluated across ten benchmarks spanning mathematical reasoning, code generation, and general-purpose reasoning, the method achieves an average performance gain of 3% and sustains over 2% improvement in subsequent post-training stages, marking the first demonstration of reasoning-driven co-optimization between pre-training and post-training.
Existing approaches to improving instruction-following capabilities in large language models often rely on costly human supervision or static instructions, limiting their ability to continuously adapt and improve. This work proposes the first self-evolutionary reinforcement learning framework that operates without external supervision. The framework establishes a closed-loop collaboration among four roles—Instructor, Filter, Follower, and Judger—to dynamically co-evolve instruction difficulty and model capability. By integrating adversarial instruction generation, a data filtering mechanism, and reward-based reinforcement learning, the approach consistently enhances instruction-following performance across diverse model scales and architectures, demonstrating strong generality and effectiveness.