Score
Designs, builds, and analyzes end-to-end post‑pretraining processes (alignment pipelines) that transform pretrained models into instruction-following or preference-aligned systems by applying supervised fine-tuning, reward modeling, reinforcement learning from human feedback, or other post-training interventions. Work includes dataset curation, training and fine-tuning schedules, evaluation and ablation studies, safety and specification updates, and producing reproducible recipes and documentation of the pipeline and its trade-offs.
This study addresses two core dimensions of large language model (LLM) alignment: value specification and data construction. Through a systematic audit of publicly available technical documentation from six mainstream LLM projects—including both proprietary and open-weight models—we employ literature review, qualitative content analysis, and cross-case comparison to extract and code alignment goal definitions, data provenance, annotation protocols, and value trade-off strategies. We introduce the novel “Values–Data Centered” dual-lens analytical framework, revealing significant disparities between proprietary and open-weight models in value objective selection and training data governance. Building on this, we develop a socio-technical critical alignment framework that identifies prevalent patterns—including ambiguous value articulation, insufficient data traceability, and opaque trade-offs—as well as structural risks. Our findings provide theoretical grounding and actionable guidance for normative, auditable, and responsible LLM alignment practices.
This work addresses the long-standing isolation among research domains such as alignment training, model organisms, and toy models, which has hindered empirical cross-pollination and led to redundant exploration and inefficiency. For the first time, it systematically transfers supervised fine-tuning (SFT) practices across these domains by integrating cross-model output training, mixed-strategy data, and benign fine-tuning to rigorously evaluate the portability of key findings. The study demonstrates three successful transfer effects: enhanced behavioral generalization, mitigation of capability degradation, and the critical insight that preserving capabilities alone is insufficient to ensure robustness in subsequent training phases. These results underscore both the efficacy and limitations of reusing methodologies across domains, thereby fostering more synergistic development across disparate research areas.
This study challenges the common assumption that models exhibiting similar performance after supervised fine-tuning (SFT) are functionally equivalent, by demonstrating that the data used in the final stage of pretraining critically influences subsequent alignment behavior. Through controlled experiments—where only the last 500 million tokens of pretraining data are varied while keeping SFT and post-training procedures identical—the authors show that ending pretraining with safety-oriented text significantly preserves a model’s ability to refuse harmful requests, an effect absent with other data types. This finding is replicated across another model family, revealing for the first time that late-stage pretraining data selectively shapes how models evolve during preference optimization and reinforcement learning. The results question evaluation practices that rely solely on post-SFT performance as a proxy for alignment capability.
To address the weak generalization of large language models (LLMs) caused by narrow instruction-tuning data distributions and misalignment with pretraining knowledge, this paper proposes a coverage-aligned instruction data adaptive synthesis framework. Methodologically, it systematically aligns instruction-tuning distributions with pretraining distributions for the first time; introduces a coverage bias detection mechanism to identify knowledge gaps; employs controllable text rewriting to transform underrepresented pretraining texts into high-quality instruction-response pairs; and designs a balanced fusion strategy for multi-stage data integration. The framework achieves significant performance gains across three fully open-source LLMs and eight benchmark datasets. Ablation studies confirm the synergistic effectiveness of all components. This work establishes a novel paradigm for preserving pretraining knowledge while enhancing task-specific adaptation in LLMs.
Current LLM alignment heavily relies on large-scale human-annotated datasets, entailing high costs, poor reproducibility, and unclear scaling relationships between data volume and performance. To address this, we propose PiKa—a highly efficient synthetic data paradigm—that constructs the high-quality alignment dataset PiKa-SFT using only 30K samples, eliminating dependence on proprietary or human-labeled data. Methodologically, PiKa integrates AI-generated data synthesis, reinforcement learning from AI feedback (RLAIF), and supervised fine-tuning (SFT) within an iterative optimization framework. We perform zero-shot post-training alignment on base models from the Llama-3 and Qwen2.5 families. Experiments show that Llama-3-8B fine-tuned on PiKa-SFT surpasses the official Llama-3-8B-Instruct on AlpacaEval 2.0 and Arena-Hard; all Qwen2.5 variants exhibit consistent improvements. These results validate the efficacy and generalizability of small-scale, high-quality synthetic data, offering a scalable, low-cost alignment pathway for resource-constrained settings.
This study addresses the limited instruction-following and mathematical reasoning capabilities of lightweight language models (e.g., Qwen2.5-0.5B). We systematically investigate the efficacy of reinforcement learning (RL)-based fine-tuning for alignment. To this end, we conduct the first comparative evaluation—on small-scale models—of RLOO, DPO, and supervised fine-tuning (SFT) for instruction alignment. We further propose a novel inference-time strategy: “synthetic data augmentation + external verifier-guided Best-of-N reasoning”, enabling tool-augmented, verification-aware reasoning. Experimental results show that RLOO with DeBERTa-based reward modeling achieves optimal instruction alignment, while DPO demonstrates superior robustness. Crucially, mathematical reasoning accuracy improves significantly, validating the synergistic benefit of combining RL-based fine-tuning with external verification at inference time. Our work establishes a reproducible, computationally efficient technical pathway for aligning small language models and enhancing their reliability in complex reasoning tasks.
This work addresses the need for post-training to enhance the accuracy and reasoning reliability of large language models on specific tasks, while the boundaries and synergies between supervised fine-tuning (SFT) and reinforcement learning (RL) remain unclear. The study proposes a unified analytical framework to systematically compare SFT and RL in terms of objective formulation, algorithmic structure, and data requirements, revealing their intrinsic connections. Building on this analysis, the authors design an integrated strategy to establish an efficient hybrid post-training paradigm. Through empirical evaluation across representative applications from 2023 to 2025, the research identifies a clear trend toward hybrid post-training approaches and distills key practical guidelines, offering both theoretical grounding and methodological guidance for scalable, effective, and generalizable post-training of large language models.
研究通过中期训练方法解决AI模型在所有可能环境中的行为泛化问题,但发现该方法在某些情况下效果有限。
Aligning a text-to-image generation flow model with a reward makes it follow objectives that the training data alone does not provide. Alignment fine-tuning delivers this by reinforcement learning (RL) or preference optimization, but it must be repeated for every checkpoint and returns a model fixed at the reward and strength it was trained with. Test-time alignment instead steers a frozen model during sampling, allowing task-specific and sample-specific guidance. Existing methods obtain this only by drawing the per-step signal from the reward function itself, through its gradient, or through a separately trained value function. We propose changing the supervision source: let a pair of weak models, not a reward function, supply the supervision. A source aligned model, kept together with its base as a source alignment pair, stores its training reward as an implicit, step-wise, KL-anchored signal expressed in the sampler's own coordinates. We explore whether this model-form supervision can cross scale, and show that it does: our method, AlignGraft, aligns a larger, frozen, never-tuned model by adding the pair's velocity difference during sampling. The transport is exact under a shared noising kernel and needs neither the reward nor its gradient at test time. The method has no schedules, only a single scalar that controls the alignment strength and can extrapolate it beyond that of the source alignment pair. Across image and video flow models (Stable Diffusion 3.5, FLUX, and Wan), the transfer lifts the frozen large model on preference, compositional, and text-rendering rewards, can exceed the source aligned model itself, and preserves the large model's fidelity at a small constant sampling overhead. Extensive experiments show that one alignment run on a weak model produces supervision that the whole model family can reuse at test time.
本文提出TailSFT方法,通过过滤已拟合序列来改进监督微调,从而提高模型在强化学习中的表现。
This work addresses the limitations of existing human feedback–based alignment methods, such as reinforcement learning from human feedback (RLHF), which rely on large-scale preference data, incur high costs, suffer from training instability, and often degrade model generalization. To overcome these challenges, the authors propose DEFT, an efficient alignment framework that introduces a novel differential distributional reward mechanism. This mechanism quantifies the divergence between the language model’s output distribution and the distribution implied by preference data, enabling the selection of a high-quality, small-scale subset for training. DEFT then integrates supervised fine-tuning with contrastive learning to guide distributional alignment. Experimental results demonstrate that DEFT significantly reduces both data requirements and training time while simultaneously improving alignment performance and model generalization, outperforming current state-of-the-art approaches across the board.