post-training alignment

Designs, builds, and analyzes end-to-end post‑pretraining processes (alignment pipelines) that transform pretrained models into instruction-following or preference-aligned systems by applying supervised fine-tuning, reward modeling, reinforcement learning from human feedback, or other post-training interventions. Work includes dataset curation, training and fine-tuning schedules, evaluation and ablation studies, safety and specification updates, and producing reproducible recipes and documentation of the pipeline and its trade-offs.

post-trainingalignment

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.36
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$201K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the long-standing isolation among research domains such as alignment training, model organisms, and toy models, which has hindered empirical cross-pollination and led to redundant exploration and inefficiency. For the first time, it systematically transfers supervised fine-tuning (SFT) practices across these domains by integrating cross-model output training, mixed-strategy data, and benign fine-tuning to rigorously evaluate the portability of key findings. The study demonstrates three successful transfer effects: enhanced behavioral generalization, mitigation of capability degradation, and the critical insight that preserving capabilities alone is insufficient to ensure robustness in subsequent training phases. These results underscore both the efficacy and limitations of reusing methodologies across domains, thereby fostering more synergistic development across disparate research areas.

alignment traininglesson transfermodel organisms

This study challenges the common assumption that models exhibiting similar performance after supervised fine-tuning (SFT) are functionally equivalent, by demonstrating that the data used in the final stage of pretraining critically influences subsequent alignment behavior. Through controlled experiments—where only the last 500 million tokens of pretraining data are varied while keeping SFT and post-training procedures identical—the authors show that ending pretraining with safety-oriented text significantly preserves a model’s ability to refuse harmful requests, an effect absent with other data types. This finding is replicated across another model family, revealing for the first time that late-stage pretraining data selectively shapes how models evolve during preference optimization and reinforcement learning. The results question evaluation practices that rely solely on post-SFT performance as a proxy for alignment capability.

alignmentmodel checkpointpost-training

Aligning Instruction Tuning with Pre-training

Jan 16, 2025
YL
Yiming Liang
🏛️ University of Chinese Academy of Sciences | Chinese Academy of Sciences | Peking University | BAAI

To address the weak generalization of large language models (LLMs) caused by narrow instruction-tuning data distributions and misalignment with pretraining knowledge, this paper proposes a coverage-aligned instruction data adaptive synthesis framework. Methodologically, it systematically aligns instruction-tuning distributions with pretraining distributions for the first time; introduces a coverage bias detection mechanism to identify knowledge gaps; employs controllable text rewriting to transform underrepresented pretraining texts into high-quality instruction-response pairs; and designs a balanced fusion strategy for multi-stage data integration. The framework achieves significant performance gains across three fully open-source LLMs and eight benchmark datasets. Ablation studies confirm the synergistic effectiveness of all components. This work establishes a novel paradigm for preserving pretraining knowledge while enhancing task-specific adaptation in LLMs.

Large Language ModelsPre-learning Knowledge UtilizationTask-specific Adaptation

PIKA: Expert-Level Synthetic Datasets for Post-Training Alignment from Scratch

Oct 08, 2025
SY
Shangjian Yin
🏛️ Microsoft | University of California, Riverside

Current LLM alignment heavily relies on large-scale human-annotated datasets, entailing high costs, poor reproducibility, and unclear scaling relationships between data volume and performance. To address this, we propose PiKa—a highly efficient synthetic data paradigm—that constructs the high-quality alignment dataset PiKa-SFT using only 30K samples, eliminating dependence on proprietary or human-labeled data. Methodologically, PiKa integrates AI-generated data synthesis, reinforcement learning from AI feedback (RLAIF), and supervised fine-tuning (SFT) within an iterative optimization framework. We perform zero-shot post-training alignment on base models from the Llama-3 and Qwen2.5 families. Experiments show that Llama-3-8B fine-tuned on PiKa-SFT surpasses the official Llama-3-8B-Instruct on AlpacaEval 2.0 and Arena-Hard; all Qwen2.5 variants exhibit consistent improvements. These results validate the efficacy and generalizability of small-scale, high-quality synthetic data, offering a scalable, low-cost alignment pathway for resource-constrained settings.

Achieving superior performance using significantly fewer training examplesAddressing data scarcity in LLM alignment with expert synthetic datasetsReducing reliance on costly human annotation for model training

This study addresses the limited instruction-following and mathematical reasoning capabilities of lightweight language models (e.g., Qwen2.5-0.5B). We systematically investigate the efficacy of reinforcement learning (RL)-based fine-tuning for alignment. To this end, we conduct the first comparative evaluation—on small-scale models—of RLOO, DPO, and supervised fine-tuning (SFT) for instruction alignment. We further propose a novel inference-time strategy: “synthetic data augmentation + external verifier-guided Best-of-N reasoning”, enabling tool-augmented, verification-aware reasoning. Experimental results show that RLOO with DeBERTa-based reward modeling achieves optimal instruction alignment, while DPO demonstrates superior robustness. Crucially, mathematical reasoning accuracy improves significantly, validating the synergistic benefit of combining RL-based fine-tuning with external verification at inference time. Our work establishes a reproducible, computationally efficient technical pathway for aligning small language models and enhancing their reliability in complex reasoning tasks.

Comparison of SFT, DPO, and RLOO techniques for model alignmentEffectiveness of RL fine-tuning for instruction following and math reasoningImproving math accuracy via synthetic data and inference-time tools

Latest Papers

What's happening recently
View more

This work addresses the need for post-training to enhance the accuracy and reasoning reliability of large language models on specific tasks, while the boundaries and synergies between supervised fine-tuning (SFT) and reinforcement learning (RL) remain unclear. The study proposes a unified analytical framework to systematically compare SFT and RL in terms of objective formulation, algorithmic structure, and data requirements, revealing their intrinsic connections. Building on this analysis, the authors design an integrated strategy to establish an efficient hybrid post-training paradigm. Through empirical evaluation across representative applications from 2023 to 2025, the research identifies a clear trend toward hybrid post-training approaches and distills key practical guidelines, offering both theoretical grounding and methodological guidance for scalable, effective, and generalizable post-training of large language models.

Large Language ModelsPost-TrainingReinforcement Learning

研究通过中期训练方法解决AI模型在所有可能环境中的行为泛化问题,但发现该方法在某些情况下效果有限。

alignment midtrainingdeployment environmentsgeneralisation

Aligning a text-to-image generation flow model with a reward makes it follow objectives that the training data alone does not provide. Alignment fine-tuning delivers this by reinforcement learning (RL) or preference optimization, but it must be repeated for every checkpoint and returns a model fixed at the reward and strength it was trained with. Test-time alignment instead steers a frozen model during sampling, allowing task-specific and sample-specific guidance. Existing methods obtain this only by drawing the per-step signal from the reward function itself, through its gradient, or through a separately trained value function. We propose changing the supervision source: let a pair of weak models, not a reward function, supply the supervision. A source aligned model, kept together with its base as a source alignment pair, stores its training reward as an implicit, step-wise, KL-anchored signal expressed in the sampler's own coordinates. We explore whether this model-form supervision can cross scale, and show that it does: our method, AlignGraft, aligns a larger, frozen, never-tuned model by adding the pair's velocity difference during sampling. The transport is exact under a shared noising kernel and needs neither the reward nor its gradient at test time. The method has no schedules, only a single scalar that controls the alignment strength and can extrapolate it beyond that of the source alignment pair. Across image and video flow models (Stable Diffusion 3.5, FLUX, and Wan), the transfer lifts the frozen large model on preference, compositional, and text-rendering rewards, can exceed the source aligned model itself, and preserves the large model's fidelity at a small constant sampling overhead. Extensive experiments show that one alignment run on a weak model produces supervision that the whole model family can reuse at test time.

alignment strengthflow modelsimplicit rewards

This work addresses the limitations of existing human feedback–based alignment methods, such as reinforcement learning from human feedback (RLHF), which rely on large-scale preference data, incur high costs, suffer from training instability, and often degrade model generalization. To overcome these challenges, the authors propose DEFT, an efficient alignment framework that introduces a novel differential distributional reward mechanism. This mechanism quantifies the divergence between the language model’s output distribution and the distribution implied by preference data, enabling the selection of a high-quality, small-scale subset for training. DEFT then integrates supervised fine-tuning with contrastive learning to guide distributional alignment. Experimental results demonstrate that DEFT significantly reduces both data requirements and training time while simultaneously improving alignment performance and model generalization, outperforming current state-of-the-art approaches across the board.

generalization abilityhuman alignmentLarge Language Models

Hot Scholars

HF

Haipeng Fang

Institute of Computer Technology, Chinese Academy of Sciences
Model AccelerationImage and Video GenerationAIGC
CY

Chengyang Ying

Tsinghua university
Machine LearningReinforcement LearningEmbodied AI
LW

Lingxuan Wu

Tsinghua University
Embodied IntelligenceAI Safety
SZ

Shaorong Zhang

Unknown affiliation
Generative ModelMachine Learning