Rufus-Air: An Open LLM Post-Training Recipe

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the opacity, poor reproducibility, and absence of standardized recipes in post-training pipelines for large language models by proposing an eight-stage sequential post-training framework based on GLM-4.5-Air. Methodologically, the framework comprehensively spans the entire pipeline from supervised fine-tuning to reinforcement learning from human feedback, integrating multi-domain reinforcement learning with agent training. It innovatively introduces a difficulty filtering mechanism and a reward reliability-guided stage ordering strategy, while incorporating engineering infrastructure as a core component of the recipe. Experimental results demonstrate that this approach surpasses official baselines, achieves state-of-the-art performance among open-source models of comparable scale, and ensures full reproducibility across the entire post-training pipeline.
📝 Abstract
Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF. We document the data, reward design, infrastructure, stage order, and stagewise results needed to reproduce the recipe. Stages progress from basic to advanced capabilities and from hard, verifiable rewards to softer judge-based signals. Training builds on open-source components and public data, much of it used as released, without new human annotation or an in-house distillation teacher. Our main findings are that (i) diverse, high-quality SFT establishes a strong capability floor; (ii) difficulty filtering keeps RL prompts within a productive learning range; (iii) reward reliability provides a practical principle for ordering stages; and (iv) infrastructure and engineering choices are part of the recipe, not just an implementation detail. Rufus-Air improves over the official GLM-4.5-Air post-trained release and is competitive with similarly sized open models.
Problem

Research questions and friction points this paper is trying to address.

LLM post-training
reproducibility
open-source recipe
reinforcement learning
multi-stage pipeline
Innovation

Methods, ideas, or system contributions that make the work stand out.

Post-Training Recipe
Reinforcement Learning
Reproducibility
Reward Design
Multi-stage Pipeline
🔎 Similar Papers
No similar papers found.