Reinforcement Learning Fine-Tuning of Language Model for Instruction Following and Math Reasoning

📅 2025-06-11
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited instruction-following and mathematical reasoning capabilities of lightweight language models (e.g., Qwen2.5-0.5B). We systematically investigate the efficacy of reinforcement learning (RL)-based fine-tuning for alignment. To this end, we conduct the first comparative evaluation—on small-scale models—of RLOO, DPO, and supervised fine-tuning (SFT) for instruction alignment. We further propose a novel inference-time strategy: “synthetic data augmentation + external verifier-guided Best-of-N reasoning”, enabling tool-augmented, verification-aware reasoning. Experimental results show that RLOO with DeBERTa-based reward modeling achieves optimal instruction alignment, while DPO demonstrates superior robustness. Crucially, mathematical reasoning accuracy improves significantly, validating the synergistic benefit of combining RL-based fine-tuning with external verification at inference time. Our work establishes a reproducible, computationally efficient technical pathway for aligning small language models and enhancing their reliability in complex reasoning tasks.

Technology Category

Natural Language Processing: Safety and RobustnessMachine Learning: Large Multimodal Models (LMMs)Knowledge Representation and Reasoning: Argumentation

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: LLM based quality controls for crowd work
📝 Abstract
This study investigates the effectiveness of reinforcement learning (RL) fine-tuning techniques on a compact language model (Qwen2.5-0.5B Base) for two challenging tasks: instruction following and mathematical reasoning. We compare supervised fine-tuning (SFT), Direct Preference Optimization (DPO) using preference-labeled data, and Reinforce Leave-One-Out (RLOO) with reward models. Our experiments show that RLOO with DeBERTa reward modeling achieves the best alignment, while DPO provides strong and consistent results. For math reasoing tasks, synthetic data augmentation and best-of-N sampling with an external verifier significantly improve accuracy, showing the potential of combining fine-tuning with inference-time tools. This study highlights key trade-offs and practical strategies for training lightweight, task-aligned small-scale language models.
Problem

Research questions and friction points this paper is trying to address.

Effectiveness of RL fine-tuning for instruction following and math reasoning
Comparison of SFT, DPO, and RLOO techniques for model alignment
Improving math accuracy via synthetic data and inference-time tools
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement learning fine-tuning for instruction following
DeBERTa reward modeling with RLOO for best alignment
Synthetic data augmentation for math reasoning tasks
🔎 Similar Papers
No similar papers found.
Y
Yifu Han
Department of Energy Science & Engineering, Stanford University
G
Geo Zhang
Department of Energy Science & Engineering, Stanford University