🤖 AI Summary
This study addresses the limited instruction-following and mathematical reasoning capabilities of lightweight language models (e.g., Qwen2.5-0.5B). We systematically investigate the efficacy of reinforcement learning (RL)-based fine-tuning for alignment. To this end, we conduct the first comparative evaluation—on small-scale models—of RLOO, DPO, and supervised fine-tuning (SFT) for instruction alignment. We further propose a novel inference-time strategy: “synthetic data augmentation + external verifier-guided Best-of-N reasoning”, enabling tool-augmented, verification-aware reasoning. Experimental results show that RLOO with DeBERTa-based reward modeling achieves optimal instruction alignment, while DPO demonstrates superior robustness. Crucially, mathematical reasoning accuracy improves significantly, validating the synergistic benefit of combining RL-based fine-tuning with external verification at inference time. Our work establishes a reproducible, computationally efficient technical pathway for aligning small language models and enhancing their reliability in complex reasoning tasks.
📝 Abstract
This study investigates the effectiveness of reinforcement learning (RL) fine-tuning techniques on a compact language model (Qwen2.5-0.5B Base) for two challenging tasks: instruction following and mathematical reasoning. We compare supervised fine-tuning (SFT), Direct Preference Optimization (DPO) using preference-labeled data, and Reinforce Leave-One-Out (RLOO) with reward models. Our experiments show that RLOO with DeBERTa reward modeling achieves the best alignment, while DPO provides strong and consistent results. For math reasoing tasks, synthetic data augmentation and best-of-N sampling with an external verifier significantly improve accuracy, showing the potential of combining fine-tuning with inference-time tools. This study highlights key trade-offs and practical strategies for training lightweight, task-aligned small-scale language models.