HuatuoGPT-3: RL-Only Domain Adaptation from Base Models

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of cold start and teacher distribution anchoring in pure reinforcement learning (RL) domain adaptation by proposing OnePO, a single-stage policy optimization framework. The method treats the teacher model's outputs as transient guidance signals, dynamically modulating instructional intensity through adaptive objective evolution and a teacher retirement mechanism. This design effectively mitigates gradient starvation while circumventing the complexity inherent in multi-stage optimization pipelines. As the first single-stage RL-only fine-tuning paradigm, OnePO achieves a score of 70.1 on the HealthBench benchmark, surpassing frontier models such as GPT-6 Astra. Furthermore, this work releases an open-source suite of 27B-parameter medical large language models to facilitate future research.
📝 Abstract
Domain adaptation aims to turn a general-purpose large language model (LLM) into an expert for a target domain. While the dominant SFT+RL pipeline offers a convenient cold start, it may reduce exploration diversity and introduces additional complexity through multi-stage optimization. These limitations motivate RL-only adaptation. However, pure on-policy RL suffers from a cold-start problem, while mixed-policy RL still falls short: informative tokens in teacher outputs are learned too slowly in early training, and stale teacher outputs can hinder later improvement. We identify these two failure modes as Gradient Starvation and Teacher-Distribution Anchoring. To address them, we propose One-stage Policy Optimization (OnePO), which treats teacher outputs as transient guidance for policy improvement. OnePO combines Adaptive Objective Evolution to strengthen learning on informative low-probability teacher tokens and Teacher Retirement to discard teacher outputs once the current policy can surpass them. On medical adaptation, OnePO achieves 67.2 on HealthBench (Total) with only 20K training samples, outperforming SFT+RL and pure RL by 2.7 and 7.4 points, respectively. We further scale OnePO to produce HuatuoGPT-3, an open-source medical LLM series whose 27B variant reaches 70.1 on HealthBench (Total) and 71.4 on HealthBench Professional, surpassing frontier models such as GPT-6 Astra. Models and code are available at https://github.com/FreedomIntelligence/HuatuoGPT-3.
Problem

Research questions and friction points this paper is trying to address.

Domain Adaptation
Reinforcement Learning
Large Language Models
Gradient Starvation
Teacher-Distribution Anchoring
Innovation

Methods, ideas, or system contributions that make the work stand out.

Domain Adaptation
Reinforcement Learning
One-stage Policy Optimization
Gradient Starvation
Teacher Retirement
🔎 Similar Papers
No similar papers found.
Junying Chen
Junying Chen
The Chinese University of Hong Kong, Shenzhen
Large Language Models
X
Xinyuan Xie
The Chinese University of Hong Kong, Shenzhen; Shenzhen Research Institute of Big Data; Shenzhen Loop Area Institute; National Health Data Institute, Shenzhen
Ziniu Li
Ziniu Li
The Chinese University of Hong Kong, Shenzhen
Machine LearningReinforcement LearningLarge Language Models
W
Wenyuan Gu
The Chinese University of Hong Kong, Shenzhen; Shenzhen Research Institute of Big Data; Shenzhen Loop Area Institute; National Health Data Institute, Shenzhen
J
Jianquan Li
The Chinese University of Hong Kong, Shenzhen; Shenzhen Research Institute of Big Data; Shenzhen Loop Area Institute; National Health Data Institute, Shenzhen
Xiang Wan
Xiang Wan
Shenzhen Research Institute of Big Data
BioinformaticsData MiningBig Data Analysis
G
Guangjun Yu
The Chinese University of Hong Kong, Shenzhen; Shenzhen Research Institute of Big Data; Shenzhen Loop Area Institute; National Health Data Institute, Shenzhen
Ruoyu Sun
Ruoyu Sun
Chinese University of Hong Kong (Shenzhen), Shenzhen Institue of Big Data
Mathematical optimizationNeural networksMachine Learning
Haizhou Li
Haizhou Li
The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), China; NUS, Singapore
Automatic Speech RecognitionSpeaker RecognitionLanguage RecognitionVoice ConversionMachine Translation
Benyou Wang
Benyou Wang
Assistant Professor, The Chinese University of Hong Kong, Shenzhen
large language modelsnatural language processinginformation retrievalapplied machine learning