HuatuoGPT-3: RL-Only Domain Adaptation from Base Models
This study addresses the challenges of cold start and teacher distribution anchoring in pure reinforcement learning (RL) domain adaptation by proposing OnePO, a single-stage policy optimization framework. The method treats the teacher model's outputs as transient guidance signals, dynamically modulating instructional intensity through adaptive objective evolution and a teacher retirement mechanism. This design effectively mitigates gradient starvation while circumventing the complexity inherent in multi-stage optimization pipelines. As the first single-stage RL-only fine-tuning paradigm, OnePO achieves a score of 70.1 on the HealthBench benchmark, surpassing frontier models such as GPT-6 Astra. Furthermore, this work releases an open-source suite of 27B-parameter medical large language models to facilitate future research.