Institution profile

National Health Data Institute

Academic institution
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

HuatuoGPT-3: RL-Only Domain Adaptation from Base Models

Oct 05, 2026

This study addresses the challenges of cold start and teacher distribution anchoring in pure reinforcement learning (RL) domain adaptation by proposing OnePO, a single-stage policy optimization framework. The method treats the teacher model's outputs as transient guidance signals, dynamically modulating instructional intensity through adaptive objective evolution and a teacher retirement mechanism. This design effectively mitigates gradient starvation while circumventing the complexity inherent in multi-stage optimization pipelines. As the first single-stage RL-only fine-tuning paradigm, OnePO achieves a score of 70.1 on the HealthBench benchmark, surpassing frontier models such as GPT-6 Astra. Furthermore, this work releases an open-source suite of 27B-parameter medical large language models to facilitate future research.

0 citationsRead paper

Why Vision Fails as a Universal Bridge: Rectifying Modality Asynchrony in Multilingual MLLMs

Aug 15, 2026

This study addresses the performance degradation of multilingual multimodal large language models in non-English visual reasoning by identifying the "Ghost Anchor" phenomenon, which induces modality asynchrony. To mitigate this, we propose ANCHOR, a framework employing an active visual anchoring mechanism that accelerates early visual semantic generation to effectively rectify cross-modal alignment discrepancies. Integrating mechanistic intervention analysis with specialized training strategies, our approach significantly outperforms existing baselines across multiple benchmarks. It achieves robust reasoning in both zero-shot and fine-tuning scenarios for non-English languages, establishing a novel paradigm for enhancing the cross-cultural generalization capabilities of multilingual MLLMs.

0 citationsRead paper
Recent publications

Latest Papers

HuatuoGPT-3: RL-Only Domain Adaptation from Base Models

Oct 05, 2026

This study addresses the challenges of cold start and teacher distribution anchoring in pure reinforcement learning (RL) domain adaptation by proposing OnePO, a single-stage policy optimization framework. The method treats the teacher model's outputs as transient guidance signals, dynamically modulating instructional intensity through adaptive objective evolution and a teacher retirement mechanism. This design effectively mitigates gradient starvation while circumventing the complexity inherent in multi-stage optimization pipelines. As the first single-stage RL-only fine-tuning paradigm, OnePO achieves a score of 70.1 on the HealthBench benchmark, surpassing frontier models such as GPT-6 Astra. Furthermore, this work releases an open-source suite of 27B-parameter medical large language models to facilitate future research.

0 citationsRead paper

Why Vision Fails as a Universal Bridge: Rectifying Modality Asynchrony in Multilingual MLLMs

Aug 15, 2026

This study addresses the performance degradation of multilingual multimodal large language models in non-English visual reasoning by identifying the "Ghost Anchor" phenomenon, which induces modality asynchrony. To mitigate this, we propose ANCHOR, a framework employing an active visual anchoring mechanism that accelerates early visual semantic generation to effectively rectify cross-modal alignment discrepancies. Integrating mechanistic intervention analysis with specialized training strategies, our approach significantly outperforms existing baselines across multiple benchmarks. It achieves robust reasoning in both zero-shot and fine-tuning scenarios for non-English languages, establishing a novel paradigm for enhancing the cross-cultural generalization capabilities of multilingual MLLMs.

0 citationsRead paper