Ontology Concept Overlap as a Training Signal: Knowledge-Grounded Reinforcement Learning for Clinical Question Answering
This study addresses the challenge of defining reinforcement learning rewards in clinical question answering, where executable verifiers are unavailable and subtle entity-level errors are prevalent. To this end, it proposes a soft verification method grounded in concept overlap within the UMLS ontology. This approach pioneers the use of controlled-vocabulary concept overlap as a graded external reward signal, integrating an entropy-normalized LLM judge with consistency penalties and optimizing model training via the GRPO algorithm to effectively capture entity substitution errors and enhance safety evaluation. Experimental results demonstrate that the proposed model achieves up to a 39% improvement in Token-F1 on MedQA and PubMedQA, significantly outperforming supervised fine-tuning baselines while exhibiting strong transferability.