🤖 AI Summary
Existing approaches to radiology report generation prioritize linguistic fluency but struggle to controllably balance critical clinical metrics—such as precision and recall—thereby limiting their clinical applicability. This work proposes a reinforcement learning–based controllable generation framework that dynamically adjusts the trade-off between precision and recall during inference via tunable parameters. To enhance training stability, the method introduces a hybrid optimization objective combining natural language and clinical rewards, along with an intra-group relative reward normalization strategy. Notably, this is the first approach to enable explicit control over clinical metrics in generated radiology reports. Evaluated on the MIMIC-CXR dataset, the proposed method significantly outperforms existing techniques, achieving the best balance between generation quality and clinical effectiveness.
📝 Abstract
Automated radiology report generation (RRG) has gained increasing attention because it can reduce the heavy workload of clinical report writing. However, most existing methods mainly optimize for natural language generation (NLG) metrics that focus on language fluency, while providing little control over clinically important factors such as precision and recall. As consequence, generated reports may be fluent but not well aligned with different clinical needs. To address this challenge, we propose a reinforcement learning framework for precision recall controllable RRG, where a control parameter explicitly adjusts the trade-off between clinical precision and recall during inference. This design allows the model to flexibly generate reports according to different clinical requirements. To ensure clinical correctness, we introduce a clinical reward into the training objective, which helps improve clinical efficacy (CE) beyond standard language-based optimization. In addition, we apply a group-relative training strategy that normalizes rewards within each training group, reducing reward variance and improving training stability. Extensive experiments on the MIMIC-CXR dataset show that our method consistently outperforms state-of-the-art approaches in both NLG and CE evaluation metrics, while providing reliable control over the CE precision recall trade-off.