Enhancing Visual Reasoning in Chest X-Ray Report Generation Using Reinforcement Learning

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of supervised fine-tuning (SFT) in chest X-ray report generation, which is prone to hallucinations and yields unverifiable reasoning processes. Building upon Qwen3-VL, this work proposes a reasoning verification framework that integrates SFT with reinforcement learning. The framework constructs a structured chain-of-thought by combining anatomical region localization with textual descriptions, and introduces a dual reward mechanism encompassing spatial and factual criteria to jointly optimize visual reasoning steps and final outputs. This approach significantly enhances both the clinical accuracy and interpretability of generated reports, outperforming purely SFT-based methods while revealing potential reward hacking patterns.
📝 Abstract
Medical report generation has made significant progress with the rise of modern vision-language models and the growing availability of large-scale medical datasets. However, hallucinations remain a major challenge, largely due to the limitations of supervised fine-tuning (SFT), which prioritizes lexical similarity to reference reports rather than clinical correctness. While reinforcement learning has shown strong performance in domains with verifiable rewards such as mathematics and code generation, its application to open-ended medical tasks remains limited. Existing work focuses on evaluating final answers, overlooking the model's reasoning, despite evidence that flawed reasoning can degrade overall performance. In this study, we propose a framework that verifies the model's reasoning process by integrating anatomical regions, bounding boxes, and region-level textual descriptions. We design spatial and factual reward mechanisms to ensure that the model's reasoning is both visually grounded and factually accurate. Starting from Qwen3-VL-8B-Instruct as our base model, we adapt it to the medical domain using supervised fine-tuning, introduce reasoning capability through a cold-start SFT stage, and refine it with reinforcement learning. We find that RL provides performance gains beyond those achievable through SFT alone, and that jointly verifying both reasoning steps and final outputs yields larger improvements than verifying either in isolation. We further identify multiple modes of reward hacking in the RL stage. Finally, the model's structured think traces enhance interpretability, making its outputs easier to audit for clinical use.
Problem

Research questions and friction points this paper is trying to address.

Chest X-Ray Report Generation
Hallucination
Visual Reasoning
Reinforcement Learning
Medical Vision-Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Visual Reasoning
Chest X-Ray Report Generation
Reward Mechanism
Hallucination Mitigation
🔎 Similar Papers
No similar papers found.