🤖 AI Summary
This study addresses the disconnection between visual features and diagnostic predictions in radiology report generation, as well as the lack of spatial discrimination inherent in global gating mechanisms. To overcome these limitations, we propose a framework integrating position-aware gating with decoder-level supervised contrastive learning. Specifically, our method leverages diagnostic predictions to dynamically modulate visual patches, enabling spatially adaptive feature fusion without requiring region-level annotations. Furthermore, it constructs structured decoder representations based on shared positive findings to enhance multimodal semantic alignment. Experimental results demonstrate that the proposed approach achieves a clinical efficacy F1 score of 0.502 on the MIMIC-CXR dataset and an F1 score of 0.226 under zero-shot transfer evaluation on IU X-Ray, significantly outperforming existing baselines.
📝 Abstract
Radiology report generation models can produce fluent text while still containing finding-level inaccuracies. Diagnosis-driven methods improve generation by conditioning on predicted findings, but these predictions do not directly modify the visual patch features provided to the decoder, and global gating applies the same modulation across spatial locations. We introduce a Position-Aware Gate (PAG) that uses predicted finding representations to modulate visual patches spatially without region supervision. We also propose a Decoder-level Supervised Contrastive Loss (DSCL) that structures decoder representations using shared positive findings rather than instance identity. On MIMIC-CXR, PAG+DSCL improves clinical efficacy (CE) F1 from 0.484 to 0.502 over a matched global-gate reference, while PAG and DSCL individually reach 0.491 and 0.495. Without additional fine-tuning, the combined model achieves 0.226 CE F1 on IU X-Ray, compared with 0.211 reported by PromptMRG.