🤖 AI Summary
This study addresses the limitation of existing multi-sequence MRI report generation methods, where fixed-rule fusion often dilutes critical lesion information and leads to suboptimal clinical recall. To overcome this, we propose a vision-language framework that employs dual parallel 3D encoders to extract volumetric features from sagittal and axial planes. Furthermore, we design a training-free gated cross-view fusion module that dynamically integrates features through channel- and token-level adaptive weight allocation, enabling high-informative views to dominate the fused representation. This representation subsequently drives a vision-language model for diagnostic report generation. Experimental results demonstrate that the proposed method significantly improves clinical F1 scores and recall across three datasets while maintaining competitive performance on natural language generation metrics.
📝 Abstract
Automated report generation can ease the burden radiolo gists face when interpreting multi-sequence MRI studies. Unlike CT, MRI examinations comprise multiple sequences and imaging planes, each con tributing complementary diagnostic information. Existing methods en code a study as a single volume and combine multiple acquisitions by fixed rules. Findings visible in only one plane are thus diluted and of ten missed, lowering recall on clinical efficacy metrics, where a missed abnormality is most costly. We propose GateSPINE, a vision-language framework that fuses sagittal T1 and T2 volumes with a training-free operator, encodes the fused sagittal and axial volumes with two parallel 3D encoders, and decodes their combined representation into a report. Its core mechanism is a gated cross view fusion module that predicts, per feature channel and token, how much of each view to admit, so the more informative view dominates at each spatial location. We evaluate GateSPINE on three lumbar MRI datasets, comprising two public bench marks and a private cohort collected from Phenikaa University Hospital, using both natural language generation (NLG) and clinical efficacy (CE) metrics. GateSPINE achieves the highest CE F1 through improved re call on all three datasets; on SPIDER, which lacks an axial sequence, this reflects the sagittal fusion component rather than the gated cross-view mechanism, which is validated on the two cohorts with both imaging planes. GateSPINE also remains competitive on standard NLG metrics.