🤖 AI Summary
This study addresses the challenge in radiology report generation where cross-modal distribution shifts and the rigid coupling of heterogeneous information dilute visual anomaly cues. To overcome this, we propose the HSA-HER framework, which introduces an explicit homogeneous distribution constraint to achieve semantic alignment for extracting purified visual features. Furthermore, it designs a disease semantic anchor-guided hierarchical expert routing mechanism to replace conventional rigid coupling paradigms, enabling the adaptive fusion and semantic reconstruction of multi-source clinical evidence through dynamic weight allocation. Extensive experiments demonstrate that our approach achieves state-of-the-art performance across three mainstream benchmark datasets, accurately characterizing complex imaging details and critical diagnostic information.
📝 Abstract
Radiology report generation (RRG) aims to convert medical images into diagnostic texts to assist in clinical decision-making and alleviate the workload of physicians. Although existing methods have made extensive progress in cross-modal interaction and the incorporation of external priors, the distribution shift of underlying representations and the undifferentiated rigid coupling of heterogeneous information cause weak visual abnormality cues to be easily diluted by massive text priors and generation inertia during decoding. To overcome this bottleneck, inspired by cognitive science, we propose a novel two-stage Homogeneous Semantic Alignment and Hierarchical Expert Routing (HSA-HER) framework. First, the model introduces an explicit homogeneous distribution constraint in the underlying latent space to effectively eliminate the cross-modal distribution shift between visual and textual features, thereby extracting purified visual features as semantic anchors that accurately align with diseases. Second, for heterogeneous clinical evidence composed of visual features, local entities, and global retrievals, we design a hierarchical expert routing mechanism guided by these disease semantic anchors. This mechanism abandons the undifferentiated rigid coupling paradigm. Specifically, it dynamically activates expert networks to perform targeted mining and semantic reconstruction on multi-source evidence, and adaptively allocates fusion weights. Extensive experiments on three mainstream benchmark datasets demonstrate that HSA-HER achieves state-of-the-art performance, accurately depicting complex imaging details and key diagnostic information.