🤖 AI Summary
This study addresses the challenges of insufficient modeling of tumor-invaded anatomical structures and the difficulty of fine-grained image-report alignment in hypopharyngeal cancer T-staging. To this end, we propose an anatomy-aware multimodal framework. Specifically, we construct an anatomically structured organ graph to integrate CT-derived organ context and design an organ-anchored cross-modal alignment mechanism for precise matching between visual and textual features. Furthermore, a report-enhanced graph refinement strategy is introduced to optimize topological representations. Technically, the framework achieves deep multimodal fusion by synergizing graph neural networks, natural language processing, and computer vision techniques. Experimental results demonstrate that the proposed method significantly outperforms existing baseline models on the hypopharyngeal cancer T-staging task.
📝 Abstract
Accurate T-staging is crucial for guiding personalized treatment strategies for laryngopharyngeal cancer. However, current clinical practice relies on invasive biopsy procedures, whereas CT-based staging remains challenging due to the complex patterns of tumor invasion. Recent computer-aided approaches face two key challenges: 1) Structural relationship modeling: existing methods underrepresent anatomically structured patterns of tumor invasion, as they either process whole CT volumes without tumor-specific anatomical constraints or rely on labor-intensive tumor segmentation. 2) Fine-grained cross-modal alignment: while radiology reports contain organ-specific invasion details, current methods that apply global feature fusion struggle to accurately align individual anatomical structures with their corresponding textual descriptions. To address these issues, we propose an anatomy-aware multimodal framework that integrates organ-level CT context and radiology reports into a unified representation for laryngopharyngeal T-staging. The framework first constructs an Anatomy-Structured Organ Graph (AOG) that captures invasion patterns between primary sites and surrounding organs, then performs Organ-Anchored Cross-Modal Alignment (OCA) so that each organ node aggregates textual evidence from the radiology report, and finally refines this graph representation by injecting organ-specific invasion cues extracted from the report via Report-Enhanced Graph-Refinement (REG), yielding a multimodal organ graph that combines spatial and textual evidence. Extensive experiments demonstrate that the proposed framework achieves superior performance in T-staging of laryngopharyngeal cancer.