🤖 AI Summary
This study addresses the challenge of fusing complementary MRI sequences and achieving structured diagnosis using vision-language models (VLMs) for knee joint imaging. To this end, we propose a novel dual-sequence whole-volume fusion modeling framework. By constructing a sequence-aware VLM that deeply integrates whole-volume data from both DESS and TSE sequences, the proposed method enables comprehensive prediction of 57 MOAKS indicators. This approach overcomes the inherent limitations of single-sequence analysis. Evaluated on a test set of one thousand cases, the model achieves an average accuracy of 72.98% and an ROC-AUC exceeding 78%, significantly outperforming existing mainstream configurations. These results demonstrate that our framework effectively enhances diagnostic precision across multiple anatomical structures in knee MRI assessment.
📝 Abstract
Vision-language models (VLMs) are increasingly being applied to three-dimensional medical imaging, but their application to knee MRI remains limited, particularly for interpreting the complementary sequences used in clinical practice. We introduce Knee3DVLM, a sequence-aware VLM that uses full-volume DESS and fluid-sensitive TSE MRI to predict 57 anatomically resolved binary diagnostic targets derived from the MRI Osteoarthritis Knee Score (MOAKS) for structured reporting. We evaluated DESS-only, TSE-only, and paired DESS-TSE configurations using subject-disjoint Osteoarthritis Initiative partitions. In a held-out cohort of 1,074 examinations, the fused model achieved 72.98% average accuracy, 71.17% balanced accuracy, 78.96% mean ROC-AUC, and 78.74% macro ROC-AUC, the highest values among the three configurations. In a secondary multiclass analysis aligned with the released 3DReasonKnee cohort, Knee3DVLM was numerically higher than the strongest reported 3DReasonKnee configuration across five pathology categories. These findings support dual-sequence full-volume modeling for comprehensive knee MRI assessment.