Knee3DVLM: Dual-Sequence Full-Volume Vision-Language Modeling for Comprehensive Knee MRI Assessment
This study addresses the challenge of fusing complementary MRI sequences and achieving structured diagnosis using vision-language models (VLMs) for knee joint imaging. To this end, we propose a novel dual-sequence whole-volume fusion modeling framework. By constructing a sequence-aware VLM that deeply integrates whole-volume data from both DESS and TSE sequences, the proposed method enables comprehensive prediction of 57 MOAKS indicators. This approach overcomes the inherent limitations of single-sequence analysis. Evaluated on a test set of one thousand cases, the model achieves an average accuracy of 72.98% and an ROC-AUC exceeding 78%, significantly outperforming existing mainstream configurations. These results demonstrate that our framework effectively enhances diagnostic precision across multiple anatomical structures in knee MRI assessment.