MDF-MLLM: Deep Fusion Through Cross-Modal Feature Alignment for Contextually Aware Fundoscopic Image Classification

📅 2025-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Current multimodal large language models (MLLMs) struggle to capture diagnostically critical low-level spatial details in fundus images, limiting classification accuracy for glaucoma, diabetic retinopathy, and inherited retinal diseases (e.g., retinitis pigmentosa). To address this, we propose a multi-depth cross-modal fusion architecture. Our method leverages a U-Net encoder to extract multiscale visual features and integrates Feature-wise Linear Modulation (FiLM) conditioning with scaled cross-attention to achieve fine-grained alignment between local image features and global textual semantics within LLaMA-3.2-11B. Evaluated on a dataset of 1,305 fundus images, the model achieves 94% classification accuracy—outperforming baseline MLLMs by 56%—and attains up to a 35% improvement in F1-score, particularly enhancing recognition performance for inherited retinal diseases.

Technology Category

Computer Vision: Multi-modal VisionMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
This study aimed to enhance disease classification accuracy from retinal fundus images by integrating fine-grained image features and global textual context using a novel multimodal deep learning architecture. Existing multimodal large language models (MLLMs) often struggle to capture low-level spatial details critical for diagnosing retinal diseases such as glaucoma, diabetic retinopathy, and retinitis pigmentosa. This model development and validation study was conducted on 1,305 fundus image-text pairs compiled from three public datasets (FIVES, HRF, and StoneRounds), covering acquired and inherited retinal diseases, and evaluated using classification accuracy and F1-score. The MDF-MLLM integrates skip features from four U-Net encoder layers into cross-attention blocks within a LLaMA 3.2 11B MLLM. Vision features are patch-wise projected and fused using scaled cross-attention and FiLM-based U-Net modulation. Baseline MLLM achieved 60% accuracy on the dual-type disease classification task. MDF-MLLM, with both U-Net and MLLM components fully fine-tuned during training, achieved a significantly higher accuracy of 94%, representing a 56% improvement. Recall and F1-scores improved by as much as 67% and 35% over baseline, respectively. Ablation studies confirmed that the multi-depth fusion approach contributed to substantial gains in spatial reasoning and classification, particularly for inherited diseases with rich clinical text. MDF-MLLM presents a generalizable, interpretable, and modular framework for fundus image classification, outperforming traditional MLLM baselines through multi-scale feature fusion. The architecture holds promise for real-world deployment in clinical decision support systems. Future work will explore synchronized training techniques, a larger pool of diseases for more generalizability, and extending the model for segmentation tasks.
Problem

Research questions and friction points this paper is trying to address.

Improving retinal disease classification accuracy from fundus images
Capturing low-level spatial details for diagnosing retinal diseases
Integrating fine-grained image features with global textual context
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fuses U-Net skip features with LLaMA MLLM via cross-attention
Projects vision patches and uses FiLM-based U-Net modulation
Integrates multi-depth fusion for spatial reasoning and classification
🔎 Similar Papers
No similar papers found.
J
Jason Jordan
School of Science and Engineering, University of Missouri-Kansas City, Kansas City, MO, USA
M
Mohammadreza Akbari Lor
School of Science and Engineering, University of Missouri-Kansas City, Kansas City, MO, USA
P
Peter Koulen
Vision Research Center, Department of Ophthalmology, School of Medicine, University of Missouri-Kansas City, Kansas City, MO, USA
Mei-Ling Shyu
Mei-Ling Shyu
Professor of Electrical and Computer Engineering, University of Missouri-Kansas City
data scienceAImachine learningdeep learningmultimedia information systems
S
Shu-Ching Chen
Data Science and Analytics Innovation Center, University of Missouri-Kansas City, Kansas City, MO, USA