🤖 AI Summary
Current AI research on color fundus photography (CFP) lacks a systematic perspective on the co-evolution of datasets, preprocessing, and modeling. This work proposes the first unified framework that synergistically optimizes data curation, preprocessing, and multimodal modeling, integrating neural data engineering, hardware-aware annotation, self-supervised electronic health record imputation, vision foundation models, state space models, and multimodal mixture-of-experts architectures. The study delineates a clear evolutionary trajectory for CFP analysis—from single-task convolutional neural networks toward multimodal, longitudinally integrated clinical systems—and provides a methodological roadmap toward robust ophthalmic AI capable of clinical deployment, cross-domain generalization, and edge intelligence.
📝 Abstract
Color Fundus Photography (CFP) is a primary non-invasive imaging modality for large-scale screening of ophthalmic and systemic diseases. Existing surveys mainly summarize task-specific algorithms, datasets, or preprocessing techniques independently, lacking a unified perspective on their co-evolution with modern artificial intelligence. This review provides an integrated overview of CFP AI through the interplay of dataset evolution, preprocessing paradigms, and modeling frameworks. We show that CFP datasets have evolved from small single-center collections with task-specific labels to large multi-center resources featuring multimodal pairings and longitudinal clinical records. Preprocessing has progressed from conventional image enhancement to neural data-engineering pipelines, hardware-aware token optimization, and self-supervised imputation for incomplete electronic health records (EHRs). Meanwhile, modeling has advanced from convolutional neural networks (CNNs) to vision foundation models, state space models (SSMs), and multimodal expert architectures. At the multimodal frontier, CFP is increasingly integrated with EHRs and longitudinal patient information, enabling more comprehensive clinical reasoning beyond isolated image analysis. We conclude that future progress depends on the collaborative optimization of datasets, preprocessing, and multimodal modeling, providing a roadmap toward robust clinical deployment, improved cross-domain generalization, and resource-efficient edge intelligence.