🤖 AI Summary
This study addresses the inherent trade-offs among resolution, field of view, and acquisition speed in biomedical imaging, along with the challenge of non-rigid deformation registration. We propose a generalizable deep learning framework based on optical flow to achieve real-time multimodal video stitching. Methodologically, an automated synthetic data generation pipeline is designed to facilitate rapid cross-modal adaptation. By integrating a fine-tuned optical flow model, synthetic deformation field training, and end-to-end pixel-level registration, the framework significantly enhances both training efficiency and generalization capability. Experimental results demonstrate that the proposed approach comprehensively outperforms existing baselines in accuracy, robustness, and speed across seven imaging modalities, successfully enabling real-time large-field-of-view visualization.
📝 Abstract
Biomedical imaging modalities often require a trade-off among resolution, field of view (FOV), and acquisition speed. Video mosaicking offers a strategy to overcome this limitation by computationally stitching sequential high-resolution frames into a wide-FOV composite. However, existing methods struggle with non-rigid deformations, and modality-specific artifacts arising in clinical and research imaging. Here, we present FloVMos, a generalizable, optical-flow-based deep learning framework for real-time video mosaicking across diverse biomedical imaging modalities. FloVMos achieves robust, pixel-level registration by fine-tuning an optical flow model on synthetic training data with ground-truth deformation fields. We introduce a pipeline for generating this training data, simulating realistic tissue motion and imaging distortions from existing mosaics or raw videos. Our automated synthetic data generation and optical flow model training based on this data allow users to adapt FloVMos to different imaging modalities. To demonstrate this, we applied FloVMos to seven diverse imaging modalities: reflection confocal microscopy, open-top light-sheet microscopy, fetoscopy, laparoscopy, dermoscopy, sparse spectral microscopy, and endoscopy. FloVMos outperforms conventional baselines in accuracy, robustness, and speed for all the tested modalities. This adaptable and training-efficient framework enables large-area visualization with real-time performance and may support broader use of video-based biomedical imaging in research and clinical workflows.