🤖 AI Summary
This work addresses the anisotropic resolution and degraded image quality in line-scanning confocal microscopy caused by varying slit widths, a challenge that typically requires separate deep learning models for each optical configuration. To overcome this limitation, the authors propose a unified resolution-conditioned image fusion framework that enables a single model to adapt to multiple slit settings. The approach employs feature-wise linear modulation (FiLM) to continuously tune the network’s response to different degradation scales and introduces an adaptive Rank Enhanced Linear Attention (RELA) module, which integrates learnable temperature parameters with multi-scale depthwise convolutions to dynamically adjust to varying degrees of degradation. The method effectively suppresses interpolation artifacts, achieving PSNR values of 34–40 dB across diverse configurations—substantially outperforming both unconditioned multi-slit training (24.3 dB) and specialized models (which suffer 4–9 dB extrapolation loss)—and demonstrates smooth generalization to unseen intermediate slit widths.
📝 Abstract
Laser line-scanning microscopy enables fast volumetric imaging but produces anisotropic lateral resolution. Orthogonal line scans provide complementary directional information that can recover near-isotropic resolution, yet existing deep-learning methods require a separate model for each optical configuration. We present a unified, resolution-conditioned fusion framework based on Rank Enhanced Linear Attention (RELA). Feature-wise Linear Modulation (FiLM) conditions the network continuously on the resolving-power ratio, enabling one model to adapt across slit widths. We further introduce Adaptive RELA, which replaces fixed-kernel rank enhancement with ratio-conditioned multi-scale depthwise convolutions and uses a learnable attention temperature to adjust selectivity with degradation severity. Training data spanning multiple slit configurations are generated using a physics-grounded separable point-spread-function model verified against measured optical data at 48.3 dB accuracy. The resulting model achieves 34-40 dB PSNR across configurations, whereas unconditioned multi-slit training collapses to 24.3 dB and per-slit specialists lose 4-9 dB outside their training setting. It also generalizes smoothly to unseen intermediate configurations without interpolation artifacts. Ablations show that FiLM resolves configuration ambiguity, global linear attention captures long-range directional correspondences, and adaptive temperature yields an additional 2 dB in the challenging near-isotropic regime, where complementary signals are weak.