🤖 AI Summary
This study addresses the mismatch between random masking and multi-scale structures, along with computational bottlenecks in self-supervised pre-training for gigapixel scientific images, by proposing the SGMA framework. Specifically, this method introduces a content-adaptive quadtree tokenizer to compress images into fixed-length sequences and a structure-guided masking strategy to focus on information-rich regions. Furthermore, it innovatively incorporates a damped accumulation mechanism that aggregates cross-scale responses to stabilize the masking process, rendering the reconstruction task compatible with standard Vision Transformer (ViT) encoders. Experimental results demonstrate that SGMA significantly outperforms baseline methods across electron microscopy, whole-slide pathology, and X-ray CT datasets, achieving improvements of up to 16.84 points in Dice score while delivering a 24.8-fold acceleration in inference speed.
📝 Abstract
Self-supervised pre-training with Vision Transformers, including Masked Autoencoders (MAE), is difficult to apply to gigapixel scientific images. Random masking is poorly matched to the structured, multi-scale morphology of scientific data, while uniform tokenization produces prohibitively long sequences that make $O(N^2)$ attention impractical. We propose SGMA, a structure-guided masked autoencoding framework for ultra-high-resolution scientific images. SGMA couples two components: a content-adaptive quadtree tokenizer that compresses gigapixel images into a fixed-length sequence, and a structure-conditioned masking process that biases reconstruction toward spatially informative regions. To stabilize this process across scales, we introduce Damped Accumulation (DA), which aggregates signal-dependent responses across the tree into a structure canvas used to guide masking. The resulting pre-training task preserves fine microstructure while remaining compatible with standard ViT encoders and MAE-style reconstruction. Across electron microscopy, whole-slide optical microscopy, and X-ray CT datasets, SGMA consistently outperforms MAE baselines. It achieves 95.68% Dice on the 8K x 8K x 28K SpringXCT dataset, improving over the same-architecture MAE baseline by +13.00 points, and 83.21% Dice on the 32K^2 WSI PAIP dataset, improving by +16.84 points, while providing up to a 24.8x inference speedup.