Structure-Guided Masked Autoencoders for Ultra-High Resolution Scientific Image Understanding

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the mismatch between random masking and multi-scale structures, along with computational bottlenecks in self-supervised pre-training for gigapixel scientific images, by proposing the SGMA framework. Specifically, this method introduces a content-adaptive quadtree tokenizer to compress images into fixed-length sequences and a structure-guided masking strategy to focus on information-rich regions. Furthermore, it innovatively incorporates a damped accumulation mechanism that aggregates cross-scale responses to stabilize the masking process, rendering the reconstruction task compatible with standard Vision Transformer (ViT) encoders. Experimental results demonstrate that SGMA significantly outperforms baseline methods across electron microscopy, whole-slide pathology, and X-ray CT datasets, achieving improvements of up to 16.84 points in Dice score while delivering a 24.8-fold acceleration in inference speed.
📝 Abstract
Self-supervised pre-training with Vision Transformers, including Masked Autoencoders (MAE), is difficult to apply to gigapixel scientific images. Random masking is poorly matched to the structured, multi-scale morphology of scientific data, while uniform tokenization produces prohibitively long sequences that make $O(N^2)$ attention impractical. We propose SGMA, a structure-guided masked autoencoding framework for ultra-high-resolution scientific images. SGMA couples two components: a content-adaptive quadtree tokenizer that compresses gigapixel images into a fixed-length sequence, and a structure-conditioned masking process that biases reconstruction toward spatially informative regions. To stabilize this process across scales, we introduce Damped Accumulation (DA), which aggregates signal-dependent responses across the tree into a structure canvas used to guide masking. The resulting pre-training task preserves fine microstructure while remaining compatible with standard ViT encoders and MAE-style reconstruction. Across electron microscopy, whole-slide optical microscopy, and X-ray CT datasets, SGMA consistently outperforms MAE baselines. It achieves 95.68% Dice on the 8K x 8K x 28K SpringXCT dataset, improving over the same-architecture MAE baseline by +13.00 points, and 83.21% Dice on the 32K^2 WSI PAIP dataset, improving by +16.84 points, while providing up to a 24.8x inference speedup.
Problem

Research questions and friction points this paper is trying to address.

ultra-high resolution scientific images
masked autoencoders
self-supervised pre-training
gigapixel images
Vision Transformers
Innovation

Methods, ideas, or system contributions that make the work stand out.

Structure-Guided Masked Autoencoder
Quadtree Tokenizer
Damped Accumulation
Self-Supervised Learning
Ultra-High Resolution Image
🔎 Similar Papers
2024-03-07arXiv.orgCitations: 2
Enzhi Zhang
Enzhi Zhang
Hokkaido University, Information Initiative Center, Specially Appointed Assistant Professor
Deep learningMeta LearningReinforcement LearningEvolutionary Algorithm.
Du Wu
Du Wu
Tokyo Institute of Technology
High Performance Computing (HPC)
Rui Zhong
Rui Zhong
Hokkaido University, Information Initiative Center, Specially Appointed Assistant Professor
Computational IntelligenceEvolutionary ComputationLarge-scale OptimizationMeta/Hyper-heuristics
C
Cong Ma
Hokkaido University, Japan
I
Isaac Lyngaas
Oak Ridge National Laboratory, USA
Amir Koushyar Ziabari
Amir Koushyar Ziabari
Senior R&D Scientist, ESTD/EEID/MSA Group, Oak Ridge National Laboratory
Data Analytics for Advanced ManufacturingPhysic-Based Computational ImagingDeep Learning
Xiao Wang
Xiao Wang
Oak Ridge National Laboratory
High Performance ComputingComputational ImagingMachine Learning
Peng Chen
Peng Chen
RIKEN Center for Computational Science (R-CCS)
HPCGPGPUMachine LearningImage Processing
T
Tao Luo
A*STAR, Singapore
T
Toshio Endo
Institute of Science Tokyo, Japan
F
Fumiyoshi Shoji
RIKEN Center for Computational Science (R-CCS), Japan
Kento Sato
Kento Sato
RIKEN Center for Computational Science (RIKEN R-CCS)
High performance computingI/O & Big dataMachine/Deep learningFault toleranceDebugging
Kentaro Uesugi
Kentaro Uesugi
Japan Synchrotron Radiation Research Institute
Computed tomographyX-ray image detectorphase contrast
T
Takayuki Nonoyama
Hokkaido University, Japan
R
Ryuji Kiyama
Hokkaido University, Japan
M
Masahiro Yoshida
Hokkaido University, Japan
M
Masaru Tezuka
Hokkaido University, Japan
T
Tetsuya Ishikawa
RIKEN SPring-8 Center, Japan
Satoshi Matsuoka
Satoshi Matsuoka
RIKEN Center for Computational Science (R-CCS) / Tokyo Institute of Technology
HPCBig DataScalable AIGreen ComputingPost-Moore Computing
Masaharu Munetomo
Masaharu Munetomo
Information Initiative Center, Hokkaido University
Evolutionary ComputationCloud Computing
Mohamed Wahib
Mohamed Wahib
RIKEN Center for Computational Science (R-CCS)
High Performance ComputingParallel/Distributed ComputingHigh-Performance AI ...