Scalable In-Domain Self-Supervised Foundation Model for Dense Representation Transfer in High-Resolution Plant Imaging

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of expensive annotation, high dimensionality, and imaging variability that hinder dense feature transfer in high-resolution plant image analysis. To overcome these limitations, this work proposes a large-scale in-domain self-supervised pretraining paradigm. Specifically, we construct a Vision Transformer (ViT)-based masked autoencoder pretrained in a distributed manner on tens of millions of multi-view plant images, incorporating multi-scale data augmentation to optimize dense representation capabilities. This approach effectively bridges the domain gap between general-purpose foundation models and scientific imaging. Experimental results demonstrate that, under limited supervision, the proposed method achieves a Mean Dice score of 0.8686, significantly outperforming natural image pretraining baselines. These findings validate the benefits of scaling up pretraining for improving dense feature transfer in specialized botanical applications.
📝 Abstract
High-resolution plant imaging enables detailed characterization of plant morphology, but dense scientific analysis remains limited by costly pixel-level annotations, large image pixel dimensions, and substantial variation in imaging conditions. This work proposes a scalable in-domain self-supervised pretrained foundation model for high-resolution, high-pixel-dimension multi-species plant imagery. A masked autoencoder with a ViT backbone is pretrained on more than 10 million multi-view plant image tiles using distributed training. Following scalable pretraining, the learned foundation-model representations are comprehensively benchmarked across fine-grained dense prediction and coarse global feature recognition, with particular emphasis on limited supervision and realistic downstream imaging conditions. This work focuses on the domain gap of existing foundation models in dense feature representation and transfer. Through extensive experiments involving limited annotations, cross-view variation, and resolution degradation, the in-domain FM achieves a Mean Dice of 0.8686 and a Pooled Dice of 0.8959, outperforming an MAE counterpart pretrained on large-scale natural-image data by 0.0694 and 0.0613, respectively. The results further indicate that increasing pretraining scale produces consistent improvements in dense feature transfer. Overall, these findings suggest that scaling in-domain self-supervised pretraining can reduce the domain gap and improve transferable dense representations for high-pixel-dimension scientific imaging.
Problem

Research questions and friction points this paper is trying to address.

plant imaging
dense representation transfer
domain gap
self-supervised learning
foundation model
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Supervised Learning
Foundation Model
Masked Autoencoder
Dense Representation Transfer
Plant Imaging
🔎 Similar Papers
No similar papers found.