🤖 AI Summary
This work addresses the challenge that existing vision foundation models struggle to effectively capture cross-scale spatial relationships among multimodal satellite images with varying spatial resolutions. To this end, the authors propose Scale-ALiBi, a novel mechanism that incorporates a linear spatial bias—derived from ground sampling distance—into Transformer attention. This is integrated within a joint representation learning framework combining triplet contrastive learning and reconstruction objectives for optical and synthetic aperture radar (SAR) imagery. The key contributions include the first extension of ALiBi to multiscale remote sensing scenarios, a tailored attention mechanism capable of modeling spatial relationships across image patches at different scales, and the creation of the first aligned multimodal, multiscale satellite image dataset. The proposed method achieves significant performance gains on GEO-Bench, and the dataset has been publicly released.
📝 Abstract
Vision foundation models have been shown to be effective at processing satellite imagery into representations fit for downstream tasks, however, creating models which operate over multiple spatial resolutions and modes is challenging. This paper presents Scale-ALiBi, a linear bias transformer attention mechanism with a spatial encoding bias to relationships between image patches at different ground sample distance scales. We provide an implementation of Scale-ALiBi over a dataset of aligned high- and low-resolution optical and low-resolution SAR satellite imagery data using a triple-contrastive and reconstructive architecture, show an improvement on the GEO-Bench benchmark, and release the newly curated dataset publicly.