Multi-modal, multi-scale representation learning for satellite imagery analysis just needs a good ALiBi

📅 2026-04-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge that existing vision foundation models struggle to effectively capture cross-scale spatial relationships among multimodal satellite images with varying spatial resolutions. To this end, the authors propose Scale-ALiBi, a novel mechanism that incorporates a linear spatial bias—derived from ground sampling distance—into Transformer attention. This is integrated within a joint representation learning framework combining triplet contrastive learning and reconstruction objectives for optical and synthetic aperture radar (SAR) imagery. The key contributions include the first extension of ALiBi to multiscale remote sensing scenarios, a tailored attention mechanism capable of modeling spatial relationships across image patches at different scales, and the creation of the first aligned multimodal, multiscale satellite image dataset. The proposed method achieves significant performance gains on GEO-Bench, and the dataset has been publicly released.

Technology Category

Computer Vision: Multi-modal VisionMachine Learning: Large Multimodal Models (LMMs)Intelligent Robots: Multimodal Perception & Sensor Fusion

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Vision foundation models have been shown to be effective at processing satellite imagery into representations fit for downstream tasks, however, creating models which operate over multiple spatial resolutions and modes is challenging. This paper presents Scale-ALiBi, a linear bias transformer attention mechanism with a spatial encoding bias to relationships between image patches at different ground sample distance scales. We provide an implementation of Scale-ALiBi over a dataset of aligned high- and low-resolution optical and low-resolution SAR satellite imagery data using a triple-contrastive and reconstructive architecture, show an improvement on the GEO-Bench benchmark, and release the newly curated dataset publicly.
Problem

Research questions and friction points this paper is trying to address.

multi-modal
multi-scale
satellite imagery
representation learning
spatial resolution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Scale-ALiBi
multi-modal representation learning
multi-scale satellite imagery
spatial encoding bias
transformer attention mechanism
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Patrick Kage
Artificial Intelligence and its Applications Institute, University of Edinburgh, 10 Crichton Street, Newington, Edinburgh EH8 9AB
P
Pavlos Andreadis
Artificial Intelligence and its Applications Institute, University of Edinburgh, 10 Crichton Street, Newington, Edinburgh EH8 9AB