GeoAlignCLIP: Enhancing Fine-Grained Vision-Language Alignment in Remote Sensing via Multi-Granular Consistency Learning

📅 2026-03-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitation of existing remote sensing vision-language models, which predominantly rely on global image-text alignment and thus struggle with fine-grained semantic understanding. To overcome this, we propose GeoAlignCLIP, a novel framework that introduces, for the first time in remote sensing, a multi-granularity consistency learning mechanism. This approach jointly optimizes region-level text alignment and intra-modal consistency through contrastive learning. We also construct RSFG-100k, a new dataset comprising scene descriptions, region annotations, and hard negative samples to provide hierarchical supervisory signals. Extensive experiments demonstrate that our method significantly outperforms current state-of-the-art models across multiple remote sensing benchmarks, exhibiting superior fine-grained alignment capability and enhanced generalization across diverse tasks.

Technology Category

Computer Vision: Remote Sensing / Geospatial AINatural Language Processing: Language Grounding & Multi-modal NLPIntelligent Robots: Multimodal Perception & Sensor Fusion

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Vision-language pretraining models have made significant progress in bridging remote sensing imagery with natural language. However, existing approaches often fail to effectively integrate multi-granular visual and textual information, relying primarily on global image-text alignment. This limitation hinders the model's ability to accurately capture fine-grained details in images, thus restricting its performance in complex, fine-grained tasks. To address this, we propose GeoAlignCLIP, a unified framework that achieves fine-grained alignment in remote sensing tasks by learning multi-granular semantic alignments and incorporating intra-modal consistency, enabling more precise visual-semantic alignment between image regions and text concepts. Additionally, we construct RSFG-100k, a fine-granular remote sensing dataset containing scene descriptions, region-level annotations, and challenging hard-negative samples, providing hierarchical supervision for model training. Extensive experiments conducted on multiple public remote-sensing benchmarks demonstrate that GeoAlignCLIP consistently outperforms existing RS-specific methods across diverse tasks, exhibiting more robust and accurate fine-grained vision-language alignment.
Problem

Research questions and friction points this paper is trying to address.

fine-grained alignment
remote sensing
vision-language alignment
multi-granular information
image-text alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

fine-grained alignment
multi-granular consistency learning
vision-language pretraining
remote sensing
GeoAlignCLIP
🔎 Similar Papers
2024-09-20IEEE Transactions on Geoscience and Remote SensingCitations: 2
💼 Related Jobs
No related jobs found.
X
Xiao Yang
College of Computer Science and Technology, Jilin University, Changchun 130012, China; Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education Jilin University
R
Ronghao Fu
College of Computer Science and Technology, Jilin University, Changchun 130012, China; Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education Jilin University
Z
Zhuoran Duan
College of Computer Science and Technology, Jilin University, Changchun 130012, China; Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education Jilin University
Z
Zhiwen Lin
College of Computer Science and Technology, Jilin University, Changchun 130012, China; Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education Jilin University
X
Xueyan Liu
College of Computer Science and Technology, Jilin University, Changchun 130012, China; Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education Jilin University
B
Bo Yang
College of Computer Science and Technology, Jilin University, Changchun 130012, China; Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education Jilin University