Cross-Modal Bidirectional Interaction Model for Referring Remote Sensing Image Segmentation

📅 2024-10-11
🏛️ arXiv.org
📈 Citations: 5
✨ Influential: 1
📄 PDF
🤖 AI Summary
Remote sensing referring image segmentation (RRSIS) faces challenges in pixel-level localization due to complex geospatial relationships, highly variable object scales, and weak visual saliency. To address these, we propose CroBIM, a cross-modal bidirectional interaction framework featuring three key innovations: (1) a novel attention-deficiency compensation mechanism, (2) context-aware prompt modulation, and (3) language-guided feature aggregation with a mutual interaction decoder. CroBIM integrates multi-scale language-guided attention and cascaded bidirectional cross-attention to enhance fine-grained alignment between linguistic expressions and remote sensing imagery. We introduce RISBench—a large-scale, manually curated benchmark comprising 52,472 triplets—and evaluate CroBIM on RISBench and two established datasets. Our method achieves significant improvements over state-of-the-art approaches across all benchmarks. The code and RISBench dataset are publicly available.

Technology Category

Computer Vision: Remote Sensing / Geospatial AIIntelligent Robots: Multimodal Perception & Sensor FusionNatural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web data
📝 Abstract
Given a natural language expression and a remote sensing image, the goal of referring remote sensing image segmentation (RRSIS) is to generate a pixel-level mask of the target object identified by the referring expression. In contrast to natural scenarios, expressions in RRSIS often involve complex geospatial relationships, with target objects of interest that vary significantly in scale and lack visual saliency, thereby increasing the difficulty of achieving precise segmentation. To address the aforementioned challenges, a novel RRSIS framework is proposed, termed the cross-modal bidirectional interaction model (CroBIM). Specifically, a context-aware prompt modulation (CAPM) module is designed to integrate spatial positional relationships and task-specific knowledge into the linguistic features, thereby enhancing the ability to capture the target object. Additionally, a language-guided feature aggregation (LGFA) module is introduced to integrate linguistic information into multi-scale visual features, incorporating an attention deficit compensation mechanism to enhance feature aggregation. Finally, a mutual-interaction decoder (MID) is designed to enhance cross-modal feature alignment through cascaded bidirectional cross-attention, thereby enabling precise segmentation mask prediction. To further forster the research of RRSIS, we also construct RISBench, a new large-scale benchmark dataset comprising 52,472 image-language-label triplets. Extensive benchmarking on RISBench and two other prevalent datasets demonstrates the superior performance of the proposed CroBIM over existing state-of-the-art (SOTA) methods. The source code for CroBIM and the RISBench dataset will be publicly available at https://github.com/HIT-SIRS/CroBIM
Problem

Research questions and friction points this paper is trying to address.

Segmenting remote sensing images using complex language expressions
Handling geospatial relationships and varying object scales
Improving cross-modal feature alignment for precise segmentation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Context-aware prompt modulation for spatial integration
Language-guided feature aggregation with attention compensation
Mutual-interaction decoder for cross-modal alignment
🔎 Similar Papers
2024-09-20IEEE Transactions on Geoscience and Remote SensingCitations: 2
Harbin Institute of Technology
Zhe Dong
Zhe Dong
Microsoft AI
Y
Yuzhe Sun
School of Electronics and Information Engineering, Harbin Institute of Technology, Harbin 150001, China
Yanfeng Gu
Yanfeng Gu
Professor of Electronics Engineering, Harbin Institute of Technology
image processingpattern recognitionmachine learning
T
Tianzhu Liu
School of Electronics and Information Engineering, Harbin Institute of Technology, Harbin 150001, China