CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges in remote sensing referring segmentation (RRSIS), where weak architectural coupling leads to localization drift and an overreliance on semantic cues at the expense of spatial reasoning. To overcome these limitations, the authors propose the CROSS framework, which establishes the first tightly integrated collaboration between vision-language models (VLMs) and the Segment Anything Model (SAM). CROSS leverages language-guided cascaded distillation to incorporate SAM’s geometric priors and introduces viewpoint-spatial contrastive learning to enhance spatial-logical consistency. Additionally, it employs a dual-constraint mechanism combining structural prior injection and adversarial suppression of semantic shortcuts to effectively mitigate shortcut learning. The method achieves state-of-the-art performance across multiple RRSIS benchmarks and demonstrates robust high-precision localization even under severe spatial description perturbations.
📝 Abstract
Referring Remote Sensing Image Segmentation (RRSIS) has achieved significant progress through the integration of VLMs and the Segment Anything Model (SAM). However, this progress largely relies on strong pre-trained capabilities, while leaving two fundamental limitations insufficiently addressed: (1) Architectural Weak-Coupling, where the unidirectional flow forces reliance on coarse VLM prompts and wastes SAM's pixel-level structural guidance, causing localization drift; and (2) Object-Centric Semantic Bias, where models overemphasize dominant object semantics while remaining insensitive to spatial reasoning crucial for RRSIS. Motivated by these observations, we propose CROSS, a tightly integrated paradigm for RRSIS. First, we introduce Linguistic-Guided Cascaded Distillation (LGCD) to bridge the architectural gap, which distills SAM's geometric affinities as soft regularizers into VLM intermediate layers, injecting dense structural priors to refine localization. Second, Perspective-Spatial Contrastive Learning (PSCL) imposes cross-anchored constraints by mining mask-filtered deceptive distractors and spatial-linguistic counterfactuals as hard negatives, explicitly shattering semantic shortcuts to enforce genuine logical consistency. Extensive experiments on RRSIS benchmarks demonstrate that CROSS achieves state-of-the-art performance and maintains precise localization even under severe spatial description perturbations, standing as a robust new paradigm for RRSIS.
Problem

Research questions and friction points this paper is trying to address.

Referring Remote Sensing Image Segmentation
Architectural Weak-Coupling
Object-Centric Semantic Bias
Spatial Reasoning
Localization Drift
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cascaded Distillation
Dual-Constraint Grounding
Referring Remote Sensing Segmentation
Spatial-Linguistic Reasoning
Vision-Language Model
🔎 Similar Papers
2024-09-20IEEE Transactions on Geoscience and Remote SensingCitations: 2