🤖 AI Summary
This study addresses the inconsistent responses of SAM3 in remote sensing image segmentation caused by geometric transformations. To this end, it proposes a robust segmentation framework grounded in D4 group equivariance. Methodologically, the framework introduces Multi-scale Harmonic-guided View Selection (MH-D4VS), Anchored Manifold Restoration (OAMR), and a Pixel Decoder Test-Time Adaptation (PD-TTA) module, which jointly enable feature-level cascaded restoration and online fine-tuning of GroupNorm layers. Experimental evaluations across eight benchmarks demonstrate that the proposed approach achieves an average mIoU of 55.6%, substantially enhancing segmentation stability and performance consistency under diverse inference conditions.
📝 Abstract
The semantic information of objects in remote sensing images is typically invariant to geometric transformations from the dihedral group D4. However, SAM3-based open-vocabulary semantic segmentation (OVSS) methods often exhibit inconsistent responses to different geometric transformations. To exploit this property and improve the stability of OVSS for remote sensing images, we propose a feature-adaptive manifold repair method based on dihedral-group geometric transformations. First, we introduce multi-scale harmonic-guided D4 view selection (MH-D4VS) to select complementary candidate views from a set of geometrically transformed views. Next, we propose original-view-anchored adaptive manifold repair (OAMR), which uses the original view as an anchor and reliable cross-view information to selectively repair locally unreliable visual features. Finally, we develop pixel decoder test-time adaptation (PD-TTA) for SAM3, which fine-tunes only the parameters of the GroupNorm layers online during inference, thereby enhancing the model's ability to adapt to sample-level distribution shifts. Experimental results show that the proposed method achieves an average mIoU of 55.6% across eight remote sensing semantic segmentation benchmarks and delivers consistent performance improvements under different SAM3-based inference frameworks.