SARATR-X-v2: Scale-Aware Structural Pre-Training for SAR Foundation Models

📅 2026-07-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing SAR pretraining methods suffer from limited transferability due to reconstruction targets lacking physical stability and multi-scale semantic compatibility. This work proposes a structured pretraining objective that, for the first time, unifies physics-driven speckle invariance with downstream task-oriented multi-scale semantics within a masked image modeling framework. Specifically, six structurally designed extractors with fixed receptive fields—such as blind-spot local aggregation and directional log-ratio regional contrast—generate multi-scale supervision signals, which are then fused via learnable weights to guide reconstruction. The proposed approach achieves state-of-the-art transfer performance across twelve SAR benchmark tasks and reduces representation drift by nearly two orders of magnitude under synthetic speckle perturbations.
📝 Abstract
Masked image modeling has become a dominant paradigm for SAR pre-training, yet the design of the reconstruction target remains fundamentally unsettled. This article argues that a SAR pre-training target should satisfy two conditions to produce transferable representations: (i) physics-grounded stability, i.e., approximate invariance of the target operator to multiplicative speckle inherent in coherent imaging; and (ii) semantic scale compatibility, i.e., coverage of the heterogeneous spatial scales that downstream tasks demand. These two conditions are individually achievable but jointly difficult: physics-grounded stability favors fixed operators, while semantic scale compatibility favors data-driven composition. To this end, SARATR-X-v2 reconciles both within a single design. The target is constructed through fixed structural extractors spanning six receptive fields, from blind-spot local aggregation to directional log-ratio region contrast, and fused via learnable weights into one unified supervision signal for masked reconstruction. On twelve SAR benchmarks across classification, detection, and segmentation, SARATR-X-v2 achieves state-of-the-art transfer performance. Under synthetic speckle variation, the proposed target reduces perturbation drift in the learned representation by nearly two orders of magnitude relative to pixel-space supervision. Taken together, these results establish physics-grounded stability and semantic scale compatibility as a principled framework for pre-training target design under coherent imaging, and suggest that effective SAR pre-training is not about reconstructing more signal, but about reconstructing the right structural target.
Problem

Research questions and friction points this paper is trying to address.

SAR pre-training
reconstruction target
physics-grounded stability
semantic scale compatibility
coherent imaging
Innovation

Methods, ideas, or system contributions that make the work stand out.

masked image modeling
physics-grounded stability
semantic scale compatibility
structural pre-training
SAR foundation models
Weijie Li
Weijie Li
National University of Defense Technology
Synthetic Aperture RadarAutomatic Target RecognitionFoundation ModelSelf­Supervised Learning
Yafei Song
Yafei Song
Alibaba Group
Computer VisionMachine LearningAugmented RealityRobotics
Yongxiang Liu
Yongxiang Liu
Professor, National University of Defense Technology
Remote SensingSynthetic Aperture RadarRadarImage ProcessingPattern Recognition
B
Bowen Peng
College of Electronic Science and Technology, National University of Defense Technology, Changsha 410073, China
Jie Zhou
Jie Zhou
National University of Defense Technology
Synthetic Aperture RadarAutomatic Target RecognitionDiffusion ModelsComputer Vision
Jingyuan Xia
Jingyuan Xia
National University of Defense Technology
Non-convex optimizationStatistical machine learningImage restoration
W
Wei Yang
College of Electronic Science and Technology, National University of Defense Technology, Changsha 410073, China
T
Tianpeng Liu
College of Electronic Science and Technology, National University of Defense Technology, Changsha 410073, China
Zhen Liu
Zhen Liu
University of Electronic Science and Technology of China
Computer Vision、Computational Photography
L
Li Liu
College of Electronic Science and Technology, National University of Defense Technology, Changsha 410073, China