Score
Designs and implements algorithms and model components that extract, represent, and refine features across multiple spatial or temporal scales—e.g., multi-resolution filters, scale-space transforms, and multi-scale refinement layers—to produce scale-aware representations. Analyzes and optimizes cross-scale interactions (aggregation, suppression of redundant information, and multi-scale inference) to improve discrimination, scale invariance, and robustness for image or signal processing tasks.
Existing pan-sharpening methods are predominantly evaluated at low resolutions and struggle to generalize to real-world high-resolution cross-scale scenarios. To address this limitation, this work introduces the PanScale dataset and the PanScale-Bench benchmark, along with ScaleFormer—the first general-purpose framework specifically designed for cross-scale pan-sharpening. ScaleFormer incorporates a Scale-Aware Patchify module and rotary positional encoding to parse images into variable-length patch sequences, enabling effective extrapolation to unseen scales. Extensive experiments demonstrate that ScaleFormer consistently outperforms state-of-the-art methods in both fusion quality and cross-scale generalization capability.
Existing single-image deraining methods are largely confined to single-domain (spatial-only) and single-scale modeling, limiting their ability to jointly exploit multi-scale external and internal features and inadequately capturing complex real-world rain streaks. To address these limitations, we propose a dual-domain collaborative multi-scale representation framework. Our approach introduces, for the first time, a parallel multi-scale representation mechanism operating simultaneously in both spatial and frequency domains. We design a hierarchical modulation and fusion module (MPSRM) and a frequency-domain scale-mixing module (FDSM) to enable cross-domain feature coupling and progressive spatial refinement. By integrating Fourier-based frequency-domain modeling, multi-scale feature pyramids, and deep learning, our method effectively balances local detail preservation and global structural dependencies. Extensive experiments demonstrate state-of-the-art performance across six benchmark datasets, with significant improvements in restoration accuracy and structural fidelity—particularly under challenging, realistic rain conditions.
This work addresses the challenge that existing vision foundation models struggle to effectively capture cross-scale spatial relationships among multimodal satellite images with varying spatial resolutions. To this end, the authors propose Scale-ALiBi, a novel mechanism that incorporates a linear spatial bias—derived from ground sampling distance—into Transformer attention. This is integrated within a joint representation learning framework combining triplet contrastive learning and reconstruction objectives for optical and synthetic aperture radar (SAR) imagery. The key contributions include the first extension of ALiBi to multiscale remote sensing scenarios, a tailored attention mechanism capable of modeling spatial relationships across image patches at different scales, and the creation of the first aligned multimodal, multiscale satellite image dataset. The proposed method achieves significant performance gains on GEO-Bench, and the dataset has been publicly released.
To address the conflict between object-scale diversity in satellite imagery and constrained computational resources, this paper proposes a scale-adaptive recognition framework. Methodologically, it introduces large language models (LLMs) for the first time to perform semantic-driven scale concept reasoning; designs an active high-resolution (HR) image sampling strategy based on model disagreement; and establishes an HR–low-resolution (LR) cross-resolution knowledge distillation mechanism. The core innovation lies in jointly modeling LLM priors, model uncertainty, and multi-scale representations to enable budget-aware, on-demand HR invocation. Under stringent resource constraints, the framework achieves a 26.3% improvement in recognition accuracy over the full-HR baseline while reducing HR image usage by 76.3%, significantly enhancing deployment efficiency and cost-effectiveness.
To address global spatial layout inconsistency in high-resolution panoramic image generation, this paper proposes a multi-scale diffusion framework that jointly models geometry and semantics via cross-resolution structural prior transfer. The core contribution is a novel multi-scale gradient-guided mechanism, which explicitly injects low-resolution layout constraints into the high-resolution generation process, integrated with multi-scale feature distillation, cross-resolution gradient backpropagation, and structure-aware loss. The method preserves seamless structural coherence and natural transitions even at 16K resolution. Quantitative experiments demonstrate a 23.6% reduction in Fréchet Inception Distance (FID) and a 31.4% improvement in layout consistency metrics, significantly outperforming state-of-the-art approaches.
This work addresses the challenge of efficient high-resolution region selection under limited observation budgets to enhance remote sensing understanding. It formulates cross-scale remote sensing interpretation as a unified cost-aware optimization problem and proposes a joint framework that integrates fine-grained importance sampling with cross-patch contextual modeling. To support this research, the authors introduce GL-10M, the first large-scale, multi-resolution remote sensing benchmark dataset with tens of millions of aligned image pairs. Experimental results demonstrate that the proposed method significantly outperforms existing approaches in both recognition and retrieval tasks, achieving superior performance at the same observational cost and effectively balancing accuracy and computational overhead.
This work addresses the challenge of multi-scale time series forecasting, where modeling cross-scale patterns and maintaining computational efficiency are often at odds. The authors propose AWGformer, a novel architecture that integrates adaptive wavelet decomposition (AWDM) with a frequency-aware multi-head attention mechanism (FAMA), enabling dynamic selection of wavelet bases guided by signal characteristics and facilitating interaction among multi-band features. Coupled with cross-scale feature fusion (CSFF) and a hierarchical prediction network (HPN), AWGformer forms an end-to-end trainable framework supported by theoretical convergence guarantees. Extensive experiments on multiple benchmark datasets demonstrate that AWGformer significantly outperforms existing methods, achieving superior prediction accuracy and robustness—particularly in multi-scale and non-stationary scenarios.
This work addresses the limitations of existing frequency-domain remote sensing image fusion methods, which rely on fixed filters and exhibit inadequate utilization of frequency information in their denoising strategies, thereby struggling to adapt to complex spectral distributions. To overcome these challenges, the authors propose CGFformer, a novel framework that introduces a K-means clustering-guided adaptive frequency separation mechanism. It further incorporates a dual-stream Transformer with cross-attention modules to jointly perform denoising and detail enhancement in both frequency and spatial domains. A dedicated frequency-spatial fusion mechanism is then employed to improve reconstruction quality. Extensive experiments on multiple remote sensing datasets demonstrate that the proposed method significantly outperforms state-of-the-art approaches, effectively preserving spectral fidelity while enhancing spatial details.
Existing deep perceptual similarity models are predominantly limited to single-scale analysis, overlooking the critical role of multi-scale structural information in image quality assessment. This work addresses this limitation by explicitly modeling spatial scale as an independent variable and introduces a minimalist multi-scale framework: DeepSSIM is computed independently at each level of a feature pyramid, and the resulting scores are aggregated via lightweight, learnable global weights. Despite introducing minimal computational overhead, the proposed method significantly outperforms single-scale baselines across multiple benchmark datasets, thereby demonstrating both the effectiveness and necessity of explicit multi-scale modeling in deep perceptual similarity estimation.