adaptive spectral-spatial fusion

Design and implement multi-scale, CNN-based fusion models that combine spectral and spatial representations of multiband data across pyramid levels. Build reliability-driven fusion mechanisms that estimate per-branch confidence and apply adaptive weighting — including rotation-equivariant consistency weighting — to suppress noisy cues, reinforce informative features, and maintain robust fusion under rotations and scale changes.

adaptivespectral-spatialfusion

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.37
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work proposes SSA, a universal framework for hyperspectral image fusion that overcomes the limitations of existing deep learning methods, which are typically constrained by fixed spectral bands and spatial scales and thus struggle to generalize across different sensors. SSA is the first model to simultaneously achieve spectral band agnosticism and fusion scale invariance. It integrates a Matryoshka Kernel with an Implicit Neural Representation (INR) backbone, enabling adaptive processing of arbitrary numbers of spectral channels and spatial resolutions through continuous signal modeling. Experimental results demonstrate that a single SSA model achieves state-of-the-art performance across multiple datasets and exhibits exceptional generalization capabilities on unseen sensors and scales.

fusion-scale agnosticismhyperspectral image fusionsensor generalization

A General Adaptive Dual-level Weighting Mechanism for Remote Sensing Pansharpening

Mar 17, 2025
JH
Jie Huang
🏛️ University of Electronic Science and Technology of China | Chinese Academy of Sciences | University of Chinese Academy of Sciences

To address insufficient modeling of feature heterogeneity and difficulty in suppressing intra-channel and inter-layer redundancy in remote sensing pan-sharpening, this paper proposes a Covariance-Driven Adaptive Dual-level Weighting Mechanism (ADWM). ADWM introduces a novel covariance-aware dual-level weighting paradigm: Intra-Feature Weighting (IFW) suppresses intra-layer redundancy via channel-wise covariance estimation, while Cross-Layer Weighting (CFW) dynamically modulates inter-layer feature contributions using cross-layer covariance. It is the first method to deeply integrate covariance matrix modeling with nonlinear weight mapping, enabling adaptive feature refinement across channels and layers. ADWM is plug-and-play and compatible with mainstream network architectures. Extensive experiments on multiple remote sensing datasets demonstrate significant improvements over state-of-the-art methods, with consistent gains in PSNR (+0.12–0.38 dB) and SSIM (+0.003–0.012). The open-sourced implementation has been widely adopted in the community.

Addresses feature heterogeneity and redundancy in remote sensing pansharpening.Introduces adaptive dual-level weighting mechanism (ADWM) to enhance deep-learning methods.Proposes Correlation-Aware Covariance Weighting (CACW) for feature adjustment.

This work addresses the challenge that existing medical image fusion methods struggle to simultaneously preserve global statistical similarity—such as correlation coefficient (CC) and mutual information (MI)—and local structural fidelity. To this end, we propose a reliability-weighted dual-expert fusion framework that integrates a spatial-domain expert for modeling global context and a wavelet-frequency-domain expert for capturing fine local details. The framework employs a dense reliability map to guide adaptive modality weighting and incorporates a soft gradient arbitration mechanism. Furthermore, it adopts a residual-mean fusion paradigm and a joint CC-MI loss function. Extensive experiments on CT-MRI, PET-MRI, and SPECT-MRI datasets demonstrate that our method significantly outperforms state-of-the-art approaches such as AdaFuse and ASFE-Fusion, achieving superior local structural preservation without compromising global statistical consistency.

correlation coefficientglobal statistical similaritylocal structural fidelity

Existing multimodal image fusion (MMIF) methods overlook architectural constraints—such as normalization schemes and convolutional kernel design—on feature representation, particularly where batch normalization attenuates sparse, salient features. Addressing fundamental disparities between natural and medical image fusion, this work proposes: (1) a hybrid normalization strategy combining instance and group normalization to enhance both feature independence and intrinsic correlation; (2) large-kernel convolutions to expand receptive fields; and (3) a multi-path adaptive fusion module for cross-scale feature recalibration. The resulting end-to-end network achieves state-of-the-art performance across diverse fusion benchmarks. Moreover, it significantly improves downstream task accuracy and fine-detail preservation in medical diagnosis and remote sensing analysis, demonstrating superior generalizability and fidelity.

Introduces adaptive fusion module for dynamic multi-scale feature calibrationProposes hybrid normalization to preserve sparse features and enhance correlationsReevaluates UNet's normalization and convolution for multimodal image fusion

This work addresses the limitations of existing approaches in remote sensing image fusion, where conventional 2D methods often introduce spectral distortion, while standard 3D convolutions incur high computational costs and struggle to capture fine local details. To overcome these challenges, the authors propose Adaptive 3D Convolution (Ada3D), which, for the first time, enables voxel-level dynamic kernel generation by adaptively producing convolutional weights and biases conditioned on both spatial and spectral content. Computational complexity is further reduced through grouped convolutions. The proposed method effectively integrates multi-source information, significantly mitigating spectral distortion across five remote sensing datasets and achieving state-of-the-art fusion performance and image quality.

3D convolutionadaptive kernelscomputational complexity

Latest Papers

What's happening recently
View more

This work addresses the challenges of insufficient multi-scale feature interaction and poor representation robustness in remote sensing urban scene classification, which arise from large intra-class variations and high inter-class similarities. To this end, the authors propose a dual-backbone multi-scale fusion architecture that enhances cross-scale feature interaction through a residual feature propagation mechanism and incorporates a spatial attention module to emphasize discriminative regions. A two-stage freeze-and-fine-tune training strategy is further introduced to improve model generalization. Evaluated on the AID dataset, the proposed method achieves an average accuracy of 97.46% ± 0.14%, and ablation studies confirm the effectiveness of each component.

feature representationinter-class similarityintra-class variability

This work addresses the longstanding challenge in multispectral and hyperspectral image fusion—balancing spatial detail enhancement with spectral fidelity—where existing methods struggle with cross-scale interaction and joint spatial-spectral modeling. The authors propose CoFusion, a novel framework featuring a three-level multiscale pyramid architecture. Each level employs a dual-branch design: SpaCAM captures multiscale contextual information through a spatial coordinate-aware mixing mechanism, while SpeCAM enhances spectral representation by integrating frequency-domain decomposition with coordinate attention. A spatial-spectral cross-fusion module (SSCFM) further enables dynamic cross-modal alignment and complementary feature integration. CoFusion is the first to jointly model cross-scale and cross-modal dependencies, achieving state-of-the-art performance across multiple benchmark datasets with superior spatial reconstruction quality and spectral fidelity.

cross-scale interactionsMultispectral and Hyperspectral Image Fusionspatial detail enhancement

This work addresses the limitation of existing RGB-infrared object detection methods, which discard spectral statistical information during cross-modal fusion, thereby precluding reliable assessment of fusion quality. To overcome this, the study introduces a novel framework that explicitly extracts a parameter-free 7-dimensional spectral reliability descriptor and reuses it throughout subsequent computations. Specifically, it proposes Spectral Reliability Fusion (SRF) and Reliability-Conditioned Expert Routing (RCER) mechanisms to enable adaptive gated fusion and sparse expert selection. Evaluated on the DroneVehicle dataset under six synthetic degradation scenarios, the method achieves an average retention rate of 95.0% and improves mAP50 by 5.2 and 5.3 points in daytime and nighttime settings, respectively, significantly outperforming content-only baseline approaches.

adaptive fusioncross-modal fusionexpert routing

Existing multispectral object detection methods are limited by discrete spectral modeling, unstable cross-scale spectral-spatial feature fusion, and a lack of rotational equivariance for arbitrarily oriented objects. This work proposes FressDet, the first fully rotation-equivariant framework for multispectral detection. It achieves continuous spectral modeling through implicit resampling that preserves spectral ordering, introduces a rotation-equivariant consistency weighting mechanism for robust multiscale feature fusion, and incorporates an orientation-aware detection head without parameter duplication. Evaluated on three public benchmarks, FressDet attains state-of-the-art performance with 93% fewer parameters, significantly enhancing robustness to rotational perturbations and generalization capability.

feature pyramidmultispectral object detectionoriented objects

Hot Scholars

BB

Benedikt Blumenstiel

Research Software Engineer, IBM Research
Computer VisionFoundation ModelsEarth Observation
FB

Fanglin Bao

Westlake University
AI physicsquantum optical sensingCasimir physics
LZ

Lefei Zhang

School of Computer Science, Wuhan University
Pattern RecognitionMachine LearningImage ProcessingRemote Sensing
RG

Rémi Giraud

Associate Professor - Bordeaux INP / Univ. Bordeaux
Image Processing
KS

Konrad Schindler

Professor of Photogrammetry and Remote Sensing, ETH Zurich
PhotogrammetryRemote SensingImage AnalysisComputer Vision