🤖 AI Summary
Existing discrete wavelet transform (DWT)-based single-image super-resolution methods struggle to model inter-scale dependencies among frequency subbands, leading to artifacts and structural inconsistencies in reconstructed images. To address this, we propose the Multi-level Wavelet Spectral Diffusion Transformer (MW-DiT), which innovatively integrates multi-level DWT, pyramid tokenization, and a dual-decoder architecture to explicitly capture cross-scale spectral correlations within joint spatial-frequency representations—thereby alleviating high-low frequency misalignment. Leveraging diffusion processes as a prior, MW-DiT employs Transformer-based long-range dependency modeling and multi-scale feature co-optimization. Extensive experiments on standard benchmarks (Set5, Set14, Urban100) demonstrate significant improvements in perceptual quality and fidelity: PSNR and SSIM scores are competitive with state-of-the-art methods, while visual results exhibit richer texture details and more natural structural consistency.
📝 Abstract
Discrete Wavelet Transform (DWT) has been widely explored to enhance the performance of image superresolution (SR). Despite some DWT-based methods improving SR by capturing fine-grained frequency signals, most existing approaches neglect the interrelations among multiscale frequency sub-bands, resulting in inconsistencies and unnatural artifacts in the reconstructed images. To address this challenge, we propose a Diffusion Transformer model based on image Wavelet spectra for SR (DTWSR).DTWSR incorporates the superiority of diffusion models and transformers to capture the interrelations among multiscale frequency sub-bands, leading to a more consistence and realistic SR image. Specifically, we use a Multi-level Discrete Wavelet Transform (MDWT) to decompose images into wavelet spectra. A pyramid tokenization method is proposed which embeds the spectra into a sequence of tokens for transformer model, facilitating to capture features from both spatial and frequency domain. A dual-decoder is designed elaborately to handle the distinct variances in lowfrequency (LF) and high-frequency (HF) sub-bands, without omitting their alignment in image generation. Extensive experiments on multiple benchmark datasets demonstrate the effectiveness of our method, with high performance on both perception quality and fidelity.