Score
Designs and implements frequency-domain transforms and filters that select or suppress signal components by spatial frequency and local orientation (frequency-domain processing and orientation-selective filtering). Builds feature-extraction and texture-suppression operators that attenuate repetitive texture interference and enhance curve-like or curvilinear structural responses.
Text-guided image editing often suffers from local detail loss and color distortion due to uniform optimization across the full frequency spectrum. To address this, we propose a frequency-aware image editing framework that introduces, for the first time, a frequency-aware denoising scoring mechanism—enabling spatially localized and band-selective editing. Our method integrates wavelet-based multi-scale decomposition, frequency-domain gradient masking, and triplane representation to achieve cross-dimensional (2D/3D) texture editing with precise frequency control. Within text-guided latent diffusion models, it enables targeted modulation of critical frequency components in user-specified regions. Quantitative evaluation and user studies demonstrate significant improvements: a 32% increase in local detail preservation and a 27% gain in color fidelity—outperforming state-of-the-art approaches.
This work addresses the limitations of existing frequency-domain remote sensing image fusion methods, which rely on fixed filters and exhibit inadequate utilization of frequency information in their denoising strategies, thereby struggling to adapt to complex spectral distributions. To overcome these challenges, the authors propose CGFformer, a novel framework that introduces a K-means clustering-guided adaptive frequency separation mechanism. It further incorporates a dual-stream Transformer with cross-attention modules to jointly perform denoising and detail enhancement in both frequency and spatial domains. A dedicated frequency-spatial fusion mechanism is then employed to improve reconstruction quality. Extensive experiments on multiple remote sensing datasets demonstrate that the proposed method significantly outperforms state-of-the-art approaches, effectively preserving spectral fidelity while enhancing spatial details.
Image inpainting remains challenging for reconstructing complex textures and restoring large occluded regions, suffering from high-frequency detail loss and excessive computational cost. To address these issues, we propose a frequency-domain enhanced Transformer framework. First, we design a wavelet-Gabor fused attention mechanism to explicitly model multi-scale structural patterns. Second, we introduce learnable FFT-based frequency-domain filters to adaptively preserve high-frequency components while suppressing noise. Third, we construct a four-stage encoder-decoder architecture, jointly optimized with a composite loss function that balances global semantic coherence and local detail fidelity. Extensive experiments demonstrate that our method achieves superior performance over state-of-the-art approaches in both quantitative metrics (PSNR/SSIM) and visual quality, with enhanced high-frequency detail preservation and approximately 23% faster inference speed—effectively reconciling restoration accuracy and computational efficiency.
This study addresses the limitations of conventional image enhancement, filtering, and pattern recognition—namely, heavy reliance on manual feature engineering and insufficient real-time performance—by proposing a theory-driven, end-to-end machine learning framework. Methodologically, it is the first to systematically integrate discrete Fourier transform (DFT), Z-transform, and continuous Fourier analysis into deep learning pipelines, synergistically coupling them with convolutional neural networks (CNNs) and classical digital filtering algorithms to enable frequency-domain-guided automated feature extraction and real-time joint signal–image processing. The key contributions include: (i) development of an extensible Python framework; (ii) average PSNR improvement of 3.2 dB in image enhancement and noise suppression tasks; and (iii) 40% acceleration in feature extraction efficiency. This work establishes a novel paradigm for AI-powered real-time computer vision that simultaneously ensures high performance and interpretability.
Single-image reflection removal (SIRR) aims to recover a reflection-free background from a single image corrupted by glass reflections, yet remains an open challenge due to the high variability in reflection intensity, morphology, and spatial distribution. This paper proposes a frequency-spatial collaborative modeling framework: it introduces, for the first time, global frequency-domain priors into SIRR via an FFT Transformer that captures periodic spectral patterns inherent to reflections; further, it integrates hierarchical Transformers with a U-Net–based multi-scale encoder-decoder architecture to enable frequency-domain disentanglement and spatially adaptive separation of reflection and background components. The method achieves state-of-the-art performance on three benchmarks—SIR2, RealBlur-J, and Reflections-Real—yielding significant PSNR and SSIM improvements. Notably, it demonstrates superior robustness under challenging conditions, including strong reflections, large reflection coverage, and non-uniform reflection distributions.
This work addresses the limitation of conventional Fourier-encoded implicit neural representations (INRs), which employ globally fixed frequencies and struggle to effectively capture spatially varying local spectra, leading to slow convergence of high-frequency details. To overcome this, the authors propose an adaptive local frequency filtering approach that introduces a spatially varying parameter α(x) to dynamically modulate Fourier components, enabling smooth, position-dependent transitions among low-pass, band-pass, and high-pass responses. This method establishes the first spatially adaptive frequency modulation mechanism within Fourier-encoded INRs and leverages neural tangent kernel (NTK) theory to reveal its spectral reshaping effect on the effective kernel, facilitating interpretable visualization of frequency preferences. Experiments demonstrate that the proposed approach significantly improves reconstruction quality and accelerates optimization across 2D image fitting, 3D shape representation, and sparse data reconstruction tasks, outperforming fixed-frequency baselines.
To address semantic information loss and low feature reliability in noisy image compression, this paper proposes a persistent homology-guided frequency-domain filtering method. First, discrete Fourier transform (DFT) is applied to input images; then, persistent homology analysis identifies topologically significant frequency components critical for classification tasks; finally, structure-preserving filtering is performed in the frequency domain to achieve noise-robust compression and reconstruction. This work is the first to embed topological data analysis into the image frequency-domain compression pipeline, explicitly preserving semantic-relevant topological structures. Evaluated on CNN-based downstream binary classification tasks, the method matches JPEG’s performance across six compression quality metrics—including PSNR and SSIM—while significantly enhancing feature discriminability and classification accuracy on noisy images. The approach establishes a novel paradigm for robust image representation learning.
This work addresses the limitation of existing object detection methods that predominantly operate on sRGB images while overlooking the superior noise characteristics and richer information inherent in RAW sensor data, particularly under low-light and adverse weather conditions. To bridge this gap, we propose FreqAdapt—a lightweight frequency-domain adaptive enhancement module—that, for the first time, decomposes and transfers in-camera ISP operations into the Fourier domain based on their physical properties, enabling principled domain separation. By jointly modeling magnitude spectra, phase spectra, and RAW features through a learnable fusion mechanism and a frequency-domain encoder, our approach achieves global context-guided adaptive enhancement. Extensive experiments demonstrate that FreqAdapt significantly improves detection performance across diverse lighting and weather conditions, offering a lightweight, efficient, and physically interpretable solution that seamlessly integrates into existing detection frameworks.
This work addresses the challenges of small object detection, which suffers from feature sparsity and the loss of high-frequency details inherent in spatial-domain approaches. To overcome these limitations, the authors propose a novel spectral-domain feature learning paradigm featuring a lightweight, plug-and-play Decomposition-Enhancement-Reconstruction (DER) operator. This operator integrates frequency-aware modulation into the backbone, neck, and detection head through three key components: a Wavelet Difference Gate (WDG), a Log-Gabor Enhancer (LGE), and a Frequency-Domain-driven Detection Head (FDHead). The resulting framework is architecture-agnostic, compatible with both CNNs and Transformers, and achieves state-of-the-art performance on benchmarks such as VisDrone2019 and UAVDT—significantly outperforming YOLOv11 at one-sixth the parameter count while enabling more precise localization of small objects.