Score
Design and implement differentiable wavelet-based encodings that map sequential signals into spectral representations. These encodings preserve temporal localization, reveal spectral fingerprints of generation, and support gradient-based analysis or downstream modeling.
Traditional position encodings are signal-agnostic and struggle to model multiscale non-stationary dynamics in time series. To address this, we propose Dynamic Wavelet Position Encoding (DWPE), the first position encoding framework for Transformers that incorporates the Discrete Wavelet Transform (DWT) to generate signal-aware, dynamically adaptive positional embeddings leveraging input-specific multiscale time-frequency features. DWPE overcomes the representational limitations of fixed sinusoidal encodings in time-frequency domains, substantially enhancing modeling capacity for complex temporal structures. Extensive experiments across 10 benchmark datasets demonstrate that DWPE achieves an average relative performance gain of 9.1% on biomedical signal tasks while maintaining computational efficiency. Our core contributions are: (i) the first DWT-based, signal-driven position encoding framework; (ii) a multiscale, dynamic, and learnable positional representation; and (iii) a favorable trade-off between accuracy and inference efficiency.
This work addresses the limitations of traditional signal modeling, which relies on discrete sampling and struggles to unify multimodal continuous signals or support analytical operations. Viewing implicit neural representations (INRs) through the lens of signal processing, the study models images, audio, 3D geometry, and other modalities as continuous coordinate-based functions, realized via differentiable neural networks. The authors systematically analyze the spectral properties, sampling theory, and multiscale mechanisms underlying INRs and propose structured representation strategies—integrating periodicity, localization, adaptive activation functions, and hash-grid encodings—to reshape the approximation space for enhanced spatial adaptivity and computational efficiency. The resulting framework demonstrates superior performance in inverse problems such as medical and radar imaging, signal compression, and 3D scene reconstruction, advancing both the theoretical understanding and practical applicability of INRs as learnable continuous signal models.
This work addresses the limitations of traditional spike coding methods, which often rely on probabilistic models and lack compatibility with mainstream signal processing theory, making it difficult to define bandwidth and guarantee reconstruction fidelity. The authors propose a novel spike coding framework grounded in causal temporal wavelets, introducing for the first time a bandwidth-controllable wavelet representation into spike coding to achieve sparse, localized, and reconstructable spiking representations of temporal signals. By integrating causal bandpass wavelet frames, spike-based quantization, and temporal discretization, the method provides rigorous theoretical bounds on reconstruction error. Experimental results demonstrate that, on ECG and audio signals, the proposed approach achieves normalized root-mean-square errors comparable to those of the continuous wavelet transform while remaining amenable to deployment on neuromorphic hardware.
Conventional short-time Fourier transform (STFT) performance critically depends on hand-crafted or heuristic hyperparameters—such as window length, hop size, and overlap ratio—while existing discrete grid-search optimization methods suffer from high computational cost and lack task-specific adaptability. Method: We propose the first fully differentiable STFT framework, modeling core STFT operations—including window function selection, overlap ratio, and discrete Fourier transform—as end-to-end trainable, differentiable signal processing modules. This enables joint optimization with downstream neural networks via gradient-based learning. Contribution/Results: By transcending the limitations of discrete parameter spaces, our approach achieves gradient-driven, adaptive time-frequency representation learning. Extensive experiments on synthetic and real-world signals demonstrate substantial improvements in time-frequency resolution and consistent performance gains across downstream tasks—including classification and denoising—validating both the effectiveness and generalizability of differentiable time-frequency analysis.
This study addresses the challenge of efficient signal representation in spiking neural networks by proposing a wavelet transform method grounded in scale-space theory. Leveraging the scale-covariant properties of leaky integrate-and-fire (LIF) neurons, the authors construct discrete mother wavelets to approximate continuous wavelets, thereby establishing— for the first time—theoretical connections between wavelet transforms and spiking neural networks. The proposed approach enables multiscale signal representation directly in the spiking domain, with reconstruction experiments confirming its feasibility. This work opens a new avenue for low-power neuromorphic signal processing, while also highlighting that current approximation errors remain amenable to further optimization.
Existing diffusion models typically employ pointwise reconstruction losses that are insensitive to signal spectra and multiscale structures, often yielding samples with imbalanced frequency content and insufficient structural detail. To address this limitation, this work proposes a lightweight, plug-and-play spectral regularization framework that introduces differentiable Fourier- and wavelet-domain losses as soft inductive biases during standard training. The approach requires no modifications to the diffusion process, model architecture, or sampling procedure, and is compatible with mainstream paradigms such as DDPM, DDIM, and EDM. Evaluated on high-resolution unconditional generation tasks for both images and audio, the method consistently enhances sample quality—particularly in terms of spectral fidelity and multiscale coherence—while incurring negligible computational overhead.
This study addresses the spectral bias in MLP-based implicit neural representations, which leads to inadequate reconstruction of high-frequency details. To overcome this limitation, we propose a spatial-frequency-aware framework that integrates MLPs with Kolmogorov–Arnold Networks (KANs) to achieve complementary frequency modeling. By incorporating the discrete wavelet transform, the input signal is decomposed into distinct frequency bands, and a band-separation regularization term is introduced to guide each branch toward specialized yet synergistic reconstruction. Extensive experiments demonstrate that the proposed method significantly improves reconstruction fidelity across multidimensional signal tasks, validating its cross-modal applicability and strong generalization capability for efficient and accurate signal representation.
This work addresses the spectral bias inherent in implicit neural representations, where periodic activations tend to overfit noise while compactly supported local activations struggle to capture low-frequency signals. The authors propose a physics-inspired spectral gating mechanism that models neuron activation as the steady-state response of a forced damped harmonic oscillator. By jointly optimizing the oscillator parameters alongside network weights, the method adaptively tunes spectral selectivity without requiring explicit regularization or task-specific hyperparameter tuning. This naturally yields a coarse-to-fine learning process, achieving state-of-the-art or competitive performance across multiple benchmarks while simultaneously preserving fine details and providing effective implicit regularization.
This work addresses the challenges of plasticity loss and catastrophic forgetting in continual learning, which stem from spectral bias and uncontrolled updates inherent in conventional activation functions. To mitigate these issues, the authors propose a learnable wavelet-based activation function that decomposes activations into high- and low-frequency components, thereby alleviating spectral bias. A hybrid wavelet architecture enables efficient L² approximation, while a decoupled learning rate mechanism restores plasticity in high-frequency information. Furthermore, a loss-driven wavelet injection strategy, combined with regularization constraints, enhances learning on new tasks without compromising previously acquired knowledge. The proposed method achieves state-of-the-art performance across multiple continual learning benchmarks, significantly improving overall trainability and generalization capability.
This work addresses the limitations of conventional spectral neural operators, which rely on fixed global bases and struggle to capture spatial heterogeneity and multiscale dynamics. The authors propose the Adaptive Basis Learning (ABLE) framework—the first approach to enable end-to-end learning of spectral bases. ABLE constructs data-driven, spatially adaptive Parseval frames that preserve invertibility and maintain O(N log N) computational complexity, effectively shifting representational capacity from spectral coefficients to the basis functions themselves. The framework leverages an FFT-based efficient implementation, incorporates learnable auxiliary density functions, and can seamlessly replace spectral layers in existing neural operators. Experiments demonstrate that ABLE significantly outperforms strong baselines across multiple PDE benchmarks, particularly excelling in scenarios with sharp gradients and multiscale features. Moreover, when integrated as a plug-in module into models such as U-FNO and HPM, it consistently enhances performance.