Score
Design and implement attention modules that compute and apply per-spectral-component weights to feature channels or spectral coefficients (e.g., frequency bands, subcarriers, or graph spectral modes), including multi-hop and hop-aware variants that aggregate information across spectral neighborhoods. Build mechanisms that drive re-weighting by statistics such as temporal or variance measures to suppress noisy or redundant channels, denoise spectra, adapt filters to spectral variability, and analyze resulting changes in channel-wise feature responses and learned filters.
The mechanistic differences between 2D convolution and self-attention—particularly in high-frequency filtering capability and shape-bias propensity—lack a unified spectral explanation. Method: This work establishes, for the first time, a unified frequency-domain modeling framework grounded in graph spectral theory to quantitatively characterize their intrinsic frequency responses. It introduces a node-connectivity-driven spectral modulation mechanism and proposes the Spectral Adaptive Modulation (SPAM) mixer, enabling dynamic rescaling and fusion of multi-scale spectral components. Contribution/Results: Based on SPAM, we design SPANetV2—a novel vision backbone—that achieves state-of-the-art performance across ImageNet-1K classification, COCO object detection, and ADE20K semantic segmentation. These results empirically validate that spectral adaptive modeling significantly enhances visual representation learning.
In multi-span optical networks, dynamic power spectral modeling faces challenges in brown-field deployments due to scarce operational data, high component-specific modeling costs, and poor generalization across spans. Method: This paper proposes a component-specific multi-decoder attention framework, introducing the first component-aware attention mechanism that decouples spectral response modeling per optical device, enabling cross-span generalization under few-shot conditions. Contribution/Results: Compared to single-decoder baselines, the method reduces training data requirements by 70% and improves model deployment efficiency by 3×. It supports rapid adaptation to complex topologies and has been validated on a real metropolitan-area network, achieving significantly higher accuracy in power spectral evolution prediction than state-of-the-art data-driven models.
This study critically examines whether spectral graph neural networks (spectral GNNs) genuinely leverage spectral properties of graphs to achieve performance gains in node classification tasks. Through theoretical analysis grounded in graph signal processing, examination of Vandermonde systems, and demonstration of equivalence with message-passing neural networks (MPNNs), the work reveals that existing spectral GNNs—such as MagNet and HoloNet—fail to correctly implement the graph Fourier transform. Their reported performance advantages are instead attributable to implicit MPNN-like architectures or implementation artifacts. Rigorous reimplementation and ablation studies confirm that when spectral GNNs strictly adhere to spectral theory, their performance deteriorates significantly, thereby challenging the theoretical foundation underpinning their use in node classification.
This work addresses the challenges in multi-behavior recommendation, where intra-behavior representation entanglement and inter-behavior reliability heterogeneity often introduce noise and deviate from the target intent. To this end, the authors propose SpectraMB, a novel model that introduces dynamic feature-level spectral filtering into multi-behavior recommendation for the first time. Operating in the frequency domain, it adaptively denoises representations without requiring predefined frequency bands. SpectraMB further incorporates a global context-aware attention mechanism that leverages the purified global representations as anchors to dynamically assess the reliability of each behavior, enabling reliability-aware fusion. Integrated with graph neural networks and a residual global backbone, SpectraMB significantly outperforms state-of-the-art methods on three real-world datasets, demonstrating enhanced recommendation accuracy and robustness to noise.
This work proposes FITMM, a novel framework that addresses the limitations of existing spatial-domain approaches to multimodal recommendation, which often neglect frequency-domain structures and suffer from modality misalignment and redundancy. FITMM is the first to introduce the information bottleneck principle into frequency-domain multimodal recommendation. It constructs item representations via graph augmentation and performs orthogonal decomposition of each modality in the frequency domain to yield lightweight intra-band components. A task-adaptive gating mechanism fuses band-specific information, while intra-band independent modeling and cross-modal spectral consistency constraints enable adaptive band selection and redundancy suppression. Extensive experiments on three real-world datasets demonstrate that FITMM significantly outperforms state-of-the-art baselines, validating the effectiveness and generalizability of frequency-domain modeling for multimodal recommendation.
This study addresses the challenges of opaque information routing and inaccurate uncertainty estimation in Transformer attention mechanisms. Grounded in spectral geometry and operator theory, this work models attention heads as functional mappings within a Hilbert space and derives token difference operators to reveal intrinsic routing structures. By constructing a token difference spectrum under probabilistic geometry, it effectively decouples the sink effect from genuine routing capacity, providing a unified explanation for attention collapse. Ultimately, this project establishes a systematic analytical framework for attention mechanisms and introduces a novel uncertainty estimator that significantly improves both estimation accuracy and model performance on long-context tasks.
This work investigates how attention mechanisms enhance the ability of sequence models to recover latent signals from noisy observations in the high-dimensional limit. Leveraging random matrix theory, the authors analyze the spectral properties of the sample covariance matrix after attention-weighted pooling, characterizing its eigenvalue distribution, outlier eigenvalues, and the alignment between eigenvectors and the true signal. The analysis reveals two BBP-type phase transitions governing signal recovery and demonstrates that optimal attention weights correspond to the leading eigenvector of a location-dependent kernel matrix. Moreover, a specific causal self-attention mechanism is shown to yield deterministic harmonic weights that outperform simple averaging. Combining free multiplicative convolution, high-dimensional spectral analysis, and Gaussian mixture embedding models, the proposed framework exhibits excellent agreement with finite-dimensional experiments, quantitatively elucidating the advantage of attention in boosting signal-to-noise ratio and recovery performance.
This study addresses the challenge of classifying neurodegenerative diseases from electroencephalography (EEG) signals, which is hindered by their inherent high noise levels and low spatial resolution, causing existing deep learning models to struggle in distinguishing between healthy and affected individuals or among different disease types. To overcome this, the authors propose a spectral prior–based feature construction approach that transforms raw EEG data into highly discriminative frequency- and time-frequency–domain representations, substantially enhancing class separability. Experiments on three resting-state and one task-based public EEG datasets demonstrate that this method enables conventional machine learning models to achieve performance comparable to or even surpassing state-of-the-art deep learning approaches. Furthermore, the work reveals a fundamental limitation of attention mechanisms in stably capturing neural activity patterns, showing negligible improvement even when combined with frequency-selected inputs.
Existing Fourier transform–based parameter-efficient fine-tuning methods assume that weight updates are spectrally sparse; however, empirical observations reveal that their spectral energy is in fact uniformly distributed, limiting performance. This work proposes S2FT, which first uncovers the inherently non-sparse nature of weight-update spectra and introduces an invertible transformation based on row–column permutation to map weight changes into a latent space exhibiting local smoothness. This structural prior induces sparsity in the transformed spectral domain. By integrating this spectral prior with a nearest-neighbor search strategy, S2FT achieves superior fine-tuning performance while updating only 0.08% of the model parameters—outperforming current state-of-the-art approaches.
Graph denoising poses a key challenge in graph learning and diffusion models, as existing linear attention mechanisms struggle to adapt to diverse graph spectral structures. This work addresses this limitation by introducing a spectral attention mechanism grounded in spectral graph theory, which is theoretically shown to outperform linear attention. Building upon this insight, the authors develop a permutation-equivariant Graph Convolutional Attention (GCA) architecture that achieves both theoretical optimality and computational efficiency without requiring explicit eigendecomposition or spectral computations. By integrating spectral graph theory, graph convolutions, stochastic block model analysis, and PEARL positional encoding, the proposed method significantly enhances denoising and diffusion performance on both synthetic and real-world datasets, with gains strongly correlated to spectral diversity. When incorporated into DiGress, it enables faster inference while preserving generation quality.