🤖 AI Summary
The mechanistic differences between 2D convolution and self-attention—particularly in high-frequency filtering capability and shape-bias propensity—lack a unified spectral explanation. Method: This work establishes, for the first time, a unified frequency-domain modeling framework grounded in graph spectral theory to quantitatively characterize their intrinsic frequency responses. It introduces a node-connectivity-driven spectral modulation mechanism and proposes the Spectral Adaptive Modulation (SPAM) mixer, enabling dynamic rescaling and fusion of multi-scale spectral components. Contribution/Results: Based on SPAM, we design SPANetV2—a novel vision backbone—that achieves state-of-the-art performance across ImageNet-1K classification, COCO object detection, and ADE20K semantic segmentation. These results empirically validate that spectral adaptive modeling significantly enhances visual representation learning.
📝 Abstract
Recent studies have shown that 2D convolution and self-attention exhibit distinct spectral behaviors, and optimizing their spectral properties can enhance vision model performance. However, theoretical analyses remain limited in explaining why 2D convolution is more effective in high-pass filtering than self-attention and why larger kernels favor shape bias, akin to self-attention. In this paper, we employ graph spectral analysis to theoretically simulate and compare the frequency responses of 2D convolution and self-attention within a unified framework. Our results corroborate previous empirical findings and reveal that node connectivity, modulated by window size, is a key factor in shaping spectral functions. Leveraging this insight, we introduce a extit{spectral-adaptive modulation} (SPAM) mixer, which processes visual features in a spectral-adaptive manner using multi-scale convolutional kernels and a spectral re-scaling mechanism to refine spectral components. Based on SPAM, we develop SPANetV2 as a novel vision backbone. Extensive experiments demonstrate that SPANetV2 outperforms state-of-the-art models across multiple vision tasks, including ImageNet-1K classification, COCO object detection, and ADE20K semantic segmentation.