Spectral-Adaptive Modulation Networks for Visual Perception

📅 2025-03-31
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
The mechanistic differences between 2D convolution and self-attention—particularly in high-frequency filtering capability and shape-bias propensity—lack a unified spectral explanation. Method: This work establishes, for the first time, a unified frequency-domain modeling framework grounded in graph spectral theory to quantitatively characterize their intrinsic frequency responses. It introduces a node-connectivity-driven spectral modulation mechanism and proposes the Spectral Adaptive Modulation (SPAM) mixer, enabling dynamic rescaling and fusion of multi-scale spectral components. Contribution/Results: Based on SPAM, we design SPANetV2—a novel vision backbone—that achieves state-of-the-art performance across ImageNet-1K classification, COCO object detection, and ADE20K semantic segmentation. These results empirically validate that spectral adaptive modeling significantly enhances visual representation learning.

Technology Category

Computer Vision: Multi-modal VisionMachine Learning: Multimodal LearningNatural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Graph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphsSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Recent studies have shown that 2D convolution and self-attention exhibit distinct spectral behaviors, and optimizing their spectral properties can enhance vision model performance. However, theoretical analyses remain limited in explaining why 2D convolution is more effective in high-pass filtering than self-attention and why larger kernels favor shape bias, akin to self-attention. In this paper, we employ graph spectral analysis to theoretically simulate and compare the frequency responses of 2D convolution and self-attention within a unified framework. Our results corroborate previous empirical findings and reveal that node connectivity, modulated by window size, is a key factor in shaping spectral functions. Leveraging this insight, we introduce a extit{spectral-adaptive modulation} (SPAM) mixer, which processes visual features in a spectral-adaptive manner using multi-scale convolutional kernels and a spectral re-scaling mechanism to refine spectral components. Based on SPAM, we develop SPANetV2 as a novel vision backbone. Extensive experiments demonstrate that SPANetV2 outperforms state-of-the-art models across multiple vision tasks, including ImageNet-1K classification, COCO object detection, and ADE20K semantic segmentation.
Problem

Research questions and friction points this paper is trying to address.

Analyzes spectral behaviors of 2D convolution vs self-attention
Proposes SPAM mixer for spectral-adaptive visual feature processing
Introduces SPANetV2 backbone for superior vision task performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Graph spectral analysis for comparing frequency responses
Spectral-adaptive modulation mixer with multi-scale kernels
SPANetV2 backbone for superior vision task performance
💼 Related Jobs
No related jobs found.
Guhnoo Yun
Guhnoo Yun
Korea University, Korea Institute of Science and Technology
Computer VisionDeep LearningVideo Understanding3D Vision
J
Juhan Yoo
Dong-A University in Busan, Korea
K
Kijung Kim
Korea University (KU) in Seoul, Korea, and Korea Institute of Science and Technology (KIST) in Seoul, Korea
J
Jeongho Lee
Korea University (KU) in Seoul, Korea, and Korea Institute of Science and Technology (KIST) in Seoul, Korea
Paul Hongsuck Seo
Paul Hongsuck Seo
Korea University
Multimodal Interactive IntelligenceVisionSpeech and Language Understanding
D
Dong Hwan Kim
Korea University (KU) in Seoul, Korea, and Korea Institute of Science and Technology (KIST) in Seoul, Korea