Score
Designs and implements convolutional neural network feature extractors whose kernels are initialized, parameterized, or constrained to follow Gabor or Gabor-like functions to enforce oriented, frequency- and time-sensitive filtering. Builds compact, Gabor-constrained CNN layers and analyzes the resulting filter parameters and representations to reduce parameter count and overfitting while extracting interpretable frequency/temporal features.
Standard convolutions, due to their fixed structure, linearity, and reliance on local averaging, struggle to capture complex image characteristics such as low-rank structures, adaptive basis representations, and non-uniform spatial dependencies. This work proposes a unified taxonomy encompassing five classes of structured operators—decomposition-based, adaptive weighting, basis-adaptive, integral/kernel-based, and attention-based—and systematically analyzes their differences along key dimensions including locality, linearity, and equivariance. By leveraging techniques such as singular value/tensor decomposition, content-adaptive weighting, learnable analysis bases, position-dependent nonlinear kernels, and attention mechanisms, the study comprehensively evaluates the performance of these operators across image-to-image and image-to-label tasks. The findings clarify the respective strengths and limitations of each operator class, offering both theoretical insights and practical guidance for future research.
The first layer of CNNs exhibits sensitivity to input translations, leading to unstable parameter learning. Method: We reveal that max-pooling can approximate complex modulus operations—and thus achieve approximate translation invariance—under specific conditions. We propose the first quantitative metric for translation invariance in subsampled convolution followed by max-pooling, and theoretically establish that filter center frequency and orientation are the key determinants of stability. Leveraging the dual-tree complex wavelet packet transform—a discrete Gabor decomposition—we construct a deterministic feature extractor for empirical validation. Contribution/Results: Both theoretical analysis and experiments consistently demonstrate that deliberately designing filters with controlled spectral orientation significantly enhances the translation robustness of pooled feature maps. Our core contribution is a novel, interpretable, and quantifiable framework for analyzing translation stability in CNN first layers, coupled with actionable, frequency-domain filter design principles to improve robustness.
Convolutional neural networks (CNNs) empirically exhibit strong frequency biases and hierarchical processing, yet the structural principles underlying their representational efficiency remain poorly understood. Method: We introduce the convolutional bottleneck (CBN)—a naturally emergent architectural motif wherein early layers compress inputs into sparse frequency-channel representations, and later layers reconstruct outputs from this compressed representation. We define the CBN rank to quantify the number and types of critical frequencies preserved in the bottleneck, and integrate Fourier-domain analysis, parameter complexity theory, stability-driven derivations of activation/weight structures, and empirical validation. Contribution/Results: We prove that parameter norm scales linearly with network depth and CBN rank, tightly linking it to the regularity of the target function. Crucially, we provide the first optimization-efficiency–driven theoretical justification for downsampling: optimally parameter-efficient CNNs must possess a CBN structure. Our framework successfully decodes learned frequency preferences and functional mechanisms in multi-task CNNs, establishing a new paradigm for understanding CNN representation learning and designing efficient architectures.
To address the high computational cost and weak modeling capability for multi-scale and multi-directional features in vision Transformers (ViTs) for dense prediction tasks, this paper proposes the Focal Vision Transformer (FViT). Methodologically, FViT replaces self-attention with learnable Gabor filters (LGFs) to explicitly encode local directional textures; introduces a biologically inspired Focal Vision (BFV) module that emulates neural mechanisms of visual attention—namely, focal enhancement and surround suppression; and constructs a lightweight pyramid architecture augmented with a multi-path feed-forward network (MPFFN) to improve feature reuse. Experimental results demonstrate that FViT significantly outperforms mainstream ViT variants on semantic segmentation and object detection benchmarks. At comparable accuracy, it reduces computational cost by 30–50%, while exhibiting strong generalization, high efficiency, and excellent scalability.
This work investigates the fundamental frequency-domain characteristics of the ReLU activation function and its impact on CNN representation learning from a signal processing perspective. We derive a rigorous closed-form spectral representation of ReLU, revealing that it inherently introduces high-frequency oscillations while simultaneously generating a dominant direct-current (DC) component. Contrary to conventional emphasis on nonlinear high-frequency effects, we establish the DC component as a critical mechanism ensuring both stability and spectral sensitivity in feature extraction. Methodologically, we integrate Taylor spectral analysis, numerical frequency-response modeling, feature visualization, and multi-scale ablation studies—including validation on real CNNs. Theoretical findings are fully corroborated by numerical experiments: the DC component significantly enhances network responsiveness to input frequency content and steers weight optimization toward stable solutions near initialization.
Standard convolutional operations cannot be directly applied to graph-structured data due to its irregular, non-Euclidean topology. Method: This paper proposes Graph Kernel-driven Learnable Structural Convolution (GK-Conv), a purely structural, end-to-end modeling framework operating directly on non-Euclidean graph domains. GK-Conv eliminates explicit graph embedding and instead constructs a parameterized, structural convolutional operator grounded in generic graph kernel functions—enabling plug-and-play integration of arbitrary graph kernels and generating CNN-style, interpretable structural masks. The model is fully differentiable and optimized via ablation-guided hyperparameter analysis. Contribution/Results: GK-Conv achieves state-of-the-art performance across multiple graph classification and regression benchmarks, empirically validating the central claim that strong generalization can be attained using topology alone—without node or edge features.
Causal convolutional neural networks (CNNs) lack interpretability in multimodal frequency–time series modeling. Method: We reveal that trained causal CNNs are mathematically equivalent to finite impulse response (FIR) filters and propose an analytical simplification leveraging the associativity of convolution to rigorously reduce deep causal CNNs to a single-layer equivalent FIR filter. Quasi-linear activation functions and least-squares optimization enable explicit frequency-domain feature extraction and interpretable mapping of filter parameters. Results: Evaluated on simulated beam dynamics and real bridge vibration data, our method accurately models sparse-spectrum physical system dynamics. Crucially, it establishes, for the first time, an analytical correspondence between causal CNN weights and the system’s frequency response function—thereby significantly enhancing the physical interpretability and spectral awareness of deep models in dynamic system identification.
This work proposes a continuous spectral-domain parameterized convolution method that overcomes the limitations of traditional convolutional neural networks, which are constrained by local receptive fields and struggle to capture global context, as well as vision Transformers, which lack spatial inductive biases and rely on fixed patch partitioning and positional encodings. By introducing direction-aware continuous spectral basis functions for the first time, the method defines smooth, shared convolutional kernels across the entire frequency domain, achieving both global receptive fields and resolution adaptability. The approach effectively combines structural priors with global modeling capacity, yielding strong robustness to geometric transformations, noise, and scale variations. It matches or surpasses the performance of existing convolutional, attention-based, and spectral methods while reducing the number of parameters by an order of magnitude across image classification, synthetic benchmarks, and 3D medical imaging tasks.
This work addresses the high parameter count and overfitting issues of conventional convolutional Kolmogorov–Arnold Networks (KANs), which employ independent learnable functions for each kernel element. The authors propose a structured KAN design paradigm that shifts learnable functions from individual weights to the convolutional architecture itself, exploring two distinct pathways: applying functions to pixel values (SV-KAN, AG-KAN) or to filter shapes (RF-KAN). Notably, RF-KAN leverages shared univariate functions, content-adaptive Gaussian gating, and Morlet wavelet–based oriented ridge filters, achieving 88.47% and 64.57% accuracy on CIFAR-10 and CIFAR-100, respectively, with only approximately 0.4 million parameters. This performance significantly surpasses both standard convolutions and existing KAN approaches at comparable scales, highlighting the critical role of localized oscillatory bases and content adaptivity.
This work proposes a learnable inter-filter connectivity mechanism that replaces the fixed pointwise nonlinear activations in conventional convolutional neural networks with a parameterized, universal connection function embedded within convolutional layers. By enabling adaptive interactions among filters, the approach overcomes the limitations of traditional fixed logical operations—such as multiplication or minimum selection—and allows the network to automatically optimize its connectivity strategy through end-to-end training. Experimental results demonstrate that this method significantly improves classification accuracy, confirming its effectiveness in enhancing both model expressivity and generalization capability.
本文通过广义Stein方法从统计角度估计卷积滤波器,提出基于奇异值分解的新方法,实现滤波器的准确学习。