Score
Designs and evaluates training-time regularizers and architectural operations that bias learned representations toward low- and mid-frequency content by attenuating or penalizing high-frequency components, and implements group-wise channel scanning/merging or skip-scanning block mechanisms to selectively combine channel groups. These techniques are used to control the spectral content of features and the behavior of group-scan-merge or skip-scanning blocks to improve robustness to real‑world degradation and related failure modes.
The physical mechanisms by which L2 regularization and dropout implicitly impose low-frequency inductive biases in CNNs remain poorly understood. Method: We propose a visual diagnostic framework and introduce the Spectral Suppression Ratio (SSR) to quantify the low-pass filtering strength of regularizers. Using discrete radial spectral analysis, weight spectral evolution tracking, and experiments on ResNet-18/CIFAR-10, we systematically characterize how regularization shapes weight frequency spectra. Contribution/Results: We find L2 regularization reduces high-frequency weight energy by over 3×, significantly improving blur robustness (+6.2%) but degrading Gaussian noise robustness—revealing an intrinsic accuracy–robustness trade-off governed by spectral bias. Furthermore, we design a small-kernel anti-aliasing model that provides interpretable theoretical grounding for understanding spectral inductive biases in deep learning.
This work investigates the implicit spectral bias regulation mechanism of gradient descent (GD) in one-dimensional shallow neural networks. Theoretically, GD is shown to act as a singular-value shrinkage operator on the network’s Jacobian matrix, where the learning rate and iteration count jointly determine the spectral bandwidth—the number of active frequency components. Three key contributions are made: (1) GD is formally modeled as an explicit frequency-domain shrinkage operation, with a quantitative relationship established between hyperparameters and bandwidth; (2) it is proven that GD’s implicit regularization effect on spectral bias is valid only under monotonic activation functions; (3) non-monotonic activations—including sinc and Gaussian—are proposed as efficient alternatives, significantly enhancing spectral control efficiency. Collectively, these results provide a novel theoretical framework for understanding implicit spectral bias in deep learning.
Adversarial training models often suffer from degraded robustness due to over-reliance on low-frequency structural components and neglect of high-frequency textural details. To address this, we propose the High-Frequency Feature Decoupling and Recalibration (HFDR) module and a Frequency Attention Regularization (FAR) method. First, we systematically identify and mitigate the low-frequency bias inherent in adversarial training. Second, we design a DCT/DFT-based frequency-domain feature decomposition mechanism, coupled with channel-frequency joint attention, to enable collaborative modeling of structure and texture. Third, we introduce a cross-spectral attention regularization loss to enhance high-frequency semantic fidelity. Extensive experiments on CIFAR-10/100 and ImageNet demonstrate significant improvements in robust accuracy under both white-box and transfer attacks. Our approach achieves superior generalization performance over existing state-of-the-art methods, with a 32% gain in high-frequency semantic fidelity.
Deep learning models often rely on spurious correlations, leading to substantial degradation in worst-group performance under subgroup distribution shifts. To address this, we propose an embedding-space regularization framework that establishes, for the first time, a theoretical connection between embedding representations and worst-group error. We prove that the direction of mean embedding differences across groups effectively separates core features from spurious ones. Leveraging this insight, we design a direction-sensitive regularization term applied directly to the embedding layer, explicitly suppressing reliance on spurious features while promoting learning of invariant core features. Our method requires no additional annotations or data augmentation and enjoys both theoretical guarantees and implementation simplicity. Extensive experiments on multiple vision and language benchmarks demonstrate significant improvements in worst-group accuracy—achieving state-of-the-art performance—and notably enhance generalization robustness, particularly for minority subgroups and out-of-distribution scenarios.
Existing reconstruction-based methods for heterogeneous subsequence anomaly detection in multivariate time series struggle to simultaneously achieve fine-grained frequency modeling and dynamic inter-channel dependency learning. Method: We propose a frequency-domain patching reconstruction framework. It introduces *frequency patching*—a novel mechanism that partitions the frequency spectrum into patches for band-level fine-grained feature representation—and a *Channel Fusion Module (CFM)* that employs patch-wise masking and masked attention to adaptively learn and cluster inter-channel dependencies. The framework integrates frequency-domain transformation, patch-based encoding, masked attention, and a two-level multi-objective optimization. Results: Evaluated on 10 real-world and 12 synthetic datasets, our method achieves state-of-the-art performance, significantly improving detection rates for subtle, localized, and cross-channel anomalies.
Existing diffusion models typically employ pointwise reconstruction losses that are insensitive to signal spectra and multiscale structures, often yielding samples with imbalanced frequency content and insufficient structural detail. To address this limitation, this work proposes a lightweight, plug-and-play spectral regularization framework that introduces differentiable Fourier- and wavelet-domain losses as soft inductive biases during standard training. The approach requires no modifications to the diffusion process, model architecture, or sampling procedure, and is compatible with mainstream paradigms such as DDPM, DDIM, and EDM. Evaluated on high-resolution unconditional generation tasks for both images and audio, the method consistently enhances sample quality—particularly in terms of spectral fidelity and multiscale coherence—while incurring negligible computational overhead.
In few-shot learning, deep networks suffer from channel bias—over-reliance on source-task-discriminative channels—leading to feature redundancy: most channels exhibit high intra-class variance and low inter-class separability, thereby hindering adaptation to novel tasks. Method: We establish, for the first time, the causal chain “channel bias → feature redundancy → performance degradation,” theoretically proving that redundancy is exacerbated under low-data regimes. We reveal the “less-is-more” principle: retaining only 1–5% of the most discriminative channels significantly improves accuracy, while >95% contribute negatively—a detrimental effect that attenuates with increasing sample size. Accordingly, we propose Augmented Feature Importance Adjustment (AFIA), a theoretically grounded method integrating data augmentation and soft channel masking to suppress redundancy. Results: AFIA achieves consistent and significant improvements across standard few-shot benchmarks, offering both a new theoretical principle and a practical tool for few-shot learning.
This work investigates whether partial data augmentation can statistically match the generalization performance and sample complexity of full-group augmentation under computational constraints. By leveraging Fourier analysis and finite group representation theory, the authors establish a unified theoretical framework that, for the first time, characterizes the conditions under which partial and full augmentations are statistically equivalent from a frequency-domain perspective. The main contributions include proving that when the augmented subset is sufficiently large, partial augmentation achieves the same minimax optimal rate as full augmentation, while also demonstrating that exact symmetry—leading to perfect invariance—can only be realized through averaging over the entire group. Consequently, the study delineates the theoretical limits of approximate symmetry and establishes an impossibility result showing that no proper subgroup can yield exact invariance.
This work addresses the inaccuracy in evaluating channel redundancy under extreme structured channel pruning by proposing a novel channel importance criterion based on complex-valued interaction fields and spectral reconstruction fidelity. Specifically, a complex interaction field is constructed between the input and each individual output channel, and a lightweight autoencoder is employed to reconstruct its Fourier spectrum; the resulting reconstruction error serves as a measure of channel compressibility. Combined with an L1-norm thresholding strategy, this approach enables threshold-driven structured pruning. To the best of our knowledge, this is the first method to incorporate complex-valued signal modeling and spectral autoencoders into channel pruning. On VGG16 trained on CIFAR-10, it achieves 90.11% FLOPs reduction and 96.30% parameter compression with only a 1.67% drop in Top-1 accuracy (from 93.44%), substantially outperforming existing extreme pruning methods.
To address the high FLOPs and training cost arising from redundant cross-channel attention computation in multi-channel images (e.g., cellular staining, satellite imagery), this paper proposes MoE-ViT—the first vision Transformer framework integrating sparse Mixture-of-Experts (MoE). It innovatively treats *channels* as experts and employs a lightweight routing network to dynamically select the most relevant channel subset for each image patch, enabling channel-wise sparse interaction. This design challenges the implicit assumption that all channel pairs must be fully connected, thereby substantially mitigating the quadratic growth of attention complexity while preserving or even improving accuracy. Extensive experiments on JUMP-CP and So2Sat demonstrate that MoE-ViT significantly reduces FLOPs and training overhead while achieving state-of-the-art performance, establishing it as a practical and efficient backbone architecture for multi-channel imaging.