frequency-biased regularization

Designs and evaluates training-time regularizers and architectural operations that bias learned representations toward low- and mid-frequency content by attenuating or penalizing high-frequency components, and implements group-wise channel scanning/merging or skip-scanning block mechanisms to selectively combine channel groups. These techniques are used to control the spectral content of features and the behavior of group-scan-merge or skip-scanning blocks to improve robustness to real‑world degradation and related failure modes.

frequency-biasedregularization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.17
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Frequency Regularization: Unveiling the Spectral Inductive Bias of Deep Neural Networks

Dec 20, 2025
JL
Jiahao Lu
🏛️ The Chinese University of Hong Kong, Shenzhen

The physical mechanisms by which L2 regularization and dropout implicitly impose low-frequency inductive biases in CNNs remain poorly understood. Method: We propose a visual diagnostic framework and introduce the Spectral Suppression Ratio (SSR) to quantify the low-pass filtering strength of regularizers. Using discrete radial spectral analysis, weight spectral evolution tracking, and experiments on ResNet-18/CIFAR-10, we systematically characterize how regularization shapes weight frequency spectra. Contribution/Results: We find L2 regularization reduces high-frequency weight energy by over 3×, significantly improving blur robustness (+6.2%) but degrading Gaussian noise robustness—revealing an intrinsic accuracy–robustness trade-off governed by spectral bias. Furthermore, we design a small-kernel anti-aliasing model that provides interpretable theoretical grounding for understanding spectral inductive biases in deep learning.

Introduces a framework to track weight frequency evolution during trainingInvestigates spectral bias and frequency selection mechanisms in CNNsReveals accuracy-robustness trade-off related to low-frequency specialization

Gradient Descent as a Shrinkage Operator for Spectral Bias

Apr 25, 2025
SL
Simon Lucey
🏛️ Australian Institute for Machine Learning | University of Adelaide

This work investigates the implicit spectral bias regulation mechanism of gradient descent (GD) in one-dimensional shallow neural networks. Theoretically, GD is shown to act as a singular-value shrinkage operator on the network’s Jacobian matrix, where the learning rate and iteration count jointly determine the spectral bandwidth—the number of active frequency components. Three key contributions are made: (1) GD is formally modeled as an explicit frequency-domain shrinkage operation, with a quantitative relationship established between hyperparameters and bandwidth; (2) it is proven that GD’s implicit regularization effect on spectral bias is valid only under monotonic activation functions; (3) non-monotonic activations—including sinc and Gaussian—are proposed as efficient alternatives, significantly enhancing spectral control efficiency. Collectively, these results provide a novel theoretical framework for understanding implicit spectral bias in deep learning.

Characterizes influence of activation functions on spectral biasDemonstrates GD as a shrinkage operator for spectral componentsProposes GD hyperparameters' impact on bandwidth selection

Mitigating Low-Frequency Bias: Feature Recalibration and Frequency Attention Regularization for Adversarial Robustness

Jul 04, 2024
KZ
Kejia Zhang
🏛️ Xiamen University | Jinan University | Minjiang University

Adversarial training models often suffer from degraded robustness due to over-reliance on low-frequency structural components and neglect of high-frequency textural details. To address this, we propose the High-Frequency Feature Decoupling and Recalibration (HFDR) module and a Frequency Attention Regularization (FAR) method. First, we systematically identify and mitigate the low-frequency bias inherent in adversarial training. Second, we design a DCT/DFT-based frequency-domain feature decomposition mechanism, coupled with channel-frequency joint attention, to enable collaborative modeling of structure and texture. Third, we introduce a cross-spectral attention regularization loss to enhance high-frequency semantic fidelity. Extensive experiments on CIFAR-10/100 and ImageNet demonstrate significant improvements in robust accuracy under both white-box and transfer attacks. Our approach achieves superior generalization performance over existing state-of-the-art methods, with a 32% gain in high-frequency semantic fidelity.

Adversarial AttacksDeep Learning RobustnessImage Processing

Spurious Correlation-Aware Embedding Regularization for Worst-Group Robustness

Nov 06, 2025
SP
Subeen Park
🏛️ Yonsei University | LG CNS

Deep learning models often rely on spurious correlations, leading to substantial degradation in worst-group performance under subgroup distribution shifts. To address this, we propose an embedding-space regularization framework that establishes, for the first time, a theoretical connection between embedding representations and worst-group error. We prove that the direction of mean embedding differences across groups effectively separates core features from spurious ones. Leveraging this insight, we design a direction-sensitive regularization term applied directly to the embedding layer, explicitly suppressing reliance on spurious features while promoting learning of invariant core features. Our method requires no additional annotations or data augmentation and enjoys both theoretical guarantees and implementation simplicity. Extensive experiments on multiple vision and language benchmarks demonstrate significant improvements in worst-group accuracy—achieving state-of-the-art performance—and notably enhance generalization robustness, particularly for minority subgroups and out-of-distribution scenarios.

Connecting embedding space representations with worst-group error theoreticallyImproving model robustness in underrepresented subpopulation groupsMitigating deep learning models' reliance on spurious correlations

CATCH: Channel-Aware multivariate Time Series Anomaly Detection via Frequency Patching

Oct 16, 2024
XW
Xingjian Wu
🏛️ East China Normal University | The Hong Kong University of Science and Technology

Existing reconstruction-based methods for heterogeneous subsequence anomaly detection in multivariate time series struggle to simultaneously achieve fine-grained frequency modeling and dynamic inter-channel dependency learning. Method: We propose a frequency-domain patching reconstruction framework. It introduces *frequency patching*—a novel mechanism that partitions the frequency spectrum into patches for band-level fine-grained feature representation—and a *Channel Fusion Module (CFM)* that employs patch-wise masking and masked attention to adaptively learn and cluster inter-channel dependencies. The framework integrates frequency-domain transformation, patch-based encoding, masked attention, and a two-level multi-objective optimization. Results: Evaluated on 10 real-world and 12 synthetic datasets, our method achieves state-of-the-art performance, significantly improving detection rates for subtle, localized, and cross-channel anomalies.

Detect anomalies in multivariate time seriesEnhance fine-grained frequency characteristics captureImprove channel correlation perception in anomaly detection

Latest Papers

What's happening recently
View more

Existing diffusion models typically employ pointwise reconstruction losses that are insensitive to signal spectra and multiscale structures, often yielding samples with imbalanced frequency content and insufficient structural detail. To address this limitation, this work proposes a lightweight, plug-and-play spectral regularization framework that introduces differentiable Fourier- and wavelet-domain losses as soft inductive biases during standard training. The approach requires no modifications to the diffusion process, model architecture, or sampling procedure, and is compatible with mainstream paradigms such as DDPM, DDIM, and EDM. Evaluated on high-resolution unconditional generation tasks for both images and audio, the method consistently enhances sample quality—particularly in terms of spectral fidelity and multiscale coherence—while incurring negligible computational overhead.

diffusion modelsfrequency balancegenerative modeling

From Channel Bias to Feature Redundancy: Uncovering the "Less is More" Principle in Few-Shot Learning

Apr 11, 2026
XL
Xu Luo
🏛️ UESTC | The University of Hong Kong | Harbin Institute of Technology Shenzhen

In few-shot learning, deep networks suffer from channel bias—over-reliance on source-task-discriminative channels—leading to feature redundancy: most channels exhibit high intra-class variance and low inter-class separability, thereby hindering adaptation to novel tasks. Method: We establish, for the first time, the causal chain “channel bias → feature redundancy → performance degradation,” theoretically proving that redundancy is exacerbated under low-data regimes. We reveal the “less-is-more” principle: retaining only 1–5% of the most discriminative channels significantly improves accuracy, while >95% contribute negatively—a detrimental effect that attenuates with increasing sample size. Accordingly, we propose Augmented Feature Importance Adjustment (AFIA), a theoretically grounded method integrating data augmentation and soft channel masking to suppress redundancy. Results: AFIA achieves consistent and significant improvements across standard few-shot benchmarks, offering both a new theoretical principle and a practical tool for few-shot learning.

Addressing channel bias in few-shot learning under distribution shiftsMitigating confounding features for robust representation transferReducing feature redundancy to improve classification with limited data

This work investigates whether partial data augmentation can statistically match the generalization performance and sample complexity of full-group augmentation under computational constraints. By leveraging Fourier analysis and finite group representation theory, the authors establish a unified theoretical framework that, for the first time, characterizes the conditions under which partial and full augmentations are statistically equivalent from a frequency-domain perspective. The main contributions include proving that when the augmented subset is sufficiently large, partial augmentation achieves the same minimax optimal rate as full augmentation, while also demonstrating that exact symmetry—leading to perfect invariance—can only be realized through averaging over the entire group. Consequently, the study delineates the theoretical limits of approximate symmetry and establishes an impossibility result showing that no proper subgroup can yield exact invariance.

computational feasibilitydata augmentationgroup invariance

This work addresses the inaccuracy in evaluating channel redundancy under extreme structured channel pruning by proposing a novel channel importance criterion based on complex-valued interaction fields and spectral reconstruction fidelity. Specifically, a complex interaction field is constructed between the input and each individual output channel, and a lightweight autoencoder is employed to reconstruct its Fourier spectrum; the resulting reconstruction error serves as a measure of channel compressibility. Combined with an L1-norm thresholding strategy, this approach enables threshold-driven structured pruning. To the best of our knowledge, this is the first method to incorporate complex-valued signal modeling and spectral autoencoders into channel pruning. On VGG16 trained on CIFAR-10, it achieves 90.11% FLOPs reduction and 96.30% parameter compression with only a 1.67% drop in Top-1 accuracy (from 93.44%), substantially outperforming existing extreme pruning methods.

channel pruningconvolutional neural networksmodel compression

Sparse Mixture-of-Experts for Multi-Channel Imaging: Are All Channel Interactions Required?

Nov 21, 2025
SY
Sukwon Yun
🏛️ University of North Carolina at Chapel Hill | Genentech

To address the high FLOPs and training cost arising from redundant cross-channel attention computation in multi-channel images (e.g., cellular staining, satellite imagery), this paper proposes MoE-ViT—the first vision Transformer framework integrating sparse Mixture-of-Experts (MoE). It innovatively treats *channels* as experts and employs a lightweight routing network to dynamically select the most relevant channel subset for each image patch, enabling channel-wise sparse interaction. This design challenges the implicit assumption that all channel pairs must be fully connected, thereby substantially mitigating the quadratic growth of attention complexity while preserving or even improving accuracy. Extensive experiments on JUMP-CP and So2Sat demonstrate that MoE-ViT significantly reduces FLOPs and training overhead while achieving state-of-the-art performance, establishing it as a practical and efficient backbone architecture for multi-channel imaging.

Addressing computational bottlenecks in cross-channel attention mechanismsOptimizing Vision Transformers for multi-channel imaging domainsReducing excessive FLOPs and training costs in channel interactions

Hot Scholars

DK

Dongjun Kim

Stanford University
Machine LearningArtificial Intelligence
XL

Xinpeng Ling

Tongji University
Federated LearningDifferential PrivacyConvex Optimization
CC

Chaofan Chen

Assistant Professor, University of Maine
Machine LearningArtificial Intelligence
ZW

Zengyi Wo

College of Intelligence and Computing, Tianjin University
Data MiningAnomaly DetectionLLM Reasoning
SF

Sina Farsiu

Anderson-Rupp Professor of Biomedical Engineering and Professor of Ophthalmology, Duke
Medical Image ProcessingVision ScienceDiaper ChangingSuper_Resolution