Score
Design training objectives, loss functions, and model components that align representations and outputs across spectral or channel bands so that semantic and textural embeddings, decoder reconstructions, and predicted channels remain mutually consistent. Implement cross-band alignment constraints, spectral-consistency losses, or guidance modules that enforce redundancy across bands and improve robustness of reconstruction and stochastic diffusion-based recovery.
本文通过自监督谱表示对齐方法优化扩散模型训练,利用谱正则化提升生成质量。
Training deep neural networks often suffers from catastrophic loss explosions, leading to costly failures. Conventional monitoring metrics—such as weight or gradient norms—are lagging and lack discriminative power for early detection. This paper proposes a spectral alignment–based early-warning mechanism: it quantifies the alignment between layer-wise input distributions and the dominant left singular vectors of weight matrices to detect incipient representational collapse. Theoretically, we show that the collapse of sign diversity in spectral alignment serves as an interpretable, pre-divergence indicator of training instability. Our method requires only lightweight SVD computation and statistical tracking, entailing minimal overhead and straightforward deployment. Empirical evaluation on language models demonstrates that our approach issues warnings significantly earlier than conventional metrics, with clearer signals and stronger generalization across architectures. This work establishes a novel paradigm for stabilizing large-model training through interpretable, spectrum-aware monitoring.
This work investigates the geometric and spectral alignment of dominant singular subspaces across residual Jacobian chains in deep neural networks, with the aim of ensuring structural stability in physical channel coordinates. Building upon Cartan coordinate rigidity and effective rank window fitting, the authors introduce a physically aligned matrix that decomposes signal propagation into core, overlap, and noise components. A static certificate radius—integrating column gaps, overlap margins, and noise bounds—is proposed to guarantee consistency between truncated and full propagation in terms of active support sets, association graphs, and mask structures, while defining SC/SA/ST-invariant channel mapping labels. Through singular subspace analysis, orthogonal decomposition, and perturbation validation, the study empirically demonstrates the alignment matrices and block energy heatmaps across CNNs, language models, and vision/diffusion backbone architectures, confirming the practical satisfiability of the proposed certificate conditions.
This work investigates the intrinsic relationship between SGD training dynamics and the spectral structure of empirical Hessian and gradient matrices in high-dimensional multiclass classification. Methodologically, it employs theoretical modeling and spectral analysis to characterize the evolution of these matrices throughout training. The key contribution is the first rigorous proof that, in deep multilayer networks, SGD trajectories align layerwise with the anomalous eigensubspaces—i.e., those spanned by top eigenvalues—of both the layer-wise Hessian and gradient matrices. Crucially, the rank of the final-layer anomalous subspace monotonically degenerates during optimization and becomes severely deficient upon convergence to a suboptimal classifier, serving as a spectral diagnostic for suboptimal convergence in overparameterized networks. This result is formally established for high-dimensional mixture models and single-/two-layer neural networks. Moreover, the work quantifies the direct link between the evolution of the final-layer anomalous subspace and classification accuracy, offering a novel spectral-geometric perspective on deep learning optimization dynamics.
Diffusion models suffer from severe image distortion under classifier-free guidance, primarily due to the misalignment between conventional mean-squared-error (MSE) loss and human visual perception. To address this, we propose a perceptually consistent self-supervised loss function, offering the first perceptual-supervision perspective on the efficacy of guidance mechanisms. Specifically, during the diffusion process, we construct self-supervised targets using deep features extracted from a pre-trained VGG network—requiring neither auxiliary classifiers nor dedicated guidance networks. Our method is fully compatible with standard diffusion frameworks and significantly improves generation quality even in the zero-guidance regime: FID scores drop substantially, human evaluation scores rise markedly, and the long-standing quality-diversity trade-off is alleviated. Generated images exhibit enhanced photorealism and richer fine-grained detail.
This work addresses the challenge of interference between primary objectives and constraint objectives when adapting foundation models under safety, privacy, or task-specific constraints. To resolve this issue, the authors propose SIFT (Spectral Interference-Free Training), a novel framework that uniquely integrates subspace orthogonalization from model merging with gradient orthogonalization. By analyzing cross-task interference in the spectral domain and introducing a localized intervention mechanism, SIFT enables selective and interference-free constrained optimization. Evaluated across four diverse tasks—machine unlearning, safety alignment, voice synthesis adaptation, and hallucination mitigation—SIFT consistently outperforms both constrained and unconstrained baselines, demonstrating robust and generalizable performance gains.
This study addresses the lack of explicit criteria in classifier-free guidance and the consequent difficulty in evaluating diffusion trajectory consistency by proposing a spectral correction guidance method. For the first time, this approach introduces intermediate-state spectral alignment as a principled criterion to adaptively rectify sampling biases, establishing a training-free trajectory correction mechanism that generalizes across diverse backbone architectures. Experimental results demonstrate that the proposed method significantly improves preference metrics and generation quality on text-to-image synthesis and ImageNet tasks. Notably, it maintains robust performance gains even under low-step sampling regimes, highlighting its effectiveness and efficiency for accelerating diffusion-based generative models without requiring additional training or architecture-specific modifications.
本文针对自监督学习中特征表示的适用性问题,通过分析任务协方差与谱对齐关系,提出了一种基于统计量的诊断方法及修正策略。
This work addresses the geometric mismatch arising from mainstream optimizers like Adam, which disregard the symmetry and equivariance structures inherent in neural network parameter spaces. The authors propose a principled framework for designing symmetry-compatible optimizers, tailoring gradient update rules to respect the equivariance of specific architectural components—such as embedding layers, language model heads, SwiGLU projections, and MoE routers—and assembling them into an end-to-end hierarchical optimizer stack. For the first time, this approach is systematically applied beyond generic matrix layers, encompassing permutation and shared translational symmetries, thereby unifying and extending equivariant optimization methods. Efficient compatibility is achieved through techniques including one-sided spectral updates, row/column-aware normalization, and centering. In pretraining both dense and sparse MoE language models, the proposed optimizer consistently outperforms AdamW, yielding lower validation loss and enhanced training stability.
This work addresses the exposure bias and error accumulation in diffusion models during inference, which stem from a mismatch between the frequency-domain distributions of training and sampling phases. The study systematically identifies that this issue arises from structural discrepancies in spectral signal-to-noise ratios and introduces Spectral Alignment (SPA), a novel method that aligns these distributions without altering the training procedure. SPA leverages offline, data-driven spectral priors and employs FFT-based gradient guidance during inference to calibrate the power spectrum of intermediate predictions. The approach is lightweight, architecture-agnostic, and fully compatible with Classifier-Free Guidance. Evaluated across diverse models—including DDPM, ADM, Stable Diffusion 2.0, SDXL, SD3.5, and FLUX—SPA consistently enhances generation quality with only a 3–4% increase in computational overhead.