Spectral Alignment as Predictor of Loss Explosion in Neural Network Training

📅 2025-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Training deep neural networks often suffers from catastrophic loss explosions, leading to costly failures. Conventional monitoring metrics—such as weight or gradient norms—are lagging and lack discriminative power for early detection. This paper proposes a spectral alignment–based early-warning mechanism: it quantifies the alignment between layer-wise input distributions and the dominant left singular vectors of weight matrices to detect incipient representational collapse. Theoretically, we show that the collapse of sign diversity in spectral alignment serves as an interpretable, pre-divergence indicator of training instability. Our method requires only lightweight SVD computation and statistical tracking, entailing minimal overhead and straightforward deployment. Empirical evaluation on language models demonstrates that our approach issues warnings significantly earlier than conventional metrics, with clearer signals and stronger generalization across architectures. This work establishes a novel paradigm for stabilizing large-model training through interpretable, spectrum-aware monitoring.

Technology Category

Machine Learning: Deep Neural Architectures and Foundation ModelsComputer Vision: Large Vision ModelsNatural Language Processing: (Large) Language Models

Application Category

Web Mining and Content Analysis: Large pretrained models with web dataGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsResponsible Web: Human-perceived consequences of algorithmic deployment on the web
📝 Abstract
Loss explosions in training deep neural networks can nullify multi-million dollar training runs. Conventional monitoring metrics like weight and gradient norms are often lagging and ambiguous predictors, as their values vary dramatically across different models and even between layers of the same model, making it difficult to establish a unified standard for detecting impending failure. We introduce Spectral Alignment (SA), a novel, theoretically-grounded metric that monitors the distributional alignment between layer inputs and the principal singular vectors of weight matrices. We show that a collapse in the sign diversity of this alignment is a powerful early predictor of representational collapse and training divergence. Empirical results on language models demonstrate that monitoring the SA distribution provides a significantly earlier and clearer warning of loss explosions than traditional scalar metrics. SA's low computational overhead makes it a practical tool for safeguarding model training.
Problem

Research questions and friction points this paper is trying to address.

Predicting loss explosions in neural network training early
Overcoming limitations of traditional weight and gradient metrics
Monitoring spectral alignment to detect representational collapse
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spectral Alignment monitors layer inputs and weight singular vectors
SA collapse predicts representational collapse before training divergence
Low computational overhead enables practical training failure detection
🔎 Similar Papers
H
Haiquan Qiu
Tsinghua University
Y
You Wu
Tsinghua University
Y
Yingjie Tan
Tsinghua University
Y
Yaqing Wang
Beijing Institute of Mathematical Sciences and Applications
Quanming Yao
Quanming Yao
Associate Professor, EE Department, Tsinghua University
Machine Learning