Force without transmission: a depth-induced rank collapse that no loss on the representation reopens

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue of rank collapse during Transformer training, which causes learning stagnation that conventional loss terms fail to remedy. Through experiments on small-scale Transformers and gradient analysis, this work proposes a "gradient reachability" theory, revealing that the disruption of gradient propagation pathways is the primary cause of such repair failures. Accordingly, a dynamic skip connection regulation mechanism is designed to restart gradient flow. The authors demonstrate that relying solely on loss terms is ineffective, whereas restoring skip connections enables unsupervised rank recovery. By distinguishing the critical factors of network reparability along the dimensions of force magnitude and transmission pathways, this research offers a novel perspective for understanding and reversing rank collapse, even though post-recovery performance remains inferior to that of healthy models.
📝 Abstract
Training can drive a transformer into a rank collapse: all token representations point in one direction, and learning stops. In a related collapse of attention, a loss term with a bounded corrective force repairs the network during the run. We ask whether such a term repairs rank collapse. We collapse small transformers by weakening their skip connection and treat copies of the collapsed network. No added loss term repaired the collapse, although the stronger kind pushed with about a tenth of the task gradient. The reason was the path, not the strength. The task gradient no longer reached the query and key weights, which decide where attention looks, and the added term's gradient faded before the blocks where the collapse forms. Restoring the skip connection, which changes no weight, reopened this path at once. The rank then recovered, but only far above the scale of collapse. After a burst of high learning rate the path stayed open and the rank recovered untreated. Registered predictions from the path ranked recovery times but did not transfer to this cause. In every case the loss stayed above that of a healthy network after the rank recovered. Whether a collapsed network can be repaired depends on whether the gradient still reaches the weights that must change, not on how strongly a loss term pushes.
Problem

Research questions and friction points this paper is trying to address.

rank collapse
transformer
gradient transmission
skip connection
loss correction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Rank Collapse
Gradient Transmission
Skip Connection
Attention Mechanism
Transformer