Flag Varieties: A Geometric Framework for Deep Network Alignment

📅 2026-05-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of a unified theoretical explanation for the alignment phenomenon observed between weight subspaces of adjacent layers in deep neural networks. Building from first principles through geometric invariant theory, the study establishes a unified geometric framework that identifies this alignment structure as a flag manifold and proves that the dimension of subspace intersections is the unique reparameterization-invariant observable. Leveraging Lie bracket analysis and a novel weight-space diagnostic tool, the theory rigorously demonstrates the mathematical necessity of subspace metrics and provides an explanation for the hierarchical structure underlying Neural Collapse. Experimental validation across multilayer perceptrons, residual networks, and pretrained language models confirms the theoretical predictions, with the proposed diagnostic metric accurately capturing internal alignment properties without requiring forward propagation.
📝 Abstract
Alignment, the tendency of adjacent weight matrices in deep networks to develop compatible subspace orientations, underlies gradient flow, Neural Collapse, and representation similarity across architectures. Despite extensive empirical documentation, these phenomena have resisted unified theoretical treatment: existing explanations are post-hoc, each fitted to a specific observation with whatever mathematics is at hand. We reverse this direction by deriving the mathematical structure that layerwise alignment inherently demands. Using geometric invariant theory, we prove that alignment geometry has a canonical closed, polystable stratum given by a flag variety, and that subspace intersection dimension is its unique reparameterization-invariant observable, establishing that subspace metrics are not empirical conventions but mathematical necessities. This unified framework yields two dynamical consequences: ridge regularization drives subspace alignment at an exponential rate set by weight decay, whereas nonlinear activations induce a commutator obstruction to exact basis alignment, generically present in nonlinear networks and absent in linear ones. Together these give a geometric explanation of the Level-2/3 hierarchy in Neural Collapse from first principles rather than post-hoc analysis. The commutator magnitude and head subspace overlap further serve as weight-space windows into internal alignment structure, requiring no forward passes. Experiments on multilayer perceptrons, residual networks, and pretrained language models support the proposed diagnostics and delineate their scope.
Problem

Research questions and friction points this paper is trying to address.

Alignment
Flag Varieties
Neural Collapse
Subspace Geometry
Deep Network
Innovation

Methods, ideas, or system contributions that make the work stand out.

flag variety
layerwise alignment
geometric invariant theory
commutator obstruction
Neural Collapse
🔎 Similar Papers
J
Jingchuan Xiao
Department of Mathematics and Computer Studies, Mary Immaculate College, Ireland
X
Xinyi Sui
Department of Computer Science and Engineering, Santa Clara University, USA
C
Cihan Ruan
Department of Computer Science and Engineering, Santa Clara University, USA