The Birkhoff Geometry of Manifold-Constrained Hyper-Connections: Two Channels, Vertex Viscosity, and Sinkhorn as a Retraction

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper elucidates the geometric mechanisms and memory decay principles underlying Sinkhorn-normalized mixers within manifold-constrained hyperconnections. Methodologically, it establishes a Birkhoff polytope geometric theory that decomposes doubly stochastic matrices into mean and difference channels. It proves that the Sinkhorn mapping constitutes a global coordinate chart, demonstrates the equivalence of straight-through updates to entropic mirror descent, and introduces the Fisher-Rao metric for differential geometric analysis. Key contributions include proving that additional width induces a finite memory horizon, deriving analytical relationships between local convergence factors and this memory horizon, and quantifying convergence rate disparities near vertices. Experimental results validate the predicted gradient flow and mirror descent rates, providing a rigorous theoretical foundation for optimizing Transformer architectures.
📝 Abstract
Hyper-connections widen the residual stream of a Transformer to $n$ parallel streams. Their manifold-constrained version (mHC) mixes the streams at each layer with a doubly stochastic matrix, which it computes by Sinkhorn normalization of exponentiated logits. We give a geometric theory of this design on the Birkhoff polytope. First, a doubly stochastic mixer splits the stream into a mean channel, on which mHC is exactly a residual network, and a difference channel, which each layer contracts by its second singular value $\sigma_2 \le 1 - n \min_{ij} H_{ij}$. Thus the extra width is a fading memory with a horizon of $1/(1-\sigma_2)$ layers, and among nonnegative mixers only the permutations do not collapse. Second, the Sinkhorn-logit map is a global chart, and its logit gradient is exactly the Fisher-Rao gradient. Thus logit gradient flow follows a squared Fisher-Rao metric, and the straight-through update is exactly entropic mirror descent. Third, under logit gradient flow the logarithm of each entry moves at a rate of at most $4n^3\|\nabla f\|_\infty \varepsilon$, where $\varepsilon$ is the distance to the nearest permutation. Thus gradient flow approaches and leaves the vertices only at rate $1/t$, but mirror descent moves at an exponential rate. Fourth, the local convergence factor of Sinkhorn is $\sigma_2^2$, so a fixed iteration budget limits the horizon. Experiments confirm the predicted rates.
Problem

Research questions and friction points this paper is trying to address.

Hyper-Connections
Birkhoff polytope
Sinkhorn normalization
doubly stochastic matrix
manifold-constrained
Innovation

Methods, ideas, or system contributions that make the work stand out.

Birkhoff polytope
Manifold-constrained Hyper-Connections
Sinkhorn normalization
Fisher-Rao metric
Entropic mirror descent
🔎 Similar Papers
No similar papers found.