Exact-Solution Volume and Length Generalization in Transformers

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of quantitative theoretical foundations for Transformer generalization under long input sequences. It proposes a normalized solution volume metric and establishes its theoretical connection to length generalization. Through asymptotic analysis and single-head attention modeling, this work derives theoretical bounds for fixed-width, single-layer Transformers across four task categories, further tightening the bound on the INDEX task via structural optimization that eliminates error sources. The findings reveal an intrinsic link between solution volume decay and length sensitivity, confirming a positive correlation between volume shrinkage and generalization difficulty. Empirically, the optimized model improves accuracy from 60% to 85% at ten times the training sequence length. Collectively, this research provides both a rigorous theoretical framework and practical strategies for enhancing Transformer generalization over long sequences.
📝 Abstract
Research on transformer expressivity shows whether a transformer is capable of solving a given task, but gives little indication of whether the solution, if learned, is generalizable to longer input lengths. We study this question through normalized exact-solution volume (NESV): the fraction of a bounded parameter region that achieves an exact solution on every input of length $n$. For fixed-width, single-layer transformers with $\log n$-scaled attention, we establish asymptotic bounds on NESV for four tasks: FIRST ($Θ(1)$), MAJORITY ($Θ(1/(n\log n))$), INDEX ($Θ(1/n^3)$), and PARITY ($0$). These results are consistent with previous empirical results: the faster the exact-solution volume decays with input length, the harder it is to length-generalize on that task. Looking deeper into INDEX, our volume analysis reveals two error sources that grow with $n$. Consequently, we study a transformer model that would structurally eliminate one of the terms, theoretically improving the NESV bound to $Θ(n^{-1})$, and empirically achieving 85% accuracy when tested at $10\times$ the training length, compared with the 60% accuracy of the original model. We conclude that volume analysis may be a useful approach to identify concrete sources of length sensitivity and thus provide insights into task-specific model refinements.
Problem

Research questions and friction points this paper is trying to address.

Transformers
length generalization
expressivity
exact-solution volume
Innovation

Methods, ideas, or system contributions that make the work stand out.

Normalized Exact-Solution Volume
Length Generalization
Transformers
Asymptotic Bounds
Volume Analysis
🔎 Similar Papers