Universality and Generalization of Causal Transformers Across Context Lengths

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the theoretical gap and generalization challenges in uniformly approximating causal mappings over arbitrarily long sequences using Transformers. To this end, this work characterizes cross-resolution causal families via α-Hölder continuity, establishing a universal approximation theory free of length-dependent parameters. By integrating masked attention mechanisms, it further derives generalization bounds within the infinite-length mean-field limit. The key contributions include achieving quantitative approximation without maximum-length factors, establishing a generalization bound of O((log log N/log N)^{β/(d+2)}), and validating the theoretical findings through experiments on physical time series data.
📝 Abstract
Long contexts are central to modern transformer systems, but most expressivity results choose a different network for each fixed sequence length. We study whether one masked transformer can approximate causal token-to-token maps uniformly over sequences of arbitrary length sampling a fixed normalized horizon. To relate sampling resolutions, we model tokens by $\alpha$-H\"older sequences or, more generally, a common modulus of continuity. Our notion of continuity across resolutions characterizes the causal families admitting uniform approximation on these compact input classes by a single transformer with length-independent parameters. The result extends to the infinite-length mean-field limit, where tokens form continuous curves and masked attention becomes a causal time integral. For bounded regression with target maps satisfying a $\beta$-smooth stability condition defined using regular test functions, quantitative approximation yields a generalization bound: exact empirical risk minimization over suitably sized bounded-weight transformers gives root mean-square prediction error $O((\log\log N/\log N)^{\beta/(d+2)})$ from $N$ iid labeled sequences. The bound holds at fixed confidence on the same sampling distribution, with $d$ the token dimension and no maximum-length factor. Finally, experiments on physical time series support the H\"older-regular token model at observed scales, with dataset-dependent fitted exponents, whereas text input embeddings provide a contrasting case. Native and dense sampling, shuffled controls, and refinement checks delimit this empirical regularity regime.
Problem

Research questions and friction points this paper is trying to address.

Causal Transformers
Context Length Generalization
Uniform Approximation
Hölder Continuity
Generalization Bound
Innovation

Methods, ideas, or system contributions that make the work stand out.

Causal Transformers
Universal Approximation
Context Length Generalization
Mean-field Limit
Hölder Continuity
🔎 Similar Papers
No similar papers found.