🤖 AI Summary
This study addresses the theoretical gap and generalization challenges in uniformly approximating causal mappings over arbitrarily long sequences using Transformers. To this end, this work characterizes cross-resolution causal families via α-Hölder continuity, establishing a universal approximation theory free of length-dependent parameters. By integrating masked attention mechanisms, it further derives generalization bounds within the infinite-length mean-field limit. The key contributions include achieving quantitative approximation without maximum-length factors, establishing a generalization bound of O((log log N/log N)^{β/(d+2)}), and validating the theoretical findings through experiments on physical time series data.
📝 Abstract
Long contexts are central to modern transformer systems, but most expressivity results choose a different network for each fixed sequence length. We study whether one masked transformer can approximate causal token-to-token maps uniformly over sequences of arbitrary length sampling a fixed normalized horizon. To relate sampling resolutions, we model tokens by $\alpha$-H\"older sequences or, more generally, a common modulus of continuity. Our notion of continuity across resolutions characterizes the causal families admitting uniform approximation on these compact input classes by a single transformer with length-independent parameters. The result extends to the infinite-length mean-field limit, where tokens form continuous curves and masked attention becomes a causal time integral. For bounded regression with target maps satisfying a $\beta$-smooth stability condition defined using regular test functions, quantitative approximation yields a generalization bound: exact empirical risk minimization over suitably sized bounded-weight transformers gives root mean-square prediction error $O((\log\log N/\log N)^{\beta/(d+2)})$ from $N$ iid labeled sequences. The bound holds at fixed confidence on the same sampling distribution, with $d$ the token dimension and no maximum-length factor. Finally, experiments on physical time series support the H\"older-regular token model at observed scales, with dataset-dependent fitted exponents, whereas text input embeddings provide a contrasting case. Native and dense sampling, shuffled controls, and refinement checks delimit this empirical regularity regime.