๐ค AI Summary
This work addresses the lack of theoretical guarantees for the generalization of Transformers under limited training samples by modeling their training dynamics as a finite-horizon Markov control problem. It pioneers an integration of optimal control theory, measure-valued dynamical systems, and Wasserstein distributionally robust optimization. By leveraging state-action space quantization, Lipschitz stability analysis, and concentration inequalities for empirical measures, the study derives an explicit finite-sample generalization bound. This bound not only uncovers the intrinsic distributional robustness of Transformers but also quantifies the impact of approximation errors on their generalization performance.
๐ Abstract
We derive finite-sample generalization bounds for Transformers trained with dynamic programming recursions. Building on the doubly lifted, measure-valued formulation of Transformer dynamics, we view data sets as probability laws on pairs of empirical input-output measures, allowing us to interpret the training problem as a finite-horizon Markovian control problem. We then analyze a quantized model, derived by quantizing the state, action, and measure-state spaces, and derive explicit finite-sample generalization bounds using concentration inequalities for empirical laws on finite metric spaces together with a Lipschitz stability estimate for the value function. These bounds are transferred to the base model at the cost of an explicit approximation error. Finally, we show that the same machinery yields a distributionally robust control formulation of the training problem, connecting Transformer generalization to Wasserstein distributionally robust optimization.