🤖 AI Summary
This study addresses the lack of theoretical foundations for in-context learning (ICL) in Transformers operating on non-Euclidean and geometrically heterogeneous data. To this end, we investigate nonparametric ICL over mixture manifolds by integrating local polynomial regression, geometric preconditioners, chartwise solvers, and near-empirical risk minimization to construct minimax-optimal local polynomial estimators alongside a two-stage Softmax Transformer architecture. Our work provides the first characterization of the adaptive mechanisms underlying Transformers when confronted with unknown local geometries. We demonstrate that approximation errors become negligible under logarithmic depth and polynomial size constraints. Furthermore, we establish an aggregated minimax lower bound and derive matching upper bounds as well as generalization guarantees, thereby bridging a critical theoretical gap in understanding Transformer-based ICL on complex geometric domains.
📝 Abstract
Transformers have become a central architecture for in-context learning (ICL), particularly through their state-of-the-art performance in large language models. This success motivates understanding how transformers exploit task-relevant structure in geometrically heterogeneous data. However, existing nonparametric ICL theory has largely focused on Euclidean domains or single-manifold models. To address this gap, we study the prediction problem under unknown local geometry, modeled by sample size-dependent mixtures of manifolds with heterogeneous dimensions, smoothness, and sampling masses. Under local separation and small-perturbation conditions, we establish a minimax lower bound capturing the aggregate difficulty of the components and construct an oracle tangent local-polynomial estimator with a matching upper bound. This estimator is connected to a structure-informed, two-stage softmax transformer with a geometric preconditioner and chartwise reduced local-polynomial solvers. The transformer achieves negligible approximation error relative to the minimax rate with logarithmic depth and polynomial size. Finally, we derive an in-context generalization bound for near empirical risk minimizers over this class. Together, these results identify conditions under which the resulting predictor exploits local geometry and attains the aggregate minimax rate.