Learning Chaos Without Seeing Chaos: Extrapolation of Global Dynamics in Autoregressive Transformers

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether autoregressive models trained exclusively on local observations can recover the global dynamical structure across unseen parameter regimes. To this end, we train small autoregressive Transformers from scratch, encoding both states and parameters as continuous tokens. By combining closed-loop evaluation with causal interventions, we elucidate how narrow-window local behaviors support global generalization and reveal the pathways through which control parameters influence closed-loop dynamics via attention mechanisms. Our approach successfully reproduces period-doubling cascades and attractor structures. Notably, it estimates the Feigenbaum constant in the logistic map with an error below 5e-4 and achieves high-fidelity long-horizon chaotic prediction. These findings demonstrate that local learning can drive the generalization of global dynamical organization.
📝 Abstract
Autoregressive models are trained to predict a system's behavior one step at a time, and recursive generation allows the learned dynamics to unfold over long horizons. To what extent can such dynamics learned from local observations recover broader organization of an underlying system that was only partially observed during training? Here we study small autoregressive transformers trained from scratch on trajectories sampled from restricted parameter regimes of several non-linear dynamical systems, including logistic and sine maps, the Lorenz system, and the generalized Hopf system, with control parameters and state trajectories represented as sequences of continuous tokens. Under closed-loop evaluation at parameters far outside the training distribution, the models can recover self-similar period-doubling cascades, chaotic dynamics, and attractor structures with remarkable visual and numerical fidelity. For the logistic map, a transformer reproduces successive period doublings up to period 128, yielding a finite-order scaling ratio of 4.6687, matching the Feigenbaum constant to within $5\times10^{-4}$. We further investigate how these structures emerge over the course of training, and reveal with causal interventions how control-parameter information is processed through attention into state prediction and shapes the resulting closed-loop dynamics. These results suggest that a surprisingly narrow window into a system's local behavior may suffice for autoregressive transformers to generalize to its unseen global dynamical organization.
Problem

Research questions and friction points this paper is trying to address.

Autoregressive Transformers
Chaotic Dynamics
Global Dynamics Extrapolation
Nonlinear Dynamical Systems
Out-of-Distribution Generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Autoregressive Transformers
Chaos Extrapolation
Nonlinear Dynamical Systems
Causal Interventions
Feigenbaum Constant