Dynamic Mode Decomposition along Depth in Vision Transformers

📅 2026-05-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates whether the deep architecture of Vision Transformers (ViTs) can be modeled as an approximately autonomous linear dynamical system. By applying Dynamic Mode Decomposition (DMD) to consecutive hidden states, the authors fit a unified recurrence operator \( K \) and use \( K^p \) to predict subsequent states, marking the first application of DMD along the depth dimension of ViTs. The study reveals that early layers admit low-rank compression and that the CLS token is more amenable to linearization. For short prediction horizons (\( p \leq 4 \)), the cosine similarity error between \( K^p \) and the true state transition remains below 0.02, enabling accurate reconstruction of intermediate activations. However, this linear fidelity does not reliably transfer to downstream tasks, and the predictive advantage vanishes in the final layers.
📝 Abstract
Recent work has shown that contiguous vision transformer (ViT) blocks (a) can be replaced by a linear map and (b) organize into recurrent phases of computation. We ask whether these observations coincide: does ViT depth implement approximately \textit{autonomous linear} dynamics, admitting a single operator $K$ applied recurrently across a contiguous span? We test this using Dynamic Mode Decomposition (DMD), which fits $K$ from selected, consecutive hidden-state pairs and predicts $p$ steps ahead via $K^p$. On four pretrained DINO ViTs, we study the regularization, rank, and calibration budget required for stable fitting. For short spans ($p \leq 4$), $K^p$ tracks an unconstrained endpoint map to within $0.02$ cosine similarity on DINOv3-H/16+, while also recovering intermediate activations at each skipped block. At early cut starts, the fitted operators compress to rank $\ll d$ with minimal calibration data, and across tokens, \texttt{cls} is most amenable to linearization; both properties decay monotonically with depth. Yet this local fidelity does not transfer downstream. At the final hidden state, after propagating through the remaining blocks, an identity baseline becomes competitive.
Problem

Research questions and friction points this paper is trying to address.

Vision Transformers
Dynamic Mode Decomposition
linear dynamics
depth-wise computation
autonomous systems
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic Mode Decomposition
Vision Transformers
Linear Dynamics
Depth-wise Recurrence
Low-rank Approximation
🔎 Similar Papers
No similar papers found.