🤖 AI Summary
This work addresses the instability of Test-Time Training (TTT) under distribution shift, which stems from its sensitivity to hyperparameters and the absence of theoretical guidance. The authors reinterpret TTT through the lens of decision theory as implicit Bayesian inference under a kernel mechanism. They propose a PAC-Bayes–guaranteed method that adaptively selects the number of update steps based on prompt evidence and characterize the Bayes-optimal update subspace within a linear Gaussian correction model to inform Transformer module selection. By integrating Gaussian processes, spectral analysis, and Bayesian inference, the study establishes a theoretical framework for TTT, revealing conditions under which fixed update strategies fail and providing principled foundations for adaptive update directions and step sizes, thereby effectively mitigating TTT’s instability.
📝 Abstract
Test-time training (TTT) adapts a pretrained model to each prompt via parameter updates, improving accuracy under pretraining-to-test distribution shifts. Yet, its performance often suffers from instability and sensitivity to hyperparameters such as update steps and subspace. We explain this behavior through a decision-theoretic lens, treating TTT as implicit Bayesian inference in the kernel regime. Under a Gaussian process benchmark, we show that TTT reduces prediction error when updates are spectrally matched to the prompt's signal-to-noise ratio and aligned with query-relevant eigen-directions. This perspective underpins the following results: (1) we show when fixed update steps and subspaces fail under distribution shifts, motivating adaptive strategies; (2) we prove that selecting update steps via prompt evidence admits a PAC-Bayes guarantee against overfitting; and (3) we characterize the Bayes-optimal update subspace under a linear-Gaussian correction model, yielding a scoring rule for selecting Transformer blocks and heads. Our theory helps explain the empirical instability of TTT, taking a step toward principled guidance for when, how far, and which directions to adapt.