ðĪ AI Summary
This study addresses the unresolved question of why language models maintain robustness under activation perturbations, as their internal error-correction mechanisms remain poorly understood. By probing the geometric structure of activation spaces through low-dimensional curve analysis, this work reveals the existence of attractor basins and passive error-correction dynamics within these models. Building on these findings, it proposes an adaptive linear steering strategy to enhance cross-lingual control capabilities. The primary contribution lies in introducing geometric analysis into model interpretability and control research. This approach significantly increases the sampling probability of target-language tokens, thereby validating the effectiveness of geometry-based model interventions. Ultimately, this research establishes a novel paradigm for both understanding and manipulating language models.
ð Abstract
Language models exhibit remarkable robustness, continuing to produce coherent text even when their activations are perturbed by interventions like linear steering. We hypothesize that this robustness is a result of passive dynamics, i.e., constraining mechanisms in the forward pass that funnel activations toward"good"regions that produce coherent outputs. To investigate these hypothesized error-correcting mechanisms, we probe the geometry of language model activation space by observing the action of model layers on low-dimensional curves. In doing so, we discover that model activations occur within a cluster of distinct attracting basins, which differentiate natural activations geometrically from distributionally similar synthetic activations. Applying this lens to language model steering, we observe feature-specific basins along semantic steering directions, and find that steering moves activations between these basins. To demonstrate the active role of this geometry in neural computation, we show that adaptively modulating steering strength to transport activations across basins improves inter-language steering, significantly increasing the probability of sampling tokens from the target language compared to fixed-strength steering. Our findings establish analysis of activation space geometry as a promising approach to interpreting and controlling language models.