Language Model Activations Inhabit Privileged Error-Correcting Basins

📅 2026-10-03
📈 Citations: 0
âœĻ Influential: 0
📄 PDF
ðŸĪ– AI Summary
This study addresses the unresolved question of why language models maintain robustness under activation perturbations, as their internal error-correction mechanisms remain poorly understood. By probing the geometric structure of activation spaces through low-dimensional curve analysis, this work reveals the existence of attractor basins and passive error-correction dynamics within these models. Building on these findings, it proposes an adaptive linear steering strategy to enhance cross-lingual control capabilities. The primary contribution lies in introducing geometric analysis into model interpretability and control research. This approach significantly increases the sampling probability of target-language tokens, thereby validating the effectiveness of geometry-based model interventions. Ultimately, this research establishes a novel paradigm for both understanding and manipulating language models.
📝 Abstract
Language models exhibit remarkable robustness, continuing to produce coherent text even when their activations are perturbed by interventions like linear steering. We hypothesize that this robustness is a result of passive dynamics, i.e., constraining mechanisms in the forward pass that funnel activations toward"good"regions that produce coherent outputs. To investigate these hypothesized error-correcting mechanisms, we probe the geometry of language model activation space by observing the action of model layers on low-dimensional curves. In doing so, we discover that model activations occur within a cluster of distinct attracting basins, which differentiate natural activations geometrically from distributionally similar synthetic activations. Applying this lens to language model steering, we observe feature-specific basins along semantic steering directions, and find that steering moves activations between these basins. To demonstrate the active role of this geometry in neural computation, we show that adaptively modulating steering strength to transport activations across basins improves inter-language steering, significantly increasing the probability of sampling tokens from the target language compared to fixed-strength steering. Our findings establish analysis of activation space geometry as a promising approach to interpreting and controlling language models.
Problem

Research questions and friction points this paper is trying to address.

Language Models
Activation Space Geometry
Error-Correcting Basins
Model Robustness
Linear Steering
Innovation

Methods, ideas, or system contributions that make the work stand out.

Activation Space Geometry
Error-Correcting Basins
Passive Dynamics
Language Model Steering
Attracting Basins
M
Matthew Finlayson
Goodfire
F
Francisco Pernice
Goodfire, Massachusetts Institute of Technology
Eric Todd
Eric Todd
PhD Student at Northeastern University
Machine LearningModel Interpretability
Amir Zur
Amir Zur
Stanford University
Natural Language ProcessingModel Interpretability
Daniel Wurgaft
Daniel Wurgaft
PhD Student, Stanford University
Cognitive ScienceGeneralizationIn-Context LearningReasoningLanguage Models
F
Fenil R. Doshi
Goodfire, Harvard University
Vasudev Shyam
Vasudev Shyam
Postdoctoral Fellow, Stanford University
High Energy Physics
Matt Feiszli
Matt Feiszli
Facebook AI Research
Machine LearningComputer VisionHarmonic AnalysisGeometry
Satchel Grant
Satchel Grant
PhD Student, Stanford University
interpretabilitynumeric reasoningalignment
Lucius Bushnaq
Lucius Bushnaq
Research Scientist, Apollo Research
Tal Haklay
Tal Haklay
PhD student, Technion
Usha Bhalla
Usha Bhalla
Ph.D. Student, Harvard University
Machine Learning Interpretability
Matthew Kowal
Matthew Kowal
FAR AI, York University, Vector Institute
Machine LearningInterpretabilityComputer Vision
Thomas Fel
Thomas Fel
Kempner Fellow, Harvard University
InterpretabilityDeep LearningComputer VisionVision InterpretabilityNeuro Interpretability
Jack Merullo
Jack Merullo
Brown University
interpretabilitylanguage modelsnatural language processingmultimodal learning
Atticus Geiger
Atticus Geiger
Pr(Ai)ÂēR Group
Artificial IntelligenceNatural LanguageMechanistic InterpretabilityCausality
Xiang Ren
Xiang Ren
ExxonMobil
Computational Mechanics
O
Owen Lewis
Goodfire
Ekdeep Singh Lubana
Ekdeep Singh Lubana
Goodfire AI
AIMachine LearningDeep Learning