🤖 AI Summary
This study addresses the factors governing the steerability of latent structures in language models, highlighting the limitations of traditional linear probes in cross-model steering. It decomposes causal influence into three components—capacity, responsiveness, and alignment—to elucidate their mechanisms as interpretability constraints and their context dependence. To overcome these bottlenecks, this work proposes a low-dimensional subspace-based causal probe that integrates activation space analysis with constrained subspace training. Empirically, the proposed probe improves cross-model steering performance by 17%–118% while incurring only marginal degradation in concept detection accuracy. These findings establish an efficient new paradigm for causal intervention and controllable generation in language models.
📝 Abstract
Localizing latent structures in the activation space of language models (LMs) is central to understanding and controlling their behavior. Yet, localized structures can differ substantially in their causal influence, raising the question of what makes a structure actionable. We tackle this question by casting causal influence as a product of three factors and showing empirically that they act as interpretable, distinct constraints: capacity, measuring the sensitivity of the model's output to movement along the structure, responsiveness, capturing how promotable the concept is given the current context, and alignment, reflecting how well the structure aligns with the context-specific representation of the concept. Across 4 LM families and 50 concepts, we observe that causal effectiveness requires all factors to be high; low capacity and responsiveness reduce it by 84% and 95%, respectively, while low alignment can reverse it, suppressing concept expression. Moreover, we find that causality is context-dependent rather than an intrinsic property of the structure, with causally effective directions forming a low-dimensional subspace that varies across contexts. By restricting the training of linear probes to this subspace, we introduce causal probes that achieve 17%-118% improvement in steering across models, with only 3% reduction in concept detection.