On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness

๐Ÿ“… 2026-07-31
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the tendency of chain-of-thought (CoT) reasoning to overlook the critical role of prompt-provided cues, often resulting in unfaithful inference. The authors propose using activation steering to explicitly encourage language models to acknowledge such cues within their CoT outputs. They systematically evaluate the generalization of this approach across diverse cue types, datasets, and methods for constructing steering vectors. Experiments on Gemma-3 4B/12B and Qwen-3.5 9B models reveal that significant improvements in cue acknowledgment occur only in Gemma-3 12B. While the intervention does not alter overall cue utilization rates, it effectively reduces implicit, undeclared reliance on cues, thereby enhancing expressive faithfulness rather than altering underlying decision logic. The consistent performance across multiple steering vector construction methods further demonstrates the robustness and generality of the proposed mechanism.
๐Ÿ“ Abstract
Model capabilities have improved in large part due to scaling chain of thought. This has been a promising development for AI safety--where models verbalize their reasoning, it is possible to monitor it. However, in some cases, models do not verbalize important steps in their reasoning process. For example, models prompted with a cue suggesting the incorrect answer may fail to acknowledge that cue, even when it appears instrumental to their conclusion. When chain of thought (CoT) fails to disclose instrumental reasoning steps, we describe it as unfaithful. Prior work has shown that activation steering can be a useful method to improve faithfulness in CoT. We extend this line of work by studying how well steering for faithfulness generalizes across cue types, datasets, and methods of constructing the steering vector for three models (Gemma-3 4B, Qwen-3.5 9B, Gemma-3 12B) in a cued question-answering setting. While steering reliably increases cue acknowledgment for only the largest model (Gemma-3 12B), we find that when steering is effective, its effect generalizes broadly across cue types and datasets--in cross-cue and cross-dataset analyses, effect size is determined primarily by the evaluation setting, rather than the vector's train setting. How the vector is built also matters little--four construction methods, including one whose optimization target mentions no specific cue, yield similar effect sizes. Finally, we consider the possibility that steering promotes the salience of the cue and causes greater cue use, rather than targeting verbalization behaviors. However, we find no evidence for this--steering leaves the rate of cue use roughly unchanged while reducing hidden cue use, i.e., cue use that is not acknowledged.
Problem

Research questions and friction points this paper is trying to address.

Chain-of-Thought Faithfulness
Activation Steering
Cue Acknowledgment
Model Interpretability
AI Safety
Innovation

Methods, ideas, or system contributions that make the work stand out.

activation steering
chain-of-thought faithfulness
cue acknowledgment
generalization
hidden cue use
๐Ÿ”Ž Similar Papers
๐Ÿ’ผ Related Jobs
No related jobs found.
M
Matthew Nguyen
University of Virginia
K
Kyle Cox
Independent
A
Austin Meek
University of Delaware
Ivรกn Arcuschin
Ivรกn Arcuschin
Independent Researcher
AI SafetyMechanistic Interpretability