Does Steering Break Your Model? A Multi-Dimensional Evaluation Suite for LLM Steering Methods

๐Ÿ“… 2026-10-06
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the fragmented evaluation landscape and unclear efficacyโ€“side-effect trade-offs in LLM activation steering by proposing SteerScope, a dual-axis, multidimensional evaluation framework. SteerScope introduces a novel paradigm that jointly characterizes outcome performance and method properties through a 15-metric benchmark suite, systematically quantifying target efficacy, side effects, and generalizability. Leveraging this framework, we benchmark 23 methods across four categories, including prompting and LoRA. Our analysis reveals coupling patterns between efficacy and side effects alongside notable disparities in sample efficiency. Crucially, we find that existing activation steering methods fail to surpass prompting baselines in overall balance, with side effects significantly exacerbated under out-of-distribution conditions.
๐Ÿ“ Abstract
Activation steering provides a lightweight and flexible way to control large language model (LLM) behavior. However, effective steering requires more than inducing the intended behavior: it should also limit unintended changes and remain robust across inputs and training data. Existing evaluations cover these dimensions only in fragments. As a result, the trade-offs between efficacy and side effects have not been systematically characterized. We introduce SteerScope, a two-axis, multi-dimensional evaluation suite that jointly characterizes steering outcomes and method properties through 15 metrics. We score target efficacy and side effects on language quality, task capabilities, and safety and reliability, and further assess generalization and data dependence through steering-specific metrics for sample efficiency and sample sensitivity. Rather than comparing methods at a single operating point, we characterize the trade-offs between efficacy and side effects. Under matched models, tasks, and evaluation protocols, we benchmark 23 methods spanning 4 families, including prompting, LoRA, and SFT as baseline methods, and release the suite as an extensible codebase. We find that current activation steering methods do not yet surpass the Prompt Steering baseline in their overall balance between steering efficacy and side effects: across both model scales, no evaluated activation steering method achieves higher efficacy without incurring greater composite side effects. We further uncover a consistent coupling between steering efficacy and side effects. Under OOD prompts, target efficacy is often preserved, whereas side effects tend to become more pronounced, particularly through declines in instruction relevance and fluency. Methods also exhibit sharply different sample-efficiency profiles.
Problem

Research questions and friction points this paper is trying to address.

Activation Steering
Large Language Models
Evaluation Suite
Side Effects
Efficacy Trade-offs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Activation Steering
Evaluation Suite
Large Language Models
Multi-dimensional Metrics
Trade-off Analysis
Haotian Yang
Haotian Yang
Kuaishou Technology
computer visioncomputer graphics
H
Huikang Jiang
School of Computer Science, Peking University; Columbia University
Y
Yucheng Wu
State Key Laboratory of Multimedia Information Processing, Peking University; School of Computer Science, Peking University
W
Wen-Jie Jiang
School of Computer Science, Peking University
C
Chenpeng Wang
YiXin-AILab, YIXIN
Y
Yibin Lou
School of Computer Science, Peking University; Southern University of Science and Technology
Liangming Pan
Liangming Pan
Assistant Professor, School of Computer Science, Peking University
Natural Language ProcessingLarge Language ModelsMachine Learning