STR: Supervised Transcoder Replacement for Reducing Steering Side Effects

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue that existing model steering methods in large language models often compromise non-target capabilities when reinforcing target behaviors, thereby inducing unintended side effects. To mitigate this, we introduce replacement training into the steering pipeline for the first time, proposing a supervised transcoder replacement method. By fitting frozen weights and training an MLP replacement module via supervised learning, our approach decouples target behavior control from the preservation of non-target capabilities. Evaluations on benchmarks such as SALAD-Bench demonstrate that this method enables even unprotected steering mechanisms to significantly reduce side effects while maintaining generalization. On the Gemma model series, out-of-distribution attack success rates decrease from 42% to 14%, effectively balancing safety alignment with target steering performance.
📝 Abstract
Model steering can strengthen a target behavior while degrading other useful behaviors. We introduce Supervised Transcoder Replacement (STR) to reduce these side effects for existing steering methods, including those fitted without a protection objective. STR learns a replacement for the multilayer perceptron (MLP) computation at the steering layer through supervision for target control, non-target preservation, and fidelity without steering. Selected steering methods then fit directions on the frozen replacement while retaining their own fitting objectives. We evaluate three steering methods across Gemma and Llama models using Corrigibility preferences and four harmful-request safety datasets. SALAD-Bench supplies protection training data and a separate in-distribution evaluation split; HarmBench, AdvBench, and StrongREJECT are reserved for out-of-distribution testing. STR substantially reduces steering side effects on the in-distribution evaluation and extends this protection to the unseen safety datasets while retaining effective target control. For target-only supervised steering vectors, pooled out-of-distribution attack success rate falls from 42.46% to 14.42% on Gemma-3-4B and from 34.97% to 12.91% on Gemma-3-12B. These results show that replacement training can benefit steering methods fitted without protection objectives.
Problem

Research questions and friction points this paper is trying to address.

Model Steering
Side Effects
Large Language Models
Safety
Behavior Preservation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Supervised Transcoder Replacement
Model Steering
Side Effect Mitigation
MLP Replacement
Safety Alignment
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Haonan Yu
Haonan Yu
Research Scientist, Skild AI
RoboticsDeep Reinforcement LearningMultimodal Learning
J
Junhao Liu
Z
Zhenyu Yan
H
Haoran Lin
X
Xin Zhang