Values as Style: Disentangling Values from Semantics with One-Way Mixing for Low-Damage LLM Steering

๐Ÿ“… 2026-09-30
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the issue that activation editing in large language models often compromises semantic and factual integrity when adjusting model values. To this end, it proposes an editable semantics-value interface over frozen residual states. The core innovation lies in constructing a unidirectional โ€œsemantics-to-valuesโ€ pathway, where stop-gradient operations block backpropagation to ensure value editing does not interfere with semantic representations. Furthermore, selective encoded representations are trained by integrating swap consistency, subject deconfounding, and decorrelation techniques. Experimental results demonstrate that the proposed method significantly enhances semantic preservation and reduces benign refusals, achieving a BERTScore of 0.938 and lowering the contradiction rate to 5.1%. Overall, it comprehensively outperforms prompt-based baselines.
๐Ÿ“ Abstract
Value steering should change an LLM's normative priorities while preserving the scenario, facts, and task constraints underlying its answer. Conventional activation edits often change both. We introduce an editable semantic-value interface on frozen residual states, with a one-way semantic-to-value pathway that grounds value recognition in context. Stop-gradient blocks feedback through this pathway; swap consistency, topic de-confounding, and decorrelation encourage selective codes. At inference, editing the value code produces a residual delta while holding the semantic code fixed. On two instruction-tuned backbones, this interface improves semantic preservation and reduces benign refusals at comparable value alignment. A matched mixing-by-gating ablation separates representation learning from selective edit activation, and dimension-matched probes establish improved code selectivity. Against validation-selected prompting on LLaMA-3.1-8B, the method achieves comparable alignment (0.750 vs. 0.748), higher BERTScore (0.938 vs. 0.923), and fewer contradictions (5.1% vs. 7.6%). Human ratings and cross-taxonomy controls provide complementary evidence for low-damage value steering.
Problem

Research questions and friction points this paper is trying to address.

value steering
semantic preservation
LLM alignment
activation editing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Value steering
Semantic-value disentanglement
One-way mixing
Activation editing
Low-damage LLM steering
J
Jiale Dai
State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University
H
Hongcan Deng
University of Chinese Academy of Sciences
L
Liuxian Ma
College of Artificial Intelligence, Tsinghua University
X
Xiaoke Niu
State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University
Guojie Song
Guojie Song
Professor (Research), Tenured of Peking University
Psychological AIAI Safe & Value AlignmentAgent Cognition & Behavioral ModelingLLM&GML