๐ค AI Summary
This study addresses the issue that activation editing in large language models often compromises semantic and factual integrity when adjusting model values. To this end, it proposes an editable semantics-value interface over frozen residual states. The core innovation lies in constructing a unidirectional โsemantics-to-valuesโ pathway, where stop-gradient operations block backpropagation to ensure value editing does not interfere with semantic representations. Furthermore, selective encoded representations are trained by integrating swap consistency, subject deconfounding, and decorrelation techniques. Experimental results demonstrate that the proposed method significantly enhances semantic preservation and reduces benign refusals, achieving a BERTScore of 0.938 and lowering the contradiction rate to 5.1%. Overall, it comprehensively outperforms prompt-based baselines.
๐ Abstract
Value steering should change an LLM's normative priorities while preserving the scenario, facts, and task constraints underlying its answer. Conventional activation edits often change both. We introduce an editable semantic-value interface on frozen residual states, with a one-way semantic-to-value pathway that grounds value recognition in context. Stop-gradient blocks feedback through this pathway; swap consistency, topic de-confounding, and decorrelation encourage selective codes. At inference, editing the value code produces a residual delta while holding the semantic code fixed. On two instruction-tuned backbones, this interface improves semantic preservation and reduces benign refusals at comparable value alignment. A matched mixing-by-gating ablation separates representation learning from selective edit activation, and dimension-matched probes establish improved code selectivity. Against validation-selected prompting on LLaMA-3.1-8B, the method achieves comparable alignment (0.750 vs. 0.748), higher BERTScore (0.938 vs. 0.923), and fewer contradictions (5.1% vs. 7.6%). Human ratings and cross-taxonomy controls provide complementary evidence for low-damage value steering.