Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders

๐Ÿ“… 2026-05-07
๐Ÿ›๏ธ arXiv.org
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the preference instability and contradictory judgments exhibited by large language model reward models under semantics-preserving perturbations. For the first time, this work localizes vulnerable features at the representation level. By leveraging sparse autoencoders and latent space analysis to isolate instability factors, it proposes two training-free mitigation strategies: feature-guided intervention and residual correction. These methods dynamically suppress anomalous activations during inference and adaptively rectify preferences. Experimental results demonstrate that the proposed approach significantly reduces erroneous preference rates on harmlessness and hallucination benchmarks while preserving the general capabilities of the underlying models.
๐Ÿ“ Abstract
Preference learning in large language models relies on reward models as proxies for human judgment. However, these models frequently exhibit preference instability, producing contradictory preference assignments in response to subtle, meaning-preserving input variations. We analyze this instability at the representation level under three semantic-preserving perturbation types: paraphrasing, pattern injection, and backdoor triggers. We attribute this instability to over-reliance on predictive yet brittle features, which we term unstable features, and isolate them via Sparse Autoencoders (SAEs) in a sparse latent space where benign and perturbed inputs activate distinctly separable patterns. Building on this separability, we propose two SAE-based instability mitigation strategies: SAE Feature Steering, which identifies and suppresses anomalously activated features at inference, and SAE Residual Correction, which learns adaptive adjustments over SAE features to restore correct preferences. Our methods substantially reduce incorrect preference assignments on harmlessness and hallucination benchmarks while preserving benign performance and general utility on other tasks, without retraining the reward model.
Problem

Research questions and friction points this paper is trying to address.

Preference Instability
Reward Models
Large Language Models
Semantic-preserving Perturbations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Autoencoders
Preference Instability
Reward Models
Feature Steering
Residual Correction
๐Ÿ”Ž Similar Papers
S
Shunchang Liu
ETH Zรผrich, EPFL
X
Xin Chen
ETH Zรผrich
B
Belen Martin Urcelay
Georgia Institute of Technology
Francesco Croce
Francesco Croce
EPFL
machine learning