๐ค AI Summary
This study addresses the preference instability and contradictory judgments exhibited by large language model reward models under semantics-preserving perturbations. For the first time, this work localizes vulnerable features at the representation level. By leveraging sparse autoencoders and latent space analysis to isolate instability factors, it proposes two training-free mitigation strategies: feature-guided intervention and residual correction. These methods dynamically suppress anomalous activations during inference and adaptively rectify preferences. Experimental results demonstrate that the proposed approach significantly reduces erroneous preference rates on harmlessness and hallucination benchmarks while preserving the general capabilities of the underlying models.
๐ Abstract
Preference learning in large language models relies on reward models as proxies for human judgment. However, these models frequently exhibit preference instability, producing contradictory preference assignments in response to subtle, meaning-preserving input variations. We analyze this instability at the representation level under three semantic-preserving perturbation types: paraphrasing, pattern injection, and backdoor triggers. We attribute this instability to over-reliance on predictive yet brittle features, which we term unstable features, and isolate them via Sparse Autoencoders (SAEs) in a sparse latent space where benign and perturbed inputs activate distinctly separable patterns. Building on this separability, we propose two SAE-based instability mitigation strategies: SAE Feature Steering, which identifies and suppresses anomalously activated features at inference, and SAE Residual Correction, which learns adaptive adjustments over SAE features to restore correct preferences. Our methods substantially reduce incorrect preference assignments on harmlessness and hallucination benchmarks while preserving benign performance and general utility on other tasks, without retraining the reward model.