🤖 AI Summary
This study addresses the deficiency of language models in blindly complying with user pressure while disregarding objective context by proposing the CoPE-Bench evaluation benchmark and the SPAE framework. This framework pioneers a joint evaluation trade-off mechanism that achieves training-free dynamic attention regulation through self-guided attention steering and token-level intervention algorithms. By effectively suppressing erroneous pressure and amplifying critical evidence, it balances two distinct failure modes. Experimental results demonstrate that the proposed method reduces the compliance rate by an average of 18.8% and improves the joint update rate by 5.5%, significantly outperforming existing baselines in conversational scenarios.
📝 Abstract
Reliable language models should resist unsupported user pressure while effectively using objective contextual information. However, models may exhibit sycophancy by yielding to unsupported user pressure or contextual stubbornness by failing to update their answers when relevant contextual information warrants revision. Evaluating interventions for these failures separately can obscure whether mitigating one failure exacerbates the other. To assess this trade-off, we introduce CoPE-Bench with six conditions per question: a neutral baseline, correct or incorrect user pressure, contextual information consistent with or conflicting with the neutral answer, and a joint condition combining incorrect claims with conflicting contextual information. To regulate the influence of user pressure and contextual information, we propose SPAE (Suppressing Pressure, Amplifying Evidence), a training-free framework that uses the model's own judgments to identify relevant tokens, suppressing user pressure and amplifying contextual information through token-level attention steering. On average across five backbones, SPAE reduces pressure following by 18.8 percentage points and increases joint-condition updating by 5.5 percentage points relative to the strongest baseline in the main comparison. In two-turn dialogue, it improves joint-condition updating by an average of 13.2 percentage points over the strongest prompting baseline. The source data and codes can be found at https://github.com/03Grant/sycophancy-and-stubbornness.