🤖 AI Summary
This study addresses a critical AI safety vulnerability in debate-based supervision, wherein human judgments are susceptible to manipulation by rhetorical strategies, causing models to optimize for persuasiveness rather than truthfulness. To investigate this, the authors construct a detective-mystery dialogue dataset incorporating interventions such as anchoring effects and logical fallacies, leveraging large language models for structured data generation. Through controlled experiments and behavioral statistical analyses involving 369 participants, this work provides the first systematic quantification of how four distinct rhetorical interventions bias debate adjudication, challenging conventional assumptions. The empirical findings demonstrate that cognitive biases significantly distort human judgment, revealing that current debate paradigms cannot effectively withstand adversarial rhetorical attacks. Ultimately, this research establishes cognitive bias as a fundamental threat to debate safety.
📝 Abstract
Reinforcement learning from human feedback (RLHF) has played a central role in making large language models responsive to human instructions. However, human evaluators often favor flattering or persuasive responses over truthful ones, creating incentives for models to appeal to evaluators at the expense of accuracy. AI safety via debate has been proposed as a way to improve the supervision of language models: in this paradigm, two agents argue opposing positions and challenge each other's claims, potentially exposing falsehoods to the adjudicator. A central premise of AI safety via debate is that truthful arguments are easier to defend than false ones under adversarial scrutiny. In this work, we investigate whether this advantage persists when debaters use rhetorical strategies that exploit biases in human judgment. Inspired by competitive debate, we construct 68 LLM-generated dialogues about detective mysteries with known culprits, spanning four interventions: anchoring, fallacy oversight, pro-jargon, and verbosity. We apply each intervention to either the side advocating for the true culprit or the side advocating for an innocent suspect, allowing us to distinguish influence on adjudication from correctness. In a study with 369 participants, we find that, pooled across bias types, these interventions significantly shift judgments toward the manipulated side. These findings expose a vulnerability in debate-based supervision: human adjudication is sensitive to manipulative rhetorical strategies.