Persuasion Dynamics in LLMs: Investigating Robustness and Adaptability in Knowledge and Safety with DuET-PD

📅 2025-08-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Large language models (LLMs) struggle to simultaneously resist misinformation and remain receptive to valid corrections in multi-turn persuasive dialogues, hindering their reliable deployment. To address this, we propose DuET-PD, the first dual-dimensional evaluation framework that systematically characterizes dynamic stance robustness across knowledge integrity and safety compliance. We further introduce Holistic DPO, a novel training method that jointly optimizes resistance to manipulation and openness to correction. Leveraging MMLU-Pro and SALAD-Bench, we construct a multi-turn adversarial persuasion dataset. Empirical results show that GPT-4o’s knowledge accuracy drops to 27.32% under sustained adversarial persuasion; in contrast, Llama-3.1-8B-Instruct trained with Holistic DPO achieves a substantial improvement in safety-scenario accuracy—from 4.21% to 76.54%—demonstrating the method’s effectiveness and generalizability in balancing reliability and adaptability.

Technology Category

Natural Language Processing: Safety and RobustnessMachine Learning: Large Multimodal Models (LMMs)Philosophy and Ethics of AI: Safety, Robustness & Trustworthiness

Application Category

Search and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsUser Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systems
📝 Abstract
Large Language Models (LLMs) can struggle to balance gullibility to misinformation and resistance to valid corrections in persuasive dialogues, a critical challenge for reliable deployment. We introduce DuET-PD (Dual Evaluation for Trust in Persuasive Dialogues), a framework evaluating multi-turn stance-change dynamics across dual dimensions: persuasion type (corrective/misleading) and domain (knowledge via MMLU-Pro, and safety via SALAD-Bench). We find that even a state-of-the-art model like GPT-4o achieves only 27.32% accuracy in MMLU-Pro under sustained misleading persuasions. Moreover, results reveal a concerning trend of increasing sycophancy in newer open-source models. To address this, we introduce Holistic DPO, a training approach balancing positive and negative persuasion examples. Unlike prompting or resist-only training, Holistic DPO enhances both robustness to misinformation and receptiveness to corrections, improving Llama-3.1-8B-Instruct's accuracy under misleading persuasion in safety contexts from 4.21% to 76.54%. These contributions offer a pathway to developing more reliable and adaptable LLMs for multi-turn dialogue. Code is available at https://github.com/Social-AI-Studio/DuET-PD.
Problem

Research questions and friction points this paper is trying to address.

Evaluating LLM robustness to persuasive misinformation and corrections
Assessing stance-change dynamics in knowledge and safety domains
Addressing sycophancy and improving persuasion resistance in dialogues
Innovation

Methods, ideas, or system contributions that make the work stand out.

Introduces DuET-PD framework for dual-dimension persuasion evaluation
Proposes Holistic DPO training balancing positive and negative examples
Enhances robustness to misinformation and receptiveness to corrections