Score
Design and implement metrics, experimental protocols, and analytic methods to quantify how agents’ or people’s beliefs shift in response to persuasive inputs, measuring magnitude and direction of change. Build comparative analyses that assess susceptibility across evidence sources, distinguish outcomes such as persuasion, backfire, and immunity, and create scalable measurement frameworks that generalize across claims and models.
This study challenges the prevailing assumption that model scale predominantly determines persuasive efficacy, instead investigating the cognitive foundations of persuasion dynamics between large language models (LLMs) and large inference models (LIMs) in multi-agent systems (MAS). We develop a controlled multi-agent experimental platform to systematically analyze how chain-of-thought sharing and reasoning depth influence persuasion outcomes. Our key contribution is the introduction of “persuasion duality”: explicit reasoning simultaneously enhances persuasiveness and strengthens individual resistance to persuasion, leading to attenuation of influence across multi-hop interactions. Empirical results demonstrate that transparent reasoning significantly improves persuasion efficacy; however, more advanced models exhibit heightened cognitive rigidity in group interactions. These findings uncover a fundamental trade-off between reasoning capability and persuasion robustness, offering a novel paradigm and empirical grounding for designing safe, interpretable MAS.
This work addresses the limited effectiveness of personalized response generation in persuasive dialogue systems. Methodologically, we propose the first generative framework integrating causal discovery, counterfactual reasoning, and variational latent variable modeling: (1) causal discovery is employed to identify strategy-level causal structures between user tactics and system responses; (2) system responses are modeled as intervenable counterfactual actions; and (3) user latent states are jointly inferred to enable dynamic personalization. Our key contribution lies in being the first to embed causal inference and counterfactual intervention mechanisms into dialogue policy learning—explicitly modeling how persuasion outcomes would change under alternative responses. Experiments on a real-world social welfare dataset demonstrate statistically significant improvements in cumulative reward (p < 0.01), validating that causally guided counterfactual modeling yields substantial gains in persuasive efficacy.
This work proposes a novel dialogue agent framework that integrates effective communication strategies from social psychology, behavioral economics, and communication studies to address the limitations of existing persuasive agents, which often rely on predefined tactics and struggle with the dynamic complexity of real-world interactions. By systematically unifying multidisciplinary persuasion mechanisms, the proposed approach significantly enhances persuasive efficacy—particularly for users with low initial willingness—and improves cross-scenario generalization. Leveraging interdisciplinary strategy modeling, dialogue system design, and advanced natural language understanding and generation techniques, the method achieves substantially higher persuasion success rates on both the Persuasion for Good and DailyPersuasion datasets, with especially strong performance on challenging cases.
To address the limitation of traditional belief revision models in capturing large-scale textual persuasion processes within social media environments, this study integrates psychological theory with large language models (LLMs) to develop an interpretable online persuasion detection framework. Methodologically, we design psychology-informed experiments to extract two key predictive factors—“cognitive emotion” and “sharing intention”—and leverage LLMs to generate fine-grained psychological feature scores, which are then fed into a random forest classifier to predict individual belief change. Compared to purely data-driven approaches, this hybrid paradigm significantly improves prediction accuracy, empirically confirming cognitive emotion and sharing intention as the strongest persuasion indicators. Our primary contribution is the first systematic incorporation of cognitive emotion mechanisms into LLM-based persuasion modeling, thereby reconciling theoretical interpretability with data-driven performance. This work provides a novel analytical tool and empirical foundation for influence assessment, misinformation detection, and narrative efficacy evaluation.
This work investigates the fundamental trade-off between rhetorical style and evidential quality in counterargument generation by large language models (LLMs). To this end, we introduce Counterfire, a dataset of 38,000 stylistically diverse counterarguments, and the first human-annotated argument triads—comprising claim, evidence, and rhetorical style—that enable joint modeling of style control and evidence integration. We evaluate six LLMs—including GPT-3.5 Turbo, GPT-4o, Claude Haiku, and LLaMA-3.1—on counterargument generation and multi-dimensional rhetorical quality assessment. Empirical results reveal that stylistic enhancement significantly improves persuasiveness (p < 0.01) but consistently reduces evidence density, confirming an intrinsic trade-off between these dimensions. Among all models, GPT-3.5 Turbo achieves the best performance yet remains substantially inferior to human-generated counterarguments. This work establishes a new benchmark for controllable and trustworthy argumentative generation and provides theoretical insights into the stylistic–evidential tension in LLM-based reasoning.
本文提出了一种新的分类测试框架,用于从定量估计中得出定性结论,以解决传统假设检验在区分竞争性假设上的不足。
This study addresses the challenges of information manipulation and aims to enhance AI safety and public health communication by proposing a 15-dimensional Persuasion Index (PI) grounded in psychological and communication theories. The framework leverages 55 lexicon- and rule-based subfeatures to construct a lightweight, interpretable, and modular system for assessing persuasiveness, featuring flexible component substitution, an open-source toolkit, and a visualization interface. Experimental evaluation across four heterogeneous datasets demonstrates that PI achieves both computational efficiency and strong predictive performance, uncovering cross-domain commonalities as well as topic-specific associations between persuasive dimensions and outcome variables. This work establishes a novel paradigm for transparent and auditable analysis of human–AI communication.
This study evaluates whether large language models (LLMs) can automatically generate questionnaires capable of effectively measuring social attitudes and rivaling established expert-designed scales. Through a within-subjects experimental design, the performance of GPT-4–generated questionnaires—elicited via structured prompts—was systematically compared against validated human-crafted scales across three domains: climate change, immigration, and diversity and inclusion. This work presents the first multi-domain, within-participant comparison between LLM-generated instruments and standard psychometric scales. Results indicate that LLM-generated questionnaires reliably capture major attitudinal divides and are suitable for exploratory, large-scale attitude assessment. However, they exhibit lower resolution in uncovering belief structures and reduced precision in differentiating subpopulations compared to expert-developed scales, suggesting promising yet supplementary utility in social science research.
研究探讨了在模型不确定性下,发送者如何根据私有信息战略性地传达叙述,以及接收者如何考虑发送者的动机调整其解读。实验验证了偏见强度对沟通的影响。
本文提出了一种框架来评估LLMs在接收新信息后是否能像人类一样更新观点,通过对比模型和人类在相同信息干预后的观点变化,发现现有模型均未能准确模拟人类的观点更新。