Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech

📅 2026-07-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing AI-generated counter-speech, which often relies on generic strategies that ignore conversational context and user-specific characteristics, thereby failing to effectively mitigate online toxicity. The authors propose a lightweight, personalized counter-speech generation approach that integrates both dialogue context and user history, and systematically evaluate the impact of various fine-tuning strategies. Employing a hybrid evaluation framework combining automatic metrics (ROUGE, BLEU, BERTScore) with a preregistered crowdsourced human assessment, the study finds that lightweight contextual modeling significantly enhances perceived counter-speech quality and persuasiveness, whereas certain more complex strategies yield diminishing returns or even degrade performance. These findings elucidate key factors for effective personalized interventions and offer a promising direction for responsible content moderation.
📝 Abstract
AI-generated counterspeech offers a scalable and effective strategy to mitigate online toxicity by promoting more constructive dialogue. Yet, existing approaches adopt a generic, one-size-fits-all paradigm, overlooking the conversational context and characteristics of the targeted users. Here, we propose and evaluate multiple strategies for generating contextualized counterspeech that is adapted to the moderation setting and personalized to the moderated user. In detail, we explore a range of configurations that integrate different forms of contextual information and fine-tuning techniques. We conduct a comprehensive evaluation combining quantitative indicators with a pre-registered, mixed-design crowdsourcing experiment. To ensure robustness, we implement algorithmic measures of counterspeech quality based on ROUGE, BLEU, and BERTScore, observing overall consistent results across metrics. Furthermore, we analyze which characteristics of both the generated counterspeech and the moderated toxic message most strongly influence perceived persuasiveness, yielding insights into how contextualized interventions can be made more effective. Our findings show that personalization can be effective, but not uniformly so. Lightweight strategies combining conversational context and user history improve perceived adequacy and persuasiveness, whereas several other contextualization strategies degrade human-perceived counterspeech quality. Taken together, these results provide actionable directions for developing more personalized, effective, and responsible counterspeech systems, ultimately advancing human-AI collaboration in online content moderation.
Problem

Research questions and friction points this paper is trying to address.

counterspeech
online toxicity
contextualization
personalization
content moderation
Innovation

Methods, ideas, or system contributions that make the work stand out.

contextualized counterspeech
personalized AI moderation
conversational context
fine-tuning strategies
persuasiveness evaluation