Counterspeech the ultimate shield! Multi-Conditioned Counterspeech Generation through Attributed Prefix Learning

πŸ“… 2025-05-17
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
To address the limited effectiveness of counter-hate speech generation caused by insufficient multi-attribute collaborative control, this paper proposes HiPPrOβ€”a novel framework integrating hierarchical prefix embedding learning with reward-free preference optimization (RFPO), enabling the first joint modeling of intent and emotion attributes. We introduce IntentCONANv2-emotion, the first large-scale counter-speech dataset with fine-grained, multi-annotator emotion labels. Evaluation employs both ROUGE metrics and human assessment. Experiments demonstrate a 38% improvement in intent adherence rate and absolute gains of 3%, 2%, and 3% in ROUGE-1, ROUGE-2, and ROUGE-L, respectively. Human evaluation further confirms significant improvements in relevance, appropriateness, and constructiveness of generated outputs. This work establishes a new paradigm for fine-grained, controllable counter-hate speech generation.

Technology Category

Natural Language Processing: GenerationHumans and AI: Learning Human Values and PreferencesCognitive Modeling & Cognitive Systems: Affective Computing

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systemsEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasets
πŸ“ Abstract
Counterspeech has proven to be a powerful tool to combat hate speech online. Previous studies have focused on generating counterspeech conditioned only on specific intents (single attributed). However, a holistic approach considering multiple attributes simultaneously can yield more nuanced and effective responses. Here, we introduce HiPPrO, Hierarchical Prefix learning with Preference Optimization, a novel two-stage framework that utilizes the effectiveness of attribute-specific prefix embedding spaces hierarchically optimized during the counterspeech generation process in the first phase. Thereafter, we incorporate both reference and reward-free preference optimization to generate more constructive counterspeech. Furthermore, we extend IntentCONANv2 by annotating all 13,973 counterspeech instances with emotion labels by five annotators. HiPPrO leverages hierarchical prefix optimization to integrate these dual attributes effectively. An extensive evaluation demonstrates that HiPPrO achieves a ~38 % improvement in intent conformity and a ~3 %, ~2 %, ~3 % improvement in Rouge-1, Rouge-2, and Rouge-L, respectively, compared to several baseline models. Human evaluations further substantiate the superiority of our approach, highlighting the enhanced relevance and appropriateness of the generated counterspeech. This work underscores the potential of multi-attribute conditioning in advancing the efficacy of counterspeech generation systems.
Problem

Research questions and friction points this paper is trying to address.

Generating counterspeech with multiple attributes simultaneously
Improving intent conformity and relevance in counterspeech
Enhancing counterspeech effectiveness through hierarchical prefix optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical prefix learning optimizes attribute embedding
Preference optimization enhances counterspeech constructiveness
Multi-attribute conditioning improves intent and emotion conformity
πŸ’Ό Related Jobs
No related jobs found.
Aswini Kumar Padhi
Aswini Kumar Padhi
Ph.D Scholar (Indian Institute of Technology, Delhi)
Natural Language Processing
A
Anil Bandhakavi
Logically.ai
T
Tanmoy Chakraborty
IIT Delhi, India