π€ AI Summary
To address the limited effectiveness of counter-hate speech generation caused by insufficient multi-attribute collaborative control, this paper proposes HiPPrOβa novel framework integrating hierarchical prefix embedding learning with reward-free preference optimization (RFPO), enabling the first joint modeling of intent and emotion attributes. We introduce IntentCONANv2-emotion, the first large-scale counter-speech dataset with fine-grained, multi-annotator emotion labels. Evaluation employs both ROUGE metrics and human assessment. Experiments demonstrate a 38% improvement in intent adherence rate and absolute gains of 3%, 2%, and 3% in ROUGE-1, ROUGE-2, and ROUGE-L, respectively. Human evaluation further confirms significant improvements in relevance, appropriateness, and constructiveness of generated outputs. This work establishes a new paradigm for fine-grained, controllable counter-hate speech generation.
π Abstract
Counterspeech has proven to be a powerful tool to combat hate speech online. Previous studies have focused on generating counterspeech conditioned only on specific intents (single attributed). However, a holistic approach considering multiple attributes simultaneously can yield more nuanced and effective responses. Here, we introduce HiPPrO, Hierarchical Prefix learning with Preference Optimization, a novel two-stage framework that utilizes the effectiveness of attribute-specific prefix embedding spaces hierarchically optimized during the counterspeech generation process in the first phase. Thereafter, we incorporate both reference and reward-free preference optimization to generate more constructive counterspeech. Furthermore, we extend IntentCONANv2 by annotating all 13,973 counterspeech instances with emotion labels by five annotators. HiPPrO leverages hierarchical prefix optimization to integrate these dual attributes effectively. An extensive evaluation demonstrates that HiPPrO achieves a ~38 % improvement in intent conformity and a ~3 %, ~2 %, ~3 % improvement in Rouge-1, Rouge-2, and Rouge-L, respectively, compared to several baseline models. Human evaluations further substantiate the superiority of our approach, highlighting the enhanced relevance and appropriateness of the generated counterspeech. This work underscores the potential of multi-attribute conditioning in advancing the efficacy of counterspeech generation systems.