Unmasking Propaganda: A Comparative Analysis of Masked and Causal Language Models

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of identifying covert and context-dependent propaganda techniques by systematically comparing masked language models, such as XLM-RoBERTa and DeBERTa V3, with causal language models on the SemEval dataset. It thoroughly evaluates the impact of basic prompting and chain-of-thought strategies, revealing architecture-specific strengths and weaknesses across distinct propaganda techniques. Experimental results demonstrate that the optimal model achieves an F1 score of 63.62, surpassing existing state-of-the-art performance. Furthermore, the findings confirm that fine-tuning and ensemble methods can substantially enhance detection capabilities. Overall, this work provides an effective paradigm for distinguishing manipulative language from legitimate persuasion in political discourse.
📝 Abstract
Propaganda detection is an essential task in natural language processing (NLP), particularly in the context of manipulative political communications. However, identifying specific propaganda techniques presents a significant challenge due to their often subtle nature and reliance on context, making them difficult to distinguish from legitimate persuasive language. Propaganda often involves highlighting certain facts while downplaying or ignoring others to create a desired perception. This biased communication aims to influence attitudes, beliefs, or behaviors towards a particular cause or position. This paper explores advances in detecting propaganda techniques through a comparative analysis of modern language models, using the SemEval-2020 Task 11 dataset. We evaluated both masked language models (based on XLM-RoBERTa or DeBERTa V3) and causal models (from OpenAI, Google, Mistral, Anthropic and Meta), employing two prompting strategies: base and chain-of-thought prompting. Our results demonstrate improvements over state-of-the-art models, with the best-performing MLM achieving an F1 score of 63.18 in technique classification and the best causal model achieving 63.62. We also observed that certain models excel in specific techniques, such as loaded language and name-calling, while struggling with others like bandwagon and black-and-white fallacy. These findings suggest that fine-tuning, ensemble modeling, and the use of larger datasets can further enhance propaganda detection capabilities.
Problem

Research questions and friction points this paper is trying to address.

Propaganda detection
Natural language processing
Language models
Technique classification
Innovation

Methods, ideas, or system contributions that make the work stand out.

Propaganda Detection
Masked Language Models
Causal Language Models
Chain-of-Thought Prompting
Comparative Analysis
💼 Related Jobs
No related jobs found.
C
Claudiu Creanga
Faculty of Mathematics and Informatics, Human Language Technologies Research Center, Interdisciplinary School of Doctoral Studies, University of Bucharest, Romania
I
Ioachim Lihor
Faculty of Mathematics and Informatics, Human Language Technologies Research Center, Interdisciplinary School of Doctoral Studies, University of Bucharest, Romania
Liviu P. Dinu
Liviu P. Dinu
Professor, University of Bucharest, Dept. of Computer Science,
Computational LinguisticsNatural Language ProcessingComputational Historical Linguistics