SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset

📅 2025-08-31
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of sentence-level sarcasm detection in Romanian news texts, where sarcastic statements are frequently misclassified as factual reporting. To this end, we introduce RoSarcasm—the first news-domain Romanian sarcasm detection dataset—comprising 13,873 manually annotated, cross-domain sentences. We design a fine-grained annotation schema and conduct systematic zero-shot and fine-tuning experiments to evaluate state-of-the-art large language models and Transformer-based baselines. Results reveal significant performance limitations across all models, underscoring the difficulty of sarcasm identification in low-resource languages. Our contribution is threefold: (1) RoSarcasm fills a critical gap in non-English sarcasm detection resources; (2) it establishes a reproducible benchmark with standardized annotation guidelines; and (3) it provides an empirical analysis framework for future research on figurative language understanding in under-resourced languages. This work lays foundational groundwork for advancing sarcasm detection in low-resource linguistic settings.

Technology Category

Natural Language Processing: Fact-Checking / Misinformation Detection (NLP Focus)Knowledge Representation and Reasoning: ArgumentationMachine Learning: Large Multimodal Models (LMMs)

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web dataEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasets
📝 Abstract
Satire, irony, and sarcasm are techniques typically used to express humor and critique, rather than deceive; however, they can occasionally be mistaken for factual reporting, akin to fake news. These techniques can be applied at a more granular level, allowing satirical information to be incorporated into news articles. In this paper, we introduce the first sentence-level dataset for Romanian satire detection for news articles, called SeLeRoSa. The dataset comprises 13,873 manually annotated sentences spanning various domains, including social issues, IT, science, and movies. With the rise and recent progress of large language models (LLMs) in the natural language processing literature, LLMs have demonstrated enhanced capabilities to tackle various tasks in zero-shot settings. We evaluate multiple baseline models based on LLMs in both zero-shot and fine-tuning settings, as well as baseline transformer-based models. Our findings reveal the current limitations of these models in the sentence-level satire detection task, paving the way for new research directions.
Problem

Research questions and friction points this paper is trying to address.

Detecting satire at sentence level in Romanian news
Addressing confusion between satire and factual reporting
Evaluating LLM capabilities for satire detection tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

First sentence-level Romanian satire dataset
Evaluated LLMs in zero-shot and fine-tuning settings
Used transformer-based models for satire detection
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Răzvan-Alexandru Smădu
Răzvan-Alexandru Smădu
National University of Science and Technology POLITEHNICA Bucharest
Adversarial LearningDomain Adaptation
A
Andreea Iuga
National University of Science and Technology POLITEHNICA Bucharest, Bucharest, Romania
Dumitru-Clementin Cercel
Dumitru-Clementin Cercel
Teaching Assistant of Computer Science, University Politehnica of Bucharest
Social Network AnalysisNatural Language ProcessingInformation RetrievalMachine Learning
F
Florin Pop
National University of Science and Technology POLITEHNICA Bucharest, Bucharest, Romania; National Institute for Research & Development in Informatics – ICI Bucharest, Bucharest, Romania