What Makes You CLIC: Detection of Croatian Clickbait Headlines

📅 2025-07-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of clickbait detection in Croatian—a low-resource language—particularly in news headlines. Method: We introduce CLIC, the first Croatian clickbait dataset spanning two decades and covering both mainstream and fringe media outlets. We systematically compare fine-tuned Croatian BERTić against multilingual large language models (e.g., mT5, LLaMA) under zero-shot and few-shot prompting in Croatian and English. Contribution/Results: Fine-tuned BERTić significantly outperforms zero-shot and few-shot LLMs, demonstrating the efficacy of domain-adaptive fine-tuning in low-resource settings. CLIC reveals that nearly 50% of Croatian news headlines exhibit clickbait characteristics. This work establishes the first benchmark dataset and reproducible methodology for content credibility assessment in low-resource languages, enabling rigorous evaluation and advancement of clickbait detection for under-resourced linguistic communities.

Technology Category

Natural Language Processing: Fact-Checking / Misinformation Detection (NLP Focus)Application Domains: Misinformation & Fake NewsMachine Learning: Large Multimodal Models (LMMs)

Application Category

Web Mining and Content Analysis: Large pretrained models with web dataSearch and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchEconomics, Online Markets and Human Computation: LLM based quality controls for crowd work
📝 Abstract
Online news outlets operate predominantly on an advertising-based revenue model, compelling journalists to create headlines that are often scandalous, intriguing, and provocative -- commonly referred to as clickbait. Automatic detection of clickbait headlines is essential for preserving information quality and reader trust in digital media and requires both contextual understanding and world knowledge. For this task, particularly in less-resourced languages, it remains unclear whether fine-tuned methods or in-context learning (ICL) yield better results. In this paper, we compile CLIC, a novel dataset for clickbait detection of Croatian news headlines spanning a 20-year period and encompassing mainstream and fringe outlets. We fine-tune the BERTić model on this task and compare its performance to LLM-based ICL methods with prompts both in Croatian and English. Finally, we analyze the linguistic properties of clickbait. We find that nearly half of the analyzed headlines contain clickbait, and that finetuned models deliver better results than general LLMs.
Problem

Research questions and friction points this paper is trying to address.

Detecting clickbait in Croatian headlines automatically
Comparing fine-tuned models vs in-context learning methods
Analyzing linguistic traits of clickbait in Croatian media
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fine-tuned BERTić model for Croatian clickbait detection
Compared performance with LLM-based in-context learning
Analyzed linguistic properties of clickbait headlines
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Marija Anđedelić
University of Zagreb, Faculty of Electrical Engineering and Computing, TakeLab
D
Dominik Šipek
University of Zagreb, Faculty of Electrical Engineering and Computing, TakeLab
Laura Majer
Laura Majer
Researcher, Faculty of Electrical Engineering and Computing, University of Zagreb
J
Jan Šnajder
University of Zagreb, Faculty of Electrical Engineering and Computing, TakeLab