MentalMARBERT: Domain-Adaptive Pre-training and Two-Stage Fine-Tuning for Arabic Mental Health Disorders Detection

📅 2026-06-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of detecting mental health disorders in Arabic social media, which include dialectal diversity, informal language usage, scarcity of annotated data, and severe class imbalance. To tackle these issues, the authors propose a two-stage framework: first applying domain- and task-adaptive pretraining (DAPT/TAPT) to AraBERT, CAMeLBERT, and MARBERT, followed by fine-tuning using either full parameter updates or LoRA within both single-stage and hierarchical two-stage classification architectures. The work introduces the first high-quality, multi-class Arabic mental health tweet dataset, comprising 50,670 labeled instances. Experimental results demonstrate that the hierarchical two-stage model based on MARBERT with full fine-tuning—dubbed MentalMARBERT—achieves state-of-the-art performance, yielding a macro-F1 score of 0.861 and an accuracy of 0.877, significantly outperforming existing baselines.
📝 Abstract
Detecting mental health disorders from Arabic social media text remains challenging due to dialectal variation, informal language, limited high-quality annotated resources, and severe class imbalance. While English mental health natural language processing (NLP) has progressed substantially, Arabic multi-class disorder classification remains insufficiently studied. This study proposes a two-phase framework for Arabic mental health text classification. In phase 1, three Arabic pre-trained language models, AraBERT, CAMeLBERT, and MARBERT, undergo Domain-Adaptive and Task-Adaptive Pretraining (DAPT and TAPT) using a large-scale corpus of unlabeled Arabic mental health tweets. The adapted models are evaluated under a unified protocol to identify the most effective backbone model. In phase 2, the selected model is assessed across four configurations combining single-stage and hierarchical two-stage classification architectures with full fine-tuning and Low-Rank Adaptation (LoRA). To support this study, we constructed a novel annotated Arabic mental health dataset comprising 50,670 tweets across six categories, with strong inter annotator agreement (Krippendorff's Alpha = 0.733, average pairwise agreement = 0.797). Experimental results show that the domain-adapted MARBERT (MentalMARBERT) achieves statistically significant improvements over baseline models in both accuracy and macro-F1. The hierarchical two-stage architecture combined with full fine-tuning achieves the best overall performance, reaching a macro-F1 of 0.861 and an accuracy of 0.877. These findings demonstrate the effectiveness of domain-specific adaptive pretraining and hierarchical classification for Arabic mental health disorder detection.
Problem

Research questions and friction points this paper is trying to address.

Arabic mental health
social media text
dialectal variation
class imbalance
multi-class classification
Innovation

Methods, ideas, or system contributions that make the work stand out.

Domain-Adaptive Pretraining
Hierarchical Two-Stage Classification
Arabic Mental Health Detection
Low-Rank Adaptation (LoRA)
MentalMARBERT
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
F
Fatimah Almalki
Department of Computer Science, Faculty of Computing and Information Technology, King Abdulaziz University, Jeddah, Saudi Arabia
Areej Alhothali
Areej Alhothali
Associate Professor of Computer Science, King Abulaziz University
Machine learningNatural language processingAffective ComputingSentiment analysis
L
Lulwah Alharigy
Department of Computer Science, Faculty of Computing and Information Technology, King Abdulaziz University, Jeddah, Saudi Arabia
A
Abdulrahman Aladeem
Department of Psychology, College of Arts and Humanities, King Abdulaziz University, Jeddah, Saudi Arabia