🤖 AI Summary
This study addresses the scarcity of Arabic argument mining resources and the inherent challenges in discourse unit identification and classification by proposing the STAR-Ar system. The proposed method unifies detection and classification into a sequence labeling task, employing a BERT-BiLSTM-CRF architecture. Furthermore, it innovatively integrates contextual embeddings with a structural transfer constraint mechanism to achieve precise span detection of discourse units. Experimental results demonstrate that the model attains F1 scores of 73.70% and 72.69% on the test and validation sets, respectively, effectively enhancing argument mining performance in low-resource scenarios. The associated code has been made publicly available.
📝 Abstract
Argument Mining (AM) is a critical NLP task that remains significantly under-resourced in Arabic. This paper presents $\testtt{STAR-Ar}$, a BERT-BiLSTM-CRF architecture for argument discourse detection and classification, as our system for Daleel 2026, the inaugural Arabic argument mining shared task. The task requires the identification and classification of argumentative discourse units (ADUs) in debate and editorial texts.We jointly model these two objectives as a token-level sequence labeling task using a BERT-BiLSTM-CRF architecture that combines contextual transformer embeddings with structural transition constraints to support accurate span detection. $\testtt{STAR-Ar}$ achieves an F1-score of 72.69 on validation and 73.7 on test data. Our domain-specific analysis shows that models trained exclusively on editorials underperform those trained on debates, a disparity we primarily attribute to the smaller size of the editorial dataset. The code for $\testtt{STAR-Ar}$ is available at ${\href{https://github.com/ENTAILab/daleel_2026_Arabic-Argumentative-Discourse-Mining}{\faGithub~TTLab at Daleel 2026}}$