Detecting Toxic Language: Ontology and BERT-based Approaches for Bulgarian Text

📅 2026-04-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the frequent misclassification of medical terminology and minority-related discourse as harmful content by online moderation systems. Focusing on Bulgarian, the work presents the first toxicity language ontology for the language and introduces a novel fine-grained annotated dataset comprising four categories designed to preserve sensitive yet non-toxic content. By integrating ontology-guided constraints with BERT fine-tuning, the authors train a model on 4,384 manually labeled sentences, achieving a macro-averaged F1 score of 0.89. The proposed approach effectively distinguishes genuinely toxic utterances from critical non-toxic information, offering a deployable solution that significantly enhances both the accuracy and inclusivity of real-world content moderation systems.

Technology Category

Natural Language Processing: Safety and RobustnessKnowledge Representation and Reasoning: OntologiesMachine Learning: Multimodal Learning

Application Category

Web Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web dataSearch and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchSemantics and Knowledge: Scalable techniques for the creation, curation, publication, maintenance, and consumption of large, Web-based, structured, reusable, knowledge graphs and ontologies
📝 Abstract
Toxic content detection in online communication remains a significant challenge, with current solutions often inadvertently blocking valuable information, including medical terms and text related to minority groups. This paper presents a more nu-anced approach to identifying toxicity in Bulgarian text while preserving access to essential information. The research explores two distinct methodologies for detecting toxic content. The developed methodologies have po-tential applications across diverse online platforms and content moderation systems. First, we propose an ontology that models the potentially toxic words in Bulgarian language. Then, we compose a dataset that comprises 4,384 manually anno-tated sentences from Bulgarian online forums across four categories: toxic language, medical terminology, non-toxic lan-guage, and terms related to minority communities. We then train a BERT-based model for toxic language classification, which reaches a 0.89 F1 macro score. The trained model is directly applicable in a real environment and can be integrated as a com-ponent of toxic content detection systems.
Innovation

Methods, ideas, or system contributions that make the work stand out.

ontology
BERT
toxic language detection
Bulgarian NLP
content moderation
🔎 Similar Papers
No similar papers found.
M
Melania Berbatova
Faculty of Mathematics and Informatics, Sofia University "St. Kliment Ohridski"
T
Tsvetoslav Vasev
Faculty of Mathematics and Informatics, Sofia University "St. Kliment Ohridski"