Voices of the Mountains: Deep Learning-Based Vocal Error Detection System for Kurdish Maqams

📅 2026-02-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing automatic singing assessment systems, which are grounded in Western twelve-tone equal temperament and thus struggle to accurately detect microtonal intervals, glissandi, and modal instability characteristic of traditional performances in the Kurdish Bayati-Kurd maqam. To bridge this gap, the work proposes the first deep learning framework specifically designed for vocal error detection in this non-Western microtonal context, focusing on pitch inaccuracies, rhythmic deviations, and modal drift. Trained on log-mel spectrograms of 221 expert-annotated error segments extracted from 50 songs performed by 13 singers, the model combines CNN and BiLSTM layers with an attention mechanism. It achieves a macro F1-score of 0.468 on the validation set, with F1-scores of 0.492 and 0.536 for pitch and rhythm errors, respectively, while modal drift detection remains a challenge. This research establishes the first dedicated evaluation system for microtonal non-Western music, moving beyond the constraints of Western musical paradigms.

Technology Category

Machine Learning: Multimodal LearningNatural Language Processing: Language Grounding & Multi-modal NLPCognitive Modeling & Cognitive Systems: Affective Computing

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsWeb Mining and Content Analysis: Robustness and generalizability of Web mining methodsGraph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphs
📝 Abstract
Maqam, a singing type, is a significant component of Kurdish music. A maqam singer receives training in a traditional face-to-face or through self-training. Automatic Singing Assessment (ASA) uses machine learning (ML) to provide the accuracy of singing styles and can help learners to improve their performance through error detection. Currently, the available ASA tools follow Western music rules. The musical composition requires all notes to stay within their expected pitch range from start to finish. The system fails to detect micro-intervals and pitch bends, so it identifies Kurdish maqam singing as incorrect even though the singer performs according to traditional rules. Kurdish maqam requires recognizing performance errors within microtonal spaces, which is beyond Western equal temperament. This research is the first attempt to address the mentioned gap. While many error types happen during singing, our focus is on pitch, rhythm, and modal stability errors in the context of Bayati-Kurd. We collected 50 songs from 13 vocalists ( 2-3 hours) and annotated 221 error spans (150 fine pitch, 46 rhythm, 25 modal drift). The data was segmented into 15,199 overlapping windows and converted to log-mel spectrograms. We developed a two-headed CNN-BiLSTM with attention mode to decide whether a window contains an error and to classify it based on the chosen errors. Trained for 20 epochs with early stopping at epoch 10, the model reached a validation macro-F1 of 0.468. On the full 50-song evaluation at a 0.750 threshold, recall was 39.4% and precision 25.8% . Within detected windows, type macro-F1 was 0.387, with F1 of 0.492 (fine pitch), 0.536 (rhythm), and 0.133 (modal drift); modal drift recall was 8.0%. The better performance on common error types shows that the method works, while the poor modal-drift recall shows that more data and balancing are needed.
Problem

Research questions and friction points this paper is trying to address.

Kurdish Maqam
Automatic Singing Assessment
microtonal music
pitch error detection
modal stability
Innovation

Methods, ideas, or system contributions that make the work stand out.

microtonal singing
automatic singing assessment
CNN-BiLSTM with attention
Kurdish maqam
modal stability error
🔎 Similar Papers
💼 Related Jobs
No related jobs found.