Continuous Saudi Sign Language Recognition: A Vision Transformer Approach

📅 2025-09-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Approximately 84,000 Deaf individuals in Saudi Arabia use Saudi Sign Language (SSL) as their primary means of communication; however, robust continuous-sentence SSL recognition remains critically underexplored, hindered by the absence of large-scale SSL datasets and the morphological complexity of Arabic. To address this gap, we introduce KAU-CSSL—the first publicly available continuous-sentence SSL dataset—comprising natural, sentence-level signing sequences. We further propose an end-to-end framework integrating ResNet-18 for spatial feature extraction, a Transformer encoder for contextual modeling, and a bidirectional LSTM for temporal sequence learning. Evaluated under speaker-dependent and speaker-independent settings, our method achieves 99.02% and 77.71% word accuracy, respectively, significantly advancing SSL spatiotemporal pattern modeling. This work establishes the first sentence-level benchmark for Arabic sign language recognition, paving the way for improved educational access and societal inclusion for the Deaf community in Arabic-speaking regions.

Technology Category

Computer Vision: Language and VisionHumans and AI: AI for AccessibilityMachine Learning: Multimodal Learning

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Algorithms and analysis for heterogeneous, signed, attributed, multi-relational, temporal, higher-order, and annotated Web-related graphs
📝 Abstract
Sign language (SL) is an essential communication form for hearing-impaired and deaf people, enabling engagement within the broader society. Despite its significance, limited public awareness of SL often leads to inequitable access to educational and professional opportunities, thereby contributing to social exclusion, particularly in Saudi Arabia, where over 84,000 individuals depend on Saudi Sign Language (SSL) as their primary form of communication. Although certain technological approaches have helped to improve communication for individuals with hearing impairments, there continues to be an urgent requirement for more precise and dependable translation techniques, especially for Arabic sign language variants like SSL. Most state-of-the-art solutions have primarily focused on non-Arabic sign languages, resulting in a considerable absence of resources dedicated to Arabic sign language, specifically SSL. The complexity of the Arabic language and the prevalence of isolated sign language datasets that concentrate on individual words instead of continuous speech contribute to this issue. To address this gap, our research represents an important step in developing SSL resources. To address this, we introduce the first continuous Saudi Sign Language dataset called KAU-CSSL, focusing on complete sentences to facilitate further research and enable sophisticated recognition systems for SSL recognition and translation. Additionally, we propose a transformer-based model, utilizing a pretrained ResNet-18 for spatial feature extraction and a Transformer Encoder with Bidirectional LSTM for temporal dependencies, achieving 99.02% accuracy at signer dependent mode and 77.71% accuracy at signer independent mode. This development leads the way to not only improving communication tools for the SSL community but also making a substantial contribution to the wider field of sign language.
Problem

Research questions and friction points this paper is trying to address.

Lack of continuous Saudi Sign Language datasets for sentences
Need for accurate translation techniques for Arabic sign languages
Addressing resource scarcity for Saudi Sign Language recognition
Innovation

Methods, ideas, or system contributions that make the work stand out.

First continuous Saudi Sign Language dataset
Transformer model with ResNet-18 feature extraction
Bidirectional LSTM for temporal dependencies handling
King Abdulaziz University
S
Soukeina Elhassen
Department of Computer Science, Faculty of Computing and Information Technology, King Abdulaziz University, Jeddah 21589, Saudi Arabia
L
Lama Al Khuzayem
Department of Computer Science, Faculty of Computing and Information Technology, King Abdulaziz University, Jeddah 21589, Saudi Arabia
Areej Alhothali
Areej Alhothali
Associate Professor of Computer Science, King Abulaziz University
Machine learningNatural language processingAffective ComputingSentiment analysis
O
Ohoud Alzamzami
Department of Computer Science, Faculty of Computing and Information Technology, King Abdulaziz University, Jeddah 21589, Saudi Arabia
N
Nahed Alowaidi
Department of Computer Science, Faculty of Computing and Information Technology, King Abdulaziz University, Jeddah 21589, Saudi Arabia