RoCoISLR: A Romanian Corpus for Isolated Sign Language Recognition

📅 2025-11-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
The scarcity of large-scale, standardized datasets hinders progress in Romanian Isolated Sign Language Recognition (RoISLR). Method: This work introduces RoCoISLR—the first high-quality, standardized corpus for RoISLR—comprising over 9,000 video samples and nearly 6,000 lexical items. It proposes a systematic methodology for constructing low-resource sign language corpora, analyzes the critical impact of long-tailed label distribution on recognition performance, and establishes the first dedicated RoISLR benchmark. Contribution/Results: Under a unified experimental protocol, seven state-of-the-art architectures—including I3D, SlowFast, Swin Transformer, and TimeSformer—are rigorously evaluated. Results demonstrate the clear superiority of Transformer-based models over CNNs: Swin Transformer achieves 34.1% Top-1 accuracy, confirming the dataset’s substantial difficulty and practical utility. RoCoISLR effectively bridges the data gap for under-resourced sign languages and advances non-dominant sign language recognition research.

Technology Category

Intelligent Robots: Multimodal Perception & Sensor FusionMachine Learning: Large Multimodal Models (LMMs)Computer Vision: Adversarial Attacks & Robustness

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsSecurity and Privacy: Large-scale security measurementsEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasets
📝 Abstract
Automatic sign language recognition plays a crucial role in bridging the communication gap between deaf communities and hearing individuals; however, most available datasets focus on American Sign Language. For Romanian Isolated Sign Language Recognition (RoISLR), no large-scale, standardized dataset exists, which limits research progress. In this work, we introduce a new corpus for RoISLR, named RoCoISLR, comprising over 9,000 video samples that span nearly 6,000 standardized glosses from multiple sources. We establish benchmark results by evaluating seven state-of-the-art video recognition models-I3D, SlowFast, Swin Transformer, TimeSformer, Uniformer, VideoMAE, and PoseConv3D-under consistent experimental setups, and compare their performance with that of the widely used WLASL2000 corpus. According to the results, transformer-based architectures outperform convolutional baselines; Swin Transformer achieved a Top-1 accuracy of 34.1%. Our benchmarks highlight the challenges associated with long-tail class distributions in low-resource sign languages, and RoCoISLR provides the initial foundation for systematic RoISLR research.
Problem

Research questions and friction points this paper is trying to address.

Addressing the lack of large-scale Romanian sign language datasets
Establishing benchmark results with seven video recognition models
Highlighting challenges of long-tail class distributions in low-resource languages
Innovation

Methods, ideas, or system contributions that make the work stand out.

Created Romanian sign language corpus with 9000 videos
Benchmarked seven video recognition models consistently
Transformer models outperformed convolutional baselines in accuracy
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Cătălin-Alexandru Rîpanu
National University of Science and Technology POLITEHNICA Bucharest, Bucharest, Romania
A
Andrei-Theodor Hotnog
National University of Science and Technology POLITEHNICA Bucharest, Bucharest, Romania
G
Giulia-Stefania Imbrea
National University of Science and Technology POLITEHNICA Bucharest, Bucharest, Romania
Dumitru-Clementin Cercel
Dumitru-Clementin Cercel
Teaching Assistant of Computer Science, University Politehnica of Bucharest
Social Network AnalysisNatural Language ProcessingInformation RetrievalMachine Learning