RunyaNER: Auxiliary Language Selection for Runyankore NER

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of unclear auxiliary language selection strategies and the lack of benchmark data in named entity recognition (NER) for low-resource languages by constructing the first Runyankore NER benchmark dataset. Methodologically, we design a semi-automatic annotation pipeline to generate high-quality data and conduct systematic experiments combining pretrained model fine-tuning with cross-lingual zero-shot transfer. Our findings reveal that embedding similarity outperforms traditional linguistic features in predicting cross-lingual transfer performance, thereby establishing an optimal strategy for auxiliary language selection. By bridging the data gap for this language, this work provides empirical evidence and practical guidance for multilingual transfer learning in low-resource scenarios.
📝 Abstract
Cross-lingual zero-shot transfer and multilingual fine-tuning are promising approaches for NLP tasks such as Named Entity Recognition (NER) in low-resource languages, but in the absence of target language benchmarks, it is unclear which auxiliary language selection strategy leads to the best transfer. We introduce RunyaNER, the first publicly available NER benchmark for the East African language Runyankore, and use it to investigate the choice of which languages to use for transfer. Created with a semi-automated pipeline and fully manually verified, RunyaNER contains over 237k annotated words across 30k sentences. We benchmark pretrained models on RunyaNER, establishing that our dataset is of sufficient quality and size to produce effective Runyankore NER models. We then use RunyaNER to investigate auxiliary language selection in cross-lingual zero-shot and multilingual fine-tuning settings. Our experiments show that while transfer performance is highly sensitive to auxiliary language selection, embedding-based measures computed from labelled training spans correlate more strongly with downstream transfer performance than traditional linguistic features based on metadata or typology. By releasing RunyaNER and providing a systematic analysis of auxiliary language selection strategies, this work contributes both a new benchmark resource and practical insights for multilingual transfer in low-resource settings.
Problem

Research questions and friction points this paper is trying to address.

Named Entity Recognition
low-resource languages
auxiliary language selection
cross-lingual transfer
Runyankore
Innovation

Methods, ideas, or system contributions that make the work stand out.

Named Entity Recognition
Cross-lingual Transfer
Low-resource Languages
Auxiliary Language Selection
Embedding-based Measures
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.