MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limited progress in Yiddish language modeling, hindered by scarce digital resources, noisy corpora, and the absence of reliable evaluation benchmarks. To overcome these challenges, we introduce MameLoshnLM, the first open-source 8B-parameter language model specifically tailored for Yiddish, built upon the Llama 3.1 8B architecture. The model is further pretrained on Oytser, a newly curated high-quality corpus comprising web texts and literary sources. We also release Kashes, a multitask evaluation benchmark covering translation, linguistic analysis, and other relevant tasks. Experimental results demonstrate that MameLoshnLM significantly outperforms comparable open-source baselines on Kashes, more accurately capturing Yiddish-specific lexical and morphological features, thereby highlighting the advantages of specialized models for low-resource yet culturally rich languages.
📝 Abstract
We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have constrained progress in Yiddish language modeling. Existing multilingual corpora and benchmarks are often poor proxies for the language, containing substantial amounts of noisy, machine-translated, and misclassified text. We address these gaps by introducing Oytser, a high-quality Yiddish pretraining corpus that combines contemporary web-native sources with literary materials, and Kashes, a multi-task benchmark spanning translation, linguistic analysis, information extraction, and language understanding. Using these resources, we continue pretraining Llama 3.1 8B to obtain MameLoshnLM. Across the tasks in the benchmark, MameLoshnLM outperforms open baselines of similar scale. Our analyses show that these gains are not only quantitative: relative to general-purpose multilingual models, MameLoshnLM better captures language-defining lexical and morphological patterns, pointing to a broader failure mode of noisy web-scale multilingual data for low-resource languages. Our results provide both a foundation for Yiddish NLP and a practical template for language model development in historically rich but digitally underrepresented languages.
Problem

Research questions and friction points this paper is trying to address.

Yiddish
low-resource languages
language modeling
evaluation benchmark
noisy data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Yiddish language model
low-resource NLP
high-quality pretraining corpus
multitask evaluation benchmark
language-specific modeling
🔎 Similar Papers