🤖 AI Summary
This study addresses the absence of dedicated benchmarks for evaluating inflectional morphology generation and recognition capabilities in modern Greek language models. To this end, we introduce and publicly release MORFES, a novel evaluation suite comprising 500 expert-validated low-frequency lexical items designed to assess models’ generalization over rote memorization of inflectional patterns. We further establish the first open benchmark specifically targeting modern Greek inflectional morphology and present Sophea-Genesis-1, an open-source model with state-of-the-art morphological competence. Experimental results demonstrate that Sophea-Genesis-1 significantly outperforms prominent open-source large language models—including LLaMA, Qwen3, and DeepSeek-R1—on inflectional tasks, while maintaining competitive general-purpose performance relative to models of comparable scale, thereby validating both the efficacy and necessity of the MORFES benchmark.
📝 Abstract
Modern Greek is a richly inflected language, yet the language models built for it are evaluated mainly on factual knowledge, and no benchmark is dedicated to their inflectional competence. We introduce MORFES (Morphological Open-class Recognition-and-Formation Evaluation Suite), a benchmark of 500 expert-verified items that tests the recognition and production of Greek inflected forms, favoring lower-frequency lemmas so that a correct answer reflects the rule rather than a memorized form. We make it publicly available at https://huggingface.co/datasets/KIEFERSA/MORFES. We evaluate a range of open language models on MORFES, situating them within the rapidly scaling open-weight ecosystem from LLaMA to Qwen3, DeepSeek-R1, Magistral, and Kimi K2, where multilingual coverage grows but grammatical competence in morphologically rich languages remains under-measured. Among them, Sophea-Genesis-1, a model we developed and release as open weights at https://huggingface.co/KIEFERSA/Sophea-Genesis-1, leads on inflectional morphology while matching similarly sized models in general capability.