🤖 AI Summary
This study addresses the absence of evaluation benchmarks for Greek large language models and the inadequacy of traditional metrics in assessing complex reasoning. We introduce two benchmarks, Prot-Ex and Pan-Ex, employing an LLM-as-a-Judge paradigm and textualized visual context techniques to systematically evaluate diverse models on academic tasks. Our analysis reveals a few-shot prompting paradox and context overload phenomena in smaller models, demonstrating that localization adaptation can effectively compensate for parameter disadvantages. Experimental results indicate that KriKri-8B achieves performance comparable to larger models on humanities tasks, validating the efficacy of domain-specific linguistic adaptation. This work establishes a novel evaluation paradigm for low-resource language models.
📝 Abstract
The rapid advancement of Large Language Models (LLMs) imposes a thorough evaluation of their linguistic and analytical capabilities as well as constraints, particularly for a language with limited benchmark coverage such as Greek. To address the limited availability of comprehensive benchmarks in this domain, we introduce Prot-Ex and Pan-Ex, two benchmarks consisting of questions from entrance exams for Greek Model and Experimental schools as well as the Panhellenic exams (the Greek national university entrance examinations). These benchmarks are employed to assess the performance of text-only LLMs-including the Greek-adapted KriKri-8B-Instruct, Llama-3.1-8B, Gemma-4-26B, and Qwen-3-32B-across diverse academic disciplines (Modern Greek, Mathematics, Physics, etc.) and task formats (closed, structured, and open-ended), including textualized visual context (i.e., image descriptions). Our findings indicate the localized KriKri-8B significantly outperforms its base model, successfully rivalling much larger LLMs in linguistically demanding humanities tasks. By leveraging an LLM-as-a-Judge methodology, we expose the inadequacy of traditional lexical metrics for evaluating complex reasoning. Crucially, we uncover a few-shot prompting paradox: while synthetic examples improve accuracy in closed-ended questions, they severely overload the context window of 8B models in structured tasks, causing significant performance degradation. Ultimately, this study suggests targeted linguistic adaptation offsets lower parameter counts in specialized domains, despite the fragility of smaller models to prompt verbosity.