LLM-as-a-Judge for Low-Resource Languages: Adapting Ragas and Comparative Ranking for Romanian

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inadequacy of traditional metrics in evaluating Retrieval-Augmented Generation (RAG) systems for low-resource languages by proposing a metric-specific evaluation strategy. By adapting the Ragas framework to the Romanian context, we construct the AdminRo-Eval benchmark dataset and conduct multi-method comparisons using the Gemini model combined with fine-grained decomposition and comparative ranking techniques. Our analysis reveals the inherent limitations of lightweight models in complex reasoning tasks under low-resource conditions. Ultimately, the Faithfulness metric achieves 96% alignment with human judgments, establishing a robust automated evaluation baseline for RAG systems in low-resource languages.
📝 Abstract
Evaluating Retrieval-Augmented Generation (RAG) systems remains a challenge for Low-Resource Languages (LRLs), where standard reference-based metrics fall short. This paper investigates the viability of the "LLM-as-a-Judge" paradigm for Romanian by adapting the Ragas framework using next-generation models (Gemini 2.5 and Gemini 3). We introduce AdminRo-Eval, a curated dataset of Romanian administrative documents annotated by native speakers, to serve as a ground truth for benchmarking automated evaluators. We compare three evaluation methodologies - direct scoring, comparative ranking, and granular decomposition - across metrics for Faithfulness, Answer Relevance, and Context Relevance. Our findings reveal that evaluation strategies must be metric-specific: granular decomposition achieves the highest human alignment for Faithfulness (96% with Gemini 2.5 Pro), while comparative ranking outperforms in Answer Relevance (90%). Furthermore, we demonstrate that while lightweight models struggle with complex reasoning in LRLs, the Gemini 2.5 Pro architecture establishes a robust, transferable baseline for automated Romanian RAG evaluation.
Problem

Research questions and friction points this paper is trying to address.

Retrieval-Augmented Generation
Low-Resource Languages
LLM-as-a-Judge
Evaluation
Romanian
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-as-a-Judge
Low-Resource Languages
RAG Evaluation
Ragas Framework
Comparative Ranking
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Claudiu Creanga
Faculty of Mathematics and Computer Science, Interdisciplinary School of Doctoral Studies, HLT Research Center, University of Bucharest, Romania
Liviu P. Dinu
Liviu P. Dinu
Professor, University of Bucharest, Dept. of Computer Science,
Computational LinguisticsNatural Language ProcessingComputational Historical Linguistics