Document-Level Text Simplification in Estonian Using Large Language Models

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of coherence and consistency in document-level text simplification for low-resource languages such as Estonian. We evaluate five multilingual large language models using three strategies: single-turn generation, pipeline processing, and guideline-enhanced prompting. Comprehensive evaluation combines automated metrics with human annotation. The primary contributions include a novel metric for document-level coherence, validation of evidence-based prompting strategies, and the release of open-source reproducible resources. Experimental results demonstrate that Gemini-2.0 and LLaMA-3.3 produce simplified texts exhibiting near-native fluency alongside strong semantic preservation.
📝 Abstract
Document-level text simplification involves transformations that go beyond sentence-internal edits, addressing discourse coherence, anaphora resolution, and cross-paragraph consistency. Despite advances in sentence-level simplification for high-resource languages, document-level simplification in morphologically rich, low-resource languages such as Estonian remains largely unexplored. This study presents a comprehensive evaluation of five state-of-the-art multilingual large language models (LLMs) for document-level simplification in Estonian. Three prompting strategies are examined: single-pass generation, pipeline-based modular agents, and guideline-augmented pipelines. The evaluation framework integrates automatic metrics assessing readability, semantic preservation, and discourse coherence, alongside a structured manual annotation protocol. The findings indicate that Gemini-2.0 and LLaMA-3.3 produce outputs with near-native fluency and strong meaning preservation, whereas other models display notable grammatical and semantic limitations. This work contributes novel document-level coherence metrics, evidence-based prompting strategies, and publicly available resources for reproducibility.
Problem

Research questions and friction points this paper is trying to address.

Document-Level Text Simplification
Low-Resource Languages
Estonian
Discourse Coherence
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Document-Level Text Simplification
Large Language Models
Low-Resource Languages
Prompting Strategies
Discourse Coherence Metrics
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Meeri-Ly Muru
National Library of Estonia, Institute of Computer Science, University of Tartu
E
Eduard Barbu
National Library of Estonia, Institute of Computer Science, University of Tartu