RoLegalGEC: Legal Domain Grammatical Error Detection and Correction Dataset for Romanian

📅 2026-04-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the longstanding scarcity of high-quality, manually annotated grammatical error correction (GEC) data for Romanian legal texts, which has significantly hindered the development of domain-specific language processing tools. To bridge this gap, the authors introduce RoLegalGEC, the first parallel dataset for grammatical error detection and correction in the Romanian legal domain, comprising 350,000 human-annotated sentence pairs. The work systematically evaluates various neural approaches: knowledge-distilled Transformer and sequence labeling models for error detection, and multiple pretrained text-to-text Transformer architectures for error correction. By providing this novel resource and establishing strong baselines, the study fills a critical void in Romanian domain-specific grammatical processing and lays a foundational benchmark for future research.

Technology Category

Natural Language Processing: Safety and RobustnessKnowledge Representation and Reasoning: Knowledge Representation LanguagesMachine Learning: Calibration & Uncertainty Quantification

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsWeb Mining and Content Analysis: Robustness and generalizability of Web mining methods
📝 Abstract
The importance of clear and correct text in legal documents cannot be understated, and, consequently, a grammatical error correction tool meant to assist a professional in the law must have the ability to understand the possible errors in the context of a legal environment, correcting them accordingly, and implicitly needs to be trained in the same environment, using realistic legal data. However, the manually annotated data required by such a process is in short supply for languages such as Romanian, much less for a niche domain. The most common approach is the synthetic generation of parallel data; however, it requires a structured understanding of the Romanian grammar. In this paper, we introduce, to our knowledge, the first Romanian-language parallel dataset for the detection and correction of grammatical errors in the legal domain, RoLegalGEC, which aggregates 350,000 examples of errors in legal passages, along with error annotations. Moreover, we evaluate several neural network models that transform the dataset into a valuable tool for both detecting and correcting grammatical errors, including knowledge-distillation Transformers, sequence tagging architectures for detection, and a variety of pre-trained text-to-text Transformer models for correction. We consider that the set of models, together with the novel RoLegalGEC dataset, will enrich the resource base for further research on Romanian.
Problem

Research questions and friction points this paper is trying to address.

Grammatical Error Correction
Legal Domain
Romanian Language
Annotated Dataset
Grammar Errors
Innovation

Methods, ideas, or system contributions that make the work stand out.

RoLegalGEC
grammatical error correction
legal domain
Romanian NLP
parallel dataset
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.