Adapting LLMs for Minimal-edit Grammatical Error Correction

📅 2025-06-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses two key challenges in minimal-edit grammatical error correction (GEC): (1) the limited adaptability of decoder-only large language models to GEC, and (2) pervasive tokenization distortions and annotation errors in mainstream GEC benchmarks (e.g., BEA, CoNLL). To tackle these, we propose an error-rate-adaptive training schedule that dynamically weights sample losses; systematically identify and rectify tokenization and annotation flaws across standard datasets, releasing a unified detokenized, naturally segmented corpus; and introduce controllable edit-constrained modeling to enhance edit precision. Our approach achieves a new single-model state-of-the-art on BEA-test (ERRANT F₀.₅ = 79.2), demonstrating that data correction significantly improves model generalization. All code and the curated, verified dataset are publicly released.

Technology Category

Natural Language Processing: Language Grounding & Multi-modal NLPMachine Learning: Large Multimodal Models (LMMs)Reasoning under Uncertainty: Relational Probabilistic Models

Application Category

Economics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metrics
📝 Abstract
Decoder-only large language models have shown superior performance in the fluency-edit English Grammatical Error Correction, but their adaptation for minimal-edit English GEC is still underexplored. To improve their effectiveness in the minimal-edit approach, we explore the error rate adaptation topic and propose a novel training schedule method. Our experiments set a new state-of-the-art result for a single-model system on the BEA-test set. We also detokenize the most common English GEC datasets to match the natural way of writing text. During the process, we find that there are errors in them. Our experiments analyze whether training on detokenized datasets impacts the results and measure the impact of the usage of the datasets with corrected erroneous examples. To facilitate reproducibility, we have released the source code used to train our models.
Problem

Research questions and friction points this paper is trying to address.

Adapting LLMs for minimal-edit grammatical error correction
Exploring error rate adaptation for improved GEC performance
Detokenizing and correcting common English GEC datasets
Innovation

Methods, ideas, or system contributions that make the work stand out.

Proposes novel training schedule method
Detokenizes common English GEC datasets
Releases source code for reproducibility
🔎 Similar Papers
No similar papers found.