Beyond Reference-Based Evaluation: Reward Models for Meta-Evaluation of Grammatical Error Correction

📅 2026-09-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文针对语法纠错评价中参考答案局限性问题,提出基于奖励模型RM-EVAL的方法进行无参考评价,并通过奖励引导文本生成改进纠错系统。
📝 Abstract
Reference-based metrics for Grammatical Error Correction (GEC) such as M$^2$ and ERRANT assume that the reference set enumerates all valid edits, and therefore often penalize corrections that are grammatical and meaning-preserving but phrased differently. We introduce RM-EVAL, a reward model trained on human preference data from SEEDA, as a reference-free meta-evaluator that predicts human-like quality judgments at both full-sequence and partial-sequence levels. Beyond evaluation, we show that the same reward model can be used as a learning signal to improve GEC generation via Reward-Guided Text Generation (RGTG), which keeps a base GEC model frozen and performs online, reward-driven decoding. Across SEEDA, RM-EVAL achieves strong agreement with human rankings, and RGTG yields consistent gains in reward and external validation, demonstrating a unified framework for both assessing and enhancing GEC systems without relying on gold references.
Problem

Research questions and friction points this paper is trying to address.

Grammatical Error Correction
Reference-based Metrics
Human Preference Data
Innovation

Methods, ideas, or system contributions that make the work stand out.

reward model
reference-free evaluation
grammatical error correction
human preference data
reward-guided text generation