🤖 AI Summary
This study addresses the over-correction tendency of large language models in Chinese semantic error correction and the unclear interaction mechanisms between chain-of-thought reasoning and self-consistency decoding. To this end, we propose a multi-stage optimization framework. Methodologically, the model is initialized via chain-of-thought distillation and supervised fine-tuning, followed by the design of a minimal-edit principle reward function. Group-level relative policy optimization (GLPO) is then introduced to align training and inference objectives, with self-consistency decoding ultimately employed to enhance output robustness. Experimental results demonstrate that the proposed approach achieves state-of-the-art performance on both the CSED-C and NaSGEC-Exam datasets, attaining F0.5 scores of 47.72% and 41.55%, respectively.
📝 Abstract
Chinese Semantic Error Correction (CSEC) targets semantic errors in Chinese text, which are typically more subtle and complex than spelling and grammatical errors but remain relatively underexplored. Existing LLM-based approaches face two recurring obstacles in this task: over-correction, and unclear interaction between Chain-of-Thought (CoT) reasoning and self-consistency decoding, such that the benefits brought by CoT cannot be reliably transferred to final corrections. We propose Vote-guided Advantage Allocation for CSEC (VAA-CSEC), a multi-stage framework that combines CoT distillation, Supervised Fine-Tuning (SFT), Reinforcement Learning (RL) and self-consistency decoding. During RL, we design a task-specific reward function that directly aligned with the minimal-editing principle of CSEC. We further introduce Group-Level Relative Policy Optimization (GLPO), which reallocates GRPO advantages according to the margin between individual rollout rewards and the vote-aggregated group reward, aligning the RL training objective with the self-consistency objective used at inference time. Experiments on CSED-C and NaSGEC-Exam show that VAA-CSEC outperforms all LLM-based baselines on CSED-C with an F0.5 of 47.72%, achieves the highest recall of 42.15% among all methods, and establishes a new state of the art of 41.55% F0.5 on NaSGEC-Exam.