VAA-CSEC: Vote-guided Advantage Allocation for Chinese Semantic Error Correction
This study addresses the over-correction tendency of large language models in Chinese semantic error correction and the unclear interaction mechanisms between chain-of-thought reasoning and self-consistency decoding. To this end, we propose a multi-stage optimization framework. Methodologically, the model is initialized via chain-of-thought distillation and supervised fine-tuning, followed by the design of a minimal-edit principle reward function. Group-level relative policy optimization (GLPO) is then introduced to align training and inference objectives, with self-consistency decoding ultimately employed to enhance output robustness. Experimental results demonstrate that the proposed approach achieves state-of-the-art performance on both the CSED-C and NaSGEC-Exam datasets, attaining F0.5 scores of 47.72% and 41.55%, respectively.