🤖 AI Summary
This study addresses the tendency of vision-language models to over-correct anomalous text into fluent expressions, thereby compromising OCR transcription fidelity. To mitigate this, we propose GAD-RL, a method that innovatively introduces gating and decay mechanisms to adaptively modulate teacher supervision intensity based on the student model's current task performance and local distribution, dynamically attenuating its guidance weight as performance improves. This approach integrates joint post-training, online distillation, sequence-level reward optimization, and probability-based KL divergence weighting. Experimental results demonstrate that GAD-RL achieves a Micro Recall of 59.92% on CHAOS-Bench, significantly outperforming baselines, and attains an overall score of 91.18 on OmniDocBench.
📝 Abstract
Vision-language models may rewrite anomalous text in images into linguistically plausible expressions, compromising OCR transcription faithfulness. Sequence-level task rewards and local teacher guidance are complementary, but guidance from the same teacher may not remain equally effective as the student improves. Offline analysis shows that supervision from a fixed teacher becomes progressively less favorable as the student improves, both across training checkpoints and across response groups with different task rewards. Motivated by this observation, we introduce GAD-RL, which adaptively regulates teacher supervision during joint post-training according to the student's current task performance and local distributions. A frozen teacher conditions on reference transcriptions and student-generated prefixes. GAD-RL disables distillation for response groups containing an output with task reward at least 0.95 and continuously attenuates distillation strength as group-mean reward increases. It also weights forward KL by the student's probability of the teacher's Top-1 token, moderating local auxiliary updates when student support for that candidate is low. On Qwen3.5-2B, GAD-RL achieves 59.92% Micro Recall on CHAOS-Bench, surpassing GRPO and GRPO+OPD (fixed-weight) by 8.45 and 4.43 percentage points, respectively, while achieving an Overall score of 91.18 on OmniDocBench v1.6.