CARM: Cancellation-Aware Response Masking for LLM Reinforcement Learning
This study addresses the issue in reinforcement learning for large language models where policy updates cause sampled responses to drift off-policy, and conventional masking obscures bidirectional drift due to the cancellation of positive and negative log-ratios. To this end, we propose Cancellation-Aware Response Masking, which introduces an absolute value mechanism that averages token-level log-ratios after taking their absolute values. This effectively prevents opposing probability shifts from canceling each other out, enabling more precise sequence-level off-policy control. We further provide theoretical proofs establishing the joint boundary conditions for accepted responses. Experimental results demonstrate that our method improves average accuracy by 3.13 percentage points on the AIME mathematical reasoning task and increases Pass@1 by 2.88 percentage points across four code benchmarks, significantly outperforming existing state-of-the-art baselines.