🤖 AI Summary
This study addresses the issue in reinforcement learning for large language models where policy updates cause sampled responses to drift off-policy, and conventional masking obscures bidirectional drift due to the cancellation of positive and negative log-ratios. To this end, we propose Cancellation-Aware Response Masking, which introduces an absolute value mechanism that averages token-level log-ratios after taking their absolute values. This effectively prevents opposing probability shifts from canceling each other out, enabling more precise sequence-level off-policy control. We further provide theoretical proofs establishing the joint boundary conditions for accepted responses. Experimental results demonstrate that our method improves average accuracy by 3.13 percentage points on the AIME mathematical reasoning task and increases Pass@1 by 2.88 percentage points across four code benchmarks, significantly outperforming existing state-of-the-art baselines.
📝 Abstract
Recent years have witnessed the rapid adoption of reinforcement learning (RL) in large language model (LLM) post-training, with substantial gains in mathematical reasoning and code generation. In practical systems, however, policy updates and differences between rollout and training engines can make sampled responses off-policy. Sequence-level masking addresses this mismatch by deciding whether an entire response should contribute to optimization. A common masking rule uses the length-normalized geometric mean of sampled token probability ratios. Its signed log-ratios can cancel across positions, concealing substantial bidirectional policy drift. We propose \emph{Cancellation-Aware Response Masking} (CARM), a sequence-level mask that takes the absolute value of each token log-ratio before averaging, preventing opposing probability changes from canceling. We prove that accepted responses satisfy a joint bound on the fraction of sampled-token ratios outside a prescribed band and their mean log-distance beyond its boundaries. Experiments on mathematical reasoning and code generation show that CARM improves mean@16 averaged over AIME 2024/2025/2026 and BeyondAIME by up to $3.13$ percentage points over geometric-mean masking, and increases average pass@1 across four code benchmarks by $2.88$ points over the strongest evaluated baseline. These findings support CARM as a theoretically grounded and effective method for response-level off-policy control in LLM reinforcement learning.