🤖 AI Summary
This study addresses the training instability and potential collapse arising from the interaction between sparse rewards and dense distillation signals in reinforcement learning. Through theoretical analysis grounded in Neural Tangent Kernel (NTK) theory, it identifies two core failure modes: magnitude swamping and local conflict. Accordingly, this work proposes the M3 optimization framework, which introduces cross-signal NTK to quantify gradient alignment and integrates policy gradients, KL-divergence constraints, and magnitude normalization to establish empirical thresholds that prevent training collapse. Experimental results demonstrate that by identifying critical thresholds in gradient norm ratios, the proposed method effectively mitigates directional conflicts and dynamic instabilities during joint training, significantly enhancing the robustness of hybrid training paradigms.
📝 Abstract
Reinforcement learning with verifiable rewards provides a sparse post-training signal: a single binary outcome evaluates the entire rollout, and every token receives the same sequence-level advantage regardless of its individual contribution. To complement this sparse supervision, a growing family of methods adds a scalar-weighted teacher KL term to the policy-gradient objective, providing dense token-level guidance that may be unreliable at some positions. Despite the benefits of combining these signals, their interaction during optimization can destabilize joint training. To understand how this instability develops, we study the learning dynamics of hybrid reward--distillation training through a neural tangent kernel (NTK) analysis. We introduce the cross-signal NTK $K_{DR}(n)$, a token-level statistic that measures the alignment between reward and distillation gradients at position n. Through this analysis, we identify two failure modes: 1 Magnitude drowning, where the reward gradient exceeds the distillation gradient by orders of magnitude, so that even weak directional conflict can cause the distillation loss to rise despite its explicit inclusion in the training objective; and 2 Localized directional conflict, where the sequence-level advantage and the teacher's position-specific distribution induce opposing updates at the same token ($K_{DR}(n)\!<\!0$). The severity of these effects depends on the optimization regime: the gradient-norm ratio $\kappa\!=\!\|\nabla\mathcal{L}_R\|/\|\nabla\mathcal{L}_D\|$ varies by roughly an order of magnitude across tasks, and our experiments reveal an empirical threshold beyond which naive mixing can lead to persistent training collapse. Motivated by these findings, we introduce the M3 family, which combines magnitude normalization with three strategies...