🤖 AI Summary
This work addresses the lack of rigorous theoretical interpretation for length-normalized importance ratios in the GSPO algorithm. We establish, for the first time, a strict equivalence between sequence-level importance weights and information-theoretic quantities—specifically, proving that such weights equal the product of perplexity ratios and exponential changes in cross-entropy. Building on this insight, we propose a novel gradient-weighting mechanism grounded in inverse perplexity ratios and geometric means in the log domain, offering a unified information-theoretic explanation for GSPO’s variance reduction and training stability. Our theoretical derivations are mathematically rigorous and empirically validated. Experiments demonstrate that this interpretation effectively accounts for GSPO’s superior performance and robustness in training mixture-of-experts models on mathematical reasoning tasks. The framework advances policy optimization by introducing an interpretable, analyzable paradigm rooted in information theory.
📝 Abstract
We provide a new perspective on GSPO's length-normalized importance ratios by establishing their connection to information-theoretic quantities. We show that GSPO's sequence-level weight $s(θ) = (π_θ/π_{θ_{ ext{old}}})^{1/|y|}$ can be equivalently expressed as the inverse perplexity ratio $ ext{PPL}_{θ_{ ext{old}}}/ ext{PPL}_θ$ and as the exponential cross-entropy change $exp(ΔH)$. While the perplexity-entropy relationship follows from standard definitions, this observation provides a useful lens for understanding GSPO: the algorithm weights policy gradient updates by perplexity ratios, offering an information-theoretic interpretation of the importance weights. This perspective helps explain GSPO's empirical properties, including log-domain variance reduction through geometric averaging and stability in training mixture-of-experts models. We validate the mathematical equivalences and variance predictions through controlled experiments on mathematical reasoning tasks.