🤖 AI Summary
This work addresses policy mode collapse, fragile exploration, and distributional shift in reinforcement learning from human feedback by proposing a geometry-driven proximal policy optimization method. The approach models the policy as a particle-based variational inference process within a mixture-of-experts architecture, updated via Stein variational gradient descent. It introduces a geometric proximal control mechanism grounded in functional kernels and an expert orthogonality loss, thereby eliminating reliance on fixed clipping or KL divergence scheduling. Evaluated on 33B/4B sparse mixture-of-experts models, the method achieves substantial performance gains: a +179 ELO improvement on Codeforces programming tasks and a 32% reduction in token consumption on AIME mathematical reasoning benchmarks.
📝 Abstract
Reinforcement Learning from Human Feedback via Proximal Policy Optimization often suffers from policy mode collapse, brittle exploration loops, and distribution drift. This paper introduces Variational Proximal Policy Optimization (\(\textsc{VP}_2\textsc{O}\)), a particle-based variational inference framework that maps policy optimization to Stein Variational Gradient Descent within a Mixture-of-Experts architecture. By leveraging functional kernels over localized expert prototypes alongside an expert orthogonalization loss, \(\textsc{VP}_2\textsc{O}\) introduces a geometry-based proximal-control mechanism that can reduce reliance on fixed clipping or KL schedules. Our results on a 33B/4B sparse Mixture-of-Experts model show several improvements across complex reasoning benchmarks, establishing a \(+\mathbf{179}\) ELO gain on Codeforces and a \(\mathbf{32\%}\) reduction in token count on AIME mathematical reasoning tasks.