MaPP: A Unified Marginalized Posterior-Predictive Framework for Data-Efficient RLVR

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the irreducible gradient estimation error and inefficient prompt selection in Reinforcement Learning with Verifiable Rewards (RLVR) caused by combinatorial noise from group normalization, proposing the MaPP framework. Theoretically, we establish an error lower bound for this noise. Methodologically, we construct a unified posterior predictive model to eliminate intra-group dependencies, derive closed-form solutions via Beta-Binomial marginalization for denoised advantage estimation, and design an uncertainty-aware online prompt selection algorithm with no additional computational overhead. Experiments demonstrate that MaPP significantly outperforms GRPO baselines on mathematical reasoning tasks, achieving up to a 2.45% improvement in average accuracy under equivalent computational budgets and establishing a new state of the art.
📝 Abstract
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models but incurs substantial costs from rollouts and policy updates. Online prompt selection improves efficiency by using per-prompt Bayesian posteriors to predict difficulty and prioritize informative prompts. However, existing methods overlook how reliably learning signals are extracted from sampled responses. In GRPO, a response's advantage depends on both its own outcome and the randomly sampled outcomes of its peers through group normalization. Our theoretical and experimental analyses show that uncertainty in group composition introduces composition noise, a non-vanishing variance component that imposes an irreducible lower bound on gradient estimation error and impairs downstream prompt selection. We propose MaPP (Marginalized Posterior-Predictive), a unified framework for data-efficient RLVR that denoises response-level advantage estimation and improves prompt selection using a shared Beta posterior. For each response, MaPP replaces the standard group-relative advantage with a composition-invariant intrinsic advantage through closed-form Beta-Binomial marginalization. The resulting posterior-predictive estimator has an error that provably diminishes as the posterior concentrates. Using the same posterior, MaPP derives an uncertainty-aware prompt selection score to improve data efficiency without additional rollout cost. Experiments on mathematics, planning, and visual geometry across five model backbones show that MaPP consistently outperforms GRPO and strong selection baselines, achieving up to +2.45 average accuracy improvement over the strongest baseline under the same rollout budget and setting a new state of the art.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning with Verifiable Rewards
Prompt Selection
Composition Noise
Advantage Estimation
Data Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Marginalized Posterior-Predictive
RLVR
Composition Noise
Beta-Binomial Marginalization
Uncertainty-aware Prompt Selection
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yangyang Ren
Beihang University
H
Haodong Zhu
Beihang University
S
Sheng Xu
Communication University of China
Y
Yanjing Li
Nanyang Technological University
N
Nikolai Yu. Zolotykh
Lobachebsky University
Wentao Zhang
Wentao Zhang
Institute of Physics, Chinese Academy of Sciences
photoemissionsuperconductivitycupratehtsctime-resolved
Baochang Zhang
Baochang Zhang
Technische Universität München
Computer assisted interventionMedical image analysisDeep learning