Reward-Driven Learning under Prompt-Level Differential Privacy

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the risk of training data privacy leakage in Reinforcement Learning with Verifiable Rewards (RLVR). To mitigate this, it proposes the first prompt-level differential privacy framework tailored for RLVR. Building upon LoRA fine-tuning, the method introduces mechanisms for prompt-level gradient aggregation, clipping, and Gaussian noise injection to achieve rigorous differential privacy guarantees, wherein the privacy budget remains independent of both the number of responses and the clipping norm. Empirical evaluations on the MATH and GSM8K benchmarks demonstrate that the proposed approach significantly outperforms supervised fine-tuning baselines in accuracy, retaining 85%–90% of the performance gains achieved by non-private methods. These results substantiate the effectiveness of reward signals under differential privacy constraints, offering a principled solution for privacy-preserving reinforcement learning.
📝 Abstract
Reinforcement learning with verifiable rewards (RLVR) trains a language model on problems that may themselves be confidential, and the trained model can reveal which problems it saw. We study RLVR under prompt-level differential privacy: the released weights must be (ε,δ)-differentially private with respect to the presence of any one training problem. Taking the group of responses to one prompt as the privacy record, our method aggregates their gradients, clips the prompt's contribution once, adds Gaussian noise, and composes the privacy loss across updates, so the budget depends on neither the number of responses per prompt nor the clipping norm; to our knowledge this is the first differential privacy guarantee for RLVR training. We train Qwen2.5-1.5B-Instruct with LoRA at a per-run budget of ε=8 and compare, on the same prompts and at the same budget, a control that removes only the reward signal and two private supervised fine-tuning recipes. The reward signal improves accuracy over the control by 2.65 points on MATH and 3.24 on GSM8K, in every seed; the improvement survives a format-robust scorer, at 1.3 points on MATH, and is not explained by response length. At the same budget the private model outperforms both supervised recipes on MATH and GSM8K by 2.3 to 3.8 points, retains 85--90% of the gain of non-private GRPO on these tasks, and on MATH the noise of an eightfold tighter budget costs at most 1.2 points. The reward effect also carries to CommonsenseQA, an exploratory non-mathematical task. Verifier feedback thus remains a usable learning signal under prompt-level privacy.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning with Verifiable Rewards
Differential Privacy
Prompt-Level Privacy
Language Model Training
Privacy Leakage
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning with Verifiable Rewards
Prompt-Level Differential Privacy
Gradient Aggregation and Clipping
Privacy Budget Composition
LoRA Fine-Tuning
🔎 Similar Papers
No similar papers found.