Persistent Negatives for Adversarial Black-Box On-Policy Distillation

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the training instability in black-box online distillation caused by the drift of the discriminator's negative sample distribution following policy updates. To mitigate this moving-target problem, we propose a persistent negative sample adversarial distillation framework that anchors the discriminator by maintaining a historical negative sample pool. Theoretically, we demonstrate that the log-density ratio between the teacher model and negative samples constitutes the Bayesian optimal reward, which is subsequently integrated with the GRPO algorithm to achieve stable policy optimization. Empirical evaluations across multiple benchmarks reveal that our approach outperforms existing methods, significantly smooths the discriminator's learning trajectory, and effectively eliminates sub-random performance fluctuations.
📝 Abstract
Black-box On-Policy Distillation (OPD) seeks to improve a student from its own generations when the teacher provides sampled responses but not token probabilities. Adversarial distillation offers one route: it learns a discriminator over prompt-matched teacher and student responses and uses its score as the policy reward. However, sampling discriminator negatives from the latest student at each step couples the learned reward to a negative distribution that changes after every policy update. We address this moving-target problem with persistent-negative adversarial distillation, a live-pool method that replaces a fraction of each discriminator batch with historical, prompt-matched teacher--student comparisons. Under matched discriminator compute, historical comparisons train the discriminator, while GRPO remains on-policy with fresh student responses. Our analysis identifies the Bayes-optimal reward as a teacher-to-negative log-density ratio and, under explicit assumptions, shows how persistent negatives anchor the discriminator and reduce reward-estimation MSE relative to fresh-negative training. Across two student families, three judges, and four judged-chat benchmarks, persistent-negative adversarial distillation consistently improves performance over current methods at matched discriminator compute. It also yields smoother fresh-policy discriminator trajectories, with fewer below-chance dips. These findings identify the discriminator's negative distribution as an important design axis in black-box on-policy distillation.
Problem

Research questions and friction points this paper is trying to address.

Black-box On-Policy Distillation
Adversarial Distillation
Moving-target Problem
Persistent Negatives
Discriminator
Innovation

Methods, ideas, or system contributions that make the work stand out.

Black-Box On-Policy Distillation
Adversarial Distillation
Persistent Negatives
Discriminator Anchoring
GRPO
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.