Revisiting On-policy Adversarial Black-Box Distillation: Calibrating Groupwise Reward Geometry for Effective Advantage Construction

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of optimization signals in black-box distillation, where advantage computation in Group Relative Policy Optimization (GRPO) suffers from reward group scale collapse. To mitigate this issue, we propose the GRGC framework, which enhances advantage construction through two-stage calibrated reward geometry. Specifically, GRGC introduces Gaussian optimal transport calibration and group power modulation to reshape reward distributions while preserving ordinal information, thereby improving the stability of policy optimization. By integrating adversarial black-box distillation with optimal transport theory, the proposed method achieves substantial improvements in both in-distribution (ID) and out-of-distribution (OOD) performance across diverse models and datasets, all with minimal computational overhead.
📝 Abstract
Black-box distillation is a practical route for transferring capabilities from API-accessible large language models that expose only text outputs into smaller student models. Recent on-policy adversarial methods such as GAD improve over SeqKD by forming an adversarial loop between a critic and a student, where the critic provides rewards for GRPO-based student policy optimization over the student's sampled responses. However, GRPO computes advantages from the within-group relative rewards of student samples for the same prompt, whereas the critic is trained primarily to distinguish teacher responses from student responses. This objective mismatch can produce reward groups with collapsed scale or fragile margins, leading to brittle grouped optimization signals. We propose Groupwise Reward Geometry Conditioning (GRGC), a two-stage framework that improves advantage construction by shaping student-side reward groups during both critic training and policy optimization. To improve critic-side conditioning, Gaussian groupwise Optimal Transport calibration regularizes the critic during training to produce reward groups with non-collapsed spread and smooth rank-wise gaps by matching sorted prompt-wise rewards to group-centered Gaussian quantiles. Building on this conditioned reward geometry, policy-side group power modulation reshapes the prompt-wise reward groups before they are converted into advantages, preserving the critic-induced ordering while increasing optimization-relevant margin separability. Extensive experiments across diverse teachers, student model families and scales, and training datasets demonstrate the effectiveness of GRGC on both in-distribution and out-of-distribution evaluations, while introducing negligible overhead over GAD. The code is available at https://github.com/2018cx/GRGC.
Problem

Research questions and friction points this paper is trying to address.

Black-box distillation
On-policy adversarial learning
Objective mismatch
Advantage construction
Reward geometry
Innovation

Methods, ideas, or system contributions that make the work stand out.

Black-box distillation
Groupwise reward geometry
Optimal transport calibration
Advantage construction
On-policy adversarial learning