🤖 AI Summary
Existing online distillation membership auditing methods fail to effectively capture teacher-guided update directions. To address this limitation, this work proposes the PAMA framework, which identifies membership cues by quantifying the deviation of student policies toward the teacher's loss descent direction. The method innovatively introduces a teacher alignment gain signal to characterize the direct contribution of membership cues to teacher updates, and integrates student drift with uncertainty alignment signals for multidimensional comprehensive evaluation. Experimental results demonstrate that PAMA achieves AUC scores ranging from 0.791 to 0.941 on the MATH benchmark, representing a substantial improvement of 14.6%–20.6% over state-of-the-art approaches.
📝 Abstract
On-policy distillation (OPD) trains a student model by aligning its policy with a teacher model on trajectories generated by the student model itself. Through this process, the student policy moves toward the teacher on the prompts used for distillation. However, these prompts are often private and costly, creating a need for prompt-level membership auditing. Existing methods mainly rely on likelihood-based confidence signals or student policy drift between checkpoints, but they do not capture the teacher-induced direction of the student update. In this paper, we propose Policy Alignment Membership Auditing (PAMA), a new auditing framework tailored for OPD. Our key observation is that a member prompt directly contributes to the teacher-guided policy update, while a non-member prompt only experiences indirect effects through cross-prompt generalization. Based on this directional trace, PAMA measures whether the student update moves toward reducing the teacher loss on a candidate prompt. Specifically, we introduce Teacher Alignment Gain (TAG) to estimate the teacher-aligned update direction from model outputs, and further combine it with student drift and uncertainty alignment signals for reliable membership auditing. We evaluate PAMA on six datasets and three teacher-student model families. On MATH, the primary evaluation benchmark, PAMA achieves AUC values of 0.791--0.941, improving AUC by 14.6--20.6% over state-of-the-art baselines.