Leaky Students: Membership Inference against On-Policy Distillation

πŸ“… 2026-09-27
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study investigates whether student models in online policy distillation inadvertently leak sensitive membership information from the teacher’s training data. Addressing the challenges of sparse fresh trajectory signals and the susceptibility of fixed-loss approaches to evasion, this work presents the first systematic membership inference attack analysis for this setting. It proposes Leaky, a detection framework that integrates dynamic trajectory comparison with ReLU noise correction, alongside log-probability evaluation and reference model matching techniques, to achieve high-precision privacy leakage detection. Experimental results demonstrate that Leaky attains an average AUROC of 0.875 across mathematical and medical tasks, significantly outperforming existing baselines. These findings substantiate the presence of privacy leakage risks in student models during online policy distillation.
πŸ“ Abstract
On-policy distillation (OPD) trains a student to match a teacher's next-token distributions on student-generated trajectories. However, privileged information supplied to the teacher for OPD training may contain sensitive data. Whether the student leaks private information about the records supplied to the teacher during distillation remains poorly understood. To the best of our knowledge, we present the first systematic study of membership inference in this setting. We find that fresh student trajectories expose sparse membership signals that fixed reference-answer losses often miss. These signals are mixed with probability changes caused by training on other records. We introduce Leaky, which samples fresh trajectories from the target model and compares its token log-probabilities with the maximum across matched reference models trained without the candidate records. It applies Leaky ReLU to the resulting gaps, preserving positive gaps and downweighting negative gaps as an approximate correction for incidental positive gaps in non-members. Across fifteen targets spanning mathematics, medical question answering, and code generation, Leaky outperforms all evaluated baselines and achieves mean AUROC 0.875, compared with 0.614 for the strongest baseline on each target in the main evaluation. On the same sampled trajectories, the strongest baseline achieves mean AUROC 0.826. These results show that students trained through OPD can expose the membership of records used for teacher supervision, even when fixed reference-answer losses provide little evidence of membership.
Problem

Research questions and friction points this paper is trying to address.

Membership Inference
On-Policy Distillation
Privacy Leakage
Knowledge Distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Membership Inference
On-Policy Distillation
Knowledge Distillation
Privacy Leakage
Leaky ReLU
πŸ”Ž Similar Papers