🤖 AI Summary
This study addresses the challenge of detecting whether large language models have illicitly appropriated proprietary capabilities through distillation attacks. We formulate distillation detection as a hypothesis testing problem and propose a novel statistical inference framework based on shadow models. By training shadow models to simulate both distillation and independent training behaviors, our approach compares discrepancies in reasoning trajectory predictions between the target model and the shadow models, computing calibrated p-values to determine provenance. Notably, this method requires only black-box access without necessitating internal model parameters. Experimental evaluations on Qwen2.5 and Llama3.2 demonstrate that the proposed framework achieves a true positive rate of 1.0 at a significance level of 0.02, effectively validating the feasibility and superiority of this black-box detection approach.
📝 Abstract
Unauthorized model distillation, in which a model is trained on the outputs of a proprietary large language model (LLM), is a growing threat to model providers. We study distillation inference: determining whether a suspect model was distilled from another model or trained independently. We formulate this problem as a hypothesis test and estimate the behavior expected under each hypothesis by training shadow models: distilled shadow models learn from the teacher's reasoning traces, whereas independent shadow models learn only from reference answers. The auditor measures how closely each model predicts the teacher's reasoning outputs and then uses the shadow models to convert the suspect's score into a calibrated p-value. In a preliminary study using Qwen2.5-7B as the teacher and Llama-3.2-3B for the suspects, our test achieves a true positive rate of 1.0 at a significance level of 0.02. These results demonstrate the feasibility of using distillation inference to detect distillation attacks.