🤖 AI Summary
This study addresses the challenge of tracing distilled models after post-training, where historical checkpoints and model weights are typically unavailable. To this end, we propose SCOUT, a framework that introduces the first output-only attribution paradigm. Without requiring access to model parameters or pre-distillation data, this method achieves weight-free and history-free distillation attribution by mining syntactic patterns in generated text, applying low-contrast filtering, and employing a normalized distance calibration algorithm. Experimental results demonstrate that syntactic signatures from teacher models persist and remain traceable throughout post-training stages, including preference optimization and reinforcement learning. Furthermore, the proposed framework supports real-time auditing and abstention decisions for publicly released models, offering a novel tool for AI safety governance.
📝 Abstract
Model distillation transfers capabilities through supervised fine-tuning (SFT) on teacher responses, often collected from commercial APIs, raising questions of model provenance. Existing distillation attribution methods have been largely evaluated on students immediately after the SFT step. However, a distilled model may undergo further SFT, preference optimization, or reinforcement learning before release, while an auditor may lack access to the pre-distillation checkpoint required by reference-based attribution. To close this gap, we propose SCOUT, an output-only method that aggregates recurring *syntactic patterns* into candidate profiles, filters low-contrast patterns, and calibrates student--candidate distances against inter-candidate distances. SCOUT supports attribution and abstention using only current texts, without model weights, token likelihoods, or historical checkpoints. Auditing publicly released descendants of distilled models spanning diverse post-training objectives, SCOUT consistently identifies the distillation source. Furthermore, tracing teacher-associated *syntactic signatures* along training trajectories reveals that they emerge during distillation and persist through subsequent preference optimization and reinforcement learning.