Retrospective Distillation Attribution via Normalized Response Similarity

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of tracing distilled models after post-training, where historical checkpoints and model weights are typically unavailable. To this end, we propose SCOUT, a framework that introduces the first output-only attribution paradigm. Without requiring access to model parameters or pre-distillation data, this method achieves weight-free and history-free distillation attribution by mining syntactic patterns in generated text, applying low-contrast filtering, and employing a normalized distance calibration algorithm. Experimental results demonstrate that syntactic signatures from teacher models persist and remain traceable throughout post-training stages, including preference optimization and reinforcement learning. Furthermore, the proposed framework supports real-time auditing and abstention decisions for publicly released models, offering a novel tool for AI safety governance.
📝 Abstract
Model distillation transfers capabilities through supervised fine-tuning (SFT) on teacher responses, often collected from commercial APIs, raising questions of model provenance. Existing distillation attribution methods have been largely evaluated on students immediately after the SFT step. However, a distilled model may undergo further SFT, preference optimization, or reinforcement learning before release, while an auditor may lack access to the pre-distillation checkpoint required by reference-based attribution. To close this gap, we propose SCOUT, an output-only method that aggregates recurring *syntactic patterns* into candidate profiles, filters low-contrast patterns, and calibrates student--candidate distances against inter-candidate distances. SCOUT supports attribution and abstention using only current texts, without model weights, token likelihoods, or historical checkpoints. Auditing publicly released descendants of distilled models spanning diverse post-training objectives, SCOUT consistently identifies the distillation source. Furthermore, tracing teacher-associated *syntactic signatures* along training trajectories reveals that they emerge during distillation and persist through subsequent preference optimization and reinforcement learning.
Problem

Research questions and friction points this paper is trying to address.

Model Distillation
Distillation Attribution
Model Provenance
Post-training Auditing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Distillation Attribution
Syntactic Patterns
Output-only Method
Model Provenance
Normalized Response Similarity
🔎 Similar Papers