🤖 AI Summary
This study addresses the challenge in knowledge distillation where varying data budgets shift the optimal teacher capacity, complicating effective sample selection. By analyzing relational ranking and score geometry, this work reveals for the first time why smaller teachers excel under low-data regimes and proposes the DVA framework. Employing a small teacher as a proxy, DVA models relational diversity through difficulty filtering and class-conditional volume maximization, establishing a training-free dynamic data selection mechanism that jointly optimizes difficulty matching and signal diversity. Without requiring training-dependent dynamic statistics, the proposed approach achieves performance comparable to dynamic methods while consistently outperforming static baselines, offering a novel paradigm for efficient knowledge distillation.
📝 Abstract
Data pruning reduces the training cost of knowledge distillation (KD). However, the preferred teacher capacity changes with the data budget: smaller teachers can outperform larger ones when limited training data are available. Understanding what drives this shift is important not only for teacher choice but also for identifying which samples are useful for distillation. We analyze teacher supervision by decomposing it into relational ordering---the ranking of classes---and score geometry---the magnitudes and margins of class probabilities---and show that the small-teacher advantage in the low-data regime arises not only from score geometry but also from relational ordering. Beyond understanding teacher capacity, our analysis reveals two properties of effective subsets: samples should match the difficulty appropriate for the available budget, and their relational signals should be diverse rather than redundant. Based on these findings, we propose DVA (Difficulty- and Volume-Aware data selection for KD), a training-dynamics-free method, which uses a small teacher as a proxy for budget-aware difficulty filtering and class-conditional relational volume maximization. Despite requiring no training dynamics statistics, our method remains competitive with training-dynamics-based methods while consistently outperforming training-dynamics-free baselines.