🤖 AI Summary
This work addresses the challenge of effectively aggregating heterogeneous probability distributions in multi-teacher knowledge distillation by proposing the first axiomatic framework for ensemble distillation. It formulates five core axioms that any valid knowledge aggregation operator should satisfy in the probability space and, leveraging operator theory and convex analysis, proves the existence—and non-uniqueness—of a family of operators fulfilling these axioms. The framework dispenses with the common assumption of teacher homogeneity, establishing generalization error and stability guarantees that are invariant to teacher heterogeneity. Theoretically, it reveals that multi-teacher aggregation simultaneously reduces both stochastic variance and systematic supervision bias, provides an upper bound on log-loss, ensures safe decay properties, and extends classical ensemble learning’s variance-reduction results to settings involving correlated errors.
📝 Abstract
Building on the probability-domain distillation framework of Sparse-KD, we develop an axiomatic, operator-theoretic framework for multi-teacher ensemble knowledge distillation. Rather than prescribing a specific aggregation formula, we define five core axioms governing valid knowledge aggregation operators, encompassing convexity, positivity, continuity, weight monotonicity, and temperature coherence. We prove the existence and non-uniqueness of operator families satisfying these axioms, establishing that multiple distinct aggregation mechanisms conform to the same foundational principles. Within this framework, we establish operator-agnostic guarantees showing that multi-teacher aggregation reduces both stochastic variance and systematic supervisory bias under heterogeneous teachers, while providing Jensen-type bounds, log-loss guarantees, and safety attenuation properties. For aggregation operators linear in teacher weights, we further establish classical ensemble variance-reduction results under standard independence assumptions, with extensions to correlated-error regimes. The framework provides theoretical grounding for multi-teacher distillation from diverse frontier models while admitting multiple valid implementation strategies.