🤖 AI Summary
This study addresses the fragmentation of capabilities across vision foundation models and the need for repetitive training in multi-teacher distillation by proposing GRAFT, a framework that constructs a unified backbone through continual multi-teacher distillation. Methodologically, this work pioneers a continual distillation paradigm, introducing task-specific readout tokens and geometry-agnostic relational losses to effectively accommodate and align heterogeneous representations. By unifying five major visual task domains within a single backbone network, GRAFT enables the acquisition of new capabilities through a single distillation step while maintaining superior performance. This approach offers an efficient pathway for integrating multi-source visual knowledge, significantly reducing computational overhead compared to conventional retraining strategies.
📝 Abstract
Vision foundation models such as DINOv2, SigLIP2, and MASt3R develop complementary capabilities from different pretraining objectives, yet their knowledge remains distributed across separate, specialized models. Multi-teacher knowledge distillation offers a path toward consolidating these capabilities into a single agglomerative backbone, but existing approaches assume a fixed set of teachers, and incorporating a new teacher requires repeating expensive joint distillation over the entire teacher set. We introduce GRAFT, a continual multi-teacher distillation framework that enables a unified backbone to progressively acquire capabilities from an open-ended sequence of foundation models. When a new teacher arrives, GRAFT treats the previously distilled model as a teacher for preserving learned capabilities, while the current student jointly learns from both the previous model and the incoming teacher. Furthermore, to reconcile the incompatible representation geometries of heterogeneous teachers, we introduce Teacher Specific Readout Tokens, which grant each teacher an independent read-out of the shared encoder, together with Geometry Agnostic Relational Loss that aligns a vision-language teacher by matching image-text similarity structures rather than raw feature values. We provide GRAFT model, which is a single, continually extensible backbone that unifies five domains, including image understanding, 2D dense prediction, 3D human pose estimation, 3D vision, and vision-language, delivering strong performance across all of them while acquiring each new capability at the cost of a single distillation rather than a full re-distillation.