π€ AI Summary
This study addresses the sensitivity of performance to learning rates in federated online distillation, which arises from the coupling between aggregation and generated feedback. Specifically, it reveals an optimization lag mechanism induced by the dual role of student models. To mitigate this issue, this work proposes FedTOPS, a method that theoretically analyzes the aforementioned coupling effect and designs an adaptive scaling strategy based on predictive change constraints to enable dynamic updates. By integrating federated learning with online knowledge distillation, FedTOPS substantially enhances multi-model collaborative training. Experimental results across six benchmarks demonstrate that the proposed approach improves macro-average scores by 4.56 to 14.57 percentage points over FedAvg, confirming its effectiveness in optimizing federated online distillation frameworks.
π Abstract
On-policy distillation (OPD) is a promising approach to language-model adaptation, aligning teacher supervision with the student's own generated trajectories. When adaptation prompts are distributed across clients, can this process benefit from federated collaboration? We study federated OPD and find that substantial collaboration gains can be obscured by learning-rate sensitivity: FedAvg can perform no better than independent local training at a small learning rate, yet recover a clear advantage at a larger rate. We explain this phenomenon through the student's dual role as learner and generator of future training data. An aggregation-induced optimization lag can delay access to useful teacher supervision, which in turn slows subsequent learning. Our theory establishes this aggregation--rollout feedback in a solvable model with a common optimum and stable updates, and identifies two coupled roles of learning rate: learning from current supervision and reaching future supervision. Guided by this analysis, we propose FedTOPS (Federated Teacher-guided On-Policy Scaling), which reuses teacher feedback on current trajectories to adapt the FedAvg update magnitude under clientwise predictive-change constraints. Across six mathematical reasoning benchmarks, FedTOPS improves macro Avg@8 over FedAvg by 4.56--14.57 percentage points across the evaluated student models and local learning rates.