Adaptive Mutual Distillation for Balanced Multi-Task Post-Training of Large Language Models

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of data imbalance and heterogeneous cross-model supervision utility in multi-task post-training of large language models. To this end, we propose an adaptive mutual distillation framework leveraging short-trained probes and validation scores. This approach introduces a novel bidirectional complementary knowledge transfer mechanism, wherein probes dynamically calibrate distillation weights. Furthermore, it integrates multi-task balancing strategies with model merging techniques to jointly optimize both models. Experimental results demonstrate that the proposed framework consistently surpasses supervised fine-tuning (SFT) baselines across six benchmarks. Notably, the merged model achieves an average improvement of 2.91 points, substantially enhancing comprehensive multi-task performance.
📝 Abstract
Multi-task post-training of large language models (LLMs) aims to improve performance across tasks with unequal amounts of training data. Existing methods focus primarily on balancing task contributions during single-model training. Different task-balancing strategies can produce models with complementary strengths, creating opportunities for mutual distillation. However, the usefulness of cross-model supervision can vary across tasks, transfer directions, and stages of training. We propose Adaptive Mutual Distillation (AMD), a collaborative post-training framework that jointly trains two models with different task-balancing strategies. AMD evaluates candidate adjustments to distillation weights through short training probes shared across tasks, then uses task-wise validation scores to select an adjustment for each task and transfer direction. Across six benchmarks and three LLM backbones, both AMD models achieve higher average benchmark scores than supervised fine-tuning (SFT) baselines trained with the same sampling strategies. They also outperform the task-balancing methods evaluated in our experiments. Merging the two trained models can further improve their average benchmark score while yielding a single model for inference. The merged models outperform multi-task SFT by an average of 2.91 points across the three backbones.
Problem

Research questions and friction points this paper is trying to address.

Multi-task post-training
Large language models
Mutual distillation
Task balancing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Mutual Distillation
Multi-Task Post-Training
Task Balancing
Model Merging
Large Language Models
🔎 Similar Papers
No similar papers found.