🤖 AI Summary
This study addresses the insufficient expert specialization in sparse Mixture-of-Experts models caused by conflicts between load balancing and gradient alignment. To mitigate this, we propose Gradient-Aligned Routing (GAR), which reformulates routing as a gradient partitioning problem. By decoupling load balancing from gradient-based routing organization, GAR rewards samples with consistent gradients within groups to optimize multi-task classification performance. We validate our method across five multi-task text classification datasets using RoBERTa, DeBERTa, and Qwen3 backbones integrated with low-rank adapter experts. Experimental results demonstrate that GAR improves accuracy by approximately 1.1 percentage points over task-loss-only routing while significantly enhancing gradient purity. These findings highlight the critical value of leveraging gradient information for strengthening expert specialization.
📝 Abstract
Sparse expert models can distribute traffic evenly while still grouping incompatible training signals within the same experts. We study routing as a gradient-partitioning problem and introduce gradient-aligned routing (GAR), whose load-normalized router objective rewards grouping observations with aligned gradients. On five multi-task text-classification mixtures, we compare GAR with task-loss-only routing, gradient-combination and gradient-conflict methods, and load-balancing losses. With a fully trainable RoBERTa backbone and classification-head experts, GAR has the highest aggregate validation accuracy, 1.07 percentage points above task-loss-only routing. With frozen DeBERTa and Qwen3-1.7B backbones and low-rank adapter experts, it again ranks first, 1.10 points above task-loss-only routing, with better-balanced expert load and higher gradient-mass purity, the share of each expert's gradient-norm mass from its dominant task; the load-balancing losses flatten load further but leave this purity near its task-loss-only level. Top-1 routing, trainable full-parameter feed-forward network (FFN) experts, and a larger backbone also show positive aggregate gains. The results distinguish expert-load balance from gradient-based routing organization and indicate the predictive value of gradient-informed routing in multi-task text classification.