Routing in Gradient Space: Balanced Usage Is Not Expert Specialization

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficient expert specialization in sparse Mixture-of-Experts models caused by conflicts between load balancing and gradient alignment. To mitigate this, we propose Gradient-Aligned Routing (GAR), which reformulates routing as a gradient partitioning problem. By decoupling load balancing from gradient-based routing organization, GAR rewards samples with consistent gradients within groups to optimize multi-task classification performance. We validate our method across five multi-task text classification datasets using RoBERTa, DeBERTa, and Qwen3 backbones integrated with low-rank adapter experts. Experimental results demonstrate that GAR improves accuracy by approximately 1.1 percentage points over task-loss-only routing while significantly enhancing gradient purity. These findings highlight the critical value of leveraging gradient information for strengthening expert specialization.
📝 Abstract
Sparse expert models can distribute traffic evenly while still grouping incompatible training signals within the same experts. We study routing as a gradient-partitioning problem and introduce gradient-aligned routing (GAR), whose load-normalized router objective rewards grouping observations with aligned gradients. On five multi-task text-classification mixtures, we compare GAR with task-loss-only routing, gradient-combination and gradient-conflict methods, and load-balancing losses. With a fully trainable RoBERTa backbone and classification-head experts, GAR has the highest aggregate validation accuracy, 1.07 percentage points above task-loss-only routing. With frozen DeBERTa and Qwen3-1.7B backbones and low-rank adapter experts, it again ranks first, 1.10 points above task-loss-only routing, with better-balanced expert load and higher gradient-mass purity, the share of each expert's gradient-norm mass from its dominant task; the load-balancing losses flatten load further but leave this purity near its task-loss-only level. Top-1 routing, trainable full-parameter feed-forward network (FFN) experts, and a larger backbone also show positive aggregate gains. The results distinguish expert-load balance from gradient-based routing organization and indicate the predictive value of gradient-informed routing in multi-task text classification.
Problem

Research questions and friction points this paper is trying to address.

Sparse Expert Models
Gradient-aligned Routing
Multi-task Learning
Expert Specialization
Text Classification
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gradient-aligned routing
Sparse expert models
Multi-task learning
Gradient partitioning
Expert specialization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yuchen Li
University of Sydney
M
Mingyu Du
University of Sydney, University of New South Wales
Z
Zongqi Fan
University of Sydney
Nguyen H. Tran
Nguyen H. Tran
The University of Sydney
Distributed compUtingoptimizAtionmachine Learning (DUAL)
Ken-Tye Yong
Ken-Tye Yong
The University of Sydney
PhysicsBiophotonicsNanophotonicsSensors and ActuatorsNanomaterials