🤖 AI Summary
This study addresses the challenges of memory constraints and high decoding latency caused by expert offloading in Mixture-of-Experts (MoE) models, as well as the difficulty of adapting experts to new routing patterns via existing fine-tuning methods. To overcome these issues, we propose a masked co-adaptive fine-tuning approach that jointly optimizes routers and experts through learnable binary masks. During inference, the trained masks are converted into soft priors for dynamic route re-ranking, reshaping inference priorities while preserving full expert accessibility. Experimental results demonstrate that our method reduces expert fetching operations by 10.1%–23.7% and generation latency by 5.5%–16.4%, while improving average accuracy by 0.53–0.92 points. These findings confirm that the proposed approach achieves synergistic optimization of both efficiency and performance in MoE systems.
📝 Abstract
Mixture-of-experts (MoE) language models often exceed the memory of a single GPU. Expert offloading keeps most experts in host memory and loads them on demand, so decoding speed depends on how many experts each token must fetch. Caching and prefetching reduce this cost only as far as the routing allows. Router-only fine-tuning can reshape the routing to reuse experts, but it keeps the experts frozen, so they cannot adapt to the tokens the new routing sends them. We propose MaskCoFT, a masked co-adaptive fine-tuning method that trains routers and experts together with the cross-entropy loss alone. During fine-tuning, a learnable binary mask restricts the Top-K routing of each layer to a subset of experts, and the experts adapt to the tokens redirected to them. At inference, the learned mask becomes a soft prior that re-ranks experts, so every expert remains selectable. We simulate a GPU cache of 4 experts per layer for Mixtral-8x7B and 12 for DeepSeek-V2-Lite. MaskCoFT cuts expert fetches per token by 23.7% and 10.1% relative to the base model. In real offloading system serving, it lowers the time per output token by up to 16.4% and 5.5%, respectively. Its average accuracy over nine benchmarks stays above the base model by 0.92 and 0.53 points.