🤖 AI Summary
This work addresses the persistent challenge of intra- and inter-modal load imbalance in vision–language mixture-of-experts models, which arises from the dynamic variation in image and text token counts and is inadequately mitigated by existing auxiliary losses. To this end, the authors propose ReBA, a novel routing mechanism that incorporates modality awareness and image boundary perception. ReBA introduces modality-separated auxiliary losses derived from geometric structure analysis, relaxes constraints within images, and enforces equal-weight routing at the image level. This approach enables stable expert load balancing across diverse input configurations—including varying resolutions, patch sizes, and prompt lengths—without compromising task accuracy. Extensive experiments across four backbone architectures demonstrate that ReBA consistently achieves superior load balance, significantly reducing both average and worst-case imbalance compared to all baseline methods.
📝 Abstract
Vision-language MoE batches contain different numbers of image and text tokens. Image resolution, image count, tiling, and prompt length all change this token mix. We call the standard token-level Switch auxiliary loss Std-Aux. Std-Aux balances only the mixed load, so large image and text load errors can cancel at one mix. On our main model, the same trained router shows more than a fivefold change in load imbalance across image resolutions. We hold the image and text load profiles fixed and derive the exact load curve as the token mix varies. The image-text load gap controls sensitivity to the token mix. Physical preprocessing can also change the conditional profiles. The fixed-profile law excludes such changes. To design a remedy, we examine the router input structure. Image and text occupy distinct regions, while visual tokens group strongly by source image. The modality boundary motivates separate image and text terms. The image boundary motivates one equal-weight routing instance per image. ReBA, or Relax Within, Balance Across, implements both choices. Across four split backbones, ReBA lowers load on every reported benchmark input while keeping mean task accuracy comparable to Std-Aux. ReBA also lowers average load over the tested range and worst physical load under resolution and tiling shifts. Code is available at https://github.com/ZiangWu-77/ReBA.