🤖 AI Summary
Existing vision Mixture-of-Experts (MoE) approaches perform routing at the image or patch level, which struggles to align with the instance-centric nature of object detection. This work proposes a hierarchical instance-conditioned MoE architecture that introduces a two-stage routing mechanism—operating at both scene and instance levels—within a DETR-style detector, achieving fine-grained expert assignment aligned with instance queries for the first time. The method employs a lightweight scene router and an instance router that jointly balance sparse computation with the heterogeneity of individual instances. Experiments demonstrate that the model outperforms dense DINO baselines and simplified routing variants on COCO, significantly enhancing small object detection performance and offering preliminary evidence of functional specialization among experts.
📝 Abstract
Mixture-of-Experts (MoE) architectures enable conditional computation by activating only a subset of model parameters for each input. Although sparse routing has been highly effective in language models and has also shown promise in vision, most vision MoE methods operate at the image or patch level. This granularity is poorly aligned with object detection, where the fundamental unit of reasoning is an object query corresponding to a candidate instance. We propose Hierarchical Instance-Conditioned Mixture-of-Experts (HI-MoE), a DETR-style detection architecture that performs routing in two stages: a lightweight scene router first selects a scene-consistent expert subset, and an instance router then assigns each object query to a small number of experts within that subset. This design aims to preserve sparse computation while better matching the heterogeneous, instance-centric structure of detection. In the current draft, experiments are concentrated on COCO with preliminary specialization analysis on LVIS. Under these settings, HI-MoE improves over a dense DINO baseline and over simpler token-level or instance-only routing variants, with especially strong gains on small objects. We also provide an initial visualization of expert specialization patterns. We present the method, ablations, and current limitations in a form intended to support further experimental validation.