HI-MoE: Hierarchical Instance-Conditioned Mixture-of-Experts for Object Detection

📅 2026-04-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing vision Mixture-of-Experts (MoE) approaches perform routing at the image or patch level, which struggles to align with the instance-centric nature of object detection. This work proposes a hierarchical instance-conditioned MoE architecture that introduces a two-stage routing mechanism—operating at both scene and instance levels—within a DETR-style detector, achieving fine-grained expert assignment aligned with instance queries for the first time. The method employs a lightweight scene router and an instance router that jointly balance sparse computation with the heterogeneity of individual instances. Experiments demonstrate that the model outperforms dense DINO baselines and simplified routing variants on COCO, significantly enhancing small object detection performance and offering preliminary evidence of functional specialization among experts.

Technology Category

Machine Learning: Mixture of Experts (MoE)Computer Vision: Multi-modal VisionSearch and Optimization: Learning to Search

Application Category

Systems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applicationsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Representation, reconstruction, and subgraph or motif discovery in Web-related graphs
📝 Abstract
Mixture-of-Experts (MoE) architectures enable conditional computation by activating only a subset of model parameters for each input. Although sparse routing has been highly effective in language models and has also shown promise in vision, most vision MoE methods operate at the image or patch level. This granularity is poorly aligned with object detection, where the fundamental unit of reasoning is an object query corresponding to a candidate instance. We propose Hierarchical Instance-Conditioned Mixture-of-Experts (HI-MoE), a DETR-style detection architecture that performs routing in two stages: a lightweight scene router first selects a scene-consistent expert subset, and an instance router then assigns each object query to a small number of experts within that subset. This design aims to preserve sparse computation while better matching the heterogeneous, instance-centric structure of detection. In the current draft, experiments are concentrated on COCO with preliminary specialization analysis on LVIS. Under these settings, HI-MoE improves over a dense DINO baseline and over simpler token-level or instance-only routing variants, with especially strong gains on small objects. We also provide an initial visualization of expert specialization patterns. We present the method, ablations, and current limitations in a form intended to support further experimental validation.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
object detection
instance-level routing
sparse computation
expert specialization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
Object Detection
Instance-Conditioned Routing
Sparse Computation
DETR-style Architecture
💼 Related Jobs
No related jobs found.
V
Vadim Vashkelis
N
Natalia Trukhina