🤖 AI Summary
This study addresses the limitation of existing sparse vision-language models (VLMs) that rely on fixed cross-modal interfaces, rendering them incapable of dynamically selecting optimal intermediate visual representations based on specific queries. To overcome this, we propose a sparse VLM coupled with visual depth routing, which integrates a Mixture-of-Experts (MoE) architecture with dynamic visual injection to adaptively select visual depths for each image-prompt pair, thereby achieving a question-aware visual access mechanism. This approach effectively couples visual depth routing with expert routing while preserving native visual token sequences. With 28B total parameters and only 9B activated, the proposed model achieves an average score of 85.9 across eight benchmarks, demonstrating the substantial architectural gains yielded by the dynamic routing interface.
📝 Abstract
Vision-language models (VLMs) increasingly use sparse mixture-of-experts (MoE) to scale language-side computation, yet visual information is typically routed only after passing through a fixed cross-modal interface. This leaves an important decision unresolved: which intermediate visual representations should be exposed to language computation for a given question? We introduce OmniMoE-VL, a sparse VLM with a coupled visual-depth routed projector. For each image-prompt pair, the projector selects a sparse set of intermediate visual depths and reuses the resulting global preference to guide both local patch fusion and dynamic visual injection into the language model. This design enables question-dependent visual access while preserving the native visual-token sequence, and complements token-level expert routing in the vision and language stacks. Across eight image-based benchmarks, OmniMoE-VL achieves an average score of 85.9 with 28B total and 9B activated parameters. Controlled comparisons show that the routed visual interface provides the dominant architectural gain, while matched route and component controls, same-image route analysis, and route interventions further support the value of coupling and question-conditioned visual access.