Mixture of Layers: Dynamic Layer Routing for Visual Reasoning

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of multimodal large language models that rely on fixed visual encoder layers, hindering query-adaptive extraction of fine-grained visual cues. To overcome this, we propose MoL, a framework introducing the first text-instruction-driven dynamic sparse layer routing mechanism. Inspired by Mixture-of-Experts architectures, MoL dynamically fuses intermediate-layer features at the visual patch level through Top-K sparse aggregation and image/patch-level mixture routing, enabling query-aware visual reasoning. This approach effectively circumvents the conventional fixed-aggregation bottleneck without incurring additional computational overhead, substantially enhancing fine-grained perceptual capabilities. Extensive evaluations demonstrate significant performance gains of 18.9%, 4.5%, and 16.3% on the V*, HRBench4K, and CharXiv benchmarks, respectively, validating the efficacy of dynamic sparse routing for adaptive visual representation learning.
📝 Abstract
Pre-trained vision encoders contain layer-wise visual representations that differ in spatial granularity, semantic abstraction, and sensitivity to local details. However, most Multimodal Large Language Models (MLLMs) rely on only the final or penultimate vision encoder representations or fixed aggregation rules, making visual abstraction largely query-agnostic and limiting access to fine-grained cues such as small objects, spatial details, text, and subtle visual attributes. In this work, we propose Mixture of Layers (MoL), an instruction-conditioned layer routing approach at the visual patch level that dynamically aggregates query-relevant latent representations from intermediate vision encoder layers. Given a text query, MoL predicts routing probabilities over vision encoder layers and performs a top-k sparse aggregation over selected hidden states at either the image level, patch level, or through a hybrid routing mechanism. In doing so, MoL enables query-adaptive access to layer-specific visual features for fine-grained visual reasoning. Our experiments across 7 fine-grained visual reasoning tasks demonstrate substantial performance improvements, especially across fine-grained visual grounding and understanding tasks such as +18.9% improvement on V* in overall accuracy, +4.5% on HRBench4K, and +16.3% on CharXiv compared to the baseline MLLMs, without resorting to multi-resolution inputs, simple interleaving of multiple vision encoders, or increasing the number of patch tokens. We study vision encoders' receptive field scales across different layers and their sampling behaviors to provide an in-depth analysis of why layer-wise sampling is helpful, demonstrating that conditional visual representations are a key step towards better visual perception and reasoning in MLLMs. Our project page is available at https://wjdghks950.github.io/mol.github.io/.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Visual Reasoning
Vision Encoder Representations
Fine-grained Visual Cues
Query-agnostic
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture of Layers
Dynamic Layer Routing
Multimodal Large Language Models
Fine-grained Visual Reasoning
Sparse Aggregation