MoLE: Mixture of Latent Experts for Complementary Visual Reasoning

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the information redundancy and lack of complementarity caused by shared projections in latent visual reasoning. To this end, we propose MoLE, a framework that introduces a novel Mixture of Latent Experts mechanism. By decoupling evidence extraction from dedicated aggregation experts, MoLE achieves computational specialization and complementary visual information mining without predefined roles. The framework employs a two-stage training paradigm, latent summary aggregation, and complementary representation isolation strategies. Experimental results demonstrate that MoLE attains an average score of 78.6 across five benchmarks, outperforming the baseline by 3.6 points. Furthermore, it significantly reduces latent state similarity while enhancing attention diversity, establishing its superiority over approaches that merely increase token counts.
📝 Abstract
Latent visual reasoning equips vision--language models with continuous intermediate states that can process visual evidence without explicit textual reasoning traces or repeated image operations. However, existing methods often allow multiple latent tokens to access the same visual evidence through shared value projections, providing no mechanism for them to extract complementary visual information; simply increasing the latent budget can therefore yield redundant latent representations. We argue that effective latent reasoning should encourage different latent tokens to extract complementary visual information, and thereby act as specialized visual experts. Based on this insight, we propose MoLE, a Mixture of Latent Experts framework that controls both what visual evidence each latent visual expert observes and how it transforms that evidence. MoLE isolates latent visual experts during evidence extraction and uses dedicated latent summary experts to aggregate the complementary representations of latent visual experts. A two-stage training pipeline first forces visual evidence through this latent pathway and then restores direct visual access, requiring neither predefined expert roles nor intermediate visual targets. Across five visual reasoning benchmarks, MoLE achieves an average score of 78.6, outperforming data-matched supervised fine-tuning by 4.9 and the strongest evaluated latent visual reasoning baseline at the same latent budget by 3.6. Representation analyses show lower latent-state similarity and more diverse visual attention, while masking the latent pathway reduces average performance by 9.2. These results demonstrate that specializing latent computation is more effective than merely increasing the number of latent tokens.
Problem

Research questions and friction points this paper is trying to address.

Latent Visual Reasoning
Vision-Language Models
Complementary Visual Information
Redundant Representations
Mixture of Experts
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture of Latent Experts
Complementary Visual Reasoning
Latent Visual Experts
Two-stage Training Pipeline
Vision-Language Models
🔎 Similar Papers