MMVistaReason: Toward Open-Data and Post-Training Recipes for Multimodal Reasoning

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of uneven data quality, inefficient supervision construction, and cross-domain interference in post-training for multimodal reasoning by proposing an open-data post-training framework. Methodologically, we design difficulty-aware cascaded distillation and Multi-teacher Online Policy Distillation (MOPD) mechanisms to achieve capacity-aware robust training. Furthermore, we employ phased data cleaning and trajectory selection strategies, combined with a "specialize-then-integrate" paradigm and expert reinforcement learning, to efficiently construct SFT and RL datasets. Experimental results demonstrate that MVR-4B achieves an average score of 72.8 across 15 benchmarks, outperforming larger models while utilizing 70% fewer samples, with the 9B variant yielding further performance improvements.
📝 Abstract
Open multimodal reasoning models have benefited from large-scale reasoning supervision, yet reliable post-training remains challenging due to uneven data quality, inefficient supervision construction, imbalanced difficulty, and cross-domain interference. We introduce MMVistaReason (MVR), an open-data post-training recipe with three components: (1) broader capability coverage across complementary Analytical and Real-World reasoning groups, emphasizing structured reasoning versus visual perception and spatial grounding; (2) efficient SFT and RL data construction, standardizing heterogeneous open data through staged cleaning and annotation, combining difficulty-aware cascaded teacher distillation with answer-likelihood-based trajectory selection to construct MVR-SFT-528K, and applying scale-specific frontier filtering for MVR-RL-63K; and (3) specialize-then-integrate training, which trains complementary RL experts and consolidates their capabilities through multi-teacher on-policy distillation (MOPD). Our analyses reveal a capacity-dependent interaction between supervision difficulty, trajectory quality, and model capacity: smaller students benefit more from selected supervision, while larger students are robust to trajectory variation and mixed-domain interference. Mixed-domain RL introduces benchmark-level negative transfer, whereas MOPD provides consistent capability integration, with the preferred KL direction varying across model scales. Across 15 multimodal benchmarks, MVR-4B achieves an average score of 72.8, outperforming Qwen3.5-9B (Instruct) and MMFineReason-8B while using about 70% fewer samples than MMFineReason. Scaling to 9B improves the average to 74.4, surpassing Qwen3.5-35B-A3B (Instruct). Overall, MMVistaReason demonstrates that systematic open-data construction and capacity-aware post-training provide a practical and scalable path toward reliable multimodal reasoning.
Problem

Research questions and friction points this paper is trying to address.

multimodal reasoning
post-training
data quality
cross-domain interference
supervision construction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Reasoning
Post-Training
Teacher Distillation
Multi-teacher On-policy Distillation
Data Construction
🔎 Similar Papers