๐ค AI Summary
Current self-evolving large models rely solely on final outcome rewards, lacking supervision over intermediate reasoning steps, which limits their visual reasoning capabilities. This work proposes an unsupervised self-evolution framework that explicitly elicits chain-of-thought reasoning on unlabeled images through a Proposer-Solver loop and introduces a trajectory-aware intrinsic reward mechanism to provide fine-grained supervision of the consistency and quality of intermediate reasoning stepsโwithout requiring ground-truth labels or external evaluators. Built upon Qwen2.5-VL-7B, the proposed method achieves, for the first time, fully unsupervised optimization of reasoning paths, yielding an average improvement of 2.1 points across multiple multimodal reasoning benchmarks and significantly enhancing the modelโs self-evolution capacity.
๐ Abstract
Recent work shows that large multimodal models (LMMs) can self-improve from unlabeled data via self-play and intrinsic feedback. Yet existing self-evolving frameworks mainly reward final outcomes, leaving intermediate reasoning weakly constrained despite its importance for visually grounded decision making. We propose iReasoner, a self-evolving framework that improves an LMM's implicit reasoning by explicitly eliciting chain-of-thought (CoT) and rewarding its internal agreement. In a Proposer--Solver loop over unlabeled images, iReasoner augments outcome-level intrinsic rewards with a trajectory-aware signal defined over intermediate reasoning steps, providing learning signals that distinguish reasoning paths leading to the same answer without ground-truth labels or external judges. Starting from Qwen2.5-VL-7B, iReasoner yields up to $+2.1$ points across diverse multimodal reasoning benchmarks under fully unsupervised post-training. We hope this work serves as a starting point for reasoning-aware self-improvement in LMMs in purely unsupervised settings.