iReasoner: Trajectory-Aware Intrinsic Reasoning Supervision for Self-Evolving Large Multimodal Models

๐Ÿ“… 2026-01-09
๐Ÿ›๏ธ arXiv.org
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Current self-evolving large models rely solely on final outcome rewards, lacking supervision over intermediate reasoning steps, which limits their visual reasoning capabilities. This work proposes an unsupervised self-evolution framework that explicitly elicits chain-of-thought reasoning on unlabeled images through a Proposer-Solver loop and introduces a trajectory-aware intrinsic reward mechanism to provide fine-grained supervision of the consistency and quality of intermediate reasoning stepsโ€”without requiring ground-truth labels or external evaluators. Built upon Qwen2.5-VL-7B, the proposed method achieves, for the first time, fully unsupervised optimization of reasoning paths, yielding an average improvement of 2.1 points across multiple multimodal reasoning benchmarks and significantly enhancing the modelโ€™s self-evolution capacity.

Technology Category

Computer Vision: Large Vision ModelsMachine Learning: Unsupervised & Self-Supervised LearningSearch and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsEconomics, Online Markets and Human Computation: LLM based quality controls for crowd work
๐Ÿ“ Abstract
Recent work shows that large multimodal models (LMMs) can self-improve from unlabeled data via self-play and intrinsic feedback. Yet existing self-evolving frameworks mainly reward final outcomes, leaving intermediate reasoning weakly constrained despite its importance for visually grounded decision making. We propose iReasoner, a self-evolving framework that improves an LMM's implicit reasoning by explicitly eliciting chain-of-thought (CoT) and rewarding its internal agreement. In a Proposer--Solver loop over unlabeled images, iReasoner augments outcome-level intrinsic rewards with a trajectory-aware signal defined over intermediate reasoning steps, providing learning signals that distinguish reasoning paths leading to the same answer without ground-truth labels or external judges. Starting from Qwen2.5-VL-7B, iReasoner yields up to $+2.1$ points across diverse multimodal reasoning benchmarks under fully unsupervised post-training. We hope this work serves as a starting point for reasoning-aware self-improvement in LMMs in purely unsupervised settings.
Problem

Research questions and friction points this paper is trying to address.

self-evolving
large multimodal models
intrinsic reasoning
trajectory-aware
unsupervised learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

trajectory-aware reasoning
intrinsic reward
self-evolving LMMs
chain-of-thought supervision
unsupervised multimodal learning
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
Meghana Sunil
Meghana Sunil
Student at VIT Chennai
Artificial IntelligenceMachine LearningComputer Vision
M
Manikandarajan Venmathimaran
Loughborough University
M
Muthu Subash Kavitha
Nagasaki University