MG-Thinker: Bi-Axial Self-Reflection for Multi-Image Reasoning Grounding

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of hierarchical reasoning and the heterogeneity of sample difficulty in multi-image reasoning grounding by proposing a reinforcement learning-based post-training framework. Methodologically, it integrates Chain-of-Thought (CoT) annotations to facilitate coarse-to-fine hierarchical reasoning and pixel-level precise localization. Furthermore, this work introduces the BiA-DAPO algorithm, which decomposes advantages along two axes—intra-group signals and inter-group capabilities—to effectively handle difficulty heterogeneity. The proposed approach achieves state-of-the-art performance on multi-image reasoning grounding tasks and significantly enhances cross-benchmark generalization capabilities.
📝 Abstract
Reinforcement learning (RL) has recently delivered substantial gains in multimodal reasoning, opening a promising route for fine-grained visual perception. Yet for multi-image reasoning grounding (MRG), reasoning over real-world multi-image contexts toward pixel-precise localization, existing RL-based approaches overlook two characteristics intrinsic to this paradigm: a coarse-to-fine hierarchical reasoning pattern, and heterogeneously distributed task--sample difficulties. In this work, we present MG-Thinker, a post-training RL framework that advances a new MRG paradigm featuring such hierarchical reasoning, supported by a curated 25K MRG dataset with task-adaptive Chain-of-Thought (CoT) annotations that elicit multi-perspective evidence before conclusion. To remedy the heterogeneous task--sample difficulties, we further propose Bi-Axial DAPO (BiA-DAPO), which decomposes rollout advantages along an intra-group signal axis and an inter-group competence axis through two complementary mechanisms, both grounded on our defined candidate pool for stable group-level statistics. Extensive experiments show that MG-Thinker achieves state-of-the-art performance on multi-image reasoning grounding while consistently improving generalization across multi-image understanding and diverse multimodal benchmarks.
Problem

Research questions and friction points this paper is trying to address.

Multi-Image Reasoning Grounding
Reinforcement Learning
Hierarchical Reasoning
Heterogeneous Difficulty
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Image Reasoning Grounding
Reinforcement Learning
Bi-Axial DAPO
Chain-of-Thought
Hierarchical Reasoning
🔎 Similar Papers
No similar papers found.