Rethinking Multi-Image Re-Representation in Multi-Image Understanding

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of effectively organizing dispersed visual evidence in multi-image understanding with multimodal large language models by proposing the Mosaic framework. It defines ten composable image operations to actively construct intermediate visual states and introduces a novel reinforcement learning approach that relies solely on accuracy-based rewards to train MosaicAgent, enabling it to autonomously acquire multi-step visual reasoning and tool chaining capabilities without demonstration trajectories. The study further reveals the task-dependent advantages of textual versus visual re-representation. Experimental results demonstrate that MosaicAgent-8B exhibits strong autonomous planning capabilities on benchmarks such as MosaicBench, yielding significant performance improvements on fine-grained comparison and orientation-sensitive tasks.
📝 Abstract
Multi-image understanding requires MLLMs not only to recognise the content of individual images, but also to organise visual evidence distributed across them. We study this problem through multi-image re-representation, viewing prompted Chain-of-Thought reasoning and agentic visual tool use as different ways of re-organising visual evidence during reasoning. We introduce Mosaic, a general-purpose multi-image visual harness that enables an MLLM to actively construct visual intermediates with ten composable image operations. We compare five re-representation settings on existing multi-image benchmarks and on MosaicBench, a new grounding-focused benchmark for fine-grained multi-image understanding. Our experiments show that the relative benefits of textual and visual re-representation are strongly task-dependent. Visual re-representation is particularly effective for tasks requiring precise visual evidence, including hypothesis testing, precision comparison, and orientation-sensitive reasoning, while tasks dominated by higher-level semantic content show smaller or less consistent gains. Building on this finding, we train MosaicAgent-8B to use Mosaic with reinforcement learning using only accuracy and format rewards. Without demonstration trajectories or rewards for specific tool-use, the agent learns to compose visual operations over multiple steps and exhibits diverse problem-solving patterns unpromptedly. Code and data will be released at https://github.com/gengyuanmax/Mosaic.
Problem

Research questions and friction points this paper is trying to address.

multi-image understanding
visual re-representation
multimodal large language models
visual evidence organization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-image understanding
Visual re-representation
Composable image operations
Reinforcement learning
Agentic tool use
🔎 Similar Papers
No similar papers found.