🤖 AI Summary
This study addresses the reasoning deviations in visual agents during image manipulation, which frequently stem from localized errors in evidence acquisition, reading, or grounding, and for which existing methods lack precise failure-stage correction. We propose ReVuE, a framework that diagnoses the initial failure stage via multi-trajectory comparison to generate reflections, pioneering the translation of trajectory-level evidence diagnosis into token-level supervision. By integrating online policy distillation with a dynamic loss reweighting mechanism, ReVuE enables staged error correction based on reflection impact. Experiments across eleven benchmarks demonstrate that our approach comprehensively outperforms baselines, significantly improving both perception and reasoning accuracy while effectively reducing redundant operations and enhancing tool invocation precision.
📝 Abstract
Visual agents solve problems by interleaving reasoning with image operations, and on-policy distillation (OPD) provides guidance from a strong teacher on student-generated interaction trajectories. However, image operations change the evidence available for subsequent reasoning, so local errors in evidence acquisition (Acquire), reading (Read), or answer grounding (Ground) can propagate through the trajectory and lead to incorrect answers. Existing multimodal OPD methods primarily construct or contrast auxiliary views of the original image to strengthen supervision, without explicitly modeling the connections between student actions, resulting observations, and subsequent reasoning. This limits their ability to provide corrections tailored to different failure stages. We introduce Reflection on Visual Evidence (ReVuE), an on-policy distillation method for visual agents. ReVuE compares multiple student-generated trajectories for the same query, summarizes the observed visual evidence, and diagnoses the first failure across the Acquire, Read, and Ground stages. The resulting reflections provide training-time context for the teacher. We group and reweight token-level distillation losses according to how strongly these reflections affect the teacher's predictions. This design translates trajectory-level evidence diagnosis into targeted token-level supervision, guiding students to improve their visual evidence acquisition and reasoning. Across 11 benchmarks spanning the Qwen2.5-VL and InternVL3.5 model families, ReVuE outperforms all evaluated OPD baselines in weighted-average scores for perception, mathematical reasoning, and general tasks. ReVuE also reduces redundancy in reasoning and tool calls while improving tool-call accuracy and task accuracy. Code is available at https://github.com/sylvain-wei/ReVuE