Same Reward, Different Skills: When Multimodal RL Learns to Look

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge in multimodal reinforcement learning where visual evidence is inadequately integrated into learning signals, resulting in models that lack genuine visual grounding capabilities. To overcome this, we propose a "visual solvability" design principle. By constructing image-dependent counterfactual coordinate scenarios and employing the GRPO algorithm, we demonstrate the decisive role of reward mechanisms in shaping learned skills, establish that text-only prompts cannot yield visual gains, and identify the necessity of visual evidence as a critical training criterion. Experimental results show that a 7B model improves its target discovery accuracy from 0.425 to 0.875 in unseen dense scenes, while also exhibiting significantly enhanced visual grounding performance in cross-task generalization evaluations.
📝 Abstract
Reinforcement learning with verifiable rewards (RLVR) improves vision-language benchmark scores even without visual information during training. With images at test, blind-trained models recover roughly half of the real-image gain at 3B and nearly four fifths at 7B. Prolonged real-image training can erode grounding while benchmark gains persist. Both findings expose the same gap: an image in the prompt is not an image in the learning signal. Our design rule, visual resolvability, asks that visual evidence be necessary for a correct answer and that the task remain learnable. We test it on counterfactual coordinate scenes in which the question stays fixed and the target is never named, so a correct answer requires finding the target in the image. With standard GRPO and correctness-and-format rewards, a 7B model raises its accuracy at finding the target (discovery) from 0.425 to 0.875 on held-out scenes denser than any it trained on, and it improves on question types it never trained on. Two controls locate the source of the gain. Replacing test images with gray canvases drops discovery to zero; training on gray canvases instead, at matched step 30 and in each of four seeds, yields essentially none of the gain even when the model is then tested with real images. The learned skill carries over to grounding tasks built independently of the training corpus. A caption that answers the training question, added to the same images, reward and budget, cuts the gain by nearly two thirds. Changing what reward requires changes what RL learns.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Reinforcement Learning
Verifiable Rewards
Visual Grounding
Vision-Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning with Verifiable Rewards
Visual Resolvability
Vision-Language Models
GRPO
Grounding
H
Haocun Ye
University of the Chinese Academy of Sciences
Xinlong Jiang
Xinlong Jiang
Institute of Computing Technology, Chinese Academy of Sciences
Q
Qile Chen
Independent Researcher
B
Bingyu Wang
Institute of Computing Technology, Chinese Academy of Sciences
T
Teng Zhang
Institute of Computing Technology, Chinese Academy of Sciences
Shubai Chen
Shubai Chen
University of the Chinese Academy of Sciences; Institute of Computing Technology, Chinese Academy of Sciences
T
Tingyu Wu
University of the Chinese Academy of Sciences; Institute of Computing Technology, Chinese Academy of Sciences
Z
Zhenkun Zheng
University of the Chinese Academy of Sciences; Institute of Computing Technology, Chinese Academy of Sciences
Yiqiang Chen
Yiqiang Chen
Professor of Computer Science, Institute of Computing Technology, CAS
Human-Computer InteractionUbiquitous ComputingWearable ComputingFederated LearningActivity Recognition