🤖 AI Summary
Existing multimodal Reinforcement Learning with Verifiable Rewards (RLVR) frameworks rely on coarse-grained, sequence-level rewards, lacking fine-grained supervision for visual grounding steps within reasoning chains. To address this limitation, this work proposes TPAE, a method that innovatively constructs a token-level perceptual grounding advantage estimation mechanism by revealing statistical regularities between visual dependence and predictive entropy. Specifically, TPAE evaluates the consistency of each token with the visual-entropy patterns of correct trajectories to generate fine-grained advantage signals, thereby refining supervision from the sequence level to the token level. Experimental results demonstrate that TPAE consistently outperforms mainstream baselines across seven benchmarks, significantly enhancing both the stability and efficiency of multimodal reasoning optimization.
📝 Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has improved the reasoning capabilities of Multimodal Large Language Models (MLLMs), yet existing frameworks rely on coarse, sequence-level reward signals that lack the fine-grained supervision over the visually-grounded steps within a multimodal reasoning chain. We investigate this gap through the lens of two token-level metrics: visual dependency (i.e. how much a token's prediction relies on the input image features) and predictive entropy. Our empirical analysis reveals two key findings: (1) correct reasoning chains exhibit a markedly sharper entropy reduction as visual grounding intensifies, compared to incorrect ones; (2) pivotal tokens, those whose misprediction triggers reasoning collapse, are statistical outliers in the joint distribution of visual dependency and predictive entropy derived from correct chains. Motivated by these findings, we propose token-level perception-grounded advantage estimation (TPAE), which estimates token-level advantages by measuring each token's statistical consistency with the vision-entropy patterns of correct rollouts. TPAE leverages this granular score to modulate the sequence-level advantage, producing a fine-grained supervision signal that can be integrated into various RLVR frameworks. Extensive experiments on seven benchmarks show that TPAE consistently outperforms leading strong baselines, yielding more stable and efficient optimization for multimodal reasoning. The code is publicly available at https://github.com/Zhihan72/TPAE.