Residual Visual Credit Optimization: Conserved Evidence Routing for Multimodal Reinforcement Learning

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the credit assignment challenge in multimodal reinforcement learning, where outcome rewards are difficult to precisely allocate across individual decision steps. To this end, it proposes a residual visual credit optimization method that formulates token-level credit as a conservative routing problem. Specifically, the approach introduces controlled visual interventions and a budgeted entropy router, combined with robust trajectory coordinates and analytical correction techniques, to achieve precise distribution of sequence utility. This ensures that positive credits effectively reinforce decisions while restoring accurate credit quality. Experimental evaluations demonstrate that the proposed method significantly improves accuracy across four model families and seven benchmarks, exhibiting superior optimization stability and robustness against perturbations.
📝 Abstract
Reinforcement learning with verifiable rewards scales multimodal reasoning, but an outcome reward says how much a trajectory is worth, not how that value should be spread over the decisions that produced it. We introduce Residual Visual Credit Optimization (RVCO), which treats token credit as a conserved routing problem. A controlled visual intervention yields a per-token evidence response; robust within-trajectory coordinates remove incidental scale; and a budgeted entropic router distributes a fixed amount of sequence utility according to perceptual dependence. A residual support path guarantees positive credit at every valid position, and an analytic correction restores the prescribed credit mass exactly. The resulting field is selective, bounded, full-support, and invariant to response-local score shifts, and recovers hard token selection as a limiting case. Across four model families and seven reasoning benchmarks, RVCO improves accuracy over strong RLVR baselines while maintaining late-stage optimization stability, corruption robustness, and competitive training cost. Rewards, rollouts, and the group-relative advantage estimator are unchanged; only the geometry of token-level credit differs.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Reinforcement Learning
Credit Assignment
Verifiable Rewards
Token-level Credit
Innovation

Methods, ideas, or system contributions that make the work stand out.

Residual Visual Credit Optimization
Conserved Evidence Routing
Multimodal Reinforcement Learning
Entropic Router
Token-level Credit Assignment
💼 Related Jobs
No related jobs found.
L
Lin Qiu
Meta Superintelligence Labs
Y
Yao Liu
Meta Superintelligence Labs
D
Diyi Hu
University of Southern California
Hanqing Zeng
Hanqing Zeng
Meta; Ph.D. in computer engineering
Graph Representation LearningHigh Performance ComputingRecommendation Systems
Onur Gungor
Onur Gungor
Unknown affiliation
C
Chujie Chen
Meta
J
Jiayi Liu
Meta Recommendation System
J
Jianyu Wang
Meta Recommendation System
X
XueLin Zheng
Meta