Momentum-Coupled Rubric Adaptation for Detailed Image Captioning

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the fragmentation of generation, scoring, and judging roles, as well as their disjointed optimization, in detailed image captioning. To this end, we propose the MoCo Rubric framework, which unifies these roles through shared-parameter multi-task fine-tuning. Furthermore, it introduces a novel momentum coupling mechanism that leverages exponential moving average to enable the rubric generator and judge to smoothly track policy changes, thereby maintaining role consistency without requiring independent reinforcement learning. Experimental results demonstrate that the proposed method achieves an average win rate of 72.83% across five benchmarks, attaining state-of-the-art performance in both blind evaluation rankings and question-answering scores.
📝 Abstract
Detailed image captioning requires accurate and comprehensive descriptions of fine-grained visual content, yet caption quality spans factual accuracy, information coverage, and clarity. Compared with conventional methods that rely mainly on high-quality supervision or holistic rewards, rubric-based reinforcement learning decomposes these requirements into explicit criteria and provides targeted, structured feedback. However, existing methods often use separate models for caption generation, rubric construction, and judging, which may lead to inconsistent interpretations across roles. Some dynamic rubric methods alternate updates between the caption policy and rubric generator while keeping the judge fixed, but staged optimization may still leave rubric construction and judging out of step with policy optimization. We propose MoCo Rubric, a two-stage framework that coordinates these roles. First, role-conditioned, shared-parameter multi-task supervised fine-tuning equips a single vision--language model to serve as the Caption Policy, Rubric Generator, and Rubric Judge. Then, the Generator constructs rubrics online from captions sampled by the current Policy, reference captions, and image evidence. The Judge provides rubric-based rewards, and only the Policy receives GRPO updates. As Policy updates change the candidates being evaluated, we use an exponential moving average of the Policy parameters to update one momentum model shared by the Generator and Judge. This gradual transfer lets both rubric roles track Policy updates without separate RL optimization while smoothing parameter changes that could disrupt their rubric capabilities under direct synchronization. Across five captioning benchmarks, MoCo Rubric achieves an average pairwise win rate of 72.83\%, the best mean rank in blind ranking, and the highest average score in caption-based question answering.
Problem

Research questions and friction points this paper is trying to address.

Detailed Image Captioning
Rubric-based Reinforcement Learning
Role Inconsistency
Staged Optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Momentum-Coupled Rubric
Detailed Image Captioning
Multi-task Supervised Fine-tuning
GRPO
Exponential Moving Average