🤖 AI Summary
This study addresses the problem of "thought drift" in vision-language models, where intermediate reasoning processes become inconsistent with final answers. To mitigate this issue, we propose Rita, a novel framework that introduces an annotation-free probabilistic reward based on thought-answer consistency. By integrating a difficulty-aware data filtering strategy to optimize training samples for reinforcement learning, our method achieves consistency modeling grounded in conditional probabilities. Extensive evaluations on benchmarks such as EgoIntention demonstrate that Rita significantly outperforms both supervised fine-tuning and conventional reinforcement learning approaches. These results confirm that the proposed framework effectively curbs reasoning deviations in visual intention grounding, offering a robust solution for aligning chain-of-thought reasoning with accurate predictions in multimodal tasks without relying on costly annotated data.
📝 Abstract
This paper studies \textbf{thinking--answer consistency} in vision-language models. We focus on Visual Intention Grounding, where a model infers a target object based on a human intention query and predicts a bounding box. We reveal that previous IoU-based reinforcement learning (RL) frameworks suffer from ``thinking drift'', where the model produces a correct bounding box, despite having an incorrect reasoning process pointing to a different target object. Thus, we propose \textbf{Rita} (\textit{ReInforcing Thinking--Answer consistency}) as a novel RL paradigm to tame the drift. Specifically, Rita introduces two reasoning-label-free RL rewards, constructed from the conditional probability of reference answers: a \textbf{thinking reward} and a \textbf{consistency reward}. It also adopts a difficulty-aware \textbf{data filtering} strategy that selects informative easy-to-medium samples for RL using rollout error rate and reward variance. Extensive experiments on EgoIntention and the new RefEgo-Int benchmarks show that Rita performs consistently superior to the supervised finetuning approaches and vanilla RL-finetuned frameworks.