🤖 AI Summary
This study addresses the lack of evaluation frameworks for historical information utilization in multimodal interactions by proposing ReTurn, a benchmark comprising 7,000 visual-audio dialogue tasks. This benchmark innovatively decouples history dependence from question difficulty, enabling precise assessment of models' selective retrieval and application of contextual information through a controlled framework. The methodology integrates behavioral probing analysis with supervised fine-tuning across both open-ended and multiple-choice evaluation settings. Experimental results reveal that existing models suffer significant accuracy degradation in multi-turn dialogues, exposing inherent deficiencies in leveraging historical context. Furthermore, the findings demonstrate that current adaptation methods yield only marginal improvements, underscoring the persistent challenges in effective context utilization within multimodal conversational systems.
📝 Abstract
Reliable multimodal interaction depends on selective use of conversational history: an earlier question may remain relevant while its previous answer is outdated, whereas a current request may depend on historical evidence despite conflicting new observations. Existing multi-turn evaluations rarely separate these history-use demands from underlying question difficulty. To address this gap, we introduce ReTurn, a benchmark of 7,000 base tasks spanning visual and audio evidence for evaluating selective history use. For task-carrying history, Reconfirm/Reground require applying a historical question to current media while varying historical agreement; for evidence-carrying history, Retrieve/Rebind require answering a current question using historical media while varying current-media competition. Each pair preserves the target question, media, and answer. Tasks support open-ended and multiple-choice evaluation, with matched single-turn counterparts serving as answerability references. Across 13 omni-modal, vision-language, and audio-language models, median model-level open-ended accuracy falls from 93.7% with direct input to 72.3% in conversation. Behavioral probes show that high question recall can coexist with weaker task application, while competing media can redirect answers away from historical targets. Supervised adaptation yields only partial gains. ReTurn provides a controlled framework for assessing whether multimodal models select and use the historical information required by each request.