🤖 AI Summary
This study addresses the performance degradation and positional bias induced by chain-of-thought (CoT) instructions during multiple-choice evaluation of vision-language models (VLMs), which cause answer decoding distortion and scoring discrepancies. It formally defines and quantifies the confounding effect of "CoT prefix scoring," identifying a misalignment between evaluation interfaces and output events as the root cause. Through conditional matching linear probes, free-form generation comparisons, and lexical-level diagnostics, the work investigates how information is preserved within hidden states. Experiments demonstrate that while CoT prompting reduces Qwen2.5-VL accuracy from 80% to 45%, correct-answer information remains encoded in deeper representations. These findings provide a critical caution for VLM evaluation paradigms, recommending against the use of mismatched CoT scoring strategies.
📝 Abstract
Chain-of-thought (CoT) instructions can distort multiple-choice VLM evaluation when a scorer appends a reasoning cue but reads answer-label logits before the model generates any rationale. We call this CoT-prefix scoring. On ScienceQA, Qwen2.5-VL-7B drops from 80.76% to 45.48%, and across five option-content permutations 93.54% of CoT-prefix predictions select the first slot. Condition-matched linear probes recover 78.94% from the same hidden states, while free generation restores 75.24%, showing that the answer often survives the prefix and the immediate readout fails. Vocabulary and layer diagnostics explain the mismatch: probability mass moves toward continuation tokens, while answer information remains linearly accessible in late layers. The effect recurs with varying severity across datasets and models, though not universally. These results show that CoT-prefix scoring can confound model knowledge with an evaluation-interface mismatch and should be avoided unless the requested and scored output events are aligned.