๐ค AI Summary
This study investigates systematic biases in vision-language description evaluation by GPT-series models. We introduce a controlled evaluation framework integrating multiple GPT variants (GPT-4o-mini, GPT-4o, GPT-5) and Gemini 2.5 Pro, leveraging semantic similarity analysis and cross-model score comparison. Our method enables fine-grained, reproducible assessment of evaluative behavior across models. Key findings reveal: (1) evaluation capability does not scale with general-purpose capability; (2) the GPT family exhibits an intrinsic โnegative evaluation bias,โ with an average negative-to-positive scoring ratio of 2:1; and (3) distinct evaluative personalities emergeโGPT-4o-mini achieves the highest inter-annotation consistency, GPT-4o demonstrates superior error detection, while GPT-5 behaves conservatively with notably higher score variance. These results provide critical empirical evidence and design insights for developing robust, fair, and model-aware AI evaluation paradigms.
๐ Abstract
As AI systems increasingly evaluate other AI outputs, understanding their assessment behavior becomes crucial for preventing cascading biases. This study analyzes vision-language descriptions generated by NVIDIA's Describe Anything Model and evaluated by three GPT variants (GPT-4o, GPT-4o-mini, GPT-5) to uncover distinct "evaluation personalities" the underlying assessment strategies and biases each model demonstrates. GPT-4o-mini exhibits systematic consistency with minimal variance, GPT-4o excels at error detection, while GPT-5 shows extreme conservatism with high variability. Controlled experiments using Gemini 2.5 Pro as an independent question generator validate that these personalities are inherent model properties rather than artifacts. Cross-family analysis through semantic similarity of generated questions reveals significant divergence: GPT models cluster together with high similarity while Gemini exhibits markedly different evaluation strategies. All GPT models demonstrate a consistent 2:1 bias favoring negative assessment over positive confirmation, though this pattern appears family-specific rather than universal across AI architectures. These findings suggest that evaluation competence does not scale with general capability and that robust AI assessment requires diverse architectural perspectives.