🤖 AI Summary
This study investigates preference alignment between multimodal large language models (MLLMs) and humans in image quality assessment, focusing on six attributes: aesthetics, artifact-freeness, anatomical accuracy, compositional correctness, object consistency, and style fidelity. To enable rigorous comparison, we construct a controlled synthetic image-pair dataset and conduct cross-task correlation analysis alongside ablation experiments with systematically varied attributes. Results reveal that human raters consistently discriminate all six dimensions and exhibit strong inter-attribute correlations; in contrast, MLLMs demonstrate substantially weaker discriminative capability—particularly for anatomical accuracy and stylistic nuance—and exhibit markedly divergent attribute sensitivity relative to human judgments. This work provides the first empirical evidence of structural deficiencies in MLLMs’ perceptual understanding of image quality, uncovering systematic misalignments in both judgment consistency and dimensional weighting. These findings offer concrete, evidence-based guidance for designing more human-aligned evaluation benchmarks and targeted alignment strategies for vision-language models.
📝 Abstract
Automated evaluation of generative text-to-image models remains a challenging problem. Recent works have proposed using multimodal LLMs to judge the quality of images, but these works offer little insight into how multimodal LLMs make use of concepts relevant to humans, such as image style or composition, to generate their overall assessment. In this work, we study what attributes of an image--specifically aesthetics, lack of artifacts, anatomical accuracy, compositional correctness, object adherence, and style--are important for both LLMs and humans to make judgments on image quality. We first curate a dataset of human preferences using synthetically generated image pairs. We use inter-task correlation between each pair of image quality attributes to understand which attributes are related in making human judgments. Repeating the same analysis with LLMs, we find that the relationships between image quality attributes are much weaker. Finally, we study individual image quality attributes by generating synthetic datasets with a high degree of control for each axis. Humans are able to easily judge the quality of an image with respect to all of the specific image quality attributes (e.g. high vs. low aesthetic image), however we find that some attributes, such as anatomical accuracy, are much more difficult for multimodal LLMs to learn to judge. Taken together, these findings reveal interesting differences between how humans and multimodal LLMs perceive images.