🤖 AI Summary
This study addresses the unclear perceptual mechanisms underlying human image quality assessment (IQA) by proposing a multidimensional observer model. The model maps images into a latent perceptual space, integrating primate neural representational constraints with large-scale behavioral data fitting to systematically characterize the low-dimensional structure of perception. Our findings reveal the extremely low-dimensional nature of this perceptual space and its task-dependent construction mechanism. While matching the predictive performance of existing IQA metrics, this work elucidates the fundamental distinctions between high- and low-level visual quality judgments. Ultimately, it provides a computable neurobehavioral framework for understanding human visual quality perception.
📝 Abstract
Judging image quality is not only ecologically relevant to everyday human tasks, but also underpins many machine vision tasks such as image generation. This paper proposes a framework to understand the inherent perceptual space underlying image quality judgment in humans. We propose a multi-dimensional observer model that represents images as distributions in a latent perceptual space and that models human judgment as comparing noisy samples. Being constrained by neural representations in the primate ventral stream and fit to large-scale behavioral data, the model enables analysis of perceptual structure while matching the predictive power of existing metrics. Using this model, we find that the perceptual spaces needed to account for image quality judgment in humans are extremely low-dimensional compared to the image space even when considering its sparsity. The exact structure of the space (e.g., dimensionalities, information encoded) varies between low-level and high-level quality judgments, suggesting that, despite a shared retinal encoding in the beginning, humans selectively construct task-dependent perceptual spaces in visual decision making.