Human-AI Perceptual Alignment by Playing Hues and Cues

📅 2026-08-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses a significant alignment gap between current vision-language models and human cognition regarding fine-grained color semantics and cultural associations, which conventional benchmarks fail to capture. The authors propose the first gamified evaluation framework inspired by the board game *Hues and Cues*, mapping 480 colors into the CIE xy chromaticity space and leveraging a 100-word vocabulary alongside color-association data from 325 human participants. Using leave-one-out cross-validation, they establish a human-consistency baseline and systematically evaluate 162 contrastive vision-language models. The study reveals a pervasive “default blue” collapse in model predictions and systematic biases in handling abstract and pop-culture concepts. It further demonstrates that carefully curated pretraining data substantially outperforms large-scale, unfiltered corpora, offering a novel paradigm and empirical benchmark for aligning human and AI perceptual understanding.
📝 Abstract
Evaluating the perceptual alignment between Contrastive Vision-Language Models (CVLMs) and humans is typically constrained by traditional benchmarks that overlook fine-grained semantic and cultural nuances. In this work, we propose a novel evaluation framework that leverages the gamified, discrete color space of the board game Hues and Cues. By mapping the board's 480 color cells to the CIE xy chromaticity diagram, we calculate empirical perceptual distances across a carefully curated 100-word vocabulary spanning seven semantic categories. To properly contextualize model performance, we establish an empirical lower bound of expected error-the Human Consistency baseline-calculated via Leave-One-Out (LOO) cross-validation on a dense dataset of color associations collected from 325 human observers through a custom digital interface. We evaluate 162 models across multiple architectural families and pre-training datasets to assess their semantic color grounding. Our results demonstrate that while CVLMs successfully replicate human cognitive biases, such as idealized memory colors for concrete physical referents (e.g., food and plants), they systematically diverge from the human baseline in abstract, subjective, and pop-culture domains. We identify two distinct failure modes in severely misaligned concepts: semantic misclassification and a systematic uncertainty collapse into a default blue coordinate. Furthermore, we reveal that highly curated pre-training datasets are significantly more effective than massive, uncurated corpora in mitigating these severe misalignments. Ultimately, this work highlights that despite their broad categorization capabilities, current CVLMs still fail to capture the nuanced, localized consensus of human color memory, emphasizing the value of gamified tasks in exposing underlying model biases. The data and code are publicly available to test other metrics.
Problem

Research questions and friction points this paper is trying to address.

perceptual alignment
color semantics
human-AI alignment
vision-language models
cultural nuance
Innovation

Methods, ideas, or system contributions that make the work stand out.

perceptual alignment
color semantics
human consistency baseline
gamified evaluation
vision-language models
🔎 Similar Papers
No similar papers found.