The Failure Is in the Readout: Fine-Grained Emotion Recognition Benchmarks Measure Elicitation, Not Perception

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the underperformance of general-purpose vision-language models (VLMs) on fine-grained emotion recognition benchmarks, which has led to the misconception that task-specific fine-tuning is indispensable. We reveal that existing benchmarks predominantly measure generative capacity rather than perceptual ability. To rectify this, we propose an evaluation paradigm that replaces text generation with logit probability reading, employing a binary query mechanism to directly verify answers. Experimental results demonstrate that, without any fine-tuning, merely shifting the answer extraction method from generation to logit verification enables all evaluated VLMs to surpass fine-tuned models, achieving Kappa coefficients of 0.507–0.586. These findings expose the systematic bias inherent in generative evaluation and establish a new paradigm for accurately assessing the perceptual capabilities of VLMs.
📝 Abstract
Fine-grained emotion recognition supports therapy tools and social robots, but it needs facial data, which raises privacy and data-protection concerns. EmoNet-Face-HQ answers that with generated portraits, expert-rated over a $40$-category taxonomy far finer than the usual six to eight basic emotions. Under the protocol it ships with, vision-language models (VLMs) score poorly on that taxonomy, and the benchmark concludes that a dedicated fine-tuned model is necessary: Empathic-Insight-Face (EIF; Small/Large). We show that off-the-shelf VLMs match or beat that fine-tuned model when the answer is not generated but read from the logits, as one binary query per category. We keep the benchmark's images, taxonomy and ratings, and change only how the answer is read. Experts agree at $κ_w = 0.468$ on the five categories they measure most reliably. Generatively, no interval among eleven open-weight VLMs lies entirely above that anchor ($κ_w=0.268$-$0.486$). Under verification all eleven clear it, each of them significantly better at $κ_w=0.507$-$0.586$. Three also significantly beat EIF sitting at $κ_w = 0.551$ (Small; $0.534$ Large). The gain comes from the graded probability and not from asking a yes/no question: as a control, thresholding those same probabilities to yes/no costs 142% of the average gains and drops binarization below generative elicitation to $κ_w=0.254$-$0.423$. A replication on real photographs (FACES) is weaker and mixed: of the ten models that pass a validity gate, six gain, three are neutral to positive and one is negative, so the effect is not confined to synthetic data.
Problem

Research questions and friction points this paper is trying to address.

fine-grained emotion recognition
benchmark evaluation
vision-language models
readout protocol
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fine-Grained Emotion Recognition
Vision-Language Models
Logit Readout Protocol
Benchmark Evaluation
Generative vs Verification
Tobias Hallmen
Tobias Hallmen
Doctoral candidate, Research assistant, University of Augsburg
Analysis of ConversationsLanguage ModelsNatural Language ProcessingMultimodal Deep Learning
Fabian Deuser
Fabian Deuser
University of the Bundeswehr Munich
deep learningmultimodal deep learninggeo localisation
R
Robin-Nico Kampa
Institute for Distributed Intelligent Systems, University of the Bundeswehr Munich
N
Norbert Oswald
Institute for Distributed Intelligent Systems, University of the Bundeswehr Munich
E
Elisabeth André
Chair for Human-Centered Artificial Intelligence, University of Augsburg