Can GPT replace human raters? Validity and reliability of machine-generated norms for metaphors

📅 2025-12-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Prior work has not systematically evaluated large language models’ (LLMs) capacity to generate psycholinguistic ratings—specifically familiarity, comprehensibility, and imageability—for multi-word metaphors (rather than isolated words). Method: We assessed GPT-3.5 and GPT-4 on 687 English–Italian bilingual metaphors (drawn from the Italian Metaphor Archive and three English studies), benchmarking their outputs against human ratings and validating them via behavioral response times and EEG neural responses. Contribution/Results: LLM-generated ratings exhibited moderate-to-strong correlations with human judgments—especially for English comprehensibility—and significantly predicted both behavioral and neural measures, performing comparably to human raters. Ratings demonstrated high cross-session stability. This study is the first to empirically demonstrate that LLMs can reliably estimate multidimensional psycholinguistic attributes of complex linguistic units (i.e., metaphors), while also revealing limitations in handling highly conventionalized metaphors and multimodal validation contexts—thereby establishing an evidence-based foundation for the credible use of LLMs in psycholinguistic research.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsCognitive Modeling & Cognitive Systems: Simulating Human Behavior

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
As Large Language Models (LLMs) are increasingly being used in scientific research, the issue of their trustworthiness becomes crucial. In psycholinguistics, LLMs have been recently employed in automatically augmenting human-rated datasets, with promising results obtained by generating ratings for single words. Yet, performance for ratings of complex items, i.e., metaphors, is still unexplored. Here, we present the first assessment of the validity and reliability of ratings of metaphors on familiarity, comprehensibility, and imageability, generated by three GPT models for a total of 687 items gathered from the Italian Figurative Archive and three English studies. We performed a thorough validation in terms of both alignment with human data and ability to predict behavioral and electrophysiological responses. We found that machine-generated ratings positively correlated with human-generated ones. Familiarity ratings reached moderate-to-strong correlations for both English and Italian metaphors, although correlations weakened for metaphors with high sensorimotor load. Imageability showed moderate correlations in English and moderate-to-strong in Italian. Comprehensibility for English metaphors exhibited the strongest correlations. Overall, larger models outperformed smaller ones and greater human-model misalignment emerged with familiarity and imageability. Machine-generated ratings significantly predicted response times and the EEG amplitude, with a strength comparable to human ratings. Moreover, GPT ratings obtained across independent sessions were highly stable. We conclude that GPT, especially larger models, can validly and reliably replace - or augment - human subjects in rating metaphor properties. Yet, LLMs align worse with humans when dealing with conventionality and multimodal aspects of metaphorical meaning, calling for careful consideration of the nature of stimuli.
Problem

Research questions and friction points this paper is trying to address.

Assessing GPT's validity and reliability in rating metaphor properties
Comparing machine-generated norms with human ratings for metaphors
Evaluating if GPT can replace human raters in psycholinguistic studies
Innovation

Methods, ideas, or system contributions that make the work stand out.

GPT models generate metaphor ratings for validity assessment
Machine ratings predict behavioral and EEG responses comparably to humans
Larger GPT models outperform smaller ones in rating reliability
🔎 Similar Papers
No similar papers found.