On the Reliability of User-Centric Evaluation of Conversational Recommender Systems

📅 2026-02-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Current evaluation practices for conversational recommender systems (CRS) often rely on third-party annotators labeling static dialogue logs along user-centric dimensions, yet the reliability of such annotations remains largely unverified. This study addresses this gap through a large-scale crowdsourcing experiment, integrating random-effects reliability modeling, correlation analysis, and the CRS-Que multidimensional evaluation framework to quantify the inter-annotator reliability of each assessment dimension for the first time. The findings reveal that utilitarian dimensions—such as accuracy and usefulness—achieve moderate reliability when aggregated, whereas social dimensions—including perceived humanness and warmth—exhibit substantially lower reliability. Moreover, most dimensions are susceptible to halo effects, collapsing into a single global quality signal. These results challenge prevailing offline evaluation paradigms that depend on single annotators or large language models.

Technology Category

Data Mining & Knowledge Management: Conversational Systems for Recommendation & RetrievalHumans and AI: Crowd Sourcing and Human ComputationCognitive Modeling & Cognitive Systems: Social Cognition And Interaction

Application Category

Economics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsUser Modeling, Personalization and Recommendation: Metrics for user behavior and evaluating successSearch and Retrieval-Augmented AI: Web evaluation methodologies and metrics
📝 Abstract
User-centric evaluation has become a key paradigm for assessing Conversational Recommender Systems (CRS), aiming to capture subjective qualities such as satisfaction, trust, and rapport. To enable scalable evaluation, recent work increasingly relies on third-party annotations of static dialogue logs by crowd workers or large language models. However, the reliability of this practice remains largely unexamined. In this paper, we present a large-scale empirical study investigating the reliability and structure of user-centric CRS evaluation on static dialogue transcripts. We collected 1,053 annotations from 124 crowd workers on 200 ReDial dialogues using the 18-dimensional CRS-Que framework. Using random-effects reliability models and correlation analysis, we quantify the stability of individual dimensions and their interdependencies. Our results show that utilitarian and outcome-oriented dimensions such as accuracy, usefulness, and satisfaction achieve moderate reliability under aggregation, whereas socially grounded constructs such as humanness and rapport are substantially less reliable. Furthermore, many dimensions collapse into a single global quality signal, revealing a strong halo effect in third-party judgments. These findings challenge the validity of single-annotator and LLM-based evaluation protocols and motivate the need for multi-rater aggregation and dimension reduction in offline CRS evaluation.
Problem

Research questions and friction points this paper is trying to address.

Conversational Recommender Systems
user-centric evaluation
reliability
third-party annotation
halo effect
Innovation

Methods, ideas, or system contributions that make the work stand out.

conversational recommender systems
user-centric evaluation
annotation reliability
halo effect
dimension reduction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Michael Müller
Department of Computer Science, University of Innsbruck, Innsbruck, Austria
A
Amir Reza Mohammadi
Department of Computer Science, University of Innsbruck, Innsbruck, Austria
A
Andreas Peintner
Department of Computer Science, University of Innsbruck, Innsbruck, Austria
B
Beatriz Barroso Gstrein
Department of Computer Science, University of Innsbruck, Innsbruck, Austria
Günther Specht
Günther Specht
Professor of Computer Science, University of Innsbruck
DatabasesInformation RetrievalRecommender SystemsPlagiarism DetectionBioinformatics
E
Eva Zangerle
Department of Computer Science, University of Innsbruck, Innsbruck, Austria