Calibrating LLM Judges for Human and AI Conversations

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of incomparable cross-model scoring and positional bias when large language models (LLMs) serve as dialogue judges. To overcome these issues, this work proposes CANDOR, an anchor-set-based calibration strategy that maps diverse LLM judges onto a unified and interpretable scale. By integrating pairwise and pointwise evaluation paradigms, the authors construct the Voice Arena dataset for systematic validation. Experimental results demonstrate that the proposed method achieves zero-shot cross-scenario transferability, effectively aligns the scoring scales of multiple judges, and reveals a significant gap in discriminative capabilities between current AI systems and humans.
📝 Abstract
Measuring how successful a conversation is remains difficult, even for humans judging spoken dialogue. We evaluate state-of-the-art LLMs as pointwise and pairwise judges of conversational success on CANDOR, finding pointwise scoring correlates moderately with human ratings, while pairwise comparison suffers from long transcripts and positional bias. Since this leaves judge scores incomparable across models, we propose a small anchor set and a calibration function that calibrates any judge onto a shared, interpretable scale. We further release the Voice Arena Goal Dataset (VA), 200 task-oriented human-AI and human-agent conversations with pairwise annotations, revealing a substantial gap between current judges and human-level discrimination. Using VA, we test whether CANDOR-fitted calibration transfers to human-AI conversations, finding it brings judges onto a shared scale despite never observing VA during fitting.
Problem

Research questions and friction points this paper is trying to address.

conversational evaluation
LLM judges
calibration
positional bias
human-AI conversations
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM Judges
Score Calibration
Anchor Set
Conversational Success
Cross-domain Transfer
🔎 Similar Papers
No similar papers found.