BayesJudge: Uncertainty-Aware Bayesian Meta-Evaluation of Human and LLM Judgments

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of human-LLM judgment conflicts and positional bias in AI evaluation by proposing an online Bayesian meta-evaluation framework. Methodologically, it introduces tie labels to explicitly model ambiguity and decouples quality effects from positional bias by alternating presentation orders. Efficient online inference is achieved through a Rao-Blackwellized sequential Monte Carlo algorithm. Experimental results demonstrate that the proposed approach effectively detects LLM order effects, precisely characterizes expert scoring behaviors, and yields uncertainty estimates that closely align with human disagreement. Collectively, this work establishes a novel paradigm for reliable AI evaluation.
📝 Abstract
AI evaluation pipelines often produce conflicting judgments rather than clean labels. In pairwise LLM evaluation, this conflict is especially visible: disagreement can arise from ambiguous items, underspecified rubrics, heterogeneous or unstable human raters, or an LLM judge whose verdict changes when the response order is swapped. We propose BayesJudge, an online Bayesian meta-evaluation layer for conflicting human-LLM judgment streams. For each comparison, BayesJudge estimates a panel-relative posterior verdict distribution over the two responses, with a tie or ambiguity state when such labels are available. At the same time, it estimates rater-specific human confusion matrices and LLM presentation-order bias. The method uses tie-open labels to keep ambiguity observable and paired order-swapped judge calls to separate response quality from presentation effects. We formulate the exact online posterior recursion and use a scalable Rao-Blackwellized assumed-density SMC approximation for streaming inference. Controlled synthetic experiments demonstrate recovery of prespecified evaluator parameters and illustrate two protocol-level identifiability mechanisms: tie-open labels expose ambiguity mass, and order-swapped paired judgments separate item preference from position bias. On a real-world SummEval dataset, BayesJudge successfully detects systematic presentation-order effects in LLM judge outputs, infers distinct expert and crowdworker behavior signatures without rater metadata, and produces posterior uncertainty estimates that correlate with human disagreement. Our code is available at https://anonymous.4open.science/r/BayesJudge-0879.
Problem

Research questions and friction points this paper is trying to address.

LLM evaluation
conflicting judgments
uncertainty quantification
presentation-order bias
meta-evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bayesian meta-evaluation
online streaming inference
Rao-Blackwellized SMC
presentation-order bias
uncertainty quantification