rater bias calibration

Designs and implements statistical models, algorithms, and procedures to estimate individual rater/judge leniency and severity, evaluate judge calibration, and detect systematic bias. Builds calibration-aware evaluation methods and automated pipelines (e.g., item-response / judge-response models) to adjust or align scores across multiple raters so aggregated ratings are comparable and bias-corrected.

raterbiascalibration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.01
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$196K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses systematic scoring biases in LLM-as-a-Judge evaluations, where existing correction methods suffer from cross-model calibration instability, often leading to directional misjudgments in model comparisons. Through theoretical analysis, Monte Carlo simulations, and experiments on MMLU-Pro, the work uncovers the hidden risks of shared calibration strategies and proposes two diagnostic metrics—judge quality (J) and cross-model calibration instability (ΔJ)—to assess the reliability of bias correction outcomes. The research successfully reproduces the sign-flipping phenomenon observed in prior work, empirically validating the effectiveness of the proposed indicators. Building on these insights, the authors establish a reporting protocol for LLM-as-a-Judge evaluations, offering both theoretical grounding and practical guidance to enhance the trustworthiness of comparative model assessments.

biascalibration instabilityLLM-as-a-Judge

This study addresses the lack of systematic analysis regarding reliability and prompt bias in LLM-as-a-Judge evaluations within software engineering, which may lead judgments to reflect prompt phrasing rather than actual code quality. It presents the first systematic quantification of LLM judgment consistency and sensitivity to minor prompt variations across code generation, repair, and test generation tasks. Through repeated runs, difficulty stratification, and controlled prompt interventions, the work isolates individual variables to precisely measure bias effects. Findings reveal that prompt bias can significantly alter—even reverse—model rankings, posing a serious threat to evaluation validity and reproducibility. The authors advocate for incorporating bias sensitivity metrics into standard evaluation protocols to mitigate these risks.

biascode evaluationLLM-as-a-Judge

This study addresses the reliability and validity of LLM-as-judge evaluations, which are susceptible to shifts in the judge model’s version even when candidate responses remain unchanged. The authors conduct a systematic audit of dense Qwen3 models (1.7B–32B) and MiniMax API iterations (M2 to M2.7) across four benchmark judgment datasets. They propose a multidimensional auditing framework incorporating multiscale judge comparisons, repeated-sampling juries, structured debate protocols, and probes for position and verbosity biases. Findings indicate that only the upgrade from Qwen3-1.7B to -4B yields consistent performance gains; stronger judges mitigate but do not eliminate systematic biases; and structured debate substantially alters verdicts, though reliable attribution requires access to detailed interaction logs.

evaluator biasjudge reliabilityLLM-as-judge

This work identifies and systematically quantifies a pervasive positional bias in large language models (LLMs) when employed as automated evaluators: their preference rankings are significantly influenced by the order in which candidate answers appear in the prompt, thereby undermining evaluation reliability. We conduct paired and listwise comparisons across 22 tasks using MT-Bench and DevBench, evaluating 15 LLM-based judges and constructing a dataset of over 150,000 instances. We introduce three novel metrics—repetition stability, positional consistency, and preference fairness—to localize bias sources at the judge-, candidate-, and task-levels, and empirically demonstrate that positional bias strongly correlates with answer quality gaps—not random noise. Results confirm the ubiquity of this bias across models and tasks, with substantial inter-model and inter-task variation. Our findings provide actionable, data-driven strategies for bias mitigation, advancing the robustness and fairness of LLM-based evaluation.

Analyzes impact of solution quality gaps on biasEvaluates position bias in LLM-as-a-Judge systemsIdentifies factors causing bias across judges and tasks

Latest Papers

What's happening recently
View more

This study addresses the persistent irreproducibility in LLM-as-judge safety evaluations, even when using greedy decoding (temperature = 0). Contrary to common assumptions, the authors demonstrate that temperature = 0 does not fully eliminate stochasticity in LLM scoring, particularly for borderline cases. Through 690 experiments across multiple models, APIs, and sampling configurations within the open-source aisev evaluation framework, they reveal that default temperature settings can induce judgment fluctuations in up to 50% of individual items, and even deterministic decoding fails to ensure reproducibility for one to two borderline cases per evaluation. To address this, the paper proposes incorporating scoring disagreement as a core health metric in evaluation frameworks, advocating for more robust and reliable practices in LLM safety assessment.

determinismLLM-as-Judgereproducibility

This study addresses the challenge of low-quality bug reports in crowdsourced testing, which impose substantial review burdens on developers and lack effective mechanisms to improve tester performance. The authors propose a large language model–based multi-agent evaluation framework that automatically assesses reports along three dimensions—textuality, sufficiency, and competitiveness—and integrates actionable feedback into human workflows. Through a four-phase controlled experiment combined with mixed-methods analysis, they provide the first empirical evidence that evaluative agents not only serve as post-hoc adjudicators but also function as in-process feedback sources, significantly enhancing the quality of report revisions, improving first-submission performance in subsequent tasks, and facilitating cross-application knowledge transfer. User studies further confirm the intelligibility and practical utility of the generated feedback.

actionable feedbackagent-human interactioncrowdsourced testing

Hot Scholars

PM

Paolo Monti

Professor in Communication Networks - Chalmers University of Technology
Optical NetworksNetwork AutomationSustainable Networking
TY

Tong Yang

Peking University, Beijing, China. PKU. 北京大学
SketchNetwork measurementBloom filterIP lookup
ES

Ethan Seefried

PhD Student Colorado State University
Computer VisionVirtual RealityNatural Language ProcessingHuman Computer Interactions
WP

Woojin Park

UNIST
Wireless communicationPhysical layerGeneralized Frequency Division MultiplexingGFDM
WM

Wentao Ma

Alibaba DAMO Academy
Dialog SystemLarge Language ModelQuestion Answering