Score
Designs and implements statistical models, algorithms, and procedures to estimate individual rater/judge leniency and severity, evaluate judge calibration, and detect systematic bias. Builds calibration-aware evaluation methods and automated pipelines (e.g., item-response / judge-response models) to adjust or align scores across multiple raters so aggregated ratings are comparable and bias-corrected.
The reliability of LLM-as-a-Judge—particularly concerning evaluation accuracy, inter-annotator consistency, and fairness—remains a critical bottleneck. Method: This paper introduces the first systematic reliability assessment framework, featuring a dedicated benchmark (ReliBench) and synthesizing three core strategies: consistency enhancement, bias mitigation, and scenario adaptation. Our methodology integrates multi-model cross-validation, adversarial testing, interpretability analysis, and structured prompt engineering, underpinned by a standardized evaluation protocol. Contributions: (1) The first comprehensive, domain-wide survey of LLM-as-a-Judge reliability; (2) An open-source, reproducible evaluation toolkit and benchmark; (3) A paradigm shift from heuristic, experience-driven LLM evaluation toward rigorous, scientific, and standardized assessment practices.
This study addresses systematic scoring biases in LLM-as-a-Judge evaluations, where existing correction methods suffer from cross-model calibration instability, often leading to directional misjudgments in model comparisons. Through theoretical analysis, Monte Carlo simulations, and experiments on MMLU-Pro, the work uncovers the hidden risks of shared calibration strategies and proposes two diagnostic metrics—judge quality (J) and cross-model calibration instability (ΔJ)—to assess the reliability of bias correction outcomes. The research successfully reproduces the sign-flipping phenomenon observed in prior work, empirically validating the effectiveness of the proposed indicators. Building on these insights, the authors establish a reporting protocol for LLM-as-a-Judge evaluations, offering both theoretical grounding and practical guidance to enhance the trustworthiness of comparative model assessments.
This study addresses the lack of systematic analysis regarding reliability and prompt bias in LLM-as-a-Judge evaluations within software engineering, which may lead judgments to reflect prompt phrasing rather than actual code quality. It presents the first systematic quantification of LLM judgment consistency and sensitivity to minor prompt variations across code generation, repair, and test generation tasks. Through repeated runs, difficulty stratification, and controlled prompt interventions, the work isolates individual variables to precisely measure bias effects. Findings reveal that prompt bias can significantly alter—even reverse—model rankings, posing a serious threat to evaluation validity and reproducibility. The authors advocate for incorporating bias sensitivity metrics into standard evaluation protocols to mitigate these risks.
This study addresses the reliability and validity of LLM-as-judge evaluations, which are susceptible to shifts in the judge model’s version even when candidate responses remain unchanged. The authors conduct a systematic audit of dense Qwen3 models (1.7B–32B) and MiniMax API iterations (M2 to M2.7) across four benchmark judgment datasets. They propose a multidimensional auditing framework incorporating multiscale judge comparisons, repeated-sampling juries, structured debate protocols, and probes for position and verbosity biases. Findings indicate that only the upgrade from Qwen3-1.7B to -4B yields consistent performance gains; stronger judges mitigate but do not eliminate systematic biases; and structured debate substantially alters verdicts, though reliable attribution requires access to detailed interaction logs.
This work identifies and systematically quantifies a pervasive positional bias in large language models (LLMs) when employed as automated evaluators: their preference rankings are significantly influenced by the order in which candidate answers appear in the prompt, thereby undermining evaluation reliability. We conduct paired and listwise comparisons across 22 tasks using MT-Bench and DevBench, evaluating 15 LLM-based judges and constructing a dataset of over 150,000 instances. We introduce three novel metrics—repetition stability, positional consistency, and preference fairness—to localize bias sources at the judge-, candidate-, and task-levels, and empirically demonstrate that positional bias strongly correlates with answer quality gaps—not random noise. Results confirm the ubiquity of this bias across models and tasks, with substantial inter-model and inter-task variation. Our findings provide actionable, data-driven strategies for bias mitigation, advancing the robustness and fairness of LLM-based evaluation.
This study addresses the persistent irreproducibility in LLM-as-judge safety evaluations, even when using greedy decoding (temperature = 0). Contrary to common assumptions, the authors demonstrate that temperature = 0 does not fully eliminate stochasticity in LLM scoring, particularly for borderline cases. Through 690 experiments across multiple models, APIs, and sampling configurations within the open-source aisev evaluation framework, they reveal that default temperature settings can induce judgment fluctuations in up to 50% of individual items, and even deterministic decoding fails to ensure reproducibility for one to two borderline cases per evaluation. To address this, the paper proposes incorporating scoring disagreement as a core health metric in evaluation frameworks, advocating for more robust and reliable practices in LLM safety assessment.
This study addresses the challenge of low-quality bug reports in crowdsourced testing, which impose substantial review burdens on developers and lack effective mechanisms to improve tester performance. The authors propose a large language model–based multi-agent evaluation framework that automatically assesses reports along three dimensions—textuality, sufficiency, and competitiveness—and integrates actionable feedback into human workflows. Through a four-phase controlled experiment combined with mixed-methods analysis, they provide the first empirical evidence that evaluative agents not only serve as post-hoc adjudicators but also function as in-process feedback sources, significantly enhancing the quality of report revisions, improving first-submission performance in subsequent tasks, and facilitating cross-application knowledge transfer. User studies further confirm the intelligibility and practical utility of the generated feedback.