Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering

📅 2026-04-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of systematic analysis regarding reliability and prompt bias in LLM-as-a-Judge evaluations within software engineering, which may lead judgments to reflect prompt phrasing rather than actual code quality. It presents the first systematic quantification of LLM judgment consistency and sensitivity to minor prompt variations across code generation, repair, and test generation tasks. Through repeated runs, difficulty stratification, and controlled prompt interventions, the work isolates individual variables to precisely measure bias effects. Findings reveal that prompt bias can significantly alter—even reverse—model rankings, posing a serious threat to evaluation validity and reproducibility. The authors advocate for incorporating bias sensitivity metrics into standard evaluation protocols to mitigate these risks.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Prompt Engineering / PromptingComputer Vision: Bias, Fairness & Privacy

Application Category

Economics, Online Markets and Human Computation: LLM based quality controls for crowd workSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
📝 Abstract
Large Language Models are increasingly used as judges to evaluate code artifacts when exhaustive human review or executable test coverage is unavailable. LLM-judge is increasingly relevant in agentic software engineering workflows, where it can help rank candidate solutions and guide patch selection. While attractive for scale, current practice lacks a principled account of reliability and bias: repeated evaluations of the same case can disagree; small prompt edits can swing outcomes; and seemingly semantics-preserving, human-equivalent perturbations may elicit divergent verdicts. This paper studies LLM-as-a-Judge for code through a measurement-first lens. We analyze two pointwise judging regimes across code generation, code repair task, and test generation, and we systematically probe prompt-induced biases. Our study considers difficulty levels for repeated runs and controlled prompt interventions that isolate one presentation cue at a time, and it evaluates judges using consistency and sensitivity to bias. We find that judge decisions are highly sensitive to prompt biases even when the underlying code snippet is unchanged. Across all three tasks, several biases systematically shift preferences toward the option favored by the prompt, improving accuracy when that option aligns with the gold answer but substantially reducing it otherwise. In some settings, these effects are large enough to change task-level conclusions and alter relative model rankings. These findings show that reported judge performance may reflect prompt artifacts rather than stable assessment ability, posing a direct threat to the validity and reproducibility of code evaluation. We therefore argue that LLM-as-a-Judge studies should report bias sensitivity alongside accuracy and incorporate explicit controls to support more trustworthy model comparison in software engineering.
Problem

Research questions and friction points this paper is trying to address.

LLM-as-a-Judge
bias
code evaluation
prompt sensitivity
reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-as-a-Judge
prompt bias
code evaluation
measurement study
reproducibility
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Z
Zixiao Zhao
University of British Columbia, Canada
A
Amirreza Esmaeili
University of British Columbia, Canada
F
Fatemeh Fard
University of British Columbia, Canada