Score
Designs and evaluates methods to detect and remove position-induced preference effects from ranked presentations or attention scores; this includes estimating per-position bias curves, learning bias from low-attention-majority signals, and applying corrections such as subtractive or divisive adjustments to raw attention. It also covers building debiased sorting and reordering procedures (for example debiased one-pass sorting), order-consistency filtering, and experimental procedures like querying both presentation orders to enable single-pass reordering and improve inter-judge reliability.
This work addresses the positional bias in long-context language models, where information at intermediate positions is underutilized. Existing attention sorting methods mitigate this bias through multi-pass re-ranking, but incur high deployment costs. The paper proposes a single-pass debiased attention sorting approach that estimates the positional bias curve from low-attention documents and corrects the original attention scores via subtraction or division, enabling effective one-round reranking. Systematic evaluation on YaRN-Llama-2-7b-64k shows that positional bias correction alone improves accuracy by 8.67 percentage points—yet accounts for only 37% of the performance gap closed by iterative methods—revealing that the benefits of repeated reranking extend beyond mere bias correction.
Large language models (LLMs) exhibit positional bias—where candidate ordering influences ranking/evaluation outcomes—and low repetition consistency—yielding unstable predictions for identical inputs—thereby undermining reliability. To address these issues, we propose a dynamic repetition strategy featuring the first confidence-driven early-stopping mechanism: for each input instance, it adaptively estimates the minimal required number of repetitions, then integrates majority voting with explicit positional bias modeling for fine-grained correction. Unlike static repetition schemes, our method eliminates the need for pre-specified repetition counts. We validate it across three LLM scales and two distinct task categories. Results show that our approach reduces average model calls by 81%–87% compared to static repetition, while preserving high ranking accuracy. This yields significant improvements in both computational efficiency and robustness without sacrificing performance.
This work identifies and systematically quantifies a pervasive positional bias in large language models (LLMs) when employed as automated evaluators: their preference rankings are significantly influenced by the order in which candidate answers appear in the prompt, thereby undermining evaluation reliability. We conduct paired and listwise comparisons across 22 tasks using MT-Bench and DevBench, evaluating 15 LLM-based judges and constructing a dataset of over 150,000 instances. We introduce three novel metrics—repetition stability, positional consistency, and preference fairness—to localize bias sources at the judge-, candidate-, and task-levels, and empirically demonstrate that positional bias strongly correlates with answer quality gaps—not random noise. Results confirm the ubiquity of this bias across models and tasks, with substantial inter-model and inter-task variation. Our findings provide actionable, data-driven strategies for bias mitigation, advancing the robustness and fairness of LLM-based evaluation.
Position bias in implicit feedback data severely degrades learning-to-rank (LTR) performance. To address this, we propose a two-stage debiasing framework grounded in control function methodology: in the first stage, ranking residuals are leveraged to construct exogenous instrumental variables; in the second stage, these instruments enable unbiased estimation of true relevance within a click model. Crucially, our approach imposes no parametric assumptions on either the click or propensity models, supports arbitrary nonlinear ranking models and state-of-the-art LTR algorithms, and—uniquely—introduces validation-set click debiasing to facilitate unbiased hyperparameter tuning. Evaluated on multiple standard LTR benchmarks, our method achieves consistent NDCG@10 improvements of 3.2–5.7% over strong baselines. Moreover, even without access to unbiased validation data, it reliably selects the optimal model, outperforming existing debiasing methods by a significant margin.
This study addresses the challenge of detecting option-position bias in multiple-choice evaluations of large language models, where content noise and stochasticity often confound analysis. The authors propose a verifiable, pre-registered framework that exhaustively permutes answer choices and employs chi-square tests, Cramér’s V statistic, and bootstrap confidence intervals to systematically diagnose the presence and mechanism of position bias. They find, for the first time, that such bias is detectable only within a model accuracy range of 60–95%, revealing two distinct mechanisms: a monotonic decline due to processing load and a non-monotonic drop specifically at option D caused by content ambiguity. Using the open-source tool inspect_permute built on the inspect_ai framework, the authors conducted 24,000 API calls across four leading models and five MMLU subjects, demonstrating that state-of-the-art models often exceed the detectable range due to ceiling effects—suggesting their apparent “unbiased” behavior reflects undetectability rather than true absence of bias.
This study addresses the vulnerability of large language models (LLMs) acting as judges to subtle prompt formatting cues, such as punctuation, which can lead to the erroneous acceptance of incorrect answers. Through paired intervention experiments and statistical calibration on the DROP and GSM8K datasets, the authors quantify the impact of presentation formats on evaluation accuracy, incorporating GPT-6 Sol for comparative analysis. The findings reveal a paradoxical coexistence of fundamental judging competence and heightened sensitivity to editorial cues. Specifically, experimental results demonstrate that format variations cause the false acceptance rate of Jev models to surge from 1.0% to 26.0%, whereas GPT-6 Sol remains unaffected by such perturbations. These observations confirm the critical role of prompt structure in LLM-based evaluation and provide empirical evidence for enhancing the robustness of automated assessment frameworks.
This work addresses the susceptibility of large language models (LLMs) to input order in list reranking, which induces inconsistent preferences and undermines recommendation reliability. The study introduces a multi-level consistency evaluation framework—spanning pairwise preferences, global preference structures, and output stability—to systematically assess the impact of position bias on LLM reranking behavior. Comprehensive experiments across multiple models and datasets reveal that merely enhancing relevance or balancing positional exposure is insufficient to ensure preference consistency. These findings expose an intrinsic instability in LLM-based rerankers that conventional evaluation metrics fail to capture, thereby underscoring the necessity of explicitly modeling preference consistency in reranking systems.
This study addresses the challenge of auditing biases in multimodal large language models (MLLMs) used for image editing evaluation, where judges are susceptible to irrelevant cues and struggle to verify quality preservation. We propose EditJudgeBias, a benchmark that systematically audits MLLM judge biases across invariance, consistency, and stability dimensions. This is achieved through counterfactual data safeguarded by calibrated validators and a noise-floor comparison mechanism. Our findings reveal that no single metric can comprehensively characterize robustness. Furthermore, we demonstrate that all evaluated judges are vulnerable to spurious extraneous cues: fabricated majority opinions inflate scores, and swapping candidate orders induces preference reversals in 60.9% of cases.
This study addresses response position bias and systematic misalignment with human preferences in LLM-based evaluation by proposing a unified debiasing and alignment framework. Methodologically, it decouples positional effects from preference structures through latent preference identification. By leveraging large-scale LLM pairwise comparison data alongside limited human annotations, the framework employs an adaptive estimation strategy to balance LLM anchoring signals with human evidence, incorporating fixed-weight uncertainty quantification to enhance robustness. Furthermore, this work constructs a multi-model, dual-order judgment dataset comprising over 410,000 entries. Extensive simulations and benchmark evaluations demonstrate that the proposed approach achieves robust positional debiasing and yields rankings highly consistent with human preferences.
This study addresses the unreliability of Vision-Language Models (VLMs) as image judges in cultural contexts, where selection bias and positional sensitivity compromise decision-making. To investigate this, we systematically audit VLM judging behavior by benchmarking against human ratings and random baselines, analyzing the effects of response reordering, and proposing an order-consistency-based filtering strategy to mitigate such biases. Our findings reveal distinct failure modes across model scales: 4B-parameter models exhibit significant positional bias and underperform CLIP, whereas 8B-parameter models avoid this deficiency yet still require targeted filtering. By elucidating these scale-dependent vulnerabilities, this work provides empirical evidence and methodological foundations for enhancing the reliability of VLMs in cross-cultural evaluation tasks.