Position Bias is Hidden Behind Ceiling Effects: A Permutation Diagnostic for LLM Benchmarks

📅 2026-07-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of detecting option-position bias in multiple-choice evaluations of large language models, where content noise and stochasticity often confound analysis. The authors propose a verifiable, pre-registered framework that exhaustively permutes answer choices and employs chi-square tests, Cramér’s V statistic, and bootstrap confidence intervals to systematically diagnose the presence and mechanism of position bias. They find, for the first time, that such bias is detectable only within a model accuracy range of 60–95%, revealing two distinct mechanisms: a monotonic decline due to processing load and a non-monotonic drop specifically at option D caused by content ambiguity. Using the open-source tool inspect_permute built on the inspect_ai framework, the authors conducted 24,000 API calls across four leading models and five MMLU subjects, demonstrating that state-of-the-art models often exceed the detectable range due to ceiling effects—suggesting their apparent “unbiased” behavior reflects undetectability rather than true absence of bias.
📝 Abstract
Position bias in multiple-choice LLM evaluation is widely cited as a confound in capability comparisons, but published measurements rely on single answer-order shuffles whose results confound the bias signal with content-level noise and sampling stochasticity. I introduce inspect_permute, an open-source extension to the inspect_ai evaluation framework that runs exhaustive answer-order permutations per question and reports the chi-squared / Cramer V signature of position bias with bootstrap confidence intervals. I apply the tool across four vendors (gpt-4o-mini, claude-haiku-4-5, gemini-2.5-flash, grok-3) on five MMLU subjects, 24,000 API calls under temperature-0 generation, with falsifier predictions pre-registered via a public SHA-256 hash before half the data was observed. Position bias turns out to be statistically detectable only within a roughly 60-95% base-accuracy Goldilocks zone. Below it, processing-load dominance swamps subject-specific signal; above it, ceiling effects compress the variance below the chi-squared test resolution. Detectable cells separate into two mechanism types: monotone A-to-D decrease (processing_load, in low-tier models) and non-monotone D-drop (content_ambiguity, in a narrow capability band). Standard MMLU places every frontier-tier model above the detection band, so absence of signal there should be read as not measurable, not unbiased. Together with the ceiling-effect characterisation in arXiv:2606.26185, this work brackets the detectable region of position-bias measurement and makes the field central question askable in a verifiable form. Package, data, preregistration under MIT.
Problem

Research questions and friction points this paper is trying to address.

position bias
ceiling effects
LLM evaluation
multiple-choice benchmarks
answer-order permutation
Innovation

Methods, ideas, or system contributions that make the work stand out.

position bias
permutation diagnostic
ceiling effects
LLM evaluation
inspect_permute
H
Hiroki Tamba
Independent researcher