Bias at the Borderline: Who Gets the Benefit of the Doubt in Peer Review?

📅 2026-07-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether area chairs at ICLR 2019–2025 exhibit systematic bias during discretionary decision-making against submissions from non-elite institutions, non-WEIRD countries, or teams lacking female authors. Focusing on 10,416 borderline-accepted papers and integrating reviewer scores with final decisions, the authors employ a causal inference framework, preregistered hypothesis tests, and a “slipped-through” analysis to disentangle statistical discrimination from genuine quality differences in the largest publicly available peer review dataset to date. They find that papers from non-elite institutions receive acceptance rates 0.5–1.6 percentage points lower than their elite counterparts at equivalent review scores—a gap concentrated among papers with public arXiv preprints, suggesting double-blind review may be compromised by preprint disclosure. However, none of 27 downstream quality metrics (e.g., citations, disruptiveness, novelty, publication venue) indicate superior performance by rejected non-elite papers, confirming statistical discrimination without evidence of its justification.
📝 Abstract
We study peer review at ICLR, a large machine-learning conference whose complete review record, including rejected submissions, is public. Reviewers score each submission; for the borderline band whose scores do not settle an outcome, an area chair makes a discretionary accept-or-reject call. We ask whether that call is even-handed: do authors from prestigious institutions, WEIRD countries, or all-male teams get the benefit of the doubt at the margin? Across ICLR 2019-2025 (31,711 submissions; 10,416 borderline), borderline papers without a top-25-institution author are accepted at a 0.5 to 1.6 percentage point lower rate at the same reviewer scores. The gap arises at the discretionary stage, reappears out-of-sample in the pre-registered ICLR 2026 cohort, and concentrates almost entirely among submissions identifiable through a pre-decision arXiv preprint (-3.4 vs. -0.2 points). Equal scores need not mean equal papers: an area chair may respond to quality the scores miss. We apply a robust outcome test, which concludes discrimination only when the group accepted at a lower rate also realizes better downstream outcomes. We measure five outcomes (citations, disruption, two forms of novelty, eventual venue) on both sides of the decision, including the first "ones that got away" test of rejected submissions. Our headline result is a null: across a pre-registered family of 27 tests, no disparity concordant with the decision-rate gap survives correction; we find no evidence that any group faced a higher bar on the outcomes we measure. That null is not an exoneration. A pre-decision preprint pierces the blind through policy-permitted means, and the acceptance gap lives almost entirely in that porosity, consistent with area chairs using revealed institutional prestige as a prior: statistical discrimination that outcome tests may not detect, and a practice double-blind review exists to prevent.
Problem

Research questions and friction points this paper is trying to address.

peer review bias
borderline decisions
institutional prestige
double-blind review
statistical discrimination
Innovation

Methods, ideas, or system contributions that make the work stand out.

peer review bias
robust outcome test
statistical discrimination
double-blind review
preprint leakage
🔎 Similar Papers
H
Hazem Ibrahim
Computer Science, New York University Abu Dhabi, UAE
Talal Rahwan
Talal Rahwan
Associate Professor of Computer Science, New York University Abu Dhabi
Artificial IntelligenceComputational Social ScienceGame Theory
Y
Yasir Zaki
Computer Science, New York University Abu Dhabi, UAE