Reviewer Scores Are Not Comparable Across Research Areas in ML Peer Review

πŸ“… 2026-04-30
πŸ“ˆ Citations: 1
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the lack of comparability in peer review scores across research topics at top machine learning conferences, which undermines fairness in paper acceptance decisions. Analyzing 50,289 submissions to ICLR from 2021 to 2026, this work provides the first systematic evidence that papers from different topics with identical review scores exhibit up to an eightfold difference in acceptance probability. The disparity stems from a fundamental flaw in the measurement design of the scoring systemβ€”not from individual reviewer bias or cultural differences in scoring practices. Through rigorous statistical modeling and attribution analysis that controls for multiple confounding factors, the authors propose calibrated review signals and topic-conditioned acceptance rates as core metrics for evaluating fairness, offering critical empirical evidence and policy guidance for reforming conference reviewing mechanisms.
πŸ“ Abstract
Peer review at ML conferences increasingly relies on reviewer scores as the primary decision instrument. As submissions have scaled from thousands to tens of thousands per year, no systematic audit has examined whether this instrument functions uniformly across research areas, or whether acceptance outcomes are in practice shaped by forces that reviewer scores neither capture nor control. This position paper argues that acceptance outcomes are shaped by forces beyond reviewer scores, and that the underlying cause is a measurement design failure, not individual bias. When a fixed numerical scale aggregates quality judgments across communities with structurally non-uniform reviewer pools, absolute scores become incomparable across areas, and area chairs must substitute community priors for score-based decisions. Using ICLR 2021--2026 data covering 50,289 papers across 219 research topics, we show that at any given reviewer score, a paper's acceptance probability varies by up to 8x depending on its topic. We rule out scoring culture, expert reviewer standards, rational area chair reweighting, and quality dilution as alternative explanations. We call on program committees to adopt inherently calibrated review signals and publish topic-stratified, score-conditional acceptance rates as a first-class fairness metric.
Problem

Research questions and friction points this paper is trying to address.

peer review
reviewer scores
research areas
acceptance fairness
score comparability
Innovation

Methods, ideas, or system contributions that make the work stand out.

peer review fairness
score calibration
research area bias
acceptance rate stratification
measurement design
B
Binyan Xu
The Chinese University of Hong Kong, Hong Kong, China
X
Xilin Dai
Zhejiang University, Hangzhou, China
F
Fan Yang
The Chinese University of Hong Kong, Hong Kong, China
Kehuan Zhang
Kehuan Zhang
The Chinese University of Hong Kong
Security of Computer systemsWebMobileCloudEmbedded System