π€ AI Summary
This study addresses the lack of comparability in peer review scores across research topics at top machine learning conferences, which undermines fairness in paper acceptance decisions. Analyzing 50,289 submissions to ICLR from 2021 to 2026, this work provides the first systematic evidence that papers from different topics with identical review scores exhibit up to an eightfold difference in acceptance probability. The disparity stems from a fundamental flaw in the measurement design of the scoring systemβnot from individual reviewer bias or cultural differences in scoring practices. Through rigorous statistical modeling and attribution analysis that controls for multiple confounding factors, the authors propose calibrated review signals and topic-conditioned acceptance rates as core metrics for evaluating fairness, offering critical empirical evidence and policy guidance for reforming conference reviewing mechanisms.
π Abstract
Peer review at ML conferences increasingly relies on reviewer scores as the primary decision instrument. As submissions have scaled from thousands to tens of thousands per year, no systematic audit has examined whether this instrument functions uniformly across research areas, or whether acceptance outcomes are in practice shaped by forces that reviewer scores neither capture nor control. This position paper argues that acceptance outcomes are shaped by forces beyond reviewer scores, and that the underlying cause is a measurement design failure, not individual bias. When a fixed numerical scale aggregates quality judgments across communities with structurally non-uniform reviewer pools, absolute scores become incomparable across areas, and area chairs must substitute community priors for score-based decisions. Using ICLR 2021--2026 data covering 50,289 papers across 219 research topics, we show that at any given reviewer score, a paper's acceptance probability varies by up to 8x depending on its topic. We rule out scoring culture, expert reviewer standards, rational area chair reweighting, and quality dilution as alternative explanations. We call on program committees to adopt inherently calibrated review signals and publish topic-stratified, score-conditional acceptance rates as a first-class fairness metric.