Aggregate-then-Calibrate for Human-centered Assessment with Theoretical Guarantees

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of human-centric evaluation—namely, the absence of verifiable ground truth, heterogeneity among expert judgments, and inconsistent rating scales—and the reliance of purely model-based evaluation on imperfect proxy metrics. To overcome these challenges, the authors propose AtC, a two-stage framework: first, it aggregates rankings by modeling annotator reliability to produce a consensus ranking; second, it calibrates arbitrary model scores onto this ranking via isotonic projection, preserving both ordinal consistency and quantitative information. This study is the first to integrate judgment aggregation with model-free calibration, providing theoretical guarantees of improved estimation efficiency, risk bounds under consensus misspecification, and asymptotic superiority over single-paradigm evaluation approaches. Experiments demonstrate that AtC significantly outperforms human-only or model-only methods on both semi-synthetic and real-world datasets.
📝 Abstract
Human-centered assessment tasks, which are essential for systematic decision-making, rely heavily on human judgment and typically lack verifiable ground truth. Existing approaches face a dilemma: methods using only human judgments suffer from heterogeneous expertise and inconsistent rating scales, while methods using only model-generated scores must learn from imperfect proxies or incomplete features. We propose Aggregate-then-Calibrate (AtC), a two-stage framework that combines these complementary sources. Stage-1 aggregates heterogeneous comparative judgments into a consensus ranking using a rank-aggregation model that accounts for annotator reliability. Stage-2 calibrates any predictive model's scores by an isotonic projection onto the order, enforcing ordinal consistency while preserving as much of the model's quantitative information as possible. Theoretically, we show: (1) modeling annotator heterogeneity yields strictly more efficient consensus estimation than homogeneity; (2) isotonic calibration enjoys risk bounds even when the consensus ranking is misspecified; and (3) AtC asymptotically outperforms model-only assessment. Across semi-synthetic and real-world datasets, AtC consistently improves accuracy and robustness over human-only or model-only assessments. Our results bridge judgment aggregation with model-free calibration, providing a principled recipe for human-centered assessment when ground truth is costly, scarce, or unverifiable.
Problem

Research questions and friction points this paper is trying to address.

human-centered assessment
ground truth
judgment aggregation
rating heterogeneity
model calibration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Aggregate-then-Calibrate
rank aggregation
isotonic calibration
human-centered assessment
annotator heterogeneity