Align Before You Combine: Reference Space Calibration for Supervision Without Ground Truth
This study addresses the challenge of multi-source evaluation in the absence of ground-truth labels and a shared annotation space, where incomparable output scales across scorers hinder the construction of effective supervision signals. To overcome this, we propose a calibration-first framework that synthesizes a universal ordinal reference space via ordered calibration features, aligning subset-specific scorers to a unified scale. This approach is further augmented by a low-resolution calibration approximation technique, enabling supervision score fusion independent of training distributions. Evaluated on three benchmark datasets, the proposed method significantly outperforms both uncalibrated averaging and the best individual scorer while substantially reducing computational costs, thereby demonstrating its effectiveness and generalizability in ground-truth-free scenarios.