🤖 AI Summary
This study addresses the challenge of multi-source evaluation in the absence of ground-truth labels and a shared annotation space, where incomparable output scales across scorers hinder the construction of effective supervision signals. To overcome this, we propose a calibration-first framework that synthesizes a universal ordinal reference space via ordered calibration features, aligning subset-specific scorers to a unified scale. This approach is further augmented by a low-resolution calibration approximation technique, enabling supervision score fusion independent of training distributions. Evaluated on three benchmark datasets, the proposed method significantly outperforms both uncalibrated averaging and the best individual scorer while substantially reducing computational costs, thereby demonstrating its effectiveness and generalizability in ground-truth-free scenarios.
📝 Abstract
We introduce a calibration-first framework that produces supervision scores without access to ground-truth labels or a shared annotation space. Our framework aligns subset-specific scorers using a synthetic ordinal reference space before fusion. This reference space is constructed from ordered calibration features that represent the latent concept, providing a common scale on which otherwise incomparable scorer outputs can be aligned. Because our calibration procedure uses the reference space rather than training samples, it is independent of the training set's empirical distribution. Across three benchmark datasets, our framework consistently outperforms uncalibrated averaging and achieves higher primary-metric point estimates on the evaluation metrics than the best individual scorer. Performance relative to sample-dependent baselines varies by domain, with absolute differences below 0.02 on Ames Housing and below 0.01 on Breast Cancer Wisconsin and Wine Quality. After Bonferroni correction, differences remain significant for all three comparisons on Ames Housing and one on Breast Cancer Wisconsin. Additionally, we show that using fewer calibration levels per feature can closely approximate higher-resolution results at substantially lower computational cost. Together, these results support our framework as a viable approach to construct supervision scores when neither ground-truth labels nor a shared annotation space is available.