π€ AI Summary
This work addresses the challenge of reliably evaluating human preferences in non-verifiable tasks, where diverse and often conflicting evaluation criteria hinder consistent assessment. To overcome this limitation, the authors propose the Constrained Shared-Private Fusion (CSPF) method, which treats multiple heterogeneous frozen reward models as complementary evaluators and fuses their hidden representations under pairwise human preference supervision. CSPF explicitly decomposes shared and expert-specific representations, preserving each evaluatorβs unique perspective while promoting alignment across experts to enhance overall preference expressiveness. Experimental results demonstrate that CSPF significantly outperforms single-expert baselines, scalar-based multi-expert fusion approaches, and conventional scoring rule ensembles on both LM-Arena domain adaptation and PPE out-of-distribution evaluation benchmarks.
π Abstract
At present, reliable evaluation of non-verifiable tasks remains challenging. Existing approaches often fail to adequately capture the diverse evaluative criteria underlying human preferences in such tasks. To this end, we propose Constrained Shared-Private Fusion (CSPF), a fusion method that treats heterogeneous frozen reward models as complementary evaluators and learns to integrate their hidden-state representations under pairwise human-preference supervision. CSPF decomposes each expert signal into shared and expert-private representations, encouraging cross-expert alignment while preserving complementary viewpoints. Across experiments on LM-Arena target-domain adaptation and PPE out-of-distribution preference evaluation, CSPF achieves the best performance on the primary metrics among the evaluated single-expert reward-model, scalar-score multi-expert, and rubric-judge baselines. Overall, CSPF suggests that fusing hidden-state representations provides a more expressive basis for preference assessment, offering a practical route toward integrated evaluative signals for non-verifiable preference tasks.