CSPF: A Constrained Shared-Private Fusion Method for Non-Verifiable Preference Evaluation

πŸ“… 2026-07-22
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenge of reliably evaluating human preferences in non-verifiable tasks, where diverse and often conflicting evaluation criteria hinder consistent assessment. To overcome this limitation, the authors propose the Constrained Shared-Private Fusion (CSPF) method, which treats multiple heterogeneous frozen reward models as complementary evaluators and fuses their hidden representations under pairwise human preference supervision. CSPF explicitly decomposes shared and expert-specific representations, preserving each evaluator’s unique perspective while promoting alignment across experts to enhance overall preference expressiveness. Experimental results demonstrate that CSPF significantly outperforms single-expert baselines, scalar-based multi-expert fusion approaches, and conventional scoring rule ensembles on both LM-Arena domain adaptation and PPE out-of-distribution evaluation benchmarks.
πŸ“ Abstract
At present, reliable evaluation of non-verifiable tasks remains challenging. Existing approaches often fail to adequately capture the diverse evaluative criteria underlying human preferences in such tasks. To this end, we propose Constrained Shared-Private Fusion (CSPF), a fusion method that treats heterogeneous frozen reward models as complementary evaluators and learns to integrate their hidden-state representations under pairwise human-preference supervision. CSPF decomposes each expert signal into shared and expert-private representations, encouraging cross-expert alignment while preserving complementary viewpoints. Across experiments on LM-Arena target-domain adaptation and PPE out-of-distribution preference evaluation, CSPF achieves the best performance on the primary metrics among the evaluated single-expert reward-model, scalar-score multi-expert, and rubric-judge baselines. Overall, CSPF suggests that fusing hidden-state representations provides a more expressive basis for preference assessment, offering a practical route toward integrated evaluative signals for non-verifiable preference tasks.
Problem

Research questions and friction points this paper is trying to address.

non-verifiable preference evaluation
reward models
preference assessment
heterogeneous evaluators
human preferences
Innovation

Methods, ideas, or system contributions that make the work stand out.

Constrained Shared-Private Fusion
hidden-state fusion
non-verifiable preference evaluation
reward model ensembling
preference alignment
πŸ”Ž Similar Papers
No similar papers found.
H
Hehao Zhang
Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100049, China
D
Danli Wang
Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100049, China
X
Xinyuan Wang
Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100049, China
Xuange Gao
Xuange Gao
Institute of Automation, Chinese Academy of Sciences
ElectroencephalogramDeep LearningBCI