Direct Judgement Preference Optimization

📅 2024-09-23
🏛️ arXiv.org
📈 Citations: 15
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited generalizability and susceptibility to positional/length biases of large language models (LLMs) when employed as generative evaluators. Methodologically, we propose a multi-source preference optimization framework integrating both positive and negative samples: (1) a three-way preference pair construction strategy; (2) a robust variant of Direct Preference Optimization (DPO); and (3) an evaluation protocol alignment mechanism enabling customizable assessment scenarios and actionable feedback generation. Our key contribution is the first explicit incorporation of negative samples into generative evaluator modeling, systematically mitigating structural biases. Empirically, the method achieves state-of-the-art (SOTA) performance on 10 out of 13 major benchmarks—surpassing GPT-4o and specialized evaluator models—while demonstrating high discriminative accuracy and strong cross-task adaptability.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsSearch and Optimization: Sampling/Simulation-based Search

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSocial Networks and Social Media: Generative AI / large language models and their impact on social systems
📝 Abstract
Auto-evaluation is crucial for assessing response quality and offering feedback for model development. Recent studies have explored training large language models (LLMs) as generative judges to evaluate and critique other models' outputs. In this work, we investigate the idea of learning from both positive and negative data with preference optimization to enhance the evaluation capabilities of LLM judges across an array of different use cases. We achieve this by employing three approaches to collect the preference pairs for different use cases, each aimed at improving our generative judge from a different perspective. Our comprehensive study over a wide range of benchmarks demonstrates the effectiveness of our method. In particular, our generative judge achieves the best performance on 10 out of 13 benchmarks, outperforming strong baselines like GPT-4o and specialized judge models. Further analysis show that our judge model robustly counters inherent biases such as position and length bias, flexibly adapts to any evaluation protocol specified by practitioners, and provides helpful language feedback for improving downstream generator models.
Problem

Research questions and friction points this paper is trying to address.

Enhancing LLM judges' evaluation capabilities via preference optimization
Improving generative judges using positive and negative data pairs
Countering evaluation biases and adapting to different protocols
Innovation

Methods, ideas, or system contributions that make the work stand out.

Preference optimization from positive and negative data
Three approaches collect preference pairs for different cases
Generative judge counters biases and adapts flexibly
🔎 Similar Papers
No similar papers found.