Comparison-based Active Preference Learning for Multi-dimensional Personalization

📅 2024-11-01
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the challenge of accurately modeling users’ implicit multi-dimensional preferences from sparse pairwise comparative feedback in large language model (LLM) response personalization, this paper proposes AMPLe—a novel framework for adaptive multi-dimensional preference learning. AMPLe innovatively refines Bayesian posterior updating to suppress noise and bias inherent in pairwise comparisons, and introduces a generalized binary search–based active querying strategy to minimize user annotation effort. By jointly modeling multi-dimensional preferences and grounding query selection in theoretical guarantees, AMPLe achieves over 40% improvement in feedback efficiency for language generation tasks, producing high-quality personalized responses with only a small number of comparative judgments. The implementation is publicly available.

Technology Category

Machine Learning: Learning Preferences or RankingsNatural Language Processing: Language Grounding & Multi-modal NLPSearch and Optimization: Learning to Search

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
Large language models (LLMs) have shown remarkable success, but aligning them with human preferences remains a core challenge. As individuals have their own, multi-dimensional preferences, recent studies have explored multi-dimensional personalization, which aims to enable models to generate responses personalized to explicit preferences. However, human preferences are often implicit and thus difficult to articulate, limiting the direct application of this approach. To bridge this gap, we propose Active Multi-dimensional Preference Learning (AMPLe), designed to capture implicit user preferences from interactively collected comparative feedback. Building on Bayesian inference, our work introduces a modified posterior update procedure to mitigate estimation bias and potential noise in comparisons. Also, inspired by generalized binary search, we employ an active query selection strategy to minimize the number of required comparisons by a user. Through theoretical analysis and experiments on language generation tasks, we demonstrate feedback efficiency and effectiveness of our framework in personalizing model responses. Our code is publicly available at https://github.com/ml-postech/AMPLe .
Problem

Research questions and friction points this paper is trying to address.

Aligning LLMs with implicit human preferences
Capturing multi-dimensional preferences from comparative feedback
Reducing user comparisons via active query selection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Active Multi-dimensional Preference Learning captures implicit preferences
Modified Bayesian posterior update reduces estimation bias
Active query selection minimizes required user comparisons
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
POSTECH