Maximizing Signal in Human-Model Preference Alignment

📅 2025-03-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the challenges of aligning large language model (LLM) outputs with end-user preferences and mitigating high noise in human feedback, this paper proposes a Noise–Signal Decoupling Framework that systematically disentangles stochastic annotation noise from genuine preference signals within labeling disagreements. Methodologically, it introduces a human-feedback-based annotation quality analysis and consistency modeling mechanism, designs a preference-driven supervised fine-tuning strategy, and incorporates dual guardrail classifiers for closed-loop evaluation. Its key innovation lies in explicitly formulating preference signal maximization as an optimization objective—integrated throughout both training and evaluation. Experiments demonstrate significant noise reduction: the approach improves accuracy and fairness on user-consensus–sensitive tasks—including toxicity detection and key-point extraction in summarization—while yielding a reusable, trustworthy evaluation practice guideline.

Technology Category

Machine Learning: Learning Preferences or RankingsNatural Language Processing: (Large) Language ModelsHumans and AI: Learning Human Values and Preferences

Application Category

Economics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labelingUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
The emergence of powerful LLMs has led to a paradigm shift in Natural Language Understanding and Natural Language Generation. The properties that make LLMs so valuable for these tasks -- creativity, ability to produce fluent speech, and ability to quickly and effectively abstract information from large corpora -- also present new challenges to evaluating their outputs. The rush to market has led teams to fall back on quick, cost-effective automatic evaluations which offer value, but do not obviate the need for human judgments in model training and evaluation. This paper argues that in cases in which end users need to agree with the decisions made by ML models -- e.g. in toxicity detection or extraction of main points for summarization -- models should be trained and evaluated on data that represent the preferences of those users. We support this argument by explicating the role of human feedback in labeling and judgment tasks for model training and evaluation. First, we propose methods for disentangling noise from signal in labeling tasks. Then we show that noise in labeling disagreement can be minimized by adhering to proven methodological best practices, while signal can be maximized to play an integral role in model training and evaluation tasks. Finally, we illustrate best practices by providing a case study in which two guardrails classifiers are evaluated using human judgments to align final model behavior to user preferences. We aim for this paper to provide researchers and professionals with guidelines to integrating human judgments into their ML and generative AI evaluation toolkit, particularly when working toward achieving accurate and unbiased features that align with users' needs and expectations.
Problem

Research questions and friction points this paper is trying to address.

Challenges in evaluating LLM outputs due to creativity and fluency.
Need for human judgments in model training and evaluation.
Maximizing signal in human feedback for user-aligned model behavior.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Disentangling noise from signal in labeling tasks
Minimizing labeling disagreement through best practices
Aligning model behavior with user preferences via human judgments
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Kelsey Kraus
Cisco Systems
M
Margaret Kroll
Cisco Systems