interactive preference elicitation

Designs and implements algorithms that interactively probe agents’ preferences by selecting informative queries (for example proposing allocations or pairwise comparisons) and maintaining a version space—typically a polytope of candidate valuation or utility vectors—then updating that set using observed responses interpreted as separating hyperplanes or ellipsoid-style refinements. Builds query-selection and inference procedures that minimize the number of queries while enabling computation or certification of desired outcomes (e.g., EF1/Prop1 allocations for additive valuations) and analyzes the convergence and query-complexity guarantees of these procedures.

interactivepreferenceelicitation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.25
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study investigates the learnability and query efficiency gap between fixed (non-adaptive) and interactive (adaptive) testing in hypothesis testing over a finite outcome space. Under the conditional sampling model, it establishes the first necessary and sufficient condition for two distribution classes to be reliably distinguishable: a positive separation in their conditional probabilities. By measuring simulation accuracy via total variation distance and combining randomized non-adaptive designs with information-theoretic lower bounds, the work demonstrates that adaptive strategies confer at most a quadratic—rather than exponential—query advantage. Furthermore, it constructs a method that $\rho$-approximates any $T$-step adaptive strategy using only $O(N^2(T + \log(1/\rho)))$ pre-specified queries, thereby proving that the worst-case query complexity gap between adaptive and non-adaptive approaches is $\Theta(N^2)$.

adaptive queriesconditional probabilityhypothesis testing

This work addresses a key limitation of existing Direct Preference Optimization (DPO) methods, which rely solely on pairwise preference signals and neglect the quantitative differences in response quality, leading to ambiguous training signals and suboptimal optimization efficiency. To overcome this, we propose a novel preference optimization algorithm that, for the first time, incorporates explicit score gaps into the DPO framework. Our approach preserves the advantage of not requiring an explicit reward model while leveraging fine-grained relative quality information to enhance alignment. By designing a loss function that accounts for score differences, the method enjoys faster theoretical statistical convergence and demonstrates robustness to scoring noise. Extensive experiments show consistent and significant improvements over current DPO variants across multiple large language models and evaluation benchmarks, with stable performance gains even when provided with inaccurate scores.

Alignment ProblemDirect Preference OptimizationFoundation Models

Algorithmic Persuasion Through Simulation: Information Design in the Age of Generative AI

Nov 29, 2023
KH
Keegan Harris
🏛️ Carnegie Mellon University | Microsoft Research

This paper studies algorithmic persuasion under generative AI, where a sender aims to influence a receiver’s binary action via signaling, knowing only limited information about the receiver’s type distribution and needing to infer the type with minimal queries to a behavioral simulation oracle. Method: We integrate Bayesian persuasion with a queryable behavioral simulation oracle, proposing a polynomial-time joint optimization algorithm for optimal querying and signal design. Contribution/Results: We fully characterize the optimal signaling strategy for arbitrary type distributions; establish robustness under approximate oracles, general query structures, and cost-sensitive constraints; and empirically demonstrate significant improvements in both utility maximization and query efficiency.

Bayesian persuasion gameoptimal messaging policysender-receiver interaction

Latest Papers

What's happening recently
View more

This work addresses the challenge of applying traditional multi-winner voting rules in large-scale or attention-constrained settings, where eliciting complete preference rankings from voters is impractical. To overcome this limitation, the authors propose a structured-query framework for multi-winner elections that approximates an optimal committee by querying voters’ preferences over subsets of candidates within a limited budget. They formally define a cognitive cost function and axiomatic evaluation criteria, and introduce a query strategy based on recursively partitioning the candidate set. Experimental results demonstrate that this approach significantly outperforms alternative querying mechanisms across various election models and multi-winner rules—such as k-Borda—achieving high committee selection accuracy while substantially reducing the information acquisition cost.

committee selectionlimited budgetmultiwinner voting

This work addresses the challenge of user preference learning, which typically relies on costly annotated data, while existing active learning approaches suffer from high computational overhead and fail to account for varying reliability in user feedback. The authors propose Info-Synth, a novel framework that uniquely integrates confidence-aware response modeling with active query synthesis in continuous space. By maximizing mutual information, Info-Synth generates highly informative preference queries and introduces two strategies—Pair M-dist and Pair Opt-dist—to effectively handle ambiguous comparisons. The method demonstrates substantially improved learning efficiency, outperforming baseline approaches across diverse tasks including synthetic preference learning, text summarization, and robot controller tuning. Furthermore, it naturally extends to practical scenarios with limited query pools.

active learningcomputational efficiencyfeedback reliability

Hot Scholars

IG

Iryna Gurevych

Full Professor, TU Darmstadt; Adjunct Professor, MBZUAI, UAE; Affiliated Professor, INSAIT, Bulgaria
Natural Language ProcessingLarge Language ModelsArtificial Intelligence
ZX

Zhuohan Xie

MBZUAI
Financial AIReasoningNatural Language ProcessingComputational Linguistics
PW

Pengfei Wan

Head of Kling Video Generation Models, Kuaishou Technology
Generative ModelsComputer VisionMultimodal AIComputer Graphics
MO

Martin Obaidi

PhD Student, Leibniz University Hannover
Software Engineering
JB

Jinbin Bai

National University of Singapore
Machine LearningContent CreationGenerative Modeling