🤖 AI Summary
This study addresses the attribute independence assumption and data coarsening issues in multivariate ordinal preference modeling by proposing a covariate-dependent joint Markov random field (MRF) model. Methodologically, it constructs a continuation-ratio MRF estimated via maximum likelihood, effectively handling the intractable normalizing constant. Theoretically, it establishes a unified framework demonstrating that existing comparison models are special cases, thereby preventing information loss. This work enables efficient inference and significantly improves preference comparison performance based on ordinal data.
📝 Abstract
Multivariate ordinal data along with covariates are commonly collected in problems ranging from alignment of language models with human preferences, as well as in recommender systems. For example, data sets such as MovieLens contain several movies rated on a scale 1--5 by human users, along with their demographic information such as age or gender. Similarly, data sets such as HelpSteer collect human feedback on several attributes such as "helpfulness" or "verbosity" of LLM response on an ordinal scale, with covariates depending on the LLM prompt--response pairs. Unfortunately, the standard approaches for modeling these data (a) look at the attributes individually rather than jointly, and (b) often convert the data into pairwise or list-wise win--loss comparisons for fitting models such as Bradley--Terry and Plackett--Luce. Both of these lead to a coarsening of what is actually observed, which we address via a joint covariate-dependent consecutive ratio Markov random field model. We also show pairwise or listwise comparison models are obtained under restrictions of our joint model, and that joint modeling improves comparisons. We also develop a maximum likelihood inference procedure even in the presence of an intractable normalizer.