π€ AI Summary
This study addresses the challenge faced by memory-augmented assistants in determining when user preferences are contextually applicable. We propose PairPref, a benchmark that introduces a novel paired evaluation paradigm focusing on preference applicability conditions. By holding preferences constant while varying contexts, it systematically assesses modelsβ ability to correctly apply preferences during both retrieval and generation. The benchmark comprises 1,227 paired samples spanning 45 preference categories and 8 contextual scenarios, organized into selection and free-form generation tracks. Experiments reveal that mainstream models achieve only 3.6%β18.3% accuracy in context-aware preference application within the free-form generation track, exposing significant deficiencies in current large language models regarding dynamic preference adaptation.
π Abstract
Memory-augmented assistants use retrieved preferences to guide their responses. A small change in the situation can change whether a preference is appropriate while barely affecting its retrieval similarity. Memory benchmarks typically test whether systems store and retrieve preferences, with less attention to when those preferences should apply. We introduce PairPref, a benchmark of contextual preference use. Each pair changes only the situation, keeping the preference, request, and four candidate replies fixed. The preference remains valid in both situations. In the selection track, models must choose the reply that applies the preference only where appropriate. In the free-generation track, they must decide when to apply it without seeing candidate replies. Both tracks use the same 1,227 pairs across 45 preferences and eight situation categories. We evaluate eight models, most of which achieve selection scores ($Ξ$) of 51 to 65 points. In free generation, however, both responses are appropriate for their respective situations in only 3.6\% to 18.3\% of pairs. Models continue to apply the preference in both situations even with fewer retrieved memories, alternative presentation formats, and a stricter prompt. These results show that models still struggle to judge when user preferences apply and respond accordingly.