🤖 AI Summary
This study addresses the challenge of dynamically selecting aggregation strategies for group recommendations that align with human perceptions of fairness, satisfaction, and consensus. To this end, the authors fine-tune large language models—Judgmental Llama and Judgmental OLMo—on human survey data to construct an evaluation module capable of simulating human group judgments in real time. Integrating social choice theory, they generate diverse recommendation candidates and enable preference-distribution-driven dynamic strategy selection. This work is the first to adapt large language models to the structure of group preferences, revealing how model judgments interact with group configurations such as minority factions or coalitions. In a user study involving 284 participants, the proposed method significantly outperformed baseline approaches in both satisfaction and group consensus metrics, with model judgments showing strong alignment with human perceptions.
📝 Abstract
Previous work in group recommender systems has demonstrated a sensitivity to the distribution of preferences within a group. Specifically, the selection of the preference aggregation strategy benefits from considering such group configurations. In this paper, we study whether LLMs are able to mimic this sensitivity and to select the ideal aggregation strategy (and corresponding recommendation) according to nuanced human perceptions of fairness, satisfaction, and consensus.
We do this by fine-tuning Large Language Models (LLMs) on human survey data to serve as real-time judgmental models within the recommendation pipeline. Using a reasoning dataset distilled from DeepSeek-V3.1 and human ground truth assessments, we develop Judgmental Llama and Judgmental OLMo to simulate group assessments. Our pipeline successfully generates multiple recommendation candidates based on social choice-based aggregation strategies and dynamically selects the one that maximizes these predicted human-like evaluations. We further validate these suggestions in a user study (n=284) and find that our methodology achieved the highest scores for satisfaction and group consensus. Furthermore, we find that LLM judgments are most aligned with human perceptions of fairness, satisfaction and consensus when we also consider interaction effects between our LLM-based method and group configuration (e.g., minority or coalition). These findings give further support for dynamically adapting aggregation strategies to specific within-group preference distributions, and highlight the advantage of using LLMs for an adaptation that is aligned with subjective human judgments.