🤖 AI Summary
This work investigates the performance degradation of large language models (LLMs) in zero-shot group recommendation when applying social choice aggregation rules (e.g., Borda count, majority voting), as group complexity—measured by number of users, items, and ratings—increases. Through systematic experiments, we first identify a critical threshold: accuracy drops significantly beyond 100 group ratings. To mitigate this, we propose an in-context learning (ICL)-based prompting strategy that improves accuracy by up to 37%, whereas explanation generation or domain-specific prompts yield no significant gains. We further demonstrate that input encoding format—user-wise versus item-wise preference representation—critically affects performance. Evaluations across Llama, Mixtral, and GPT-series models show that, under moderate complexity, smaller-parameter LLMs with optimized prompting achieve high accuracy comparable to larger models, validating their viability as computationally efficient alternatives.
📝 Abstract
Large Language Models (LLMs) are increasingly applied in recommender systems aimed at both individuals and groups. Previously, Group Recommender Systems (GRS) often used social choice-based aggregation strategies to derive a single recommendation based on the preferences of multiple people. In this paper, we investigate under which conditions language models can perform these strategies correctly based on zero-shot learning and analyse whether the formatting of the group scenario in the prompt affects accuracy. We specifically focused on the impact of group complexity (number of users and items), different LLMs, different prompting conditions, including In-Context learning or generating explanations, and the formatting of group preferences. Our results show that performance starts to deteriorate when considering more than 100 ratings. However, not all language models were equally sensitive to growing group complexity. Additionally, we showed that In-Context Learning (ICL) can significantly increase the performance at higher degrees of group complexity, while adding other prompt modifications, specifying domain cues or prompting for explanations, did not impact accuracy. We conclude that future research should include group complexity as a factor in GRS evaluation due to its effect on LLM performance. Furthermore, we showed that formatting the group scenarios differently, such as rating lists per user or per item, affected accuracy. All in all, our study implies that smaller LLMs are capable of generating group recommendations under the right conditions, making the case for using smaller models that require less computing power and costs.