🤖 AI Summary
This work investigates the reliability and consistency of large language models (LLMs) serving dual roles—joint decision-making and explanation generation—in group recommendation systems (GRS). Grounded in social choice theory, we systematically evaluate LLM-generated recommendations and their natural-language explanations against classical aggregation baselines (e.g., additive utilitarianism) across diverse group structures. We find that while LLM recommendations empirically approximate additive utilitarianism, their explanations frequently misattribute decisions to rating averaging and introduce undefined implicit criteria—such as similarity, diversity, or popularity—without justification. Moreover, LLM recommendations exhibit low sensitivity to intra-group disagreement, undermining both explainability and transparency. To our knowledge, this is the first study to identify and characterize the “decision–explanation misalignment” phenomenon in LLM-based GRS. Our findings provide a critical diagnostic framework for trustworthy group recommendation and inform concrete directions for improving alignment between LLM reasoning and behavior.
📝 Abstract
Large Language Models (LLMs) are increasingly being implemented as joint decision-makers and explanation generators for Group Recommender Systems (GRS). In this paper, we evaluate these recommendations and explanations by comparing them to social choice-based aggregation strategies. Our results indicate that LLM-generated recommendations often resembled those produced by Additive Utilitarian (ADD) aggregation. However, the explanations typically referred to averaging ratings (resembling but not identical to ADD aggregation). Group structure, uniform or divergent, did not impact the recommendations. Furthermore, LLMs regularly claimed additional criteria such as user or item similarity, diversity, or used undefined popularity metrics or thresholds. Our findings have important implications for LLMs in the GRS pipeline as well as standard aggregation strategies. Additional criteria in explanations were dependent on the number of ratings in the group scenario, indicating potential inefficiency of standard aggregation methods at larger item set sizes. Additionally, inconsistent and ambiguous explanations undermine transparency and explainability, which are key motivations behind the use of LLMs for GRS.