Less is More: Benchmarking LLM Based Recommendation Agents

๐Ÿ“… 2026-01-28
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study investigates the impact of user history length on recommendation quality in large language model (LLM)-based recommender systems, challenging the common assumption that more context yields better performance. Using four state-of-the-art modelsโ€”GPT-4o-mini, DeepSeek-V3, Qwen2.5-72B, and Gemini 2.5 Flashโ€”we conduct within-user experiments on the REGEN dataset to evaluate the effectiveness of contextual histories ranging from 5 to 50 items, while also measuring inference latency. Our results demonstrate that increasing context length from 5 to 50 items does not significantly improve recommendation performance, with NDCG scores remaining stable between 0.17 and 0.23. Notably, as few as 5โ€“10 historical items suffice to maintain recommendation quality while reducing inference cost by approximately 88%. This work provides the first empirical evidence of the efficiency of short-context inputs in LLM-based recommendation, offering a practical foundation for low-cost deployment.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Data Mining & Knowledge Management: Conversational Systems for Recommendation & RetrievalNatural Language Processing: (Large) Language Models

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSearch and Retrieval-Augmented AI: Large language models for search
๐Ÿ“ Abstract
Large Language Models (LLMs) are increasingly deployed for personalized product recommendations, with practitioners commonly assuming that longer user purchase histories lead to better predictions. We challenge this assumption through a systematic benchmark of four state of the art LLMs GPT-4o-mini, DeepSeek-V3, Qwen2.5-72B, and Gemini 2.5 Flash across context lengths ranging from 5 to 50 items using the REGEN dataset. Surprisingly, our experiments with 50 users in a within subject design reveal no significant quality improvement with increased context length. Quality scores remain flat across all conditions (0.17--0.23). Our findings have significant practical implications: practitioners can reduce inference costs by approximately 88\% by using context (5--10 items) instead of longer histories (50 items), without sacrificing recommendation quality. We also analyze latency patterns across providers and find model specific behaviors that inform deployment decisions. This work challenges the existing ``more context is better'' paradigm and provides actionable guidelines for cost effective LLM based recommendation systems.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Recommendation Systems
Context Length
Personalized Recommendation
Inference Cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

context length
LLM-based recommendation
cost efficiency
latency analysis
REGEN benchmark
๐Ÿ”Ž Similar Papers
No similar papers found.
K
Kargi Chauhan
University of California, Santa Cruz
M
Mahalakshmi Venkateswarlu
Georgia Institute of Technology