Score
Designs and builds systems that combine retrieval of relevant items or documents with recommendation models and LLM-based generators to produce accurate, context-aware suggestions; this includes constructing retrieval pipelines to fetch candidate items, reranking and filtering to ensure factual consistency, and formatting outputs as human-readable recommendations. Also develops RAG-enabled architectures that condition LLM generation on retrieved evidence to constrain suggestions to verified entries and to generate explainable or counterfactual recommendation alternatives.
To address inherent limitations in recommender systems—including data sparsity, cold-start challenges, insufficient personalization, and weak semantic understanding—this paper proposes a novel LLM-based recommendation paradigm, positioning large language models as the foundational architecture rather than merely auxiliary components. Methodologically, it integrates prompt-driven retrieval, language-native ranking, retrieval-augmented generation (RAG), and conversational interaction to enable zero-shot and few-shot cross-task generalization; multi-stage candidate generation coupled with external knowledge injection further enhances semantic alignment and interpretability. The key contribution is the first structured, LLM-enhanced recommendation framework, which significantly improves personalization accuracy and semantic comprehension—particularly under cold-start and long-tail conditions. Additionally, the work systematically investigates synergistic optimization pathways balancing accuracy, scalability, and real-time responsiveness.
To address the challenges of cold-start item recommendation and weak modeling of co-purchase patterns in recommender systems, this paper proposes ItemRAG—the first LLM-based recommendation framework leveraging item-level Retrieval-Augmented Generation (RAG). Methodologically, it replaces conventional user-level retrieval with a novel item-level semantic-co-purchase joint retrieval mechanism, which jointly encodes semantic similarity and frequency-weighted co-purchase signals to precisely capture item-level collaborative relationships. Furthermore, it constructs co-purchase sequences as retrieval contexts to enhance LLMs’ zero-shot recommendation capability. Empirically, ItemRAG achieves up to a 43% improvement in Hit-Ratio@1 under zero-shot settings across multiple benchmark datasets, significantly outperforming user-centric baselines. Crucially, it maintains consistent superiority in both standard and cold-start scenarios, demonstrating robust generalization without task-specific fine-tuning.
This work addresses the challenge of generating personalized, factually accurate, and topically relevant reading materials tailored to user-specified queries and target readability levels. To this end, the authors propose a four-module system that integrates retrieval-augmented generation (RAG) with large language models (LLMs), uniquely combining RAG with multiple prompting strategies—including Chain-of-Thought, zero-shot, and few-shot prompting—for personalized reading recommendations. The system further incorporates an LLM-as-a-Judge mechanism to automatically evaluate the factual accuracy, relevance, and readability alignment of generated content. Experimental results demonstrate that RAG consistently enhances the performance of diverse models—including LLaMA 4 Scout, LLaMA 3.1 8B, and Gemma2 9B—across all prompting strategies, yielding improvements of up to 26–35 percentage points in relevance and factual accuracy, thereby enabling high-quality customized reading material generation.
To address critical challenges in LLM-based recommender systems—including severe hallucination, outdated knowledge, and insufficient structured modeling—this paper proposes K-RagRec, a knowledge graph (KG)-enhanced retrieval-augmented generation framework. Methodologically, it pioneers the deep integration of structure-aware KG retrieval into the LLM recommendation pipeline, introducing KG-aware dense retrieval and graph-structure-guided generation to mitigate noise and weak relational modeling inherent in conventional RAG approaches for recommendation. The framework jointly incorporates KG embeddings, graph-context injection, and end-to-end retrieval-generation fine-tuning. Extensive experiments on multiple public benchmarks demonstrate substantial improvements: Recall@10 increases by 12.7%, hallucination rate decreases by 38.5%, outperforming state-of-the-art methods. These results empirically validate that structured external knowledge significantly enhances recommendation accuracy, interpretability, and robustness.
To address the dual challenges of semantic mismatch in retrieval-augmented generation (RAG)-based recommendation and the lack of interpretable reasoning during generation, this paper proposes an LLM-based recommendation framework integrating representation learning with explicit chain-of-thought (CoT) reasoning. Methodologically: (1) it pioneers the incorporation of CoT into recommendation generation to enhance decision transparency; (2) it constructs a multi-source joint representation fusing textual semantics and collaborative signals; (3) it introduces a lightweight temporal re-ranking module to model user interest evolution; and (4) it strengthens retrieval-generation synergy via knowledge-injected prompting and consistency-aware fusion. Evaluated on three real-world datasets, the framework achieves up to 12.7% improvement in Recall@10 over state-of-the-art methods, while simultaneously enhancing reasoning interpretability and long-tail item coverage.
This work addresses the bottleneck in RAG system evaluation—its heavy reliance on human-annotated ground-truth answers—by proposing RAGAs, a reference-free automated evaluation framework. Methodologically, it introduces a computable, three-dimensional metric suite covering retrieval relevance, context faithfulness, and generation quality, integrating BERTScore for semantic similarity, NLI-based models for factual consistency, and self-supervised prompting strategies. Its key contribution is the first end-to-end, multidimensional, reference-free evaluation paradigm, enabling quantitative, pipeline-level diagnostics of RAG systems. Experiments demonstrate strong agreement between automated metrics and human judgments (average Spearman ρ > 0.82) across multiple benchmarks, validating efficacy and robustness. The open-source RAGAs toolkit has been widely adopted in industry for iterative RAG system optimization.
Existing RAG-based recommender systems struggle to effectively leverage the Web—a dynamic, noisy external knowledge source—due to two key challenges: (1) a semantic gap between recommendation tasks and web retrieval, hindering precise user preference query generation; and (2) high noise and information sparsity in web content, impeding reliable signal extraction. To address these, we propose WebRec, the first framework that employs LLMs for end-to-end generation of structured, web-adapted queries. It introduces the MP-Head mechanism, which enhances long-range token interaction via message passing to improve attention modeling robustness against noise. Experiments on multiple recommendation benchmarks demonstrate that WebRec significantly improves both accuracy and information utilization efficiency—especially under low signal-to-noise ratio conditions. WebRec establishes a scalable, retrieval-augmented paradigm for open-domain recommendation.
This work proposes a human-feedback-driven dual-RAG architecture designed to continuously enhance the accuracy, relevance, and overall quality of Retrieval-Augmented Generation (RAG) systems through a human-in-the-loop mechanism. The approach introduces an auxiliary feedback RAG module that automatically collects, categorizes, and integrates user feedback into the primary RAG inference pipeline, enabling autonomous iterative refinement without requiring explicit supervisory signals. Leveraging an LLM-as-a-Judge evaluation strategy, the method demonstrates significant improvements in response quality across three benchmark datasets encompassing both general and domain-specific knowledge, thereby advancing RAG systems toward self-optimizing capabilities.
This work addresses the challenge of hallucination in tool-augmented retrieval-augmented generation (agentic RAG) systems, where user queries often fall outside the system’s knowledge scope. While existing approaches merely filter out invalid questions, they lack mechanisms to guide users toward answerable formulations. To bridge this gap, this study introduces query suggestion into the agentic RAG paradigm for the first time, proposing an answerability-aware dynamic in-context learning method. By retrieving relevant workflow examples from historical interactions, the approach constructs few-shot prompts that enable a large language model to generate semantically similar yet answerable reformulations of the original query. The method supports multi-step workflows and incorporates self-learning capabilities. Evaluated on three real-world user query benchmarks, it significantly outperforms conventional few-shot and pure retrieval baselines, yielding suggestions that exhibit both high relevance and strong answerability.
This study addresses critical limitations of traditional Retrieval-Augmented Generation (RAG) systems—such as retrieval noise, misuse of retrieved content, weak query-document alignment, and high generation costs—and presents the first large-scale empirical comparison between enhanced RAG and agentic RAG paradigms. Leveraging a large language model–driven agent control flow, modular RAG components, and a multidimensional evaluation framework encompassing accuracy, robustness, and computational cost, the work systematically assesses performance and efficiency across diverse scenarios. Findings reveal that agentic RAG demonstrates superior adaptability in complex tasks, whereas enhanced RAG achieves higher efficiency in simpler settings. These results provide clear guidance on the trade-offs between performance and cost for real-world deployment and underscore a pathway toward more autonomous, agent-based RAG architectures.
This work addresses the overreliance of existing agent-based RAG systems on complex retrieval backends, which constrains the ability of large language models (LLMs) to precisely articulate information needs. The authors propose a novel agent RAG framework that restores retrieval control to the LLM by enabling it to explicitly formulate query intent through generated logical expressions, coupled with a lightweight inverted index for structured retrieval. By eschewing sophisticated embedding mechanisms, the approach substantially simplifies system architecture while maintaining performance comparable to strong baselines. This design significantly reduces both construction and serving costs and effectively mitigates hallucination in generated outputs.