Score
Designs and implements retrieval-augmented generation systems that incorporate real-time inventory signals to retrieve, rank, and condition on inventory-relevant passages and items. This competence covers building retrieval pipelines and generative conditioning that bias candidate rewrites and outputs toward available products or ads, incorporate historically successful queries, and handle inventory freshness and availability constraints.
Empirical understanding of Retrieval-Augmented Generation (RAG) in industrial practice remains scarce, particularly regarding real-world usage patterns, stakeholder requirements, operational challenges, and evaluation methodologies. Method: We conducted semi-structured interviews with 13 industry practitioners and performed qualitative thematic analysis to systematically characterize RAG deployment in enterprise settings. Contribution/Results: This study is the first to empirically identify prevalent industrial use cases—primarily vertical-domain question answering—as well as core requirements: data security, content quality assurance, and model interpretability. Key bottlenecks include labor-intensive data preprocessing, reliance on manual evaluation, and absence of robust automated evaluation frameworks. We further find that most deployed RAG systems remain at the prototype stage, with critical concerns—including ethical risks, bias mitigation, and scalability—largely unaddressed. Our findings provide an evidence-based foundation and actionable guidance for transitioning RAG from academic research to production-grade industrial adoption.
This work addresses the challenge of achieving end-to-end coordination between retrieval and generation in traditional retrieval-augmented generation (RAG) systems. The authors propose GRIP, a novel framework that integrates retrieval into autoregressive generation through control tokens, dynamically determining when to retrieve, how to reformulate queries, and when to terminate retrieval. Its core innovation is the Self-Triggered Information Planning mechanism, which unifies multi-step reasoning and real-time evidence integration within a single generation trajectory. GRIP employs structured supervision signals to train retrieval behaviors across diverse scenarios—including answerable, partially answerable, and multi-hop queries. Experimental results demonstrate that GRIP significantly outperforms strong RAG baselines on five question-answering benchmarks, achieving performance comparable to GPT-4o with substantially fewer parameters.
This work addresses the lack of a systematic taxonomy for retrieval-augmented generation (RAG) applications. We propose the first comprehensive, lifecycle-spanning classification framework for RAG applications. Methodologically, we introduce a novel four-stage iterative construction paradigm—comprising multi-round expert collaboration, systematic literature review, dimensional abstraction, and empirical validation—thereby filling a critical gap in classification research beyond the ACL community. The framework comprises five meta-dimensions and sixteen fine-grained dimensions, balancing structural clarity with extensibility. Empirical validation across education, healthcare, and legal domains demonstrates its effectiveness in supporting design decisions, technical evaluation, and cross-domain understanding of RAG applications. By providing a foundational taxonomic infrastructure, this work advances the engineering-oriented deployment and standardization of RAG systems.
The impact of key design decisions—RAG activation, retrieval granularity, and knowledge integration strategy—on RAG system performance remains poorly understood. Method: We conduct systematic ablation studies across three code/qa benchmarks and two state-of-the-art LLMs, quantitatively evaluating how document type, recall rate, document selection strategy, and prompt engineering jointly affect answer correctness and confidence via multi-dimensional analysis, cross-model/dataset comparison, and joint prompt-retrieval analysis. Contribution/Results: We identify precise interaction patterns and operational boundaries among these factors and propose nine actionable, empirically grounded guidelines for diagnosing and optimizing RAG failures. Our findings significantly improve RAG system stability, debuggability, and reliability, offering rigorous empirical evidence and a principled methodology to support the engineering deployment of LLM-augmented systems.
This study addresses the prevalent issue in current advertising systems where semantic mismatches between user queries and ad keywords lead to failed ad fill, resulting in lost revenue and engagement opportunities. To mitigate this, the authors propose InvAwr-RAG, an inventory-aware retrieval-augmented generation framework that uniquely integrates real-time ad inventory information into the RAG architecture. By leveraging generative AI to dynamically rewrite user queries—combining semantic retrieval with historically high-conversion query patterns—the approach enhances coverage diversity while preserving relevance. Experimental results demonstrate that InvAwr-RAG improves ad fill rate by 68%, substantially boosting both advertising revenue potential and user experience.
Existing asynchronous retrieval-augmented generation (RAG) systems rely on heuristic coordination strategies that struggle to adapt to dynamically evolving information needs across diverse domains, resulting in limited efficiency and flexibility. To address this, this work proposes a novel asynchronous retrieval framework that leverages semantic precursors emerging early in the generation process to explicitly predict both the optimal timing and content for retrieval, enabling intelligent prefetching aligned with dynamic user demands. The framework integrates a retrieval predictor, a context monitor, and a query generator, jointly modeling semantic precursors and evolving information requirements. Experimental results demonstrate that the approach reduces end-to-end latency by up to 43.5% and accelerates first-token output speed by up to 62.4% across multiple benchmarks, while maintaining answer quality comparable to that of synchronous RAG systems.
This study addresses the challenge of inefficient cross-site and cross-organizational reuse of industrial spare parts, hindered by decentralized storage, inconsistent naming conventions, and missing information. To overcome these issues, the authors propose PhRAG, a novel framework that integrates multitask generative language modeling with named entity recognition and hybrid retrieval-augmented generation (RAG) to construct a unified virtual spare parts pool (VSPool) from heterogeneous unstructured data. The approach achieves robust structured information extraction under data-scarce conditions and enables natural language querying with generative, interpretable retrieval. Experimental results demonstrate that PhRAG outperforms conventional NER methods in technical specification extraction and significantly enhances both the efficiency of spare parts reuse across organizations and the transparency of the overall system.
This work addresses the limitation of existing generative retrieval systems in e-commerce search, which often overlook item commercial value, leading to high-value products being either omitted or ranked suboptimally during retrieval. To overcome this, we propose TSGR, a unified generative retrieval framework that introduces a novel query-aware parallel Semantic ID (SID) encoding mechanism to integrate item value signals into representation learning. TSGR further incorporates an end-to-end jointly optimized, value-aware ranking module, eliminating the need for a separate pre-ranking stage. By combining an autoregressive generative model, statistically driven SID encoding, and a progressive training strategy, TSGR achieves a 9.16% improvement in offline HR@1000 and demonstrates significant online gains in A/B tests, including +0.43% in item page views (IPV), +1.12% in transaction count, and +1.64% in gross merchandise value (GMV).