🤖 AI Summary
Empirical understanding of Retrieval-Augmented Generation (RAG) in industrial practice remains scarce, particularly regarding real-world usage patterns, stakeholder requirements, operational challenges, and evaluation methodologies. Method: We conducted semi-structured interviews with 13 industry practitioners and performed qualitative thematic analysis to systematically characterize RAG deployment in enterprise settings. Contribution/Results: This study is the first to empirically identify prevalent industrial use cases—primarily vertical-domain question answering—as well as core requirements: data security, content quality assurance, and model interpretability. Key bottlenecks include labor-intensive data preprocessing, reliance on manual evaluation, and absence of robust automated evaluation frameworks. We further find that most deployed RAG systems remain at the prototype stage, with critical concerns—including ethical risks, bias mitigation, and scalability—largely unaddressed. Our findings provide an evidence-based foundation and actionable guidance for transitioning RAG from academic research to production-grade industrial adoption.
📝 Abstract
Retrieval-Augmented Generation (RAG) is a well-established and rapidly evolving field within AI that enhances the outputs of large language models by integrating relevant information retrieved from external knowledge sources. While industry adoption of RAG is now beginning, there is a significant lack of research on its practical application in industrial contexts. To address this gap, we conducted a semistructured interview study with 13 industry practitioners to explore the current state of RAG adoption in real-world settings. Our study investigates how companies apply RAG in practice, providing (1) an overview of industry use cases, (2) a consolidated list of system requirements, (3) key challenges and lessons learned from practical experiences, and (4) an analysis of current industry evaluation methods. Our main findings show that current RAG applications are mostly limited to domain-specific QA tasks, with systems still in prototype stages; industry requirements focus primarily on data protection, security, and quality, while issues such as ethics, bias, and scalability receive less attention; data preprocessing remains a key challenge, and system evaluation is predominantly conducted by humans rather than automated methods.