Score
Designs and builds systems that perform retrieval-augmented generation entirely on end-user devices, including local indexing and retrieval from private or offline corpora and fusion of retrieved context into the prompt conditioning of on-device language models. Work covers engineering to run mid-sized models under device resource and offline constraints, implementing local/offline retrieval and context-fusion, and analyzing correctness, latency, and privacy trade-offs.
This work addresses the limitations of traditional retrieval-augmented generation (RAG) systems, which rely on server-side computation and suffer from privacy risks, high latency, and substantial storage overhead, while on-device deployment is hindered by resource constraints that compromise both retrieval efficiency and generation quality. To overcome these challenges, the authors propose a unified on-device RAG model that jointly optimizes retrieval and context compression for the first time. By sharing document representations across both tasks, the model enables efficient retrieval and compact context generation within a single architecture, eliminating redundancy from multiple separate models. The approach achieves generation performance comparable to conventional RAG using only approximately one-tenth of the context length, while maintaining embedding storage costs no higher than existing multi-vector retrieval methods, thereby significantly enhancing the practicality and efficiency of on-device RAG.
To address the limitations of small language models (SLMs) on edge devices—namely, degraded reasoning performance due to insufficient local knowledge—and the privacy risks of cloud-based retrieval-augmented generation (RAG) that may expose users’ private documents, this paper proposes DRAGON, a distributed RAG framework. DRAGON enables privacy-preserving retrieval by jointly leveraging a cloud-side general knowledge base and edge-side private documents, without uploading raw private data. Its core contributions are: (1) the first edge-oriented distributed multi-document RAG architecture; (2) a bilateral speculative aggregation algorithm supporting asynchronous, cloud-edge collaborative token generation; and (3) a dynamic aggregation-side scheduling mechanism adaptive to real-time network conditions. Evaluation on a real-world edge platform demonstrates that DRAGON achieves 1.9× higher end-to-end throughput than centralized RAG, significantly reduces per-token latency, and drives time-to-first-token (TTFT) near zero.
This study systematically investigates the impact mechanisms of individual components in Retrieval-Augmented Generation (RAG) systems on complex question answering and cross-domain tasks. Addressing key challenges—including low retrieval precision, weak contextual relevance, and poor multilingual adaptability—we propose three core innovations: (1) a Contrastive In-Context Learning (CICL) RAG paradigm to improve generation accuracy; (2) sentence-granularity focused retrieval (“Focus Mode”) combined with multi-granularity chunking to enhance retrieval relevance; and (3) a multilingual knowledge base integration framework that balances retrieval–generation efficiency. Through large-scale hyperparameter analysis, we quantitatively characterize the influence of critical factors—including language model scale, chunk size, and retrieval stride—on end-to-end performance. The findings yield a reproducible best-practice guideline for RAG system design and deployment, accompanied by open-sourced, fully implemented code.
Current retrieval-augmented generation systems require uploading sensitive documents to remote servers, posing significant privacy risks. This work proposes a “local-first information retrieval” paradigm that fully deploys indexing, models, and inference on the user’s device, with optional remote service invocation. The authors formally define this design paradigm for the first time and establish a system framework centered on three dimensions: privacy control, capability, and accessibility. They identify search scope—not quality—as the primary trade-off in local-first systems. Experimental results demonstrate that, on consumer-grade hardware, a local system combining dense retrieval, BM25, and HNSW indexing achieves over 91% nDCG@10 on collections of up to 100K documents, with only a 2% drop when scaled to 1M documents. Furthermore, question-answering quality using a 7B-parameter language model lags behind cloud-based baselines by merely 4 points.
The impact of key design decisions—RAG activation, retrieval granularity, and knowledge integration strategy—on RAG system performance remains poorly understood. Method: We conduct systematic ablation studies across three code/qa benchmarks and two state-of-the-art LLMs, quantitatively evaluating how document type, recall rate, document selection strategy, and prompt engineering jointly affect answer correctness and confidence via multi-dimensional analysis, cross-model/dataset comparison, and joint prompt-retrieval analysis. Contribution/Results: We identify precise interaction patterns and operational boundaries among these factors and propose nine actionable, empirically grounded guidelines for diagnosing and optimizing RAG failures. Our findings significantly improve RAG system stability, debuggability, and reliability, offering rigorous empirical evidence and a principled methodology to support the engineering deployment of LLM-augmented systems.
This work addresses the security and privacy risks inherent in Retrieval-Augmented Generation (RAG) systems—such as sensitive data leakage and knowledge base tampering—by proposing the first unified threat taxonomy that spans the entire RAG pipeline, including retrieval, context construction, and generation. It systematically analyzes attack surfaces across centralized, on-device (Micro-RAG), federated, and hybrid deployment paradigms, and integrates corresponding defense mechanisms leveraging techniques like differential privacy, secure aggregation, and encrypted retrieval. The study provides a comprehensive survey of existing research, clarifies the trade-offs between privacy and utility, highlights practical deployment challenges, and identifies critical open problems, thereby establishing a theoretical foundation and charting future directions for developing trustworthy, secure, and robust RAG systems.
This study investigates the capacity of small language models (7B parameters or fewer) to effectively leverage external information in retrieval-augmented generation (RAG). Through systematic evaluation on models such as SmolLM2, Qwen2.5, and Llama 3.1—combined with BM25, E5-large-v2, and oracle retrievers across multiple prompt templates—the work introduces a novel parameterized knowledge partitioning framework that cleanly disentangles retrieval failure from context utilization failure for the first time. The findings reveal a fundamental bottleneck in small models’ ability to use retrieved content: even under oracle retrieval conditions, 85%–100% of samples fail to correctly incorporate the relevant answer, and 42%–100% of the model’s original knowledge is disrupted by the retrieved context. The dominant error mode is generation entirely unrelated to the provided context, indicating a pervasive inability to attend to or integrate external information.
This study addresses the lack of systematic evaluation of small language models (SLMs) in retrieval-augmented generation (RAG) systems, particularly their potential for deployment on resource-constrained devices. It presents the first comprehensive assessment of SLMs in the RAG generation phase, benchmarking performance across diverse domains using both open-source and proprietary datasets. The work further demonstrates end-side inference entirely on CPU-based hardware without GPU acceleration. Experimental results show that SLMs can operate efficiently under such constraints, substantially reducing computational overhead while maintaining reasonable response times and generating high-quality outputs. These findings establish a viable pathway for deploying lightweight, edge-compatible AI systems leveraging RAG architectures.
This work addresses key challenges in deploying large language models for retrieval-augmented generation (RAG), including high computational overhead, rapid knowledge obsolescence, and manual dependency in component selection. The authors propose a modular evaluation framework that, for the first time, directly links hardware constraints to RAG performance. By integrating resource telemetry with an automated recommendation mechanism, the framework efficiently identifies optimal combinations of components—including document chunking strategies, embedding models, vector databases, and retrievers—for domain-specific datasets. This approach maintains high generation quality while substantially reducing resource consumption. Designed to support rapid prototyping on consumer-grade hardware, the framework enables automatic, domain-tailored RAG configuration, achieving a favorable trade-off among accuracy, efficiency, and scalability.
Traditional information retrieval prioritizes topical relevance, which often fails to meet the practical utility demands of downstream large language model (LLM) tasks. This work proposes a utility-centered retrieval paradigm that shifts the objective from relevance to the actual contribution of retrieved content to LLM generation quality. It introduces the first unified framework encompassing diverse utility forms—spanning LLM-agnostic and LLM-aware, as well as context-independent and context-dependent settings—and explicitly links LLM information needs with agent-based RAG mechanisms. By integrating retrieval-augmented generation, information need modeling, utility-oriented evaluation metrics, and intelligent retrieval strategies, this study redefines retrieval evaluation criteria and establishes a synergistic optimization pathway between retrieval and generation, offering both theoretical foundations and practical guidance for information retrieval in the LLM era.