Score
Designs, builds, and evaluates systems that retrieve relevant textual context or exemplar prompts and assemble that context into inputs for language models (retrieval-augmented or RAG-style pipelines), including components for vector search, encoder integration, context selection, and prompt templating; analyzes how retrieval strategies, encoder choices, and assembled context affect few-shot, multilingual, or cross-lingual model behavior and output quality.
Large language models (LLMs) excel at complex contextual understanding but exhibit pronounced capability asymmetry—struggling to stably generate long, equally sophisticated texts. Method: We systematically establish a unified “context engineering” framework, proposing a four-dimensional taxonomy encompassing retrieval, generation, processing, and management. Based on a systematic review and architectural analysis of 1,300+ papers, we construct the first comprehensive context engineering technology map; identify the intrinsic mechanisms underlying the understanding–generation capability mismatch; and delineate architectural integration pathways for four key application paradigms: retrieval-augmented generation, memory modeling, tool integration, and multi-agent coordination. Contribution/Results: The work delivers a standardized conceptual framework, a strategic technology roadmap, and identified critical breakthrough directions—providing both theoretical foundations and practical guidance for developing advanced context-aware AI systems.
Large language models are constrained by static knowledge, limited context windows, and weak causal reasoning capabilities, hindering their effective use of external information. This work systematically reviews techniques that enhance model performance during inference through structured contextual augmentation, including in-context learning, prompt engineering, retrieval-augmented generation (RAG), graph-augmented RAG, and causal RAG. We propose a unified analytical framework that integrates literature synthesis, cross-study evidence aggregation, and claim auditing to distinguish high-confidence findings from emerging results. Building on this framework, we develop a deployment-oriented decision guide and release a prioritized research agenda for trustworthy retrieval-augmented NLP, offering systematic support for future research and practical applications.
To address core limitations of large language models (LLMs)—including hallucination, knowledge obsolescence, and poor domain adaptability—this work systematically advances the Retrieval-Augmented Structured (RAS) generation paradigm. We propose a multi-granularity knowledge acquisition mechanism integrating sparse, dense, and hybrid retrieval, coupled with text structuralization, taxonomy construction, knowledge embedding, and prompt-driven reasoning to enable efficient external knowledge retrieval, semantic alignment, and controllable integration. Crucially, we deeply embed structured modeling into the augmentation pipeline, enhancing factual accuracy, temporal freshness, and domain-specific competence of generated outputs. Our contributions include: (1) a unified methodological framework for RAS generation; (2) principled pathways toward multimodal, cross-lingual, and interactive augmented generation; and (3) empirically validated improvements in reliability and specialization across diverse domains. This work establishes foundational design principles and future research directions for next-generation RAS systems.
This study systematically investigates the impact mechanisms of individual components in Retrieval-Augmented Generation (RAG) systems on complex question answering and cross-domain tasks. Addressing key challenges—including low retrieval precision, weak contextual relevance, and poor multilingual adaptability—we propose three core innovations: (1) a Contrastive In-Context Learning (CICL) RAG paradigm to improve generation accuracy; (2) sentence-granularity focused retrieval (“Focus Mode”) combined with multi-granularity chunking to enhance retrieval relevance; and (3) a multilingual knowledge base integration framework that balances retrieval–generation efficiency. Through large-scale hyperparameter analysis, we quantitatively characterize the influence of critical factors—including language model scale, chunk size, and retrieval stride—on end-to-end performance. The findings yield a reproducible best-practice guideline for RAG system design and deployment, accompanied by open-sourced, fully implemented code.
The impact of key design decisions—RAG activation, retrieval granularity, and knowledge integration strategy—on RAG system performance remains poorly understood. Method: We conduct systematic ablation studies across three code/qa benchmarks and two state-of-the-art LLMs, quantitatively evaluating how document type, recall rate, document selection strategy, and prompt engineering jointly affect answer correctness and confidence via multi-dimensional analysis, cross-model/dataset comparison, and joint prompt-retrieval analysis. Contribution/Results: We identify precise interaction patterns and operational boundaries among these factors and propose nine actionable, empirically grounded guidelines for diagnosing and optimizing RAG failures. Our findings significantly improve RAG system stability, debuggability, and reliability, offering rigorous empirical evidence and a principled methodology to support the engineering deployment of LLM-augmented systems.
Existing RAG research suffers from the absence of a unified, lightweight, and scalable standardized framework, hindering efficient method reproduction, comparative analysis, and evaluation. To address this, we propose RAGFlow: an open-source, modular, and highly customizable RAG research toolkit. Its key contributions are threefold: (1) a novel fine-grained modular architecture that decouples core components—including retrieval, re-ranking, prompt engineering, and evaluation—enabling flexible composition and independent optimization; (2) integrated support for 16 state-of-the-art RAG methods and 38 standardized benchmarks spanning textual and multimodal scenarios; and (3) a unified interface abstraction for LLMs and multimodal LLMs (MLLMs), coupled with efficient preprocessing and evaluation pipelines. Implemented in Python, RAGFlow significantly lowers the barrier to algorithmic experimentation and has been widely adopted for validating novel RAG methods and supporting pedagogical practice.
RAG system performance critically depends on the retriever-reader configuration, retrieval depth, and context quality; suboptimal settings often degrade performance rather than improve it. Method: We propose RAGGED, a systematic evaluation framework that—through multidimensional controlled experiments, controlled noise injection, and response attribution analysis—quantifies language models’ sensitivity spectra to contextual signals versus noise, and establishes a behavior-driven diagnostic paradigm for RAG configuration. Contribution/Results: We identify two canonical performance patterns—monotonic improvement and inverted-U—as context quality varies. Crucially, we reveal fundamental disparities across models in noise robustness and signal utilization capacity. Based on these insights, we distill reusable, model-aware configuration principles, validated across multiple DBQA benchmarks for both effectiveness and generalizability.
To address key bottlenecks in multimodal RAG systems—including inaccurate user intent understanding, monolithic retrieval strategies, and weak inappropriate response filtering—this paper proposes an end-to-end, multi-stage framework. First, it introduces an image-context-enhanced intent refinement module to improve query semantic accuracy. Second, it designs an intent-driven, context-aware query generation mechanism coupled with heterogeneous API collaborative retrieval. Third, it pioneers a dynamic, organization-aware three-tier joint filtering mechanism—operating over images, text, and multimodal representations—to enable fine-grained relevance and safety control. The method integrates multimodal large language models (MLLMs), cross-modal classifiers, policy-aware dynamic filtering, and API-integrated architecture. Extensive evaluation on multiple public benchmarks—including knowledge-intensive visual question answering (VQA) and safety-oriented datasets—as well as real-world data demonstrates consistent superiority over state-of-the-art methods, setting new records across several key metrics.
This study investigates how the relevance and utility of retrieved documents in Retrieval-Augmented Generation (RAG) influence the internal representations and generation behavior of large language models (LLMs). Through controlled experiments, hidden state analyses, and cross-dataset, multi-model comparisons, the authors systematically evaluate representation shifts across model layers under single- and multi-document settings on four question-answering benchmarks and three LLMs. They provide the first internal-representation-level evidence that the relevance of retrieval context and its interaction with layer depth critically shape the model’s information integration mechanisms. The findings demonstrate a strong correlation between representational changes and generation quality, offering a theoretical foundation for understanding RAG output behavior and informing the design of more effective RAG systems.
This study addresses the confounding effects of metadata, structured representations, and retrieval mechanisms in current RAG systems, which often combine multiple context-augmentation strategies, obscuring their individual contributions to answer quality. Through controlled experiments across six benchmarks, four models, and five augmentation levels—totaling over 24,000 evaluations—the work reveals that increased contextual richness does not necessarily improve accuracy. It introduces the “tractability hierarchy” theory, emphasizing that context must align with model capacity. The findings demonstrate that most augmentation strategies actually degrade performance; however, when metadata and retrieval strategies are carefully matched to a model’s capabilities, smaller models can outperform state-of-the-art large models by up to 19 F1 points on specific tasks, challenging the prevailing RAG design paradigm centered on stacking metadata.
This study investigates the capacity of small language models (7B parameters or fewer) to effectively leverage external information in retrieval-augmented generation (RAG). Through systematic evaluation on models such as SmolLM2, Qwen2.5, and Llama 3.1—combined with BM25, E5-large-v2, and oracle retrievers across multiple prompt templates—the work introduces a novel parameterized knowledge partitioning framework that cleanly disentangles retrieval failure from context utilization failure for the first time. The findings reveal a fundamental bottleneck in small models’ ability to use retrieved content: even under oracle retrieval conditions, 85%–100% of samples fail to correctly incorporate the relevant answer, and 42%–100% of the model’s original knowledge is disrupted by the retrieved context. The dominant error mode is generation entirely unrelated to the provided context, indicating a pervasive inability to attend to or integrate external information.
This study addresses the limited reasoning capabilities of small language models (SLMs) in multi-hop question answering by systematically evaluating 24 prompt templates on the HotpotQA dataset. The evaluation encompasses standard RAG prompts, nine existing prompting strategies, and fourteen newly designed hybrid prompts, with experiments conducted on Qwen2.5-3B and Gemma3-4B-It. The work introduces the first efficient hybrid prompting template tailored for SLMs, significantly enhancing multi-hop reasoning performance under resource-constrained conditions. On a test set of 18,720 samples, the proposed approach achieves up to 83% and 84.5% relative accuracy improvements over standard RAG prompts for the two models, respectively, corresponding to an absolute accuracy gain of up to 6%. The paper also provides a reproducible guideline for effective prompt design in SLM-based multi-hop QA systems.