build cross-chart rag benchmarks

Design and build evaluation datasets, test suites, and scoring protocols that assess retrieval-augmented generation (RAG) systems on tasks requiring reasoning across multiple charts and multi‑hop retrieval steps; this includes curating diverse chart types and query templates, creating ground-truth retrieval targets and answers, and implementing metrics and experiment pipelines that reveal performance gaps across RAG paradigms.

buildcross-chartragbenchmarks

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

This study addresses the lack of systematic, up-to-date synthesis of retrieval-augmented generation (RAG) research amid rapid methodological diversification and evaluation fragmentation. Adopting the PRISMA 2020 framework, we systematically curated 128 highly cited papers (2020–2025) from ACM, IEEE, and other authoritative sources, introducing a dynamic citation threshold to mitigate temporal bias. Our analysis maps evolutionary trajectories across three dimensions: architectural design, benchmark datasets, and evaluation metrics—revealing critical methodological gaps, particularly in non-parametric memory augmentation and neural retrieval–generation co-adaptation. We construct a structured knowledge graph of RAG research and propose a prioritized roadmap that jointly optimizes robustness, interpretability, and generalization. The findings provide empirically grounded guidance for both foundational RAG theory development and practical deployment.

Analyzes retrieval-augmented generation techniques and challengesEvaluates datasets, architectures, and metrics in RAG studiesIdentifies research gaps and future directions for RAG

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the inefficiency and low accuracy of agent-based RAG systems in multi-hop question answering, which stem from redundant retrieval and insufficient context integration. Without modifying the training procedure, the authors propose a test-time optimization strategy built upon the Search-R1 framework. Their approach introduces an LLM-driven contextualization module to enhance the fusion of retrieved information and incorporates a deduplication mechanism that dynamically replaces redundant documents. Evaluated on the HotpotQA and Natural Questions benchmarks, the method achieves a 5.6% absolute improvement in Exact Match (EM) score and reduces the average number of retrieval rounds by 10.5%, substantially enhancing both reasoning accuracy and computational efficiency.

contextualizationmultihop questionsrepetitive retrieval

RAGAs: Automated Evaluation of Retrieval Augmented Generation

Sep 26, 2023
ES
ES Shahul
🏛️ Cardiff University | AMPLYFI

This work addresses the bottleneck in RAG system evaluation—its heavy reliance on human-annotated ground-truth answers—by proposing RAGAs, a reference-free automated evaluation framework. Methodologically, it introduces a computable, three-dimensional metric suite covering retrieval relevance, context faithfulness, and generation quality, integrating BERTScore for semantic similarity, NLI-based models for factual consistency, and self-supervised prompting strategies. Its key contribution is the first end-to-end, multidimensional, reference-free evaluation paradigm, enabling quantitative, pipeline-level diagnostics of RAG systems. Experiments demonstrate strong agreement between automated metrics and human judgments (average Spearman ρ > 0.82) across multiple benchmarks, validating efficacy and robustness. The open-source RAGAs toolkit has been widely adopted in industry for iterative RAG system optimization.

Assessing retrieval and generation module performanceEvaluating RAG pipelines without ground truth annotationsReducing hallucinations in LLM-based knowledge systems

This work addresses the lack of an integrated Retrieval-Augmented Generation (RAG) development and evaluation toolkit in the R programming language, where existing solutions predominantly rely on the Python ecosystem. The authors propose ragR, the first native R framework enabling end-to-end RAG workflows—including document ingestion, vector storage, similarity-based retrieval, evidence synthesis, and structured question-answering logging—and faithfully reimplements the four core RAGAS evaluation metrics: context precision, context recall, faithfulness, and answer relevance. Experimental results demonstrate that ragR yields evaluation outcomes highly consistent with those produced by the original Python-based RAGAS implementation. This provides R users with a lightweight, reproducible, and self-contained environment for RAG research and education without requiring language switching.

LLM-based scoringR languageRAG evaluation

Towards Understanding Retrieval Accuracy and Prompt Quality in RAG Systems

Nov 29, 2024
SZ
Shengming Zhao
🏛️ University of Alberta | The University of Tokyo | East China Normal University

The impact of key design decisions—RAG activation, retrieval granularity, and knowledge integration strategy—on RAG system performance remains poorly understood. Method: We conduct systematic ablation studies across three code/qa benchmarks and two state-of-the-art LLMs, quantitatively evaluating how document type, recall rate, document selection strategy, and prompt engineering jointly affect answer correctness and confidence via multi-dimensional analysis, cross-model/dataset comparison, and joint prompt-retrieval analysis. Contribution/Results: We identify precise interaction patterns and operational boundaries among these factors and propose nine actionable, empirically grounded guidelines for diagnosing and optimizing RAG failures. Our findings significantly improve RAG system stability, debuggability, and reliability, offering rigorous empirical evidence and a principled methodology to support the engineering deployment of LLM-augmented systems.

Analyzing key engineering trade-offs in RAG deployment decisionsDetermining optimal retrieval volume for different task typesEvaluating effective knowledge integration methods across tasks

Latest Papers

What's happening recently
View more

This work addresses the limitations of existing RAG system evaluations, which rely on query-level annotations or reference answers and thus fail to comprehensively assess the behavioral coverage of retrieval components across test suites. To overcome this, we propose Chunk Coverage (CC), a novel metric that—without requiring test oracles—quantifies retrieval coverage by measuring the proportion of corpus chunks retrieved at least once during testing, thereby characterizing coverage at the structural level. Building upon CC, we develop an unsupervised framework for evaluating retrieval behavior and guiding test case selection and prioritization. Empirical results in clinical and financial domains demonstrate that CC-guided testing achieves 50% coverage 1.7× faster than random strategies and 4.2× faster than redundancy-based approaches, while improving fault detection capability (measured by APFD) by 10%–25%.

chunk coverageoracle-independent evaluationretrieval coverage

This work proposes a retrieval-augmented generation (RAG) evaluation framework tailored to deep informational needs, addressing the challenges of reasoning and factual consistency posed by multi-sentence narrative queries in complex real-world scenarios. The framework introduces narrative-style multi-sentence queries and a multi-level evaluation protocol, enabling comprehensive assessment of system performance in terms of completeness, attributability, and consistency over the MS MARCO V2.1 corpus. Aggregating over 150 system submissions, the study demonstrates the feasibility of simultaneously ensuring factual accuracy, transparency, and high-quality generation under complex querying conditions, thereby establishing a new benchmark for trustworthy, context-aware RAG systems.

complex information needsfactual groundingnarrative queries

This study investigates whether retrieval quality can serve as a reliable early indicator of information coverage in responses generated by Retrieval-Augmented Generation (RAG) systems. Through systematic experiments across three benchmarks—TREC NeuCLIR 2024, TREC RAG 2024, and WikiVideo—the authors evaluate 15 text-based and 10 multimodal retrieval systems using the Auto-ARGUE and MiRAGE assessment frameworks. The work provides the first empirical evidence of a strong correlation between coverage-oriented retrieval metrics and the informational coverage of generated outputs. Findings reveal that, at both topic and system levels, such metrics effectively predict RAG output coverage when retrieval and generation objectives are aligned, underscoring the critical role of goal consistency in optimizing RAG performance.

information coveragenugget coverageRAG performance

This work addresses data staleness, tenant leakage, and combinatorial query explosion in production-grade retrieval-augmented generation (RAG) systems caused by decoupled data layers. To resolve these issues, the authors propose a unified data layer architecture built on PostgreSQL that, for the first time, integrates vector retrieval and structured filtering within a single database. By leveraging pgvector with HNSW indexing and a hybrid hierarchical design, the system eliminates cross-system synchronization overhead while guaranteeing strict tenant isolation and strong data consistency. Experimental results on a dataset of 50,000 documents demonstrate a 92% reduction in latency for date-filtered queries and a 74% reduction for tenant-scoped queries, alongside a 93% decrease in synchronization code, achieving zero data inconsistency and enabling efficient, secure RAG.

data stalenessproduction RAG systemsquery composition explosion

Current evaluation methods for retrieval-augmented generation (RAG) systems predominantly focus on single-hop queries, failing to accurately assess retriever performance in multi-hop reasoning scenarios. To address this limitation, this work proposes Context-Aware Retriever Evaluation (CARE), the first framework to systematically incorporate contextual relevance into multi-hop retrieval evaluation. Leveraging RAG simulation environments built on HotPotQA, MuSiQue, and SQuAD, and employing large language models from OpenAI, Meta, and Google as judges, CARE establishes an LLM-as-judge automated evaluation paradigm. Experimental results demonstrate that CARE significantly outperforms existing evaluation approaches on multi-hop queries, with particularly pronounced gains when using large-parameter models with extended context windows, whereas single-hop settings exhibit minimal sensitivity to contextual awareness.

context-aware evaluationLLM-as-judgemulti-hop reasoning