🤖 AI Summary
This study addresses the lack of theoretical guidance for budget allocation among questions, trajectories, and repeated reads in agentic RAG evaluation. By quantifying precision-cost frontiers across sampling dimensions using datasets such as HotpotQA, and integrating retrieval feedback comparisons, nested predictive models, and generalization-theoretic analysis, this work reveals that expanding question coverage is statistically more efficient than sampling multiple trajectories or conducting repeated reads. Furthermore, a zero-temperature strategy is proposed to mitigate answer divergence. The primary contribution lies in establishing best practices that prioritize increasing the number of questions under constrained search budgets, which reduces standard error by 33% while maintaining prediction bias below 4%.
📝 Abstract
Evaluation budgets in agentic retrieval-augmented generation span questions, search trajectories, and repeated answers. We measure allocation precision, reading efficiency, and cost boundaries using a retrieval-feedback comparison on HotpotQA and MuSiQue. At 34.14--34.39M model tokens, broader question coverage lowers standard error by 33\% versus five reads and 12.6\% versus three trajectories. Archived nested and Q-only forecasts predict these allocations within 4.0\% and 3.5\%, respectively. Depth subsets establish no clear forecasting advantage beyond the two-trajectory audit. One-read variance penalties relative to the fitted optimum at the same token budget are 0--9.9\%, with substantial Pro uncertainty. Under recorded model fees, more questions beat more trajectories at search prices of \$0--1 per 1,000 requests; question-versus-read fee rankings remain unresolved. Temperature zero cuts answer disagreement from 14.3\% to 3.4\% while comparison precision stays similar. \par\medskip\noindent\textbf{Keywords:} Agentic RAG; Evaluation budget; Generalizability theory; Repeated sampling.