Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research

📅 2026-04-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the tendency of general-purpose large language models to propagate misinformation and exhibit insufficient epistemic humility in policy and development research. To counter this, the authors develop AVA, a multilingual generative AI platform grounded in over 4,000 World Bank reports. AVA employs a multi-agent architecture with retrieval-augmented generation and integrates mechanisms for citation traceability, page-level anchoring, and principled refusal to answer, embodying an “ecosystem-aware” approach to humble AI. In a field evaluation involving over 2,200 users across 116 countries, consistent users saved 2.4–3.9 hours per week. Users regarded AVA as a trustworthy “evidence engine,” with trust established through authoritative provenance and precise citation practices.

Technology Category

Philosophy and Ethics of AI: Safety, Robustness & TrustworthinessHumans and AI: AI for AccessibilityNatural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Economics, Online Markets and Human Computation: Trust and reliance of crowd workers and data experts on GenAISemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Agentic search
📝 Abstract
General-purpose LLMs pose misinformation risks for development and policy experts, lacking epistemic humility for verifiable outputs. We present AVA (AI + Verified Analysis), a GenAI platform built on a curated library of over 4,000 World Bank Reports with multilingual capabilities. AVA's multi-agent pipeline enables users to query and receive evidence-based syntheses. It operationalizes epistemic humility through two mechanisms: citation verifiability (tracing claims to sources) and reasoned abstention (declining unsupported queries with justification and redirection). We conducted an in-the-wild evaluation with over 2,200 individuals from heterogeneous organisations and roles in 116 countries, via log analysis, surveys, and 20 interviews. Difference-in-Differences estimates associate sustained engagement with 2.4-3.9 hours saved weekly. Qualitatively, participants used AVA as a specialized "evidence engine"; reasoned abstention clarified scope boundaries, and trust was calibrated through institutional provenance and page-anchored citations. We contribute design guidelines for specialized AI and articulate a vision for "ecosystem-aware" Humble AI.
Problem

Research questions and friction points this paper is trying to address.

misinformation
epistemic humility
generative AI
policy research
development research
Innovation

Methods, ideas, or system contributions that make the work stand out.

epistemic humility
verified generative AI
citation verifiability
reasoned abstention
evidence-based synthesis
🔎 Similar Papers
No similar papers found.