🤖 AI Summary
This work addresses the tendency of general-purpose large language models to propagate misinformation and exhibit insufficient epistemic humility in policy and development research. To counter this, the authors develop AVA, a multilingual generative AI platform grounded in over 4,000 World Bank reports. AVA employs a multi-agent architecture with retrieval-augmented generation and integrates mechanisms for citation traceability, page-level anchoring, and principled refusal to answer, embodying an “ecosystem-aware” approach to humble AI. In a field evaluation involving over 2,200 users across 116 countries, consistent users saved 2.4–3.9 hours per week. Users regarded AVA as a trustworthy “evidence engine,” with trust established through authoritative provenance and precise citation practices.
📝 Abstract
General-purpose LLMs pose misinformation risks for development and policy experts, lacking epistemic humility for verifiable outputs. We present AVA (AI + Verified Analysis), a GenAI platform built on a curated library of over 4,000 World Bank Reports with multilingual capabilities. AVA's multi-agent pipeline enables users to query and receive evidence-based syntheses. It operationalizes epistemic humility through two mechanisms: citation verifiability (tracing claims to sources) and reasoned abstention (declining unsupported queries with justification and redirection). We conducted an in-the-wild evaluation with over 2,200 individuals from heterogeneous organisations and roles in 116 countries, via log analysis, surveys, and 20 interviews. Difference-in-Differences estimates associate sustained engagement with 2.4-3.9 hours saved weekly. Qualitatively, participants used AVA as a specialized "evidence engine"; reasoned abstention clarified scope boundaries, and trust was calibrated through institutional provenance and page-anchored citations. We contribute design guidelines for specialized AI and articulate a vision for "ecosystem-aware" Humble AI.