A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench

📅 2026-08-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Although general-purpose large language models demonstrate strong performance on medical benchmarks, they often fail to align with the clinical realities of low- and middle-income countries (LMICs). This work proposes VITA, a retrieval-augmented generation (RAG) system tailored for LMIC settings such as India, which integrates locally relevant resources—including national disease guidelines, antimicrobial resistance data, the national formulary, and resource-constrained care protocols—into a domain-specific corpus. Experimental results show that VITA achieves a score of 51.9% on HealthBench, outperforming mainstream models, and performs comparably to GPT-5.5 in neutral evaluations while surpassing it in both weighted scoring and the number of questions answered more accurately. These findings underscore the critical role of corpus specificity in enhancing the accuracy and clinical applicability of AI-generated medical responses in resource-limited contexts.
📝 Abstract
General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed largely in high-income settings. We evaluate VITA, a retrieval-augmented generation (RAG) system purpose-built for contextual knowledge retrieval in India and other low- and middle-income (LMIC) settings. VITA retrieves from a curated corpus of disease-specific guidelines, India-specific antimicrobial resistance data, national formulary constraints, and resource-limited care protocols; its architecture and corpus are proprietary, but the benchmark, the physician-written rubrics, and our full response and scoring outputs are public for independent verification. On 4,023 English-language HealthBench questions (80.5% of the benchmark), scored with a GPT-4.1 judge, VITA ranked first with 51.9% of possible rubric points, ahead of GPT-5.4 (46.1%), o4-mini (44.3%), Gemini 3.1 Pro (42.6%), and Claude Sonnet 4.6 (37.3%), and scored highest on 45.4% of questions. To test robustness to newer models and judge lineage, a 500-question subset was re-run against current-generation models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Pro, Grok 4.3) and graded by a neutral open-weight judge (DeepSeek-V4-Pro) sharing no lineage with any system tested. Here the gap narrowed to parity: VITA and GPT-5.5 were statistically indistinguishable on mean per-question score, while VITA led on points-weighted score and won the most questions. VITA's advantages in accuracy and completeness persisted under the neutral judge; its communication scores were lower. These results indicate that a purpose-built clinical RAG system remains competitive with frontier LLMs on an open benchmark, consistent with corpus specificity as a design variable that improves grounding at some cost to communication polish.
Problem

Research questions and friction points this paper is trying to address.

clinical AI
low- and middle-income countries
medical benchmarking
retrieval-augmented generation
corpus specificity
Innovation

Methods, ideas, or system contributions that make the work stand out.

retrieval-augmented generation
corpus-specific RAG
low- and middle-income countries
clinical AI benchmarking
HealthBench
🔎 Similar Papers
2024-05-27International Conference on Information and Knowledge ManagementCitations: 4
💼 Related Jobs
No related jobs found.
P
Praveen Reddy
C
Charuta Mandke
S
Suvrankar Datta
S
Sarah Khan
S
Siddharth Reddy Anthireddy
S
Shitij Arora
Vishal Singh
Vishal Singh
Stern School of Business, NYU