GraphRareBench: An Auditable Graph-Evidence Benchmark for Phenotype-Driven Rare-Disease Diagnosis

📅 2026-07-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of transparent evaluation in existing rare disease diagnosis benchmarks regarding model decision rationales and the ranking of confounding diseases. To this end, the authors introduce an auditable graph-based evidence benchmark comprising 2,365 ontology-generated cases and 18,093 target–confounder pairs, which— for the first time—incorporates hard confounders that preserve provenance information and an observable evidence mechanism. This framework enables multidimensional assessment of models’ capabilities in evidence retrieval, discrimination among confounding diseases, and evidence utilization. Leveraging ontology-driven case generation, graph-structured confounder modeling, and a 21-dimensional feature-supervised ranker within an agent-based reasoning framework (Agents-A1, DeepSeek-V4-Flash), the approach achieves MRR scores of 0.640–0.746 and target-over-confounder accuracy of 0.898–0.916 on a 237-case gene-component disentanglement test set, revealing that current models remain significantly vulnerable to diagnostic confusion.
📝 Abstract
Phenotype-driven diagnostic benchmarks usually report the rank of the reference disease, but they rarely reveal which plausible alternatives are ranked above it or what evidence a tool-using model examines before making its decision. We introduce GraphRareBench, a provenance-preserving benchmark containing 2,365 ontology-derived cases and 18,093 target-confounder pairs. Each case includes a coarsened HPO query, a fixed candidate pool, graph-defined hard confounders, and source-linked evidence records. On the 237-case gene-component-disjoint test split, supervised rankers using a shared 21-feature interface achieved MRRs ranging from 0.640 to 0.740 and case-averaged target-over-confounder accuracies ranging from 0.898 to 0.916. Agents instantiated with Agents-A1 and DeepSeek-V4-Flash achieved MRRs of 0.746 and 0.718, respectively. Their paired MRR difference was not statistically significant, whereas their target-evidence coverage differed by 0.561. Together with the observation that 22.1% to 43.7% of selected Hit@10 successes still ranked at least one graph-defined hard confounder above the target, these results indicate that full-pool retrieval, hard-confounder discrimination, and observable evidence access capture complementary aspects of model behavior. GraphRareBench therefore provides a foundation for more transparent and evidence-aware evaluation of phenotype-driven diagnostic systems. Code and data are available at https://github.com/GUI0609/GraphRareBench.
Problem

Research questions and friction points this paper is trying to address.

rare disease diagnosis
phenotype-driven benchmark
auditable evidence
hard confounders
graph-based evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

graph-based benchmark
provenance-preserving evaluation
hard confounder discrimination
evidence-aware diagnosis
phenotype-driven rare disease