🤖 AI Summary
This study addresses the challenges of missing external knowledge, high inference costs, and untraceable evidence in discharge risk prediction from electronic health records by proposing a budget-aware large language model framework. The framework performs constrained reasoning and evidence citation over medical knowledge graphs via a "Plan–Navigate–Verify" loop. It introduces quality-annotated evidence graphs, designs a budget-constrained iterative verification mechanism, and develops a reinforcement learning strategy incorporating citation integrity. Experimental results demonstrate that the proposed method improves AUPRC by 3.4 points and achieves a citation precision of 77.9% while consuming only 62%–65% of the inference budget, thereby realizing synergistic optimization across predictive performance, interpretability, and computational efficiency.
📝 Abstract
Post-discharge risk prediction from electronic health records (EHRs) is difficult because many dependencies that link discharge-time observations to downstream complications, such as comorbidity cascades and drug-disease interactions, are absent from the record. External medical knowledge graphs (KGs) can supply these missing dependencies, but tracing them demands three properties: KG exploration must remain cost-bounded, retrieved evidence must be differentiated by source quality, and the resulting rationale must be citable for retrospective review. Large language models (LLMs) can plan and verify over structured evidence, making them natural candidates for KG reasoning, but existing LLM-based methods do not satisfy these three properties jointly. In this paper, we propose BAR, a Budget-Aware LLM Reasoning framework over medical KGs with three contributions. First, BAR refines the raw KG into disease-specific evidence graphs whose edges carry support scores and provenance records, turning the KG into a quality-annotated reasoning space rather than a static feature source. Second, an LLM then reasons over this graph through a plan-navigate-verify loop that decomposes the question into steps, retrieves evidence under a patient-specific budget, and revises when verification fails. Third, a reasoning policy is trained with a reward that compares predictions with and without acquired evidence, combined with acquisition cost and citation-integrity terms. Across 8 diseases and 3 prediction horizons on MIMIC-III and MIMIC-IV, BAR improves AUPRC by 3.4 points over the strongest baseline, raises citation precision from 59.8% to 77.9%, and consumes only 62-65% of the budget cap.