Agentic discovery of blood biomarker from distilled private health records

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that privacy constraints impede medical data sharing, thereby limiting the application of advanced language models in biomarker discovery. To overcome this, we propose a “distillation scoring” mechanism that compresses 5.4 million patient records via graph attention networks into a scoring tool for evaluating complete blood count (CBC) expression performance. By releasing only the trained weights rather than raw data, this approach enables external large language model agents to perform closed-loop iterative optimization under strict privacy protection. External validation demonstrates that expressions identified by these agents achieve a median AUC improvement of 4.18 percentage points over literature baselines and outperform existing preferred tools across three independent cohorts. This work effectively unifies privacy preservation with efficient biomarker discovery.
📝 Abstract
Routine complete blood counts (CBCs) could yield new biomarkers, but the private records needed to evaluate candidates cannot be shared with frontier language model agents that excel at discovery. We distilled the evidence held in the Clalit Health Services panel of over 5.4 million patients into a released scoring tool: for each of 13 immune-mediated diseases, a graph attention network was trained inside the data boundary to predict the case-control AUC of candidate CBC expressions, and only the trained weights were released. The tool grounds an agent's propose-score-refine loop in real-world data without exposing any patient data. In external validation, agent-discovered expressions improved on their literature-seeded starting points by a median of 4.18 AUC percentage points, and across three independent cohorts, reranking the candidates of three frontier research tools improved on their first choices in most comparisons, with gains that varied by cohort. The released scorer supports privacy-preserving biomarker hypothesis generation.
Problem

Research questions and friction points this paper is trying to address.

blood biomarker discovery
privacy-preserving
health records
complete blood count
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic discovery
Privacy-preserving distillation
Graph attention network
Blood biomarkers
Propose-score-refine loop
🔎 Similar Papers
No similar papers found.