Answering clinicians' questions over trial evidence tables with verifiable, feedback-driven language models

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of querying unrecorded attributes, such as drug targets, that cannot be directly retrieved from clinical databases. To this end, it proposes FD-SCoPE, a framework that leverages large language models to generate structured queries and inference rules, combined with code-based verification to answer open-ended queries and support trial selection tasks. Furthermore, this work introduces the first closed-loop optimization mechanism integrating executable queries, programmatic verification, and expert feedback, thereby enabling fully auditable access to clinical evidence. In experiments conducted on oncology evidence tables, the framework achieves an F1 score of 77.7% for derived values, which improves to 84.9% following optimization via reinforcement learning from human feedback, significantly outperforming baseline methods.
📝 Abstract
Systematic reviews condense clinical trials into evidence tables, yet clinicians can interrogate these tables only through database queries, and many questions concern attributes that the table does not record, such as a drug's target class or a harmonised endpoint. Here we introduce FD-SCoPE, a language-model framework that answers both kinds of question, exposes the query, the selected trials and the derivation rule behind every answer, and learns from expert corrections. On an oncology evidence table of 159 immune checkpoint inhibitor trial records, FD-SCoPE completed all 140 clinician-style tasks (alternatives, 90.7-97.9%). For questions needing derived attributes it retrieved 99.3% of relevant trial records at a positive predictive value of 89.8% and outperformed four alternative approaches (derived-value F1 77.7% versus 64.8-73.4%). Corrections on 299 questions, simulated from reference answers, raised F1 on 1,201 unseen questions from 77.9% to 84.9%. Language models coupled with executable queries, verified programs and expert feedback can give clinicians auditable access to trial evidence.
Problem

Research questions and friction points this paper is trying to address.

clinical evidence tables
systematic reviews
derived attributes
question answering
verifiable language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Language Models
Evidence Tables
Verifiable Reasoning
Feedback-driven Learning
Auditable AI
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Manan Roy Choudhury
Manan Roy Choudhury
Ph.D. Student, Arizona State University
Natural Language ProcessingAnomaly DetectionGenAILLM Analysis and Evaluation
S
Suparno Roy Chowdhury
Arizona State University, Tempe, Arizona, USA
S
Swastik Sahoo
Arizona State University, Tempe, Arizona, USA
M
Muhammad Ali Khan
Mayo Clinic, Phoenix, Arizona, USA
K
Kaneez Zahra Rubab Khakwani
Mayo Clinic, Phoenix, Arizona, USA
M
Mohamad Bassam Sonbol
Mayo Clinic, Phoenix, Arizona, USA
Irbaz Bin Riaz
Irbaz Bin Riaz
Mayo Clinic
Vivek Gupta
Vivek Gupta
Assistant Professor of Computer Science, Arizona State University
Artificial IntelligenceNatural Language ProcessingLarge Language ModelsInformation Retrieval