From Given to Gathered Evidence: Agentic Learning for Longitudinal Medical Reasoning

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inability of foundation models to actively retrieve longitudinal multimodal clinical evidence by proposing the CASE agent framework alongside a longitudinal multimodal benchmark derived from the UK Biobank. Methodologically, the framework employs a vision-language policy model that achieves autonomous evidence seeking through supervised fine-tuning, tool-sequence-agnostic agent reinforcement learning, privileged self-distillation, and large-model feedback. Experimental results demonstrate that the CASE agent, built upon Qwen3-VL-8B, surpasses GPT-5.4 and Claude Opus 4.8 in accuracy by over 16% and 10%, respectively. These findings validate the substantial advantages of the proposed framework for clinical evidence reasoning.
📝 Abstract
Foundation models can serve as clinical agents through tool-use harnesses. However, conventional medical benchmarks assess reasoning over preselected evidence rather than the ability to seek it across clinical records and longitudinal imaging. We propose CASE: a series of role-specific Clinical Agents for Seeking Evidence, together with a tool-use harness and an agentic post-training framework for compact vision-language policy models. We further introduce a longitudinal multimodal benchmark built on UK Biobank, comprising 50,401 clinical questions derived from real-world ICD-10-coded diagnoses of 4,739 participants. Each question links to a patient-specific environment containing clinical context and multi-sequence MRI from baseline and follow-up visits, where agents autonomously select which visits, organs, modalities, slices, and specialist tools to inspect and compare. Supervised fine-tuning transfers evidence-seeking workflows from 14,734 frontier-model interaction trajectories, followed by agentic reinforcement learning on the learner's own environment interactions. Privileged on-policy self-distillation and rubric-based LLM feedback refine evidence-to-conclusion reasoning without prescribing tool sequences. Experiments show that CASE moves beyond question-answer imitation toward transferable investigation policies, strengthening evidence-grounded longitudinal reasoning. Under matched evaluation conditions, our Qwen3-VL-8B based agent achieves over 16% and 10% relative improvements in answer accuracy over GPT-5.4 and Claude Opus 4.8. Code will be available at https://github.com/VinyehShaw/CASE.
Problem

Research questions and friction points this paper is trying to address.

longitudinal medical reasoning
evidence seeking
clinical agents
multimodal benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Learning
Longitudinal Medical Reasoning
Vision-Language Policy Models
Tool-use Harness
Reinforcement Learning
🔎 Similar Papers
2024-05-27International Conference on Information and Knowledge ManagementCitations: 4