Fathom-Vaidya: Advancing Medical Reasoning with Rubric-Based Rewards

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文针对医疗领域大语言模型在诊断推理和临床交互中的不足,通过合成数据和基于评分标准的强化学习方法提升了模型在这两方面的能力。
📝 Abstract
Deploying Large Language Models (LLMs) in healthcare requires robust performance across two complementary dimensions - diagnostic reasoning: the convergent, evidence-driven task of inferring a patient's condition from clinical data to produce a diagnosis, and clinical healthcare reasoning: the broader, navigational judgment required to communicate, plan, and adapt across multi-turn clinical interactions where a single correct answer may not exist. Recent benchmarks such as HealthBench and MedXpertQA reveal persistent weaknesses in both areas, exposing failures in complex diagnostic scenarios and limitations in contextual, patient-centered dialogue. We introduce a sequential training framework that targets these facets using synthetic data and rubric-based reinforcement learning. First, we improve diagnostic reasoning using MedBullets-derived questions with rule- and rubric-guided Reinforcement Learning (RL). We then shift to clinical reasoning by generating 5.3k synthetic multi-turn scenarios, each paired with multi-dimensional rubrics to comprehensively assess the response. This approach yields over 10% improvement on MedXpertQA, and our 30B model achieves 50.1% accuracy on HealthBench-Hard, surpassing proprietary baselines including GPT-5 (thinking). Our results show that targeted synthetic datasets and rubric-based training can systematically improve both diagnostic and interactive clinical reasoning in medical LLMs.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Diagnostic Reasoning
Clinical Healthcare Reasoning
HealthBench
MedXpertQA
Innovation

Methods, ideas, or system contributions that make the work stand out.

sequential training framework
synthetic data
rubric-based reinforcement learning
diagnostic reasoning
clinical healthcare reasoning
🔎 Similar Papers
No similar papers found.
K
Kalash Shah
Fractal AI Research
K
Kunal Singh
Fractal AI Research
S
Snehan J
Fractal AI Research
Shreyas Singh
Shreyas Singh
Indian Institute of Technology Madras
Computer VisionDeep LearningComputational Imaging