🤖 AI Summary
This study addresses the disconnect between visual recognition and clinical reasoning, as well as evidence acquisition bottlenecks, in video-based diagnostics. Leveraging 71 neurological consultation cases, we employ vision-language models to parse the “video–hypothesis–questioning–examination” reasoning chain. Trajectory replay analysis reveals that video-derived gains stem from investigative cues rather than mere visual recognition. Accordingly, we propose a novel paradigm for examination selection that integrates optimized descriptions with literature retrieval. Results demonstrate that video input improves diagnostic accuracy by 9.9%–22.5%, reaching 73.2%–93.0% when decisive evidence is available, thereby substantially narrowing the decision-making gap between computational models and clinicians.
📝 Abstract
Diagnosing a patient from video requires more than recognizing the sign: a vision-language model must turn what it sees into hypotheses, questions and tests. DynamicDx evaluates each step in 71 neurological consultations across 11 sign categories, linking authentic patient videos to confirmed diagnoses and fixed charts built from the same case reports, so that every model queries the same evidence. Across five such models, video improves accuracy by 9.9-22.5 percentage points over blind input, but neither recognition alone nor temporal order explains the gain: the cause is usually missing from the model's video-only differential diagnosis even when the sign is recognized, and shuffling the frames produces no reliable accuracy loss. Instead, a trajectory replay traces most of the gain to the investigation results the video prompts. Evidence acquisition is the bottleneck: supplying the decisive investigations raises accuracy to 73.2-93.0%. Two interventions act on it. A post-trained 4B video describer improves sign descriptions, especially from a short, densely sampled segment, and source-clean literature retrieval expands initial hypotheses; both bring the tests a model orders closer to those the treating clinicians documented and, through them, raise accuracy. For video-based diagnosis, seeing better helps when it leads to asking better.