🤖 AI Summary
This study addresses the high costs and information overload arising from indiscriminate information acquisition in multimodal medical diagnosis by proposing a cost-aware framework. Built upon a frozen vision-language model with probability distribution tracking, the method employs trajectory-level offline reinforcement learning to train a lightweight MLP controller. By incorporating clinical stepwise escalation logic, the controller intelligently determines when to acquire additional evidence and when to submit a diagnosis. Crucially, this work identifies and mitigates the "information overload effect," demonstrating that learning to disregard redundant information is as essential as effective acquisition. Experiments across three benchmarks show that the proposed framework consistently outperforms both unguided and full-modality baselines, achieving high diagnostic accuracy while utilizing fewer modality channels.
📝 Abstract
Clinical diagnosis is inherently sequential: clinicians escalate from cheap to costly tests only when additional evidence is expected to resolve diagnostic uncertainty. We present ActiveMedAgent, a framework that brings this cost-aware sequential logic to multimodal medical AI. Given a frozen, API-accessed vision-language model, ActiveMedAgent tracks probability distributions over candidate diagnoses and scores each acquisition by its per-step diagnostic utility minus cost. A lightweight MLP controller is then trained offline on these scored trajectories, learning when to request additional evidence and when to commit. Across three commonly used benchmarks, trajectory-based policy learning consistently outperforms both unguided acquisition and full-modality baselines. Notably, we identify an information overload effect. In 175 cases, the agent produces a correct diagnosis with fewer channels while the full-modality baseline fails, showing that learning what to omit can be as important as learning what to acquire.