🤖 AI Summary
This work addresses the limitations of current clinical natural language processing evaluations, which predominantly rely on multiple-choice question answering and assess only answer correctness without probing whether models reason based on relevant, existent, and consistent evidence. To bridge this gap, the authors introduce MIRA-Ev, a multilingual argument mining benchmark tailored to clinical examinations. Built upon Spanish MIR exam cases and annotated by clinical experts at the sentence level for premises, claims, and their support or attack relations, MIRA-Ev is released in Spanish, English, and Basque. It presents the first fine-grained annotation of clinical argument structures and establishes a three-tier evaluation framework encompassing evidence retrieval, claim extraction, and relation classification. Notably, it also delivers the first Basque-language clinical argumentation resource, enabling nuanced assessment of models’ evidence utilization and relational reasoning capabilities.
📝 Abstract
Clinical NLP evaluation remains dominated by multiple-choice question answering (MCQA), which scores only final-answer accuracy and cannot detect when a model reaches the correct diagnosis while grounding it in irrelevant, absent, or contradictory evidence. We introduce MIRA-Ev, a clinical argument mining benchmark built on Spanish Médico Interno Residente (MIR) licensing-exam cases, re-annotated by expert clinicians with span-level premises, claims, and directed support/attack relations, and released in parallel Spanish (native), English, and Basque versions, the first clinical argumentation resource in Basque. MIRA-Ev organizes evaluation into a three-tier task hierarchy: evidence sentence retrieval, argumentative component extraction, and relation classification.