🤖 AI Summary
This study addresses the challenge of aligning visual evidence with clinical guidelines in Tactical Combat Casualty Care (TCCC) by proposing TC3-VQA, the first doctrine-based visual question answering dataset. Methodologically, the authors employ visual annotation, passage retrieval, textual entailment verification, and multi-model validation to precisely map video observations to authoritative medical documents, thereby enabling traceable supervision of clinical reasoning. The resulting high-quality dataset comprises 1,860 question-answer pairs and has undergone rigorous review by domain experts to ensure accuracy. By establishing a robust connection between visual data and doctrinal medical knowledge, this work provides a critical benchmark resource for trustworthy visual question answering in tactical casualty care scenarios.
📝 Abstract
Tactical Combat Casualty Care (TC3) requires responders to connect visual observations of injuries and interventions with established clinical guidance. Developing vision-language models to support this process requires supervision that links visible evidence to traceable doctrine. We present TC3-VQA, a dataset constructed from public instructional and field TC3 videos and authoritative TC3 documents. It contains 581 items spanning 11 concepts, with 1,860 questions covering intervention recognition, doctrine, clinical reasoning, procedural guidance, and refusal when visual information is insufficient. Doctrine-based answers preserve verbatim source passages and character offsets. Construction combines visual annotation, passage retrieval, entailment checks, and verification across model families. Equipment boxes, anatomical labels, temporal segments, and source metadata accompany the question-answer pairs. Automated audits and ratings by two physicians and two medical students characterize annotation quality, with human ratings available for 88 retained items. The dataset provides a resource for adapting vision-language models to TC3, studying the connection between visual evidence and clinical knowledge, and evaluating recognition, doctrine recall, and abstention.