🤖 AI Summary
This study addresses the accessibility and interaction inefficiencies faced by both visually impaired and sighted scientists when querying multimodal large language models—such as ChatGPT and Gemini—about scientific papers containing figures and charts. Through in-depth interviews with ten scientists across diverse STEM disciplines, the research systematically compares their behaviors and core challenges in conducting question-answering tasks on multimodal scientific literature within authentic research contexts. The findings reveal critical accessibility shortcomings in general-purpose AI tools, including vague image descriptions and inaccurate responses. Building on these insights, the work articulates cross-visual-ability and cross-disciplinary design principles and contributes an open-source dataset comprising 115 domain-specific queries and corresponding model responses, thereby providing empirical grounding and practical resources to advance inclusive interaction with scientific literature.
📝 Abstract
Visual diagrams, figures, and tables are central to scientific papers, and convey information beyond what is captured in text. While blind or low-vision (BLV) scientists have traditionally relied on static alternative text to access figures in papers, the rise of artificial intelligence (AI) has made interactive question-answering (QA) a feasible paradigm for visual exploration; yet little is known about how scientists use visual QA in practice or how to improve its accessibility. In this work, we interview five BLV and five sighted scientists across different STEM fields to understand how they use two AI tools, ChatGPT and Gemini, to query multimodal scientific documents. Our findings characterize how scientists review multimodal content, including existing practices (along with accessibility workarounds) for engaging with visuals, and feedback on the suitability of AI-generated responses to multimodal queries. We further find that vague or incomplete image descriptions, as well as incorrect AI outputs more broadly, can cause both BLV and sighted scientists to abandon AI workflows. To support future research, we additionally contribute a dataset of 115 queries and responses from our participants' interactions with the AI tools for papers in their field. We close by discussing implications for AI-powered scientific QA systems, emphasizing considerations for access across abilities and domains.