ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of extracting critical structured information from scientific images in atomic layer deposition and etching (ALD/E) research, which is often inaccessible through text alone. To this end, we introduce Sci-ImageMiner—the first multi-task benchmark specifically designed for ALD/E scientific image understanding—encompassing four end-to-end tasks: image classification, captioning, data extraction, and visual question answering. Built upon expert annotations and validated through a community challenge that attracted 68 participating teams submitting 1,263 results, the benchmark reveals that while state-of-the-art multimodal models perform well on classification and captioning, they exhibit significant limitations in data extraction and scientific reasoning. These findings underscore the necessity of integrating visual perception with domain-specific knowledge to advance scientific image understanding.
📝 Abstract
Scientific figure comprehension and reasoning using multimodal AI requires integrating visual perception with domain-specific reasoning to extract meaningful knowledge, often not presented in the text of a research publication. The Sci-ImageMiner benchmark dataset, accompanied by a community-driven competition, raises the bar over prior scientific competitions by curating a comprehensive, expert-annotated dataset across four end-to-end complementary tasks. The competition attracted 68 active participants and 1,263 public/private submissions from 9th January 2026 to 8th April 2026. Our results show that state-of-the-art multimodal models perform well on classification and summarization tasks but struggle with data extraction and scientific reasoning, particularly in visual question-answering. These findings reveal key limitations and highlight challenges and opportunities for improving domain-aware multimodal AI systems. Overall, the Sci-ImageMiner benchmark and competition establish a rigorous platform for advancing research in scientific figure comprehension and reasoning and demonstrate the potential of state-of-the-art approaches for a challenging and complex research area.
Problem

Research questions and friction points this paper is trying to address.

scientific figure comprehension
information extraction
multimodal AI
visual question answering
domain-specific reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

multimodal AI
scientific figure comprehension
domain-specific reasoning
benchmark dataset
visual question answering
🔎 Similar Papers
2024-06-08Annual Meeting of the Association for Computational LinguisticsCitations: 2