🤖 AI Summary
This study addresses the limitations of existing vision-language models in comprehending the complex multimodal information present in real-world construction drawings and the absence of a domain-specific, multi-level reasoning benchmark for architecture and civil engineering. To bridge this gap, the authors introduce the first evaluation benchmark based on authentic “Issued for Construction” as-built drawings, featuring three tiers of visual-textual reasoning: perceptual understanding, contextual interpretation, and expert-level inference. The benchmark comprises 33 drawings and 92 expert-annotated question-answer pairs. Innovatively, it explicitly maps engineering workflows to AI capability dimensions through a dual-taxonomy framework encompassing seven engineering aspects and four model capability axes. Experimental results demonstrate that current multimodal large language models significantly underperform human experts on high-level reasoning tasks, thereby validating the benchmark’s efficacy and laying the groundwork for domain-specialized multimodal AI development.
📝 Abstract
We introduce DrawingVQA, the first benchmark designed to evaluate multimodal large language models (MLLMs) on real-world construction drawings -- a core media in architecture, civil, and many other engineering practices. Unlike natural images or schematic floor plans, construction drawings fuse abstract geometry, symbolic notation, tabular data, annotations, and domain-specific text, forming a uniquely complex visual-textual domain core to engineering workflows. DrawingVQA bridges this gap with 33 "Issued for Construction" drawings and 92 expertly curated question-answer pairs, spanning three reasoning depths: perceptual understanding, contextual interpretation, and domain-expert reasoning. To evaluate model capabilities, we present a dual categorization framework to jointly analyze performance across seven construction-engineering and four MLLM capability dimensions -- the first to explicitly map engineering workflows to AI reasoning competencies. Evaluations of state-of-the-art MLLMs reveal a substantial gap between model and expert performance, particularly at higher reasoning depths. This benchmark lays a foundation for domain-specialized multimodal reasoning to allow for advancement on integration of AI-driven understanding and real-world engineering workflows.