🤖 AI Summary
This study systematically evaluates the alignment between state-of-the-art vision-language models (VLMs)—including GPT-4o, Claude 3.5 Sonnet, and Gemini 2.0 Flash—and human experts in inferring building attributes (structure type, functional use, and number of stories) from street-view imagery. By comparing model outputs with annotations from civil engineers and architects, and employing Chain-of-Thought prompting alongside keyword probability analysis, the work reveals that VLMs primarily rely on explicit visual cues, whereas human experts integrate contextual information and domain knowledge in their reasoning. Experimental results demonstrate that VLMs achieve an average accuracy of 70% on this task, highlighting their potential for city-scale automated building characterization and utility as expert-assistive tools. This research presents the first systematic comparison of the divergent inference mechanisms underlying VLMs and domain specialists in architectural semantic understanding.
📝 Abstract
This research investigates the potential of Vision-Language Models (VLMs) to infer building typologies: Construction, Current Use, and Storeys from Google Street View (GSV) images. Predictions generated by VLMs are compared with inference by human experts (civil engineers and architects) as a source of manually labelled ground-truth data. We evaluate several state-of-the-art VLMs, including GPT-4o, Claude 3.5 Sonnet, and Gemini 2.0 Flash. By applying different scaling strategies and prompting techniques, we found that Chain-of-Thought prompts provide an overall more stable model performance. We also investigate the reasoning behind VLMs' building-typology predictions by examining the probabilities of keywords appearing in AI explanations. This enabled us to analyse patterns in these reasonings and identify key themes driving both agreements and disagreements between VLM and expert labels. We find that AI tends to focus on visual indicators, whereas human experts place greater emphasis on broader contextual cues and domain knowledge, in addition to visual cues. Overall, VLM can approximate experts' capability in building-typology classification at scale, with an average accuracy of approximately 70%. The study demonstrates the VLM's potential for AI automation in tasks that require pattern recognition and object identification in an urban context. AI have the potential to serve as complementary and collaborative tools for urban analysis, leveraging their strengths in understanding visual patterns. This study contributes to the exploration of the efficiency and scalability of AI visual prediction and provides insights into the reasoning processes that could support automation processes in urban analysis and prediction.