π€ AI Summary
This study addresses the reliance on domain experts and limited scalability in urban diagnostics by evaluating the end-to-end diagnostic capabilities of large language model agents driven by multimodal data. Methodologically, we construct a hierarchical, interactive benchmark grounded in multimodal data from twelve cities, formalizing the diagnostic task as a multi-stage workflow encompassing both atomic tasks and complete pipelines. Our findings reveal a significant gap between modelsβ isolated analytical abilities and their end-to-end diagnostic performance, while validating the critical roles of adaptive planning, evidence integration, and feedback-driven correction mechanisms. Furthermore, we identify limitations such as the cross-stage propagation of evidence gaps. Ultimately, this work establishes a novel paradigm for systematically evaluating the complex analytical capabilities of intelligent agents.
π Abstract
Urban diagnosis integrates heterogeneous observations to identify urban problems, localize affected areas, and investigate contributing factors, informing evidence-based urban planning and management. However, its reliance on labor-intensive, case-specific expert workflows limits scalability and reuse, motivating the exploration of agent-based execution. To evaluate this capability, we introduce DUDA-Bench, a hierarchical and interactive benchmark that formalizes data-driven urban diagnosis as a multi-stage agent workflow. It comprises 86 atomic and 22 workflow tasks spanning four analytical stages, grounded in multimodal data from 12 cities covering five urban problem types. Evaluations of seven backbone models and five agent systems reveal a substantial gap between isolated analytical competence and end-to-end diagnosis, with system benefits varying across backbones. Trajectory analysis shows that unresolved evidence gaps propagate across stages, while successful recovery involves revising assumptions and actions using feedback. These findings highlight limitations in coordinating analytical capabilities across stages, particularly adaptive planning, evidence integration, and verification. More broadly, DUDA-Bench provides a framework for translating expert analytical workflows into hierarchical agent tasks and process-aware evaluation, supporting systematic assessment of end-to-end analytical capabilities.