DUDA-Bench: Benchmarking LLM Agents on Multimodal Data-Driven Urban Diagnosis

πŸ“… 2026-10-06
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the reliance on domain experts and limited scalability in urban diagnostics by evaluating the end-to-end diagnostic capabilities of large language model agents driven by multimodal data. Methodologically, we construct a hierarchical, interactive benchmark grounded in multimodal data from twelve cities, formalizing the diagnostic task as a multi-stage workflow encompassing both atomic tasks and complete pipelines. Our findings reveal a significant gap between models’ isolated analytical abilities and their end-to-end diagnostic performance, while validating the critical roles of adaptive planning, evidence integration, and feedback-driven correction mechanisms. Furthermore, we identify limitations such as the cross-stage propagation of evidence gaps. Ultimately, this work establishes a novel paradigm for systematically evaluating the complex analytical capabilities of intelligent agents.
πŸ“ Abstract
Urban diagnosis integrates heterogeneous observations to identify urban problems, localize affected areas, and investigate contributing factors, informing evidence-based urban planning and management. However, its reliance on labor-intensive, case-specific expert workflows limits scalability and reuse, motivating the exploration of agent-based execution. To evaluate this capability, we introduce DUDA-Bench, a hierarchical and interactive benchmark that formalizes data-driven urban diagnosis as a multi-stage agent workflow. It comprises 86 atomic and 22 workflow tasks spanning four analytical stages, grounded in multimodal data from 12 cities covering five urban problem types. Evaluations of seven backbone models and five agent systems reveal a substantial gap between isolated analytical competence and end-to-end diagnosis, with system benefits varying across backbones. Trajectory analysis shows that unresolved evidence gaps propagate across stages, while successful recovery involves revising assumptions and actions using feedback. These findings highlight limitations in coordinating analytical capabilities across stages, particularly adaptive planning, evidence integration, and verification. More broadly, DUDA-Bench provides a framework for translating expert analytical workflows into hierarchical agent tasks and process-aware evaluation, supporting systematic assessment of end-to-end analytical capabilities.
Problem

Research questions and friction points this paper is trying to address.

Urban Diagnosis
LLM Agents
Multimodal Data
Benchmark
Agent Workflow
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM Agents
Multimodal Benchmark
Urban Diagnosis
Agent Workflow
Process-aware Evaluation
Yizhi Song
Yizhi Song
Research Scientist, Bytedance / Tiktok
Image generationGenerative AIMLLMDiffusion
Hang Ni
Hang Ni
HKUST(GZ)
Spatiotemporal Data MiningUrban Intelligence
W
Weijia Zhang
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China
H
Hao Liu
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China