🤖 AI Summary
This study addresses the high cost of urban intervention evaluation and the limitation of existing methods that merely monitor indicators without generating actionable improvement plans. To this end, we propose VIDA-Geo, a multi-agent system that shifts the paradigm from identifying to discovering interventions. By leveraging black-box indicator models alongside diffusion-based image inpainting, VIDA-Geo constructs a digital twin environment in which segmentation, inpainting, and scoring agents collaboratively explore the image intervention space. The system further incorporates perceptual realism and policy alignment to support expert-in-the-loop decision-making. Experimental results demonstrate that our approach outperforms baselines across eight metrics, achieving a twofold improvement in both perceptual quality and policy alignment scores, thereby effectively generating diverse candidate intervention strategies.
📝 Abstract
Urban environments are shaped by design choices with long-term implications for health, safety, and quality of life, yet evaluating proposed interventions remains costly, time-consuming, and often impractical. Existing geospatial vision methods largely focus on monitoring urban indicators from aerial and street-view imagery, rather than proposing interventions and estimating their effects on such indicators. Moving beyond recognition, we introduce the problem of discovering interventions that improve target indicators for a given aerial or street-view image. We argue that a black-box indicator model, combined with a generative editing model, can serve as an implicit digital twin for testing intervention hypotheses. We present VIDA-Geo , a multi-agent system that explores this intervention space by coordinating segmentation, diffusion-based inpainting, and indicator scoring models to produce interventions that are both perceptually realistic and aligned with real-world policies. We evaluate our system on 8 indicators across aerial and street-view imagery, measuring changes in factors such as perceived safety and greenery. Our approach outperforms existing baselines in many cases, achieving up to 2X higher perceptual quality and policy alignment scores. Finally, our model provides users with multiple candidate interventions, supporting an expert city-planner-in-the-loop workflow.