🤖 AI Summary
This study addresses the unclear incremental value of street view imagery over existing data sources for urban perception using vision-language models (VLMs), as well as the accuracy-centric bias in conventional evaluations. By employing multi-source data fusion and spatial distance sensitivity analysis, this work compares image-based predictions with established urban datasets across seven attributes, systematically investigating how image visibility and data coverage influence predictive performance. The findings reveal that the utility of imagery depends critically on attribute visibility and the coverage rate of existing data; specifically, images demonstrate advantages only for certain attributes such as building typology, while non-image data outperform them in other scenarios. Furthermore, this research contributes OpenFACADES, a newly annotated dataset released to support future investigations in VLM-driven urban analytics.
📝 Abstract
Street-view imagery is increasingly analysed with vision-language models (VLMs) to infer urban attributes, but predictive accuracy alone does not show how much a photograph contributes beyond data already available for the same place. Using three VLMs, we compare image-based predictions with existing urban data for seven attributes drawn from five public resources. Each urban unit is evaluated with imagery, with location or text context, and against non-image predictions from nearby observations or public records. We also replace images and add conflicting records to test which source the predictions follow. Nearby observations or public records matched or exceeded image-only predictions for road damage, curb ramps, population, and house price. Images were more informative for building type, building function, and the floor count of low-rise buildings. For floor count, the advantage of images over nearby OpenStreetMap labels increased by 5.7 percentage points per doubling of distance to the nearest labelled building and declined for tall buildings whose rooflines often fell outside the frame. When images and records disagreed, predictions usually moved towards the supplied record. Released OpenFACADES floor annotations, generated with OpenStreetMap floor values as input, showed the same dependence: their agreement with the reference increased with building height, whereas that of image-only reruns decreased. The value of street-view imagery therefore depends on whether an attribute is visible and how well the place is already covered by existing data. Because machine-derived labels are often reused as references, these comparisons also bear on how urban datasets are documented and evaluated.