Mapping the City Through the Lens of Language Models

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the implicit and unexamined assumptions about urban scale, form, and infrastructure embedded in language models when referencing cities, for which no systematic quantitative framework previously existed. The authors propose the first empirically verifiable framework for assessing urban typicality by anonymizing real city data and leveraging multiple open-source language models to assign probabilistically constrained scores across 40 urban indicators. The methodology integrates reliability filtering, phylogeny-aware aggregation, multi-population weighting, and validation via independently replicated samples. Results reveal a consistent model preference for cities characterized by large developed area, rapid growth, dense infrastructure, and compact morphology. Notably, after controlling for size and development level, geographic variation diminishes substantially, and replication outcomes remain highly consistent, demonstrating a strong association between urban typicality and model preferences.
📝 Abstract
Language models often complete an underspecified reference to a city with unstated assumptions about urban size, form, infrastructure, environment, and function. We measure those assumptions without naming places. Ten open-weight checkpoints rate anonymized profiles derived from real morphological urban centres across 40 audited indicators and seven domains. The design combines constrained probability-based ratings, prespecified reliability screens, lineage-aware aggregation, multiple population weightings, an independent replication sample, and whole-profile validation. The clearest shared tendency favours urban profiles with larger developed area, faster recent growth, greater mapped infrastructure and non-residential capacity, and less sparse form. Most eligible directions recur in the replication data, and direct ratings of complete profiles show moderate agreement with the indicator-wise construction. Geographic differences shrink after accounting for city scale and development, while reliably measured paired tasks indicate that typicality and desirability are often closely aligned. The framework makes an otherwise vague notion of what models regard as an ordinary city empirically traceable. The resulting evidence delineates a shared yet model-dependent portrait of the city through the lens of language models.
Problem

Research questions and friction points this paper is trying to address.

language models
urban representation
city typology
model bias
urban morphology
Innovation

Methods, ideas, or system contributions that make the work stand out.

language models
urban morphology
anonymized profiling
model bias
empirical validation