🤖 AI Summary
This study investigates whether richer visual representations necessarily improve the alignment between computational assessments of urban environmental engagement and human perception. Leveraging 61 first-person urban walking videos segmented into over 50,000 ten-second clips, we evaluate four modalities—full video, temporally averaged images (TAIs), audio embeddings, and textual semantics—using Spearman correlation analysis, binary classification models, and two-alternative forced-choice experiments on Amazon MTurk. Results demonstrate that TAIs perform comparably to, and often better than, full video across multiple classifiers and thresholds. Human subject experiments further reveal that TAIs achieve recognition accuracy on par with full video, while text performs poorly and audio approaches chance-level performance. These findings suggest that perception-driven temporal compression can effectively substitute for full video encoding, highlighting a functional divergence between dynamic and static scene representations in urban perception tasks.
📝 Abstract
We examine whether richer visual representations yield more human-aligned measures of urban engagement, using 61 first-person city-walk videos from YouTube segmented into over 50,000 ten-second clips and represented across four modalities: spatiotemporal video features, temporally averaged images (TAIs), audio embeddings, and text-based semantic descriptions. Spearman correlation analysis reveals the expected ordering along the temporal-richness continuum, with video features showing the strongest continuous alignment. However, this ordering breaks down under binary classification of high- versus low-engagement moments (the paradigm most commonly used to train perceptual scoring models), where TAIs consistently match or outperform video across most classifiers and quantile thresholds. An independent two-alternative forced-choice study on Amazon Mechanical Turk confirms that this parity reflects human judgment: participants identified engaging moments with comparable accuracy from TAIs and full video clips, while text performed substantially worse and audio remained near chance. Gap analysis reveals a functional dissociation: video features are advantaged in activity-driven scenes with dynamic content, whereas TAIs better align with human judgments in composition-driven scenes dominated by stable spatial structure. These findings challenge the assumption that richer representations are inherently more human-aligned, and suggest that perceptually grounded temporal compression can be a principled alternative to full video encoding.