🤖 AI Summary
This study addresses the lack of systematic evaluation of spatial intelligence and audio-visual reasoning capabilities at the urban scale. Hosted as an ECCV challenge comprising four tracks, this work introduces two novel city-level video benchmarks, KilometerAudio and KilometerVision, and conducts comprehensive evaluations utilizing multimodal models, agent-based pipelines, and video question-answering techniques. By consolidating the winning solutions, the project reveals that current standalone models remain inadequate for complex spatiotemporal reasoning tasks, demonstrating that high-performance inference still necessitates computationally expensive agent-based pipelines. Ultimately, this work establishes a critical benchmark and identifies promising new directions for advancing research in urban-scale multimodal understanding.
📝 Abstract
Continuing the Perception Test challenge series, we organised the fourth edition as a workshop at the European Conference on Computer Vision (ECCV) 2026 in Malmö, Sweden. This edition focused on spatial intelligence and featured four different tracks: unified multiple-choice videoQA and grounded videoQA from the original Perception Test benchmark, alongside two new tracks based on city-scale walking-tour videos (KilometerAudio and KilometerVision). In this report, we describe the new benchmarks used for the city-scale tracks and summarise the winning solutions across all tracks, including a generalist model that competed across all tracks with satisfactory performance. The winning solutions in the newly added city-scale tracks demonstrated that complex spatial and multimodal reasoning can be solved by expensive agentic pipelines, but remains difficult for multimodal models used standalone.